Targeting deaminases and base editing using the same

CN122521744APending Publication Date: 2026-08-07INST FOR BASIC SCI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INST FOR BASIC SCI
Filing Date
2021-09-17
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

然而,这些核酸酶不能诱导或恢复线粒体中的特定突变:与细胞核中的DNA双链断裂不同,线粒体中的DNA双链断裂不能通过非同源末端连接或同源重组有效修复

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The present invention relates to targeted deaminases and base editing using the same, including isolated cytosine or adenine deaminases or variants thereof, non-toxic full-length cytosine deaminases or variants thereof, further including fusion proteins of the deaminases or variants thereof, compositions for base editing, and methods for editing bases by using the deaminases or variants thereof, the fusion proteins, the compositions.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of Chinese Patent Application No. 202180076508.8, filed on September 17, 2021, entitled "Targeted Deaminase and Base Editing Using the Same". Technical Field

[0002] This invention relates to isolated forms of cytosine or adenine deaminases or variants thereof, non-toxic full-length cytosine deaminases or variants thereof, fusion proteins comprising said deaminases or variants thereof, compositions for base editing, and methods for editing bases using said deaminases or variants thereof, said fusion proteins, or said compositions. Background Technology

[0003] Fusion proteins, in which DNA-binding proteins and deaminases are fused together, enable targeted nucleotide substitutions or base editing in the genome without generating DNA double-strand breaks (DSBs), correcting point mutations that cause genetic disorders, or performing targeted single nucleotide conversions to introduce desired single nucleotide mutations into prokaryotic cells as well as human and other eukaryotic cells.

[0004] Unlike nucleases such as CRISPR-Cas9 that induce small insertions or deletions (insertion-deletion) at the target site, deaminase fusion proteins convert single bases within a window of several nucleotides at the target site. Therefore, point mutations that cause genetic diseases or single nucleotide polymorphisms (SNPs) can be edited in cultured cells, animals, and plants.

[0005] Examples of fusion proteins in which DNA-binding proteins and deaminases are fused together can include: 1) base editors (BEs), which include catalytically deficient Cas9 (dCas9) or D10A Cas9 nickase (nCas9) derived from Streptococcus pyogenes and rAPOBEC1 as a cytosine deaminase derived from rats; 2) target-AIDs, which include dCas9 or nCas9 and PmCDA1 (an activation-induced cytidine deaminase (AID) ortholog derived from sea lamprey) or human AID; 3) CRISPR-X, which includes MS2 RNA hairpin-linked sgRNA and dCas9 to recruit overactive AID variants fused to MS2 binding proteins, and so on.

[0006] Programmable genome editing tools, such as ZFN (zinc finger nucleases), TALEN (transcription activator-like effector nucleases), CRISPR (regularly spaced clustered short palindromic repeats) systems, and base editors composed of CRISPR-associated protein 9 (Cas9) variants and nucleobase deaminase proteins, have been developed for plant genetic research and for improving crop traits by altering base sequences. However, these tools are not suitable for editing the DNA sequences of plant organelles, including mitochondria and chloroplasts, primarily because it is difficult to deliver guide RNA to organelles or co-express two compounds within them. Plant organelles encode essential genes required for photosynthesis. Methods or tools for editing organelle genes are indispensable for functional studies of organelle genes or for improving crop yield and traits. For example, targeted mutations in the mitochondrial atp6 gene can lead to male sterility, a useful trait for seed production, and specific point mutations in the 16S rRNA gene of the chloroplast genome can lead to antibiotic resistance.

[0007] Bacterial toxin DddA tox It is an enzyme domain derived from a bacterial toxin of *Burkholderia cenocepacia*, and it is capable of deaminating cytosine in double-stranded DNA. As an example of a deaminase, DddA... tox It is cytotoxic; therefore, to avoid toxicity in host cells, DddA is... tox It splits into two inactive halves, each of which fuses with a DNA-binding protein in a DddA-derived cytosine base editor (DdCBE). As the DNA-binding protein binds the two inactive halves together, the functional deaminase reassembles at the target DNA site.

[0008] In principle, this deaminase reaction is only activated when the two inactive halves are in close proximity to the target DNA via a DNA-binding protein. Therefore, cytosine-to-thymine (C-to-T) base editing is induced in the spacer region between the binding sites of the two DNA-binding proteins. The two inactive forms fused with the TALE (transcription activator-like effector) DNA-binding array become functional when they aggregate together via TALE-DNA interactions. C-to-T editing is typically induced in a 14-18 base region between the two TALE binding sites. However, DddA... tox Split-body systems have many limitations in experiments.

[0009] Full-length encoding DddA tox The gene cannot be cloned in E. coli due to its toxicity. Cloning is only possible when the DddA inhibitor gene is co-expressed in E. coli.

[0010] On the other hand, mitochondrial DNA plays a crucial role in cellular respiration, facilitated by the mitochondrial oxidative phosphorylation (OXPHOS) mechanism. Because the OXPHOS mechanism is essential for survival, mutations in mitochondrial DNA can cause severe functional impairment in many organs and muscles, particularly in tissues with high energy demands. In many human mitochondrial diseases, wild-type mitochondrial DNA coexists with mutant mitochondrial DNA containing single-base mutations, resulting in a heterogeneous state of mitochondrial DNA. The balance between mutant and wild-type mitochondrial DNA determines the development of clinically symptomatic mitochondrial diseases. In vitro and in vivo, programmable nucleases have been used to cleave and thereby remove mutant mitochondrial DNA without cleaving wild-type mitochondrial DNA. However, these nucleases cannot induce or restore specific mutations in mitochondria: unlike DNA double-strand breaks in the cell nucleus, DNA double-strand breaks in mitochondria cannot be effectively repaired by non-homologous end joining or homologous recombination.

[0011] Mitochondrial base editing can be used to create models of various diseases or to produce therapeutic agents for treating these diseases. In this regard, there is an increasing need to develop highly efficient mitochondrial base editing enzymes.

[0012] In this technical context, we have accomplished the present invention by demonstrating that DNA can be corrected by using a desired CBE (cytosine base editor) or ABE (adenine base editor) or by using a novel full-length deaminase that is non-cytotoxic, wherein the CBE or ABE is generated by substituting deaminase residues to reduce non-selective base editing. Summary of the Invention

[0013] One object of the present invention is to provide a fusion protein comprising a DNA-binding protein and an isolated form of cytosine or adenine deaminase or a variant thereof, or a non-toxic full-length cytosine deaminase or a variant thereof.

[0014] Another object of the present invention is to provide a nucleic acid encoding a fusion protein.

[0015] Another object of the present invention is to provide a composition for base editing, said composition comprising a fusion protein or nucleic acid.

[0016] Another object of the present invention is to provide a base editing method comprising treating cells with the composition.

[0017] To achieve the above objectives, the present invention provides a fusion protein comprising (i) a DNA-binding protein and (ii) a first cleavage and a second cleavage derived from cytosine deaminase or a variant thereof, wherein each of the first cleavage and the second cleavage is fused to the DNA-binding protein.

[0018] Furthermore, the present invention provides a fusion protein comprising (i) a DNA-binding protein and (ii) a non-toxic full-length cytosine deaminase derived from cytosine deaminase or a variant thereof.

[0019] Furthermore, the present invention provides a fusion protein comprising (i) a DNA-binding protein, (ii) a cytosine deaminase or a variant thereof, and (iii) an adenine deaminase, wherein the cytosine deaminase or a variant thereof comprises (a) a non-toxic full-length cytosine deaminase or (b) a first cleavage and a second cleavage derived from a cytosine deaminase or a variant thereof, each of the first cleavage and the second cleavage being fused to the DNA-binding protein.

[0020] In addition, the present invention provides a nucleic acid encoding a fusion protein.

[0021] Furthermore, the present invention provides a composition for base editing, the composition comprising the fusion protein or the nucleic acid.

[0022] Furthermore, the present invention provides a composition for base editing in eukaryotic cells, the composition comprising the fusion protein or the nucleic acid.

[0023] Furthermore, the present invention provides a composition for base editing in plant cells, the composition comprising the fusion protein or the nucleic acid and a nuclear localization signal (NLS) peptide or a nucleic acid encoding thereon.

[0024] Furthermore, the present invention provides a composition for base editing in plant cells, the composition comprising the fusion protein or the nucleic acid and a chloroplast transport peptide or a nucleic acid encoding the fusion protein or the nucleic acid.

[0025] Furthermore, the present invention provides a composition for base editing in plant cells, the composition comprising the fusion protein or the nucleic acid and a mitochondrial targeting signal (MTS) or a nucleic acid encoding the mitochondrial targeting signal (MTS).

[0026] In some cases, the present invention also provides a composition for base editing in plant cells, said composition further comprising a nuclear output signal or nucleic acid encoding thereon.

[0027] Furthermore, the present invention provides a method for base editing in plant cells, the method comprising treating the plant cells with the composition.

[0028] Furthermore, the present invention provides a method for base editing in plant cells, the method comprising treating the plant cells with the fusion protein or the nucleic acid containing a nuclear localization signal (NLS) peptide or encoding the same.

[0029] Furthermore, the present invention provides a method for base editing in plant cells, the method comprising treating the plant cells with the fusion protein or the nucleic acid containing a chloroplast transport peptide or encoding the chloroplast transport peptide.

[0030] Furthermore, the present invention provides a method for base editing in plant cells, the method comprising treating the plant cells with the fusion protein or the nucleic acid containing a mitochondrial targeting signal (MTS) or encoding the MTS.

[0031] Furthermore, the present invention provides a composition for base editing in animal cells, the composition comprising the fusion protein or the nucleic acid containing a nuclear localization signal (NLS) peptide or encoding the same.

[0032] Furthermore, the present invention provides a composition for base editing in animal cells, the composition comprising the fusion protein or the nucleic acid and a mitochondrial targeting signal (MTS) or a nucleic acid encoding the mitochondrial targeting signal (MTS).

[0033] In some cases, the present invention also provides a composition for base editing in animal cells, said composition further comprising a nuclear output signal or nucleic acid encoding thereon.

[0034] Furthermore, the present invention provides a method for base editing in animal cells, the method comprising treating the animal cells with the composition.

[0035] Furthermore, the present invention provides a method for base editing in animal cells, the method comprising treating the animal cells with the fusion protein or the nucleic acid containing a nuclear localization signal (NLS) peptide or encoding the nucleic acid.

[0036] Furthermore, the present invention provides a method for base editing in animal cells, the method comprising treating the animal cells with the fusion protein or the nucleic acid containing a mitochondrial targeting signal (MTS) or encoding the MTS.

[0037] Furthermore, the present invention provides a composition for A-to-G base editing in prokaryotic or eukaryotic cells, the composition comprising the fusion protein or nucleic acid encoding it, wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-associated protein, and a cytosine deaminase or a variant thereof derived from bacteria and specific for double-stranded DNA.

[0038] Furthermore, the present invention provides a composition for A-to-G base editing in prokaryotic or eukaryotic cells, the composition comprising the fusion protein or nucleic acid encoding it, wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-associated protein, and the fusion protein contains a cytosine deaminase or a variant thereof derived from bacteria and specific for double-stranded DNA. The DNA-binding protein is fused to the N-terminus and C-terminus of the cytosine deaminase or a variant thereof. Similarly, the DNA-binding protein is also fused to the N-terminus and C-terminus of the adenine deaminase of the fusion protein. In the context of a fusion protein comprising a DNA-binding protein, a cytosine deaminase or a variant thereof, and adenine deaminase, the adenine deaminase may be located at the N-terminus or C-terminus of the cytosine deaminase within the fusion protein, or may exist as a separate protein independent of other DNA-binding proteins.

[0039] Furthermore, the present invention provides a composition for C-to-T base editing in prokaryotic or eukaryotic cells, the composition comprising the fusion protein or the nucleic acid encoding it and a uracil glycosylase inhibitor (UGI), wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-associated protein, and the cytosine deaminase or a variant thereof is a non-toxic full-length cytosine deaminase, and the cytosine deaminase or a variant thereof in the fusion protein is derived from bacteria and is specific for double-stranded DNA.

[0040] Furthermore, the present invention provides a composition for C-to-T base editing in prokaryotic or eukaryotic cells, the composition comprising a fusion protein or nucleic acid encoding thereon and a UGI, wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-associated nuclease, and the cytosine deaminase of the fusion protein or a variant thereof is a mitotic cytosine deaminase comprising a first mitotic fragment and a second mitotic fragment, and the cytosine deaminase of the fusion protein or a variant thereof is derived from bacteria and is specific for double-stranded DNA.

[0041] Furthermore, the present invention provides a method for performing A to G base editing in prokaryotic or eukaryotic cells, the method comprising treating the prokaryotic or eukaryotic cells with the fusion protein or the nucleic acid encoding it, wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-associated nuclease, and the fusion protein is a cytosine deaminase or a variant thereof derived from bacteria and specific for double-stranded DNA.

[0042] Furthermore, this invention provides a method for performing A-to-G base editing in prokaryotic or eukaryotic cells, the method comprising treating the prokaryotic or eukaryotic cells with a fusion protein or a nucleic acid encoding thereon, wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-associated nuclease.

[0043] The fusion protein's cytosine deaminase or a variant thereof is derived from bacteria and is specific for double-stranded DNA.

[0044] The fusion protein contains a cytosine deaminase or a variant thereof derived from bacteria and specific for double-stranded DNA. The DNA-binding protein is fused to the N-terminus and C-terminus of the cytosine deaminase or a variant thereof. Similarly, the DNA-binding protein is also fused to the N-terminus and C-terminus of the adenine deaminase of the fusion protein. In the context of a fusion protein comprising a DNA-binding protein, a cytosine deaminase or a variant thereof, and adenine deaminase, the adenine deaminase may be located at the N-terminus or C-terminus of the cytosine deaminase within the fusion protein, or it may exist as a separate protein independent of the other DNA-binding proteins.

[0045] Furthermore, the present invention provides a method for C-to-T base editing in prokaryotic or eukaryotic cells, the method comprising treating the prokaryotic or eukaryotic cells with a fusion protein or a nucleic acid encoding thereon and a UGI, wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-associated nuclease, and the fusion protein is a cytosine deaminase or a variant thereof derived from bacteria and specific for double-stranded DNA. Attached Figure Description

[0046] Figure 1 The results of ZFD optimization using the pTarget plasmid are shown.

[0047] Figure 1 Inset a shows the ZFD construct in which the split DddAtox half is fused to the C-terminus of ZFP (zinc finger protein: type C). Figure 1Inset Figure b shows the optimization of the ZFD platform using the pTarget library, where the pTarget plasmid contains spacer regions ranging from 1 to 24 bp in size (indicated in red) and ZFP DNA binding sites (indicated in green), and the ZFD construct contains AA adapters of various lengths (indicated in yellow and orange) and different DddAtox splitting sites and orientations (indicated in blue), as well as Figure 1 Small image c and Figure 1 Inset d shows ZFD activity measured at the target site in the pTarget library to examine Figure 1 The role of the variables described in b, where ZFD pairs with the same (c) or different (d) lengths of adapters in the left and right ZFDs were tested, the base editing frequency was measured by targeted depth sequencing of the relevant regions of the pTarget plasmid, and the data are expressed as the mean ± standard error of the mean (sem) from n = 2 biologically independent samples.

[0048] Figure 2 The results show the ZFD efficiency in the pTarget plasmid with various adapters.

[0049] Figure 2 Inset a shows the edit frequencies from C / G to non-C / G for various pTarget spacers of length 1–24 bp depicted in the heatmap, where various ZFD configurations were tested, including various types of joints between ZFP and split DddAtox regions where DddAtox is the split region. Figure 2 Inset chart b shows the overall activity for each ZFD pair, where the nomenclature used for the x-axis indicates the left ZFD at the bottom and the right ZFD at the top, and Figure 2 Inset c shows the base editing efficiency depending on the spacer length, where “AA” represents the number of amino acids in the linker, and the data are expressed as the mean ± standard error of the mean (sem) from n = 2 biologically independent samples.

[0050] Figure 3 Results demonstrating the efficiency of ZFD with 24AA connectors and different connectors are shown.

[0051] This shows the effect of ZFD connector length on editing efficiency from C / G to non-C / G in the heatmap, where the left ZFD of the ZFD pair is fixed with a 24AA connector and the right ZFD contains a connector of variable length, or vice versa. The error bars are the standard error (sem) of the mean of n = 2 biologically independent samples.

[0052] Figure 4 The results demonstrate the activity of ZFD in targeting the cell nucleus in vivo.

[0053] Figure 4 Figure a shows the conformation of the ZFD targeting nuclear DNA, where the splitting DddAtox half fuses to the C-terminus (C-type) or N-terminus (N-type) of the ZFP. The ZFD pair is designed in CC or NC conformation, consisting of a C-type left ZFD and a C-type right ZFD, or an N-type left ZFD and a C-type right ZFD, respectively. Figure 4 Inset b shows the frequency of base editing induced by ZFD at nuclear DNA target sites in HEK 293T cells. Data are expressed as mean ± sem from n = 3 biologically independent samples. Figure 4 Small image c- Figure 4 Inset f shows the ZFD-induced base editing efficiency at each base position within the spacer at the target sites NUMBL (c), INPP5D-2 (d), TRAC-CC (e), and TRAC-NC (f) in HEK 293T cells. Data are expressed as mean ± sem from n = 3 biologically independent samples. Figure 4 Inset g shows the frequency of ZFD-induced base editing in K562 cells after electroporation or direct delivery of ZFD protein or ZFD-encoding plasmid, where ZFD protein with one or four NLS was tested, and left and right ZFD were used in equal molar amounts, electroporation was performed using Amaxa 4D-Nucleofector, and for direct delivery, K562 cells were incubated with cell culture medium containing left and right ZFD protein and treated in the same manner once (1x) or twice (2x), and data are presented as mean ± sem from n = 2 biologically independent samples;

[0054] Figure 5 The conformation of ZFD targeting nuclear DNA is schematically shown. Figure 5 Small image a- Figure 5 Inset d shows four possible ZFD configurations, where the NC and CN configurations are structurally identical, but the left and right ZFD configurations are of different types;

[0055] Figure 6 Results confirming the insertion / deletion rate of ZFDs targeting the cell nucleus in vivo are shown, with all tested ZFDs producing insertions / deletions at a frequency of less than 0.4%, and data are presented as mean ± sem from n = 3 biologically independent samples;

[0056] Figure 7 The results of testing the activity of recombinant ZFD protein in vitro are shown.

[0057] Figure 7Figure a shows the purification of ZFD pairs targeting the TRAC site. GST-labeled proteins were purified from *E. coli* cell lysates using glutathione agarose beads. The purification process was monitored using polyacrylamide gel electrophoresis, and the gels were stained with Coomassie blue. Lane 1 shows the molecular weight markers; lane 2 shows the cell sample (whose protein expression was not IPTG-induced); lane 3 shows the cell sample (whose protein expression was IPTG-induced); lane 4 shows the soluble fraction after sonication; lane 5 shows the insoluble fraction after sonication; lane 6 shows the column flow fraction; lane 7 shows the wash fraction; and lane 8 shows the elution fraction. The sizes of representative markers are shown on the left, and the red boxes indicate the ZFD proteins. Figure 7 Inset image b shows the left and right ZFD binding sites, with red arrows indicating possible sites for ZFD-induced deamination. Figure 7 Inset figure c shows the ZFD activity for PCR amplicon containing the TRAC site, where the TRAC-NC ZFD deamination of cytosine to produce uracil (indicated in red), followed by USER enzyme cleavage of uracil to form a cleft (indicated by a red triangle), and... Figure 7 Inset d shows untreated PCR amplicon (left) and PCR amplicon treated with ZFD (right) as analyzed by agarose gel electrophoresis.

[0058] Figure 8 The various conformations of ZFD targeting mitochondrial DNA are schematically shown. Figure 8 Small image a- Figure 8 Inset d shows four possible mito ZFD configurations, where the existing ZFD's NLS is replaced by MTS and NES, and the NC and CN configurations are structurally identical, but the left and right ZFD configurations are of different types;

[0059] Figure 9 The results demonstrate the efficiency of mitoZFD in mitochondrial gene base editing.

[0060] Figure 9 Inset a shows the frequency of base editing in mtDNA induced by mitoZFD and TALE-DdCBE in HEK 293T cells. Data are expressed as mean ± standard error of mean (sem) from n = 2 biologically independent samples. Figure 9 Small image b- Figure 9Inset g shows the base editing efficiency induced by mitoZFD at each base position within the spacer at the target sites ND2(b), ND4L(c), COX2(d), ND6(e), and ND1(f) in HEK 293T cells, and the base editing efficiency induced by TALE-DdCBE at the target site ND1(g). Data are expressed as mean ± standard error of mean (sem) from n = 2 biologically independent samples. Figure 9 Inset h shows a comparison of DNA and amino acid changes in the ND1 gene introduced by mitoZFD and TALE-DdCBE, where the frequency (%) of sequencing reads of each mutant allele was measured by targeted deep sequencing, and the spacer subregions of the ZFD pair and the TALE-DdCBE pair are represented by blue dashed lines.

[0061] Figure 10 Results confirming the base editing efficiency of a single-cell-derived clonal population isolated from HEK293T cells treated with MT-ZFD are shown.

[0062] Single-cell-derived clones were obtained for allele analysis, and the C / G to non-C / G editing frequencies in individual single-cell-derived clones were determined by targeted deep sequencing. Figure 10 Figure a shows single-cell-derived clones of a HEK 293T cell population treated with mitoZFD targeting ND1. Figure 10 Inset Figure b shows single-cell-derived clones of a HEK293T cell population treated with mitoZFD targeting ND2, and Figure 10 Inset Figure c shows single-cell-derived clones of an untreated HEK 293T cell population, with ZFP binding sites indicated in green and high editing frequencies of clones subjected to mitoZFD-induced editing indicated in red.

[0063] Figure 11 Results confirming the base editing efficiency of a single-cell-derived clonal population isolated from HEK293T cells treated with MT-ZFD are shown.

[0064] It provides allelic analysis of single-cell-derived clones with high base editing frequency. The table shows the amino acids changed by base editing in ND1, and in the top reference sequence, red letters indicate spacers, and in the alleles, red letters indicate changes in the amino acid sequence (* indicates stop codons).

[0065] Figure 12a and Figure 12bResults confirming the base editing efficiency of single-cell-derived clonal populations isolated from HEK293T cells treated with MT-ZFD are shown, with allele analysis of single-cell-derived clones having high base editing frequencies provided. The table shows the amino acids altered by base editing in ND2, and in the top reference sequence, red letters indicate spacers, and in the alleles, red letters indicate changes in amino acid sequences.

[0066] Figure 13 Results confirming the base editing efficiency achieved through the combination of ZFD and TALE-DdCBE are shown.

[0067] Figure 13 Figure a shows the DNA sequence of the binding region of the mitoZFD and TALE-DdCBE pair, with the site recognized by TALE-DdCBE highlighted in green and the site recognized by mitoZFD highlighted in blue. The upper sequence represents the mtDNA heavy chain, and the lower sequence represents the mtDNA light chain. Figure 13 Inset b shows the frequencies of cytosine edited via ZFD, TALE-DdCBE, and ZFD / DdCBE hybridization pairs. Data were obtained using targeted deep sequencing and are expressed as mean ± standard error of mean (sem) from n = 2 biologically independent samples. Figure 13 Inset c shows a heatmap of base editing activity at each base position, with red boxes indicating spacer regions for each conformation and blue arrows indicating the location of mtDNA;

[0068] Figure 14 The results demonstrate the editing efficiency of mRNA relative to plasmid at different ZFD concentrations, where the mitochondrial genome-wide targeting specificity of mitoZFD targeting ND1 varies with the concentration of the mRNA or plasmid encoding the ZFD. The results show the on-target and off-target base editing frequencies determined by whole-mtDNA sequencing. HEK 293T cells transfected with plasmids or mRNA encoding mitoZFD targeting ND1 at specified concentrations are plotted as dots in the figure. Red arrows indicate target sites, red dots represent the base editing frequency at the target sites, and gray dots represent SNPs also present in the controls. Data are expressed as mean ± standard error (sem) of the mean from n = 2 biologically independent samples.

[0069] Figure 15 The results demonstrate that the editing efficiency depends on mRNA relative to plasmid at different ZFD concentrations. Figure 15Inset a shows ZFD binding at the top ND1 site, where the ZFD binding site is indicated in green, the target cytosine within the spacer is indicated in red, and the sequence from the ND1 site is shown. Figure 14 The mid-target activity was determined from whole-mtDNA sequencing data, and this activity decreased with decreasing amounts of plasmid or mRNA encoding mitoZFD. Figure 15 Inset figure b shows the number of C / G sites edited at a frequency of >1% for each plasmid or mRNA amount, and Figure 15 Inset Figure c shows the average C / G to T / A editing frequency of all C / G in the mitochondrial genome depending on the concentration of each plasmid or mRNA. The data are expressed as mean ± standard error of mean (sem) from n = 2 biologically independent samples.

[0070] Figure 16 The results demonstrate the editing efficiency of mRNA relative to plasmid at different ZFD concentrations, showing the mid- and off-target base editing frequencies determined by whole-mtDNA sequencing. HEK 293T cells transfected with plasmids or mRNA encoding mitoZFD targeting ND2 at specified concentrations are plotted as dots in the graph. Red arrows indicate target sites, red dots represent the base editing frequency at the target sites, and gray dots represent SNPs also present in the controls. Data are expressed as mean ± standard error (sem) of the mean from n = 2 biologically independent samples.

[0071] Figure 17 The results demonstrate that the editing efficiency depends on mRNA relative to plasmid at different ZFD concentrations.

[0072] Figure 17 Inset a shows ZFD binding at the top ND2 site, where the ZFD binding site is indicated in green, and the target cytosine within the spacer is indicated in red, showing the effect from... Figure 16 The mid-target activity was determined from whole-mtDNA sequencing data, and the activity decreased with decreasing amounts of plasmid or mRNA encoding mitoZFD. Figure 17 Inset figure b shows the number of C / G sites that undergo base editing at a frequency of >1% for each plasmid or mRNA amount. Figure 17 Inset Figure c shows the average C / G to T / A editing frequency of all C / G in the mitochondrial genome depending on the concentration of each plasmid or mRNA. The data are expressed as mean ± standard error of mean (sem) from n = 2 biologically independent samples.

[0073] Figure 18 The results demonstrate the editing efficiency after constructing the whole mitochondrial sequencing / QQ variant.

[0074] Figure 18 Figure a shows the QQ mitoZFD variant, which contains an R(-5)Q mutation in each zinc finger of the ZFD to eliminate nonspecific DNA contact (if there is no R at position -5 in the zinc finger frame, the nearby K or R is converted to Q). Figure 18 Inset b shows the whole mtDNA sequencing of mitoZFD-treated cells, where in-site and out-of-site editing frequencies are represented by red and black dots, respectively. Data are expressed as mean ± standard error (SEM) of the mean from n = 2 biologically independent samples, showing all base edits from C / G to T / A with efficiencies > 1%. Figure 18 Small image c and Figure 18 Inset d shows the editing efficiency and specificity as a function of the capacity of the delivered ZFD-encoding mRNA. Figure 18 Inset figure c shows the average C / G to T / A editing frequency for all C / G pairs in the mitochondrial genome, and Figure 18 Inset d shows the number of C / G segments edited at a base editing frequency of >1%;

[0075] Figure 19 The Golden Gate assembly system of base editors in plants is shown, and the Golden Gate assembly of cp-DdCBE and mt-DdCBE constructs is schematically shown, wherein for each position in the target sequence, the TALE subarray plasmid is selected from a total set of 424 sequences (= 6 × 64 triplets + 2 × 16 bipartites + 2 × 4 monopartites) and mixed with the desired vector to obtain a plasmid encoding a DdCBE targeting a specific sequence;

[0076] Figure 20 It shows base editing in plant chloroplasts and mitochondria. Figure 20 Small image a, Figure 20 Small image b, Figure 20 Small image c and Figure 20 Inset d shows the frequency and pattern of chloroplast base editing induced by cp-DdCBE in 16S rDNA (a, b) and psbA (c, d), where splitting DdCBE G1333 and G1397 pairs were transfected into lettuce and rapeseed protoplasts. Figure 20 Small picture e and Figure 20 Figure f shows the efficiency and pattern of mitochondrial base editing induced by mt-DdCBE in the ATP6 gene, in which the splitting of DdCBE G1333 and G1397 pairs was transfected into lettuce and rapeseed protoplasts, in Figure 20 Small image a, Figure 20 Small image c and Figure 20In small figure e, the TALE binding region is shown in blue, and cytosine in the spacer is shown in orange. Error bars in all figures represent the mean ± standard deviation of three independent biological replicates. Figure 20 Small image b, Figure 20 Small image d and Figure 20 In figure f, the transformed nucleotides are shown in red, and the percentage of edited alleles (mean ± standard deviation) was obtained from three independent experiments;

[0077] Figure 21 This demonstrates plant organelle DNA editing via DdCBE. Figure 21 Inset a schematically illustrates the mutagenesis of plant organelles. Figure 21 Inset Figure b shows the C∙G to T∙A conversion efficiency in cp-DdCBE-transfected callus cultured in the absence of spectinomycin, including a representative Sanger sequencing chromatogram, where the converted nucleotides are indicated in red on the left, and the arrows indicate the substituted nucleotides in the chromatogram. Figure 21 Inset figure c shows DdCBE-driven plant organelle mutagenesis, where mutant callus exhibits a much higher editing frequency than in simulated callus. Figure 21 Inset d shows the C-to-T conversion frequency induced after transfection of lettuce protoplasts with mRNA encoding cp-DdCBE targeting 16S rDNA. Error bars represent the mean ± sd of n = 3 independent biological replicates. Figure 21 Inset e shows the editing frequency and pattern of spectinomycin-resistant callus at 2.5 months. Figure 21 Inset f shows the C∙G to T∙A conversion efficiency in streptomycin-resistant plants transfected with DdCBE mRNA, obtained using a representative Sanger sequencing chromatogram. Arrows indicate substituted nucleotides in the chromatogram. Scale bar: 1 mm.

[0078] Figure 22 This study compares off-target activities near the target site in lettuce protoplasts transfected with DdCBE plasmids or DdCBE mRNA. Lettuce protoplasts were transfected with plasmids or mRNA encoding the cp-DdCBE pair targeting the chloroplast 16S rRNA gene. Off-target TC to TT editing was detected near the target site. Editing efficiency was measured by targeted depth sequencing seven days post-transfection. Frequencies (mean ± sd) were obtained from three independent experiments. Student's unpaired two-tailed t-test was applied. **P < 0.01; *P < 0.05; NS, not significant (P > 0.05);

[0079] Figure 23This study demonstrates a chloroplast and mitochondrial base editing strategy, in which the preproteins cp-DdCBE and mt-DdCBE each contain either a chloroplast transport peptide (CTP) or a mitochondrial targeting signal (MTS), and are thus translated in plant cells and then transported to chloroplasts and mitochondria. The preproteins cross the outer and inner membranes of the organelles. The CTP and MTS are cleaved by interstitial processing peptidase and mitochondrial processing peptidase, respectively, and then cp-DdCBE and mt-DdCBE (mature proteins) form their final conformations.

[0080] Figure 24 Editing via DdCBE plasmid in lettuce protoplasts over time is shown, with transfected protoplasts collected at each time point and editing efficiency analyzed by targeted deep sequencing. The frequencies (mean ± sd) were obtained from three independent experiments.

[0081] Figure 25 The base editing frequency of the psbB gene is shown, in which the plasmid encoding the cp-DdCBE gene targeting the chloroplast psbB gene (left-G1333-N + right-G1333-C) was transfected into rapeseed protoplasts, and the base editing efficiency in the spacers was analyzed by targeted deep sequencing. The TALE binding region, target cytosine, and transformed nucleotides are represented in blue, orange, and red, respectively, and the frequencies (mean ± standard deviation) were calculated from n = 3 independent experiments.

[0082] Figure 26 The base editing frequency of the mitochondrial RPS14 gene is shown. The plasmid encoding mt-DdCBE targeting the RPS14 gene (left-G1333-N + right-G1333-C) was transfected into rapeseed protoplasts, and the C-to-T conversion efficiency was analyzed by targeted deep sequencing. The TALE binding region, target cytosine, and converted nucleotides are represented in blue, orange, and red, respectively. The frequencies (mean ± standard deviation) were calculated from n = 3 independent experiments.

[0083] Figure 27 (a) shows the base editing efficiency of DdCBE targeting the chloroplast genome in callus, the base editing frequency and pattern at the target sites of 16SrDNA and psbA in lettuce and rapeseed callus after 4 weeks of culture, with the nucleotides converted in the spacers indicated in red. Figure 27 (b) shows the base editing efficiency of the mitochondrial genome in the target callus, where the base editing frequency and pattern of DdCBE at the target sites of the ATP6 and RPS14 genes in rapeseed callus were confirmed by targeted deep sequencing, and the nucleotides converted in the target spacer are shown in red.

[0084] Figure 28The results show the frequency and pattern of chloroplast base editing at the target site of 16S rDNA after DdCBE mRNA was transfected into lettuce protoplasts without DNA base editing, wherein the protoplasts were cultured for 7 days and then targeted deep sequencing was performed, and the nucleotides transformed in the target spacer are shown in red.

[0085] Figure 29 The results of gel electrophoresis show that DdCBE mRNA or DNA sequences (where M is a marker) are not present in protoplasts and callus.

[0086] Figure 30 The 16S rDNA mutation screening is shown, with red arrows indicating streptomycin-resistant green callus.

[0087] Figure 31 This indicates that there are no off-target mutations near the DdCBE target site in antibiotic-resistant callus or buds. Figure 31 (a) and Figure 31 (b) shows the off-target activity analyzed by target depth sequencing, with TALE binding sites and spacer regions underlined in green and red, respectively. Figure 31 (a) Shows spectinomycin-resistant callus produced from a culture of lettuce protoplasts transfected with the DdCBE plasmid, and Figure 31 (b) shows buds obtained from streptomycin-resistant buds;

[0088] Figure 32 The analysis results of off-target activity at the five sites most homologous to the target site are shown. The top five candidate off-target sites for DdCBE gene-specific 16S rRNA in the lettuce chloroplast genome were selected, including up to nine mismatches in the TALE binding site. The TALE binding sequence and mismatched nucleotides are shown in blue and red, respectively. The off-target mutation frequency in protoplasts and drug-resistant callus or shoots transfected with DdCBE plasmid or DdCBE mRNA was measured using targeted depth sequencing. The frequencies (mean ± standard deviation) were obtained from three independent experiments.

[0089] Figure 33 The diagram schematically illustrates DdCBE assembly and mitochondrial DNA editing. Figure 33 a shows a one-pot Golden Gate assembly for efficient DdCBE construction, in which a total of 424 sequences (64 three-part arrays × 6 + 16 two-part arrays × 2 + 4 single-part arrays × 2) are mixed with the expression vector to construct the left and right modules for the final plasmid construction, and Figure 33b schematically shows the interaction between DdCBE and the target gene ND5 in mouse mitochondrial DNA, where the TALE binding site is shown in gray, the base editing site is shown in black, and the corresponding repeating variable double residue modules are shown in orange, blue, green, and yellow: “NI”, “NG”, “NN”, and “HD” are used to identify adenine, thymine, guanine, and cytosine, respectively.

[0090] Figure 34 The image shows a mouse mitochondrial ND5 point mutation induced by DdCBE base editing. Figure 34 a shows the efficiency of DdCBE deaminase-mediated cytosine-to-thymine base editing of target sequences and in NIH3T3 cells, where translation codons in the target sequence are underlined, editable sites are indicated in red, and DdCBE transfection combinations are represented as left or right, -G1333 or -G1397, and -N or -C, with left-G1333-N + right-G1333-C, left-G1333-C + right-G1333-N, left-G1397-N + right-G1397-C, and left-G1397-C + The p-values ​​for the C10 mutation of right-G1397-N were 0.0012, 0.0003, 0.0014 and 0.0009, respectively, and the p-values ​​for the C13 mutation were 0.0116, 0.0076, 0.0030 and 0.0003, respectively (*p < 0.05 and **p < 0.01, Student's two-tailed t-test). Figure 34 b shows the base editing efficiency in mouse blastocysts, where sequencing data were obtained from blastocysts developing from fertilized eggs microinjected with left-G1397-N and right-G1397-C DdCBE mRNA. Figure 34 c shows the alignment of mutant sequences from newborn pups, where targeted deep sequencing was performed using genomic DNA extracted from tissues obtained immediately after birth from the tail and from tissues obtained from the toes at 7 and 14 days after birth. Edited bases are indicated in red, and the editing frequency of the mutant mitochondrial genome is also indicated. Figure 34 Figure d shows the editing efficiency in various tissues of adult F0 mice (sipup-1), where sequencing data were obtained from each tissue at 50 days after birth. In all figures, dark gray and light gray bars represent the corresponding editing frequencies of the m.C12539T (C10) and m.G12542A (C13) mutations, and error bars are the standard error (sem) of the mean of n = 3 biological independent samples.

[0091] Figure 35 This demonstrates the transfer of mutant mitochondrial DNA to germ cells. Figure 35Figure a shows the results of targeted deep sequencing performed after crossing female F0 (sipup-3) mice with wild-type C57BL6 / J males to obtain F1 offspring (101, 102) to observe the germline transmission of mtDNA mutations. Edited bases are indicated in red, and the editing frequency of the mutant mitochondrial genome is also indicated. Figure 35 Inset b shows the base editing efficiency in various tissues of F1 pups (101) obtained using targeted deep sequencing of genomic DNA, where dark gray and light gray bars represent the corresponding frequencies of the m.C12539T (C10) and m.G12542A (C13) mutations, and error bars are the standard error (sem) of the mean of n = 3 biological independent samples.

[0092] Figure 36 The DdCBE-induced mouse mitochondrial ND5 G12918A mutation was shown. Figure 36 a shows the DdCBE target for generating the m.G12918A point mutation in the ND5 protein, thereby causing the D393N change, with the target codon underlined and the editable site indicated in red. Figure 36 b shows the cytosine-thymine base editing efficiency obtained using DdCBE in NIH3T3 cells, where the combination of transfected DdCBE pairs is indicated. Error bars are SEM values ​​for n = 3 biologically independent samples (ns: not significant, *p < 0.05, **p < 0.01, using Student's two-tailed t-test). The P-values ​​for C6 mutations of left-G1333-N + right-G1333-C, left-G1333-C + right-G1333-N, left-G1397-N + right-G1397-C, and left-G1397-C + right-G1397-N were 0.0052, 0.0099, 0.0027, and 0.0040, respectively, and the P-value for ns was 0.4971. Figure 36 c shows the efficiency of point mutation base editing in m.G12918A mouse blastocysts, where sequencing data were obtained from blastocysts developed by microinjecting mRNAs encoding left-G1397-C and right-G1397-N DdCBE into 1-cell stage embryos and then culturing them. Figure 36 d shows mice with the ND5 point mutation (F0), which shows F0 pups with the ND5 point mutation generated after microinjection of DdCBE mRNA and the mutant sequence array identified in newborn pups. Edited bases are shown in red, and the editing frequency of mutant mitochondrial genes is shown on the right.

[0093] Figure 37The study showed a mouse mitochondrial ND5 nonsense mutation generated by cytosine deaminase-mediated base editing. Figure 37 Figure a shows the DdCBE target sequences that produce the m.C12336T nonsense mutation and the m.G12341A silencing mutation. The m.C12336T (C9) mutation produces the Q199 termination mutation in the ND5 protein, while the m.G12341A (C14) causes the silencing Q200Q mutation. The transcriptional triad is underlined, and editable sites are indicated in red. Figure 37 Inset b shows the efficiency of cytosine-thymine base editing in generating nonsense mutations in NIH3T3 cells, indicating the combination of transfected DdCBE pairs. Dark and light gray bars represent the corresponding frequencies of m.C12336T (C9) and m.G12341A (C14) mutations. Error bars represent SEM values ​​for n = 3 biologically independent samples (ns: not significant, *p < 0.05, **p < 0.01, using Student's two-tailed t-test), left-G1333-N + right-G1333-C, left-G1333-C + right-G1333-N, left-G1397-N + right-G1397-C, and left-G1397-C + The p-values ​​for the C9 mutation of right-G1397-N were 0.0065, 0.1143, 0.0266, and 0.0037, and the corresponding p-values ​​for the C14 mutation were 0.0077, 0.0144, 0.0406, and 0.0214. Figure 37 Inset Figure c shows the editing efficiency in mouse blastocysts, where sequencing data were obtained from blastocysts developed after microinjection of mRNA encoding left-G1333-N and right-G1333-C DdCBE into fertilized eggs, and dark gray and light gray bars represent the frequencies of C9 and C14 mutations, respectively. Figure 37 Inset d shows the mutant sequence array of newborn pups, with edited bases indicated in red, and the editing frequency of the mutant mitochondrial genome is shown on the right. Figure 37 Figure e shows the Sanger sequencing chromatograms of wild-type and edited mice, with red arrows indicating substituted nucleotides;

[0094] Figure 38The Golden Gate clone used to generate the DdCBE construct is schematically shown, in which all reactions occur simultaneously in one tube. Arrows do not indicate sequential reactions. Empty expression vectors and module vectors are cut with BsaI enzyme to remove the linearized backbone and TALE module inserts, including compatible sticky ends. The backbone and six module inserts are ligated by T4 DNA ligase to generate the final DdCBE construct. Eight DdCBE clone backbone plasmids are used, and for SOD2MTS, left-G1333-N, left-G1333-C, left-G1397-N, and left-G1397-C are provided, and for COX8A MTS, right-G1333-N, right-G1333-C, right-G1397-N, and right-G1397-C are provided.

[0095] Figure 39 The ND5 mutant mouse (F0) is shown. Figure 39 Inset figure a shows the ND5 silencing mutant mouse. Figure 39 Inset b shows the ND5 G12918A mutant mouse, and Figure 39 Inset Figure c shows ND5 nonsense mutant mice generated by microinjection of DdCBE mRNA;

[0096] Figure 40 : Figure 40 Inset figure a schematically shows vectors containing DdCBE-NES and NES sequences. Figure 40 Figure b shows the sequences of the mouse m.G12918 ND5 gene and ND5-like gene on chromosome 4 of the cell nucleus, the sequence of mitochondrial TrnA on chromosome 5 of the cell nucleus, and the sequence of mitochondrial Rnr2 on chromosome 6 of the cell nucleus. Figure 40 Inset Figure c shows the editing efficiency of the ND5 gene achieved using the NIH3T3 cell line via DdCBE and DdCBE-NES. Figure 40 Inset figure d shows the editing efficiency of the TrnA gene achieved using the NIH3T3 cell line via DdCBE and DdCBE-NES. Figure 40 Inset e shows the editing efficiency of the Rnr2 gene achieved using the NIH3T3 cell line via DdCBE and DdCBE-NES, with the orange plot representing the editing efficiency of DdCBE and the gray plot representing the editing efficiency of DdCBE-NES. Figure 40 Inset f shows the DNA recognition sequence of the mitoTALEN TALE array, and Figure 40 Inset g shows the DdCBE base editing efficiency in the experimental groups treated with or without mitoTALEN. All plots are n = 2 and the error bars are the standard error of the mean.

[0097] Figure 41 The improved editing efficiency achieved using DdCBE-NES and mitoTALEN in mouse embryos and mice is demonstrated. Figure 41 Inset figure a shows the efficiency of base editing of various mitochondrial DNA targets (mtND5, mtTrnA, and mtRNR2) in the blastocyst using DdCBE and DdCBE-NES. Figure 41 Inset Figure b shows a comparison of the base editing efficiency of m.G12918A achieved using DdCBE and DdCBE-NES with and without mitoTALEN, and Figure 41 Inset c shows a comparison of m.G12918A base editing efficiency in mice. All plots are n>=3, and the error bars are the standard error of the mean (ns: not significant, *p < 0.05, **p < 0.01, ***p < 0.001, obtained using Student's two-tailed t-test).

[0098] Figure 42 : Figure 42 Inset figure a schematically shows the improvements to the DdCBE protein. Figure 42 Small image b and Figure 42 Figure c shows the crystal structure of DddAtox deaminase, where residues at the interface of the cleavage dimer are represented as rods. Figure 42 Inset figure b shows the G1397-N and G1397-C fission products, represented in purple and light blue, respectively. Figure 42 Inset figure c shows the G1333-N and G1333-C fission products, represented in orange and green respectively, and... Figure 42 Small image d and Figure 42 Inset e shows the amino acid sequences of G1397-N and G1397-C (d) and G1333-N and G1333-C (e), with interface residues indicated in red;

[0099] Figure 43 : Figure 43 Inset a shows the base editing efficiency of the G1397 interface mutant, with the editing range and target cytosine indicated at the top. The mutant was co-transfected with wild-type / TALE-free DddAtox protein as shown below, and for left-DdCBE, the TALE-free DddAtox protein was G1397-N, and for right-DdCBE, the TALE-free DddAtox protein was G1397-C. Figure 43 Inset b shows a heatmap of the target cytosine-thymine (guanine-adenine) base editing efficiency of DdCBE and mutants;

[0100] Figure 44 : Figure 44 Inset a shows the base editing efficiency of the G1333 interface mutant, with the editing range and target cytosine indicated at the top. The mutant was co-transfected with wild-type / TALE-free DddAtox protein as shown below. For left-DdCBE, the TALE-free DddAtox protein was G1333-N, and for right-DdCBE, the TALE-free DddAtox protein was G1333-C. Figure 44 Inset b shows a heatmap of the target cytosine-thymine (guanine-adenine) base editing efficiency of DdCBE and mutants;

[0101] Figure 45 The results show a comparison of the amino acid sequences of wild-type and novel full-length DddA;

[0102] Figure 46 This shows the conformation in which full-length DddA is delivered into animal or plant cells;

[0103] Figure 47 The results demonstrate the activity of cytosine to thymine conversion in the TC motif at the ROR1 (a), HEK3 (b), and TYRO3 (c) sites in the human cell genome background;

[0104] Figure 48 This demonstrates the advantages of full-length DddA;

[0105] Figure 49 The results show the measurement of full-length DddA activity in human cell genome background TRAC site 1 (a), TRAC site 2 (b), FANCF (c), and HBB (d);

[0106] Figure 50 The results show the results of measuring DddA activity in human cell genomic background TYRO3 (a), ROR1 (b), HEK3 (c), EMX1 site 2 (d), TRAC site 1 (e) and HBB (f) using DddA-dCas9(D10A, H840A)-UGI.

[0107] Figure 51 This demonstrates the base editing efficiency of full-length DddAtox in HEK293T cells. Figure 51 Figure a schematically shows the screening of full-length DddAtox in a structure-based manner, with red alanine indicating the substitution of positively charged amino acid residues with alanine. Figure 51 Inset Figure b shows an E. coli transformant with the DddA variant substituted with alanine, with E1347A used as the active site mutant in the control. Figure 51Inset c shows the edit frequency and insertion / deletion of DddA AAAAA and CBE at the TYRO3 site, as well as Figure 51 Inset d shows the allele frequency at the TYRO3 locus, with C-to-T conversion in red, prototype spacers in blue, and prototype spacer neighbor motifs (PAMs) in orange.

[0108] Figure 52 It showed non-toxic DddA GSVG. Figure 52 Figure a schematically illustrates the screening of non-toxic full-length DddAtox variants based on error-prone PCR, and Figure 52 Small image b and Figure 52 Inset figure c shows the editing frequencies of genes (b) and alleles (c) fused to the N-terminus and C-terminus of Cas9, nCas9(D10A), nCas9(H840A), and dCas9(D10A, H840A), with prototype spacers shown in blue and prototype spacer neighbor motifs (PAMs) shown in orange.

[0109] Figure 53 The editing frequency of the DddAtox variant at the N-terminus of nCas9 (D10A) is shown, in which positively charged amino acid residues at TYRO3 (a), ROR1 (b), and HEK3 (c) are replaced with alanine, the prototype spacer is shown in blue, and the prototype spacer neighbor motif (PAM) is shown in orange.

[0110] Figure 54 shows the editing frequency at several sites. Figure 54a , Figure 54b , Figure 54c , Figure 54d , Figure 54e , Figure 54f , Figure 54g and Figures 54h to 54j ROR1 site 1, ROR1 site 2, ROR1 site 3, FANCF site, HBB site, HEK3 site, TRAC5 site 1 and EMX1 site are shown respectively, with the prototype spacer in blue, the prototype spacer neighbor motif (PAM) in orange, C to T conversion in red, and the target window of DddA indicated by counting the 5' upstream of the prototype spacer as negative;

[0111] Figure 55a , Figure 55b , Figure 55c , Figure 55d , Figure 55e , Figure 55f , Figure 55g , Figure 55h , Figure 55i and Figure 55jEditing frequencies at TYRO3, ROR11, ROR12, ROR13, FANCF, HBB, HEK3, TRAC51, TRAC52, and EMX12 sites in AAAAA and E1347A cells of HeLa cells are shown, with prototype spacers in blue, prototype spacer neighbor motifs (PAMs) in orange, the target window of DddA indicated by counting the 5' upstream of the prototype spacer as negative, and target cytosine in red.

[0112] Figure 56 The time-dependent base editing and insertion / deletion rates of AAAAA and E1347A at TYRO3 site (a) and ROR1 site 1 (b) are shown.

[0113] Figure 57a , Figure 57b , Figure 57c , Figure 57d , Figure 57e , Figure 57f , Figure 57g and Figure 57h Editing, insertion-deletion, and allele frequencies of GSVGs fused to the N-terminus of nCas9 (D10A), nCas9 (H840A), and dCas9 at EMX1 site 2, FANCF site, TRAC5 site 1, TRAC5 site 2, ROR1 site 1, ROR1 site 2, ROR1 site 3, and HBB site are shown, with prototype spacers indicated in blue, prototype spacer neighbor motifs (PAMs) in orange, C-to-T conversions in red, and the target window of the GSVG indicated by counting the 5' upstream of the prototype spacer as negative.

[0114] Figure 58a , Figure 58b , Figure 58c and Figure 58d Editing, insertion-deletion, and allele frequencies of GSVGs fused to the C-terminus of nCas9 (D10A), nCas9 (H840A), and dCas9 at EMX1 site 2, EMX1 site 4, ROR1 site 2, and HBB site are shown, with the prototype spacer in blue, the prototype spacer neighbor motif (PAM) in orange, G to A conversion in red, and the target window of the GSVG indicated by counting from the 3' downstream of the prototype spacer at position 1.

[0115] Figure 59 The time-dependent edit and insertion-deletion frequencies of E1347A, GSVG, SSVG, GSAG, and GSVS fused to the C-terminus of nCas9 (H840A) at TYRO3 (a) and EMX1 site 2 (b) are shown.

[0116] Figure 60 This demonstrates mitochondrial base editing of mDdCBE in HEK293T cells. Figure 60 Small picture a and Figure 60 Inset b shows the editing efficiency of ND4 and ND6, with target cytosine and TALE binding sites indicated in red and gray, respectively. Figure 60 c to Figure 60 f shows the editing efficiency of ND4 (c, d) and ND6 (e, f) when only half of the DddAtox arrays are fused with the TALE arrays and the remaining half do not contain TALEs, where the left and right TALE arrays are represented by L and R, respectively, and mismatches between the ND6 TALE array and the reference genome are underlined in purple.

[0117] Figure 61 : Figure 61 Figure a schematically shows a zinc finger cytosine deaminase (ZFD) using a conventional ZFP DNA-binding protein.

[0118] Figure 61 Inset b shows the location where adenine deaminase is inserted into the ZFD (where the red arrow indicates the insertion site).

[0119] Figure 61 Inset Figure c shows the base editing efficiency (C to T) of the constructed ZF-DdABE at the nuclear DNA Trac site.

[0120] Figure 61 Inset d shows the base editing efficiency (A to G) of the constructed ZF-DdABE at the nuclear DNA Trac site.

[0121] (Where WT-ZFD is a C-to-T deaminase that cleaves DddAtox in the absence of adenine deaminase.)

[0122] Figure 61 Inset e shows the efficiency (C to T) of ZF-DdABE targeting mitochondrial DNA at the ND1 site, and

[0123] Figure 61 Inset f shows the efficiency of ZF-DdABE targeting mitochondrial DNA at the ND1 site (A to G).

[0124] Figure 62 : Figure 62 Inset a schematically shows DdABE using TALE and split DddAtox (where the components include split DddAtox, adenine deaminase, and TALE array).

[0125] Figure 62Inset Figure b shows the base editing efficiency when a single adenine deaminase is linked to TALE targeting the mitochondrial ND4 site.

[0126] Figure 62 Inset Figure c shows the base editing efficiency when adenine deaminase is linked to TALE-cleaved DddAtox targeting the mitochondrial ND1 site.

[0127] Figure 62 Inset d shows the base editing efficiency in a single nucleotide unit when using the DdCBE pair on the left and linking adenine deaminase to TALE-Cleft DddAtox on the right (where the green box represents the part linked by TALE), and

[0128] Figure 62 Inset e shows the base editing efficiency in a single nucleotide unit when adenine deaminase is linked to TALE-split DddAtox on the left and DdCBE pairs are used on the right (where the green box is the part linked by TALE).

[0129] Figure 63 : Figure 63 Inset a shows the C-to-T and A-to-G base editing efficiency of DdABE targeting the mitochondrial ND1 site in the absence or presence of UGI (where the red box indicates adenine deaminase).

[0130] Figure 63 Inset b shows the C-to-T and A-to-G base editing efficiency of DdABE targeting the mitochondrial ND4 site in the absence or presence of UGI (where the red box indicates adenine deaminase).

[0131] Figure 63 Inset figure c shows the configuration with the highest efficiency in a single nucleotide unit among the DdABE configurations targeting the mitochondrial ND1 site (where the green box represents the portion connected by TALE), and

[0132] Figure 63 Inset d shows the configuration with the highest efficiency in a single nucleotide unit among the DdABE configurations targeting the mitochondrial ND4 site (where the green box is the part connected by TALE).

[0133] Figure 64 : Figure 64 Figure a at the top schematically shows a single TALE module containing all the constructs within a TALE module (where the components include full-length DddAtox, adenine deaminase, and a TALE array), and

[0134] At the bottom, a dual-TALE module using two TALE modules is also shown (the components include a full-length DddAtox and TALE array on one side and an adenine deaminase and TALE array on the other side).

[0135] Figure 64 Inset Figure b shows the base editing efficiency of single-module and dual-module DdABE targeting the mitochondrial ND1 site, and

[0136] Figure 64 Inset Figure c shows the base editing efficiency of single-module and dual-module DdABE targeting the mitochondrial ND4 site;

[0137] Figure 65 Results show that confirm the base editing efficiency of a single module targeting the ND1 site (components include a TALE array, adenine deaminase (AD), and full-length DddAtox (GSVG, AAAA, and E1347A are variants)).

[0138] Figure 66 Results confirming the base editing efficiency of the bimodule targeting the ND1 site (components include a TALE array, adenine deaminase (AD), and full-length DddAtox (GSVG, AAAA, and E1347A are variants)) are shown.

[0139] Figure 67 : Figure 67 Inset figure a shows the base editing efficiency when TadA(AD) adenine deaminase is linked to a TALE-binding protein targeting the ND1 site, and

[0140] Figure 67 Inset Figure b shows the base editing efficiency when TadA(AD) adenine deaminase is ligated to a TALE-binding protein targeting the ND4 site; and

[0141] Figure 68 The efficiency of adenine and cytosine base editing in dual-module, single-module, and split DddA-AD structures targeting the ND1 site is shown. Among them (from bottom to top), when the UGI is attached to both sides in the absence of AD, referred to as DdCBE, only cytosine base editing occurs; or when AD is replaced by either side of the UGI, both cytosine and adenine base editing occur; or when the UGI is absent, only adenine base editing occurs selectively; and similarly, even in dual-module and single-module structures, only adenine base editing occurs selectively. Detailed Implementation

[0142] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. Generally, the nomenclature used herein is well-known and typical in the art.

[0143] As used herein, the term “editing” is used interchangeably with “correction” and refers to a method of altering the nucleic acid sequence at a specific genomic target site in a cell. Such a specific genomic target includes, but is not limited to, chromosomal regions, genes, promoters, open reading frames, or any nucleic acid sequence.

[0144] As used herein, the term "single base" refers to only one nucleotide in a nucleic acid sequence. When used in the context of single base editing, it refers to the substitution of a base at a specific position in a nucleic acid sequence by a different base. Such substitution can occur through a variety of mechanisms, including, but not limited to, substitution or modification.

[0145] As used herein, the term “target” or “target site” refers to any previously identified nucleic acid sequence of any composition and / or length. These target sites include, but are not limited to, chromosomal regions, genes, promoters, open reading frames, or any nucleic acid sequence.

[0146] As used in this article, the term “mid-target” refers to a subsequence of a specific genomic target that is bound by a programmable DNA-binding protein or can be fully complementary to a single guide RNA sequence.

[0147] As used herein, the term “off-target” refers to a subsequence of a specific genomic target that is partially complementary to a mid-target sequence and / or a single guide RNA sequence that can be identified by a programmable DNA binding region.

[0148] 1. Cytosine deaminase

[0149] According to one aspect of the invention, the fusion protein comprises cytosine deaminase or a variant thereof, wherein the cytosine deaminase or a variant thereof comprises a first cleavage and a second cleavage derived from cytosine deaminase or a variant thereof, and each of the first cleavage and the second cleavage is fused to a DNA-binding protein.

[0150] Cytosine deaminase is an enzyme that removes the amino group from the cytosine base and converts cytosine (C) into uridine (U).

[0151] It can be a cytosine deaminase. Examples of cytosine deaminases can include APOBEC1 (apolipoprotein B editing complex 1) and AID (activation-induced deaminase), but most DNA deaminases may only act on single-stranded DNA and may not be well-suited for base editing via linkage with DNA-binding proteins. Specifically, cytosine deaminases can be derived from double-stranded DNA deaminases (DddA) or their orthologs. More specifically, cytosine deaminases can be double-stranded DNA-specific bacterial cytosine deaminases.

[0152] Cytosine deaminase is provided in a fission form, comprising a first fission and a second fission, each of which is not deaminase activity.

[0153] It may include the sequence corresponding to the DddAtox cleavage fragment of a full-length cytosine deaminase, SEQ ID NO: 1. The cytosine deaminase comprises a first cleavage fragment and a second cleavage fragment, each of which lacks deaminase activity.

[0154] [SEQ ID NO: 1]

[0155] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC

[0156] In one embodiment, the first or second cleavage product of the cytosine deaminase may comprise a sequence from the N-terminus to at least one selected from G33, G44, A54, N68, G82, N98, and G108 in SEQ ID NO: 1. The first or second cleavage product of the cytosine deaminase may comprise a sequence from at least one selected from G34, P45, G55, N69, T83, A99, and A109 to the C-terminus in SEQ ID NO: 1.

[0157] Specifically, the cytosine deaminase may comprise the first cleavage product (G1333-N) of SEQ ID NO: 23 and the second cleavage product (G1333-C) of SEQ ID NO: 24, the first cleavage product (G1397-N) of SEQ ID NO: 25 and the second cleavage product (G1397-C) of SEQ ID NO: 26, the first cleavage product (G1333-N) of SEQ ID NO: 23 and the second cleavage product (G1397-C) of SEQ ID NO: 26, or the first cleavage product (G1397-N) of SEQ ID NO: 25 and the second cleavage product (G1333-C) of SEQ ID NO: 24.

[0158] (SEQ ID NO: 23) Wild-type DddAtox G1333-N

[0159] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGG

[0160] (SEQ ID NO: 27)

[0161] GGCTCTGGTTCCTACGCCCTGGGTCCATATCAGATTAGTGCTCCCCAACTCCCCGCCTACAACGGTCAGACAGTGGGGACCTTTTACTATGTCAACGACGCCGGGGGATTGGAATCCAAGGTTTTTCTCTAGCGGTGGG

[0162] (SEQ ID NO: 24) Wild type G1333-C

[0163] PTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC

[0164] (SEQ ID NO: 28)

[0165] CCAACACCTTATCCTAACTACGCTAACGCCGGGCACGTCGAGGGGCAGTCAGCTCTTTTTATGAGAGATAACGGCATTAGCGAAGGGCTTGTGTTCCATAATAATCCTGAGGGCACCTGTGGCTTCTGTGTAAATATGACCGAAACACTTCTGCCTGAGAACGCTAAAATGACTGTCGTACCACCCGAAGGCGCAATCCCAGTTAAACGGGGCGCAACCGGCGAAACCAAAGTATTCACCGGAAACAGCAATAGTCCAAAGTCCCCCACCAAGGGAGGTTGC

[0166] (SEQ ID NO: 25)Wild-type DddAtox G1397-N

[0167] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEG

[0168] (SEQ ID NO: 29)

[0169] GGTAGCTACGCACTTGGTCCTTACCAGATTAGCGCACCCCAACTCCCCGCCTATAATGGTCAAACCGTCGGGACCTTTTACTACGTAAACGATGCTGGTGGGCTGGAATCCAAAGTATTCTCCTCAGGGGGCCCTACACCCTACCCCAACTACGCCAATGCTGGTCATGTAGAAGGGCAGTCAGCACTGTTTATGCGCGATAATGGTATAAGCGAGGGGTTGGTCTTCCATAACAACCCAGAGGGTACTTGTGGCTTCTGTGTGAATATGACTGAAACCCTTCTGCCCGAAAATGCCAAGATGACTGTCGTCCCACCTGAAGGC

[0170] (SEQ ID NO: 26)Wild-type DddAtox G1397-C

[0171] AIPVKRGATGETKVFTGNSNSPKSPTKGGC

[0172] (SEQ ID NO: 30)

[0173] GCCATACCTGTGAAGCGGGGAGCAACAGGGGAGACAAAGGTGTTCACAGGCAACTCTAACAGTCCAAAGAGCCCCACCAAAGGCGGGTGT

[0174] The combinations of G1333N, G1333C, G1397N, and G1397C can be used as cleavage forms of deaminase. Specifically, they can be used in the forms of left-G1333-N + right-G133-C, left-G1397-N + right-G1397-C, left-G1397-N + right-G1333-C, or left-G1333-N + right-G1397-C.

[0175] 2. Variations

[0176] The inventors of this application sought to suppress unwanted base editing through DdCBE mutations in which amino acid residues are substituted. A high-precision DddA-derived cytosine base editor capable of reducing the off-target effects of DdCBE is proposed. This off-target base editing effect is caused by DddA… tox The phenomenon is caused by the spontaneous assembly of deaminase cleavage products and is unrelated to the interaction between TALE and DNA.

[0177] Here, the mutated amino acid residue is located at DddA. tox Contact sites on the surfaces where the splitting dimers interact with each other. This is achieved by replacing the DddA site with alanine. tox High-fidelity DdCBE (HF-DdCBE) is constructed by using amino acid residues on the surface between the mitochondria. HF-DdCBE prevents the two deaminase hemiparts of the mitochondria linked to TALE from functioning normally when not bound to DNA. Whole-mitochondrial genome analysis confirmed that HF-DdCBE is highly efficient and precise, unlike conventional DdCBE, which induces many undesirable off-target C-to-T conversions in human mitochondrial DNA.

[0178] For DddAtox, base editing should theoretically only be induced when both fission dimers are recruited to the target site of the DNA. Based on actual experimental results, even when using DdCBE, targeted base editing occurs, with one half binding to DNA and the other half not. To address this issue, DdCBE pairs are prevented from binding at unwanted sites by replacing residues on the protein surface where the fission dimers interact with each other.

[0179] Based on this, the present invention relates to novel variants that reduce non-selective base editing by substituting amino acid residues of the cytosine deaminase DddAtox.

[0180] Cytosine deaminase comprises the first cleavage product (G1333-N) of SEQ ID NO: 23 and the second cleavage product (G1333-C) of SEQ ID NO: 24, or the first cleavage product (G1397-N) of SEQ ID NO: 25 and the second cleavage product (G1397-C) of SEQ ID NO: 26.

[0181] The variant of cytosine deaminase can be configured such that at least one amino acid selected from positions 3, 5, 10, 11, 13, 14, 15, 16, 17, 18, 19, 28, 30, and 31 in the first cleavage of SEQ ID NO: 23 is substituted with a different amino acid, or further such that at least one amino acid selected from positions 13, 16, 17, 20, 21, 28, 29, 30, 31, 32, 33, 56, 57, 58, and 60 in the second cleavage of SEQ ID NO: 24 is substituted with a different amino acid.

[0182] The variants according to the invention can be configured such that at least one amino acid selected from positions 87, 88, 91, 92, 95, 100, 101, 102 and 103 in the first split of SEQ ID NO: 25 or at least one amino acid selected from positions 13, 14, 15 and 16 in the second split of SEQ ID NO: 26 is replaced by a different amino acid.

[0183] Here, "different amino acid" can be alanine, isoleucine, leucine, methionine, phenylalanine, proline, tryptophan, valine, asparagine, cysteine, glutamine, glycine, serine, threonine, tyrosine, aspartic acid, glutamic acid, arginine, histidine, or lysine, and can refer to an amino acid selected from all known variants of the above-mentioned amino acids, excluding the amino acid at the original mutation site in the wild-type protein. In one exemplary embodiment, "different amino acid" can be alanine.

[0184] Specifically, it may include an amino acid substitution of at least one of Y3A, L5A, I10A, S11A, V13A, G14A, T15A, F16A, Y17A, Y18A, V19A, K28A, F30A and S31A in the first cleavage of SEQ ID NO: 23 (corresponding to Y1292A, L1294A, I1299A, S1300A, V1312A, G1313A, T1314A, F1315A, Y1316A, Y1317A, V1318A, K1327A, F1329A and S1330A, respectively).

[0185] In addition, it may include an amino acid substitution of at least one of V13A, Q16A, S17A, F20A, M21A, E28A, G29A, L30A, V31A, F32A, H33A, K56A, M57A, T58A and V60A from the second split product of SEQ ID NO: 24 (corresponding to V1346A, Q1349A, S1350A, F1353A, M1354A, E1361A, G1362A, L1363A, V1364A, F1365A, H1366A, K1389A, M1390A, T1391A and V1393A, respectively).

[0186] Specifically, it may include an amino acid substitution of at least one of C87A, V88A, T91A, E92A, L95A, K100A, M101A, T102A and V103A in the first cleavage of SEQ ID NO: 25 (corresponding to C1376A, V1377A, T1380A, E1381A, L1384A, K1389A, M1390A, T1391A and V1392A, respectively).

[0187] In addition, it may include an amino acid substitution of at least one of K13A, V14A, F15A and T16A in the second cleavage of SEQ ID NO: 26 (corresponding to K1410A, V1411A, F1412A and T1413A, respectively).

[0188] [DddAtox G1333-N variant]

[0189]

[0190]

[0191] [DddAtox G1333-C variant]

[0192]

[0193]

[0194]

[0195] [DddAtox G1397-N variant]

[0196]

[0197]

[0198] [DddAtox G1397-C variant]

[0199]

[0200] The cytosine deaminase variants according to the invention may comprise at least one sequence selected from the amino acid sequences described in the table above. The cytosine deaminase variants according to the invention exhibit the potential for unwanted editing by reducing various bases at nonspecific target sites.

[0201] 3. Full-length cytosine deaminase

[0202] The inventors of this application have developed a novel programmable cytosine deaminase using full-length DddA, which modifies wild-type cytosine deaminase DddA. tox Prepared from positively charged amino acids, the novel programmable cytosine deaminase is used in cleavage form due to its cytotoxicity.

[0203] The present invention relates to a fusion protein comprising (i) a DNA-binding protein and (ii) a cytosine deaminase or a variant thereof, wherein the cytosine deaminase or a variant thereof is a non-toxic, full-length cytosine deaminase.

[0204] In DddA tox The C-terminus of DNA contains a specific cluster of positively charged amino acids (KRKKK). Since DNA is negatively charged, it binds to the positively charged amino acids of proteins. Replacing positively charged amino acids with uncharged amino acids weakens the binding of DddA. tox The binding force with DNA is reduced, thereby decreasing or eliminating cytotoxicity. In particular, the non-toxic combination produced by replacing positively charged amino acids with nonpolar amino acids enables cloning using E. coli to provide full-length DddA.

[0205] Wild-type DddA toxIts use in two fission forms due to cytotoxicity presents several limitations in experiments. Specifically, when using Cas9, orthogonal Cas9 variants that recognize other PAMs are employed. Therefore, due to the limited presence of PAMs, it is often difficult to accurately deamination of cytosine to thymine at the desired location. Furthermore, the target window with the highest activity is a 40 bp region between two Cas9 variants that bind to each other, and unwanted cytosine in this region is also deaminated. However, full-length DddA is not constrained by PAMs because it is not dissociative. Moreover, it is most active within 10 bp of the TC motif, resulting in high precision.

[0206] It has been confirmed that Cas9 deaminates cytosine in the ACA, GC, and CC motifs of the R ring formed by binding to the target site to thymine. This activity has not yet been identified in isolated form.

[0207] In full-length DddA, TALE modules or zinc finger proteins and Cas9 can be used to replace cytosine with thymine at the desired location. Existing DddA tox It must be delivered in pairs in a separate manner, but full-length DddA can be delivered using only one module of the TALE module or zinc finger protein, allowing for unlimited selection of target sites. Furthermore, by targeting DNA in mitochondria, plant chloroplasts, or plastids, as well as genomic sites, cytosine in specific DNA can be converted to thymine.

[0208] Furthermore, all constructs can be inserted into AAVs, which are viral vectors used for gene therapy due to their small size. Existing CBEs (cytosine base editors) replace cytosine with thymine in the R-loop formed by Cas9 binding to the target site, but the full-length DddA of this invention deaminates the cytosine outside the R-loop. Therefore, cytosine can be converted to thymine at locations where editing is restricted to conventional CBEs.

[0209] Based on this, in non-toxic full-length cytosine deaminases, at least one, at least two, at least three, at least four, or at least five amino acids of the wild-type deaminase of SEQ ID NO: 1 can be replaced by different amino acids.

[0210] Here, "different amino acid" can be alanine, isoleucine, leucine, methionine, phenylalanine, proline, tryptophan, valine, asparagine, cysteine, glutamine, glycine, serine, threonine, tyrosine, aspartic acid, glutamic acid, arginine, histidine, or lysine, and can refer to an amino acid selected from all known variants of the above-mentioned amino acids, excluding the amino acid at the original mutation site in the wild-type protein. In one exemplary embodiment, "different amino acid" can be alanine.

[0211] Depending on its type, the non-toxic full-length DddA may include a sequence selected from SEQ ID NO: 12 to SEQ ID NO: 18.

[0212] Wild type (SEQ ID No: 1)

[0213] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC

[0214] A1341D KRKKA (SEQ ID No: 12)

[0215] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYDNAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTAGGC

[0216] [[ID=·17]]AAAAA (SEQ ID No: 13)

[0217] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVAAGATGETAVFTGNSNSPASPTAGGC

[0218] AAAAK (SEQ ID No: 14)

[0219] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVAAGATGETAVFTGNSNSPASPTKGGC

[0220] AAKAA (SEQ ID No: 15)

[0221] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVAAGATGETKVFTGNSNSPASPTAGGC

[0222] AAKAK (SEQ ID No: 16)

[0223] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVAAGATGETKVFTGNSNSPASPTKGGC

[0224] KAAAA (SEQ ID No: 17)

[0225] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKAGATGETAVFTGNSNSPASPTAGGC

[0226] E1347A (SEQ ID No: 18)

[0227] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVAGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC

[0228] The full-length deaminase variant may include at least one substitution selected from the following in the amino acid sequence of SEQ ID NO: 1:

[0229] Replace S with G at position 37;

[0230] Replace G with S at position 59;

[0231] Replace A with V at position 109; and

[0232] Replace S with G at position 129.

[0233] In one embodiment, the full-length deaminase variant may include the sequence of SEQ ID NO: 19, comprising replacing S with G at position 37, replacing G with S at position 59, replacing A with V at position 109, and replacing S with G at position 129 in the amino acid sequence of SEQ ID NO: 1.

[0234] GSVG (SEQ ID No: 19)

[0235] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLEGKVFSSGGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGVIPVKRGATGETKVFTGNSNGPKSPTKGGC

[0236] Full-length DddA GSVG can be cloned using universal *E. coli*. In a human genomic context, full-length DddA GSVG has been shown to deaminate cytosine to thymine at the target site of the TC motif. Full-length DddA GSVG can be cloned into each of the N-terminus and C-terminus of Cas9. DddA GSVG linked to the N-terminus of Cas9 can substitute cytosine for thymine at the same target site. In human cells, DddA GSVG linked to the C-terminus of Cas9 has been shown to induce cytosine-to-thymine substitution in the TC motif (guanine-to-adenine substitution in the complementary sequence).

[0237] In one embodiment, the full-length deaminase variant may include a sequence selected from SEQ ID NO: 20 to 22.

[0238] SSVG (SEQ ID No: 20)

[0239] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGVIPVKRGATGETKVFTGNSNGPKSPTKGGC

[0240] GSAG (SEQ ID No: 21)

[0241] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLEGKVFSSGGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNGPKSPTKGGC

[0242] GSVS (SEQ ID No: 22)

[0243] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLEGKVFSSGGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGVIPVKRGATGETKVFTGNSNSPKSPTKGGC

[0244] 4. DNA-binding proteins

[0245] DNA-binding proteins can be, for example, zinc finger proteins, TALE (transcription activator-like effector) proteins, CRISPR-associated nucleases, or combinations of two or more of them.

[0246] Zinc finger proteins possess DNA-binding domains in their zinc finger motifs, and the C-terminal portion of the finger specifically recognizes DNA sequences. DNA-binding proteins containing 3 to 6 zinc finger motifs recognize DNA sequences.

[0247] In one embodiment, each of the first and second cleavage products of cytosine deaminase may be fused to the N-terminus or C-terminus of a zinc finger protein.

[0248] The C-terminus (ZF-left) of the zinc finger protein is fused with the N-terminus of the first cleavage of cytosine deaminase, and the C-terminus (ZF-right) of the zinc finger protein is fused with the N-terminus of the second cleavage of cytosine deaminase (CC configuration).

[0249] The N-terminus (ZF-left) of the zinc finger protein is fused with the C-terminus of the first cleavage of cytosine deaminase, and the C-terminus (ZF-right) of the zinc finger protein is fused with the N-terminus of the second cleavage of cytosine deaminase (NC configuration).

[0250] The C-terminus (ZF-left) of the zinc finger protein is fused with the N-terminus of the first cleavage of cytosine deaminase, and the N-terminus (ZF-right) of the zinc finger protein is fused with the C-terminus of the second cleavage of cytosine deaminase (CN configuration).

[0251] The N-terminus (ZF-left) of the zinc finger protein is fused with the C-terminus of the first cleavage of cytosine deaminase, and the N-terminus (ZF-right) of the zinc finger protein is fused with the C-terminus of the second cleavage of cytosine deaminase (NN configuration).

[0252] ZF-left may include the sequence of SEQ ID NO: 2:

[0253] [SEQ ID NO: 2]

[0254] GIHGVPAAMAERPFQCRICMRNFSDRSNLSRHIRTHTGEKPFACDICGRKFAISSNLNSHTKIHTGSQKPFQCRICMRNFSRSDNLARHIRTHTGEKPFACDICGRKFATSGNLTRHTKIHLR.

[0255] ZF-right may include the sequence of SEQ ID NO: 3:

[0256] [SEQ ID NO: 3]

[0257] GIHGVPAAMAERPFQCRICMRNFSRSDNLSVHIRTHTGEKPFACDICGRKFAQKINLQVHTKIHTGEKPFQCRICMRNFSRSDVLSEHIRTHTGEKPFACDICGRKFAQRNHRTTHTKIHLR.

[0258] The sequence of a ZF can vary depending on the DNA target. ZFs can be customized based on the target DNA sequence. Since ZFs recognize 3-bp DNA, combinations of ZFs recognizing 9-18 bp DNA can be constructed by linking 3 to 6 ZFs. For example, it can be generated using libraries that include modules recognizing GNNs, TNNs, CNNs, or ANNs.

[0259] In some cases, zinc finger proteins can be linked to deaminases via linkers. Linkers can be peptide linkers comprising 2 to 40 amino acid residues. Linkers can be, for example, linkers of length 2, 5, 10, 16, 24, or 32 amino acids, but are not limited to these.

[0260] In one implementation, the connector may include:

[0261] 2a.a connector: GS

[0262] 5a.a Connector: TGEKQ (SEQ ID NO: 8)

[0263] 10a.a Connector: SGAQGSTLDF (SEQ ID NO: 9)

[0264] 16a.a Connector: SGSETPGTSESATPES (SEQ ID NO: 10);

[0265] 24a.a connector: SGTPHEVGVYTLSGTPHEVGVYTL (SEQ ID NO: 115); or

[0266] 32a.a Connector: GSGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID NO: 11).

[0267] In one specific embodiment of the invention, the cleavage deaminase and zinc finger protein can be linked via a connector, wherein the zinc finger protein is fused to the N-terminus of the cleavage hemiaminase comprising the first cleavage, and the zinc finger protein is fused to the N-terminus of the hemiaminase comprising the second cleavage. Here, C-to-T base conversion can occur in the spacer between the left and right ZFP binding sites. When linked via a 24a.a connector to the hemiaminase comprising the first cleavage and the hemiaminase comprising the second cleavage, respectively, both left and right ZFPs have been shown to exhibit high editing efficiency.

[0268]

[0269] TAL effectors (TALEs) are configured to repeat a 33-34 amino acid sequence, with approximately nine RVD (repetitive variant domains) repeated. Based on the 12th-13th amino acid sequence (HD->cytosine, NI->adenine, NG->thymine, NN->guanine), it recognizes one nucleotide per domain and can bind to a specific DNA sequence. TAL effectors (TALEs) recognize single-stranded DNA within the target site. The distance between target sites can be 12-14 nucleotides.

[0270] A TALE domain is a protein domain that binds to nucleotides in a sequence-specific manner through a combination of at least one TALE-repetitive sequence. It includes, but is not limited to, at least one TALE-repetitive sequence, particularly 1 to 30 TALE-repetitives. The TALE-repetitive sequence is a domain that recognizes a specific nucleotide sequence within the TALE domain.

[0271] The TALE domain comprises a region containing the N-terminus of the TALE and a region containing the C-terminus of the TALE as a skeleton structure. The first TALE, including the N-terminus of the TALE, can be encoded by SEQ ID NO: 4 or 5. The second TALE, including the C-terminus of the TALE, can be encoded by SEQ ID NO: 6 or 7.

[0272] Depending on the location of the TALE domain binding based on the cleavage site, a single TALE array or each of the first and second TALE arrays can be bound to it.

[0273] The first TALE (left TALE) can fuse with the first cleavage product of cytosine deaminase, and the second TALE (right TALE) can fuse with the second cleavage product of cytosine deaminase. The corresponding constructs can be described as N'-TALE-first cleavage product-C' and N'-TALE-second cleavage product-C'.

[0274] When the cytosine deaminase is full-length, a single TALE module can bind to the N-terminus of the cytosine deaminase. A single TALE module and the cytosine deaminase are included in the NC direction. A dual-module TALE can be included, wherein a first TALE can be fused to the N-terminus of the full-length cytosine deaminase, and a second TALE can be included separately. A first TALE module and the cytosine deaminase are included in the NC direction, and constructs of N'-TALE-cytosine deaminase-C' and N'-TALE-C' are provided.

[0275] TALE arrays can be customized based on target DNA sequences. TALE arrays are configured such that modules consisting of 33 to 35 amino acid residues are repeatedly arranged. These are derived from the plant pathogen genus *Xanthomonas*, and the modules recognize each of the bases A, C, G, and T, then bind to DNA. The base specificity of each module is determined by the 12th and 13th amino acid residues, a so-called repeating variable double residue (RVD). For example, a module with RVD NN recognizes G, NI recognizes A, HD recognizes C, and NG recognizes T. TALE arrays can consist of at least 14 to 18 modules and can be designed to recognize target DNA sequences 15–20 bp in length.

[0276] Regarding CRISPR-related nucleases, two types of RNA are encoded in the CRISPR array: crRNA (CRISPRRNA) and tracrRNA (trans-activating CRISPRRNA). Furthermore, crRNA is transcribed at the prototypical spacer site and binds to tracrRNA to form a tertiary structure. Both types of RNA facilitate the recognition and cleavage of foreign DNA.

[0277] Cas proteins may include, but are not limited to, Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9, Cas10, Cas12a, Cas12b, Cas12c, Cas12d, Cas12e, Cas12g, Cas12h, Cas12i, Cas12j, Cas13a, Cas13b, Cas13c, Cas13d, Cas14, Csy1, and Cs. y2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, CsMT2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, or Csf4 endonucleases.

[0278] Cas proteins can be derived from microbial genera containing Cas protein orthologs, selected from genera such as Corynebacter, Sutterella, Legionella, Treponema, Filifactor, Eubacterium, Streptococcus (Streptococcus pyogenes), Lactobacillus, Mycoplasma, Bacteroides, Flaviivola, Flavobacterium, Azospirillum, Gluconacetobacter, Neisseria, Roseburia, Parvibaculum, and Staphylococcus (Staphylococcus aureus). (Aureus), Nitratifractor, Corynebacterium, and Campylobacter, and can be easily isolated or recombined from them.

[0279] Cas proteins can be included in a mutated form that may lose endonuclease activity. Examples include at least one mutant target-specific nuclease selected from those mutated to lose endonuclease activity but retain nickase activity, and those mutated to lose both endonuclease and nickase activity.

[0280] When possessing nicking enzyme activity, a nick can be introduced into the strand in which the base conversion occurs or the opposing strand (e.g., the strand opposite to the strand in which the base conversion occurs) simultaneously with or regardless of the order of base conversion by cytosine deaminase (e.g., the conversion of cytosine to uridine). For example, on the strand opposite to the strand containing PAM, a nick can be introduced at the position between the 3rd and 4th nucleotides in the direction of the 5' end of the PAM sequence. Such mutations (e.g., amino acid substitutions) can occur in catalytically active domains (e.g., the RuvC catalytic domain in Cas9). Additionally, Cas9 derived from *Streptococcus pyogenes* may include mutations in which at least one of the following residues is replaced by any different amino acid: catalytically active aspartic acid residue (such as aspartic acid at position 10 (D10)), glutamic acid at position 762 (E762), histidine at position 840 (H840), asparagine at position 854 (N854), asparagine at position 863 (N863), or aspartic acid at position 986 (D986). Here, any different amino acid substituted may be alanine, but is not limited to this.

[0281] In some cases, the Cas9 protein derived from Streptococcus pyogenes can be mutated to recognize an NGA (where N is any base selected from A, T, G, and C) different from the PAM sequence (NGG) of wild-type Cas9. This is achieved by substituting at least, for example, all three of the following amino acids with different amino acids: aspartic acid at position 1135 (D1135), arginine at position 1335 (R1335), and threonine at position 1337 (T1337).

[0282] For example, in the amino acid sequence of the Cas9 protein derived from Streptococcus pyogenes, amino acid substitutions can occur in:

[0283] (1) D10, H840 or D10 + H840;

[0284] (2) D1135, R1335, T1337 or D1135 + R1335 + T1337; or

[0285] (3) Residues (1) and (2).

[0286] Here, "different amino acids" can be alanine, isoleucine, leucine, methionine, phenylalanine, proline, tryptophan, valine, asparagine, cysteine, glutamine, glycine, serine, threonine, tyrosine, aspartic acid, glutamic acid, arginine, histidine, or lysine, and can refer to amino acids selected from all known variants of the above-mentioned amino acids, excluding the amino acid at the original mutation site in the wild-type protein. In one exemplary embodiment, "different amino acids" can be alanine, valine, glutamine, or arginine.

[0287] In some cases, a guide RNA may be further included. The guide RNA may be, for example, at least one selected from CRISPR RNA (crRNA), trans-activating crRNA (tracrRNA), and single guide RNA (sgRNA). Specifically, it may be a double-stranded crRNA:tracrRNA complex in which crRNA and tracrRNA are bound to each other, or a single-stranded guide RNA (sgRNA) in which crRNA or a portion thereof and tracrRNA or a portion thereof are linked by an oligonucleotide linker.

[0288] 5. Addition of adenine deaminase

[0289] This invention relates to a fusion protein comprising three components: a DNA-binding protein, a cytosine deaminase or a variant thereof, and an adenine deaminase. The cytosine deaminase or a variant thereof is split into two parts, referred to as "splits," which are derived from a non-toxic full-length cytosine deaminase or a variant thereof. Both splits are fused with the DNA-binding protein.

[0290] The inventors of this application constructed a base editor capable of editing base A by linking adenine deaminase, which can cause A-to-G conversion, and DddAtox cytosine deaminase to DNA-binding TALE or ZFP proteins.

[0291] The deaminase using existing DddAtox (DdCBE) is a cytosine deaminase that uses TALE repeat sequences as DNA-binding modules. Unlike DdCBE, which only causes C to T conversion, DdABE can induce A to G conversion, thus allowing for other mutational patterns.

[0292] Because DdABE itself recognizes double-stranded DNA and induces deamination, it lacks additional components such as RNA. The RNA delivery mechanism in mitochondria or chloroplasts is unclear, therefore the CRISPR system cannot be applied. However, DdABE, lacking this component, can target not only genomic DNA in cells but also DNA in organelles such as mitochondria or chloroplasts, inducing A-to-G conversion of specific DNA.

[0293] Currently, DdCBE is only a gene editing technology targeting mitochondria or organelles. Therefore, the mutations that can be introduced through all conventional techniques are limited to C-to-T conversions, but DdABE can induce A-to-G conversions, thus making the range of mutations that can be introduced much more diverse. This makes it possible to generate or treat mitochondrial disease models that were previously impossible.

[0294] The existing DdCBE requires two TALE modules (connected to the left and right sides), therefore it cannot be loaded onto AAV, a viral vector with low gene capacity used in gene therapy. However, since DdABE can be used as a single module capable of using only one TALE module, it can be loaded onto AAV and used in gene therapy.

[0295] DdABE is highly compatible because it can use either split DddAtox or full-length DddAtox variants as needed.

[0296] Adenine deaminases can be selected from, for example, APOBEC1 (apolipoprotein B editing complex 1), AID (activation-induced deaminase), and tadA (tRNA-specific adenosine deaminase), and can be particularly tadA (tRNA-specific adenosine deaminase). Adenine deaminases can be, for example, deoxyadenine deaminases, variants of Escherichia coli TadA.

[0297] In a construct (NC configuration) in which cytosine deaminase is cleaved, the DNA-binding protein is a zinc finger protein, the N-terminus (ZF-left) of the zinc finger protein is fused to the C-terminus of the first cleavage of cytosine deaminase, and the C-terminus (ZF-right) of the zinc finger protein is fused to the N-terminus of the second cleavage of cytosine deaminase, adenine deaminase can be fused to the C-terminus (ZF-left) of the zinc finger protein, the N-terminus or C-terminus of the first cleavage of cytosine deaminase, the N-terminus (ZF-right) of the zinc finger protein, or the N-terminus or C-terminus of the second cleavage of cytosine deaminase.

[0298] Furthermore, adenine deaminase can be fused with the C-terminus (ZF-left) of a zinc finger protein, the N-terminus or C-terminus of the first cleavage of cytosine deaminase, the N-terminus (ZF-right) of a zinc finger protein, or the N-terminus or C-terminus of the second cleavage of cytosine deaminase, even in constructs in which the C-terminus (ZF-left) of the zinc finger protein is fused with the N-terminus of the first cleavage of cytosine deaminase and the C-terminus (ZF-right) of the zinc finger protein is fused with the N-terminus of the second cleavage of cytosine deaminase (CC configuration); the C-terminus (ZF-left) of the zinc finger protein is fused with the N-terminus of the first cleavage of cytosine deaminase and the N-terminus (ZF-right) of the zinc finger protein is fused with the C-terminus of the second cleavage of cytosine deaminase (CN configuration); or the N-terminus (ZF-left) of the zinc finger protein is fused with the C-terminus of the first cleavage of cytosine deaminase and the N-terminus (ZF-right) of the zinc finger protein is fused with the C-terminus of the second cleavage of cytosine deaminase (NN configuration).

[0299] When the enzyme includes a mitotic form of cytosine deaminase and the DNA-binding protein is a TALE, a first TALE can fuse with the first mitotic product of the cytosine deaminase, and a second TALE can fuse with the second mitotic product of the cytosine deaminase. The corresponding constructs can be described as N'-TALE-first mitotic product DddA-C' and N'-TALE-second mitotic product DddA-C'. Adenine deaminase can fuse with the N-terminus or C-terminus of the first mitotic product of the cytosine deaminase or with the N-terminus or C-terminus of the second mitotic product of the cytosine deaminase.

[0300] When the full-length form of cytosine deaminase is included and the DNA-binding protein is TALE, a single TALE module can be N'-TALE-full-length DDDA-C'. Here, the adenine deaminase can be fused to either the N-terminus or the C-terminus of the cytosine deaminase. Alternatively, the adenine deaminase can be fused to either the C-terminus of the single TALE module or to either the N-terminus or the C-terminus of the cytosine deaminase.

[0301] When a full-length cytosine deaminase is included and the DNA-binding protein is a TALE, a dual-TALE module may be included, comprising a first TALE module and a cytosine deaminase (N'-TALE-full-length DDDA-C') in the NC direction, and may also include an adenine deaminase and a second cleavage containing a second TALE (N'-TALE-adenine deaminase-C'). Here, the adenine deaminase may be fused to the N-terminus or C-terminus of the TALE, for example in the constructs of N'-TALE-cytosine deaminase-C' and N'-TALE-adenine deaminase-C'.

[0302] In some cases, UGI (uracil DNA glycosyltransferase inhibitor) can be added to increase base editing efficiency. UGI increases base editing efficiency by inhibiting the activity of UDG (uracil DNA glycosyltransferase), an enzyme that repairs mutated DNA by removing the umba radical from the DNA.

[0303] This invention relates to a method for A-G base editing in prokaryotic or eukaryotic cells, the method comprising a fusion protein or a nucleic acid encoding thereon, wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-associated nuclease, and a cytosine deaminase of the fusion protein or a variant thereof derived from bacteria and specific for double-stranded DNA.

[0304] This invention relates to a composition for A-G base editing in prokaryotic or eukaryotic cells, comprising a fusion protein or nucleic acid encoding it, wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-associated protein.

[0305] The fusion protein contains a cytosine deaminase or a variant thereof derived from bacteria and specific for double-stranded DNA. The DNA-binding protein is fused to the N-terminus and C-terminus of the cytosine deaminase or a variant thereof. Similarly, the DNA-binding protein is also fused to the N-terminus and C-terminus of the adenine deaminase of the fusion protein. In the context of a fusion protein comprising a DNA-binding protein, a cytosine deaminase or a variant thereof, and adenine deaminase, the adenine deaminase may be located at the N-terminus or C-terminus of the cytosine deaminase within the fusion protein, or it may exist as a separate protein independent of the other DNA-binding proteins.

[0306] This invention relates to a composition for C-to-T base editing in prokaryotic or eukaryotic cells, the composition comprising the fusion protein or the nucleic acid encoding it and UGI (uracil glycosylase inhibitor), wherein the DNA-binding protein is a zinc finger protein, TALE protein, or CRISPR-associated nuclease, and the fusion protein is a cytosine deaminase or a variant thereof derived from bacteria and specific for double-stranded DNA.

[0307] Specifically, the present invention relates to a composition for A-to-G base editing (UGI-free) in prokaryotic and eukaryotic cells, the composition comprising 1) a DNA-binding protein, 2) a full-length double-stranded DNA-specific bacterial cytosine deaminase or a variant thereof, and 3) a deoxyadenine deaminase derived from Escherichia coli TadA, wherein the DNA-binding protein is a zinc finger protein (ZFP), a transcription activator-like effector (TALE) array, or a catalytically deficient CRISPR-Cas9 (nCas9 or dCas9) or Cas12a, and the full-length double-stranded DNA-specific bacterial cytosine deaminase is DddAtox derived from Burkholderia cenocepacia.

[0308] This invention relates to a composition for A-to-G base editing (UGI-free) in prokaryotic and eukaryotic cells, the composition comprising 1) a left DNA-binding protein operatively linked to a full-length double-stranded DNA-specific bacterial cytosine deaminase or a variant thereof, and 2) a right DNA-binding protein operatively linked to a deoxyadenine deaminase derived from *Escherichia coli* TadA, wherein the left or right DNA-binding protein is a zinc finger protein (ZFP), a transcription activator-like effector (TALE) array, or a catalytically deficient CRISPR-Cas9 (nCas9 or dCas9) or Cas12a, and the full-length double-stranded DNA-specific bacterial cytosine deaminase is DddAtox derived from *Burkholderia cepacia*. The order of the left and right components in the fusion protein may be interchanged.

[0309] The present invention also relates to a composition for A-to-G and C-to-T base editing in prokaryotic and eukaryotic cells, the composition comprising 1) a DNA-binding protein, 2) a full-length double-stranded DNA-specific bacterial cytosine deaminase or a variant thereof, 3) a deoxyadenine deaminase derived from Escherichia coli TadA, and 4) a UGI (uracil glycosylase inhibitor), wherein the DNA-binding protein is a zinc finger protein (ZFP), a transcription activator-like effector (TALE) array, or a catalytically deficient CRISPR-Cas9 (nCas9 or dCas9) or Cas12a, and the full-length double-stranded DNA-specific bacterial cytosine deaminase is DddAtox derived from Burkholderia cepacia.

[0310] The present invention also relates to a composition for A-to-G base editing (UGI-free) in prokaryotic and eukaryotic cells, the composition comprising 1) a DNA-binding protein, 2) a cleavage double-stranded DNA-specific bacterial cytosine deaminase or a variant thereof, and 3) a deoxyadenine deaminase derived from *Escherichia coli* TadA, wherein the DNA-binding protein is a zinc finger protein (ZFP), a transcription activator-like effector (TALE) array, or a catalytically deficient CRISPR-Cas9 (nCas9 or dCas9) or Cas12a, and the cleavage double-stranded DNA-specific bacterial cytosine deaminase is DddAtox derived from *Burkholderia cepacia*.

[0311] The present invention also relates to a composition for A-to-G and C-to-T base editing in prokaryotic and eukaryotic cells, the composition comprising 1) a DNA-binding protein, 2) a cleavage double-stranded DNA-specific bacterial cytosine deaminase or a variant thereof, 3) a deoxyadenine deaminase derived from Escherichia coli TadA, and 4) UGI (uracil glycosylase inhibitor), wherein the DNA-binding protein is a zinc finger protein (ZFP), a transcription activator-like effector (TALE) array, or a catalytically deficient CRISPR-Cas9 (nCas9 or dCas9) or Cas12a, and the cleavage double-stranded DNA-specific bacterial cytosine deaminase is DddAtox derived from Burkholderia cepacia.

[0312] This invention relates to a method for performing A-to-G base editing in prokaryotic or eukaryotic cells, the method comprising treating prokaryotic or eukaryotic cells with a fusion protein or a nucleic acid encoding thereon, wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-associated nuclease, and the fusion protein is a cytosine deaminase or a variant thereof derived from bacteria and specific for double-stranded DNA.

[0313] This invention relates to a method for A-G base editing in prokaryotic or eukaryotic cells, the method comprising treating prokaryotic or eukaryotic cells with a fusion protein or nucleic acid encoding thereon, wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-associated nuclease, and the fusion protein contains a cytosine deaminase or a variant thereof derived from bacteria and specific for double-stranded DNA. The DNA-binding protein is fused to the N-terminus and C-terminus of the cytosine deaminase or a variant thereof. Similarly, the DNA-binding protein is also fused to the N-terminus and C-terminus of the adenine deaminase of the fusion protein. In the context of a fusion protein comprising a DNA-binding protein, a cytosine deaminase or a variant thereof, and an adenine deaminase, the adenine deaminase may be located at the N-terminus or C-terminus of the cytosine deaminase within the fusion protein, or may exist as a separate protein independent of other DNA-binding proteins.

[0314] This invention relates to a method for C-to-T base editing in prokaryotic or eukaryotic cells, the method comprising treating prokaryotic or eukaryotic cells with a fusion protein or a nucleic acid encoding thereon and a UGI (uracil glycosylase inhibitor), wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-associated nuclease, and the fusion protein is a cytosine deaminase or a variant thereof derived from bacteria and specific for double-stranded DNA.

[0315] The specific sequence of the components included in the composition or method according to the present invention is as follows.

[0316] ND1-ZFP-Right-1397C-AD ( Figure 61 Small image f- Figure 61 Small image g: SEQ ID No: 410

[0317] MLGFVGRVAAAPASGALRRLTPSASLPPAQLLLRAAPTAVHPVRDYAAQDYKDDDDKVDEMTKKFGTLTIHDTEKAAEFGIRIPGEKPFQCRICMRNFSDSGNLRVHI RTHTGEKPYKCPDCGKSFSQSSSLIRHQRTHTGEKPYECDHCGKSFSQSSHLNVHKRTHTGEKPYRCKYCDRSFSISSNLQRHVRNIHLRSGTPHEVGVYTLSGTPHEV GVYTLAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN

[0318] ND1-ZFP-Left-1397C-UGI ( Figure 61 Small image f- Figure 61 Small image g: SEQ ID No: 411

[0319] MLGFVGRVAAAPASGALRRLTPSASLPPAQLLLRAAPTAVHPVRDYAAQDYKDDDDKVDEMTKKFGTLTIHDTEKAAEFGIRIPGEKPFQCRICMRNFSDSGNLRVHIRTHTGEKPYKCPDCGKSFSQSSSLIRHQRTHTGEKPYECDHCGKSFSQSSHLNVHKRTHTGEKPYRCKYCDRSFSISSNLQRHVRNIHLRSGTPHEVGVYTLSGTPHEVGVYTLAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML

[0320] ND1-ZFP-Right-1397N-UGI ( Figure 61 Subfigure f- Figure 61 Subfigure g: SEQ ID No: 412)

[0321] MLGFVGRVAAAPASGALRRLTPSASLPPAQLLLRAAPTAVHPVRDYAAQYPYDVPDYAVDEMTKKFGTLTIHDTEKAAEFGIHGVPAAMGGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGSGTPHEVGVYTLSGTPHEVGVYTLYKCPECGKSFSTKNSLTEHQRTHTGEKPYKCPECGKSFSSKKALTEHQRTHTGEKPYECNYCGKTFSVSSTLIRHQRIHTGEKPYRCKYCDRSFSISSNLQRHVRNIHLRSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML

[0322] Trac-ZFP-Left-1397N-UGI ( Figure 61 Subfigure d: SEQ ID No: 413)

[0323] MAPKKKRKVGIHGVPAAMGGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGSGTPHEVGVYTLSGTPHEVGVYTLFQCRICMRKFATSGSLTRHTKIHTGEKPFQCRICMRNFSRSDHLSTHIRTHTGEKPFACDICGRKFATSSNRTKHTKIHTHPRAPIPKPFQCRICMRNFSRSDNLSEHIRTHTGEKPFACDICGRKFAWHSSLRVHTKIHLRSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML

[0324] Trac-ZFP-Right-1397C-AD ( Figure 61 Small figure d: SEQ ID No: 414)

[0325] MAPKKKRKVGIHGVPAAMAERPFQCRICMRNFSRSDHLSTHIRTHTGEKPFACDICGRKFADRSHLARHTKIHTGSQKPFQCRICMRKFALKQHLNEHTKIHTGEKPFQCRICMRNFSQSGNLARHIRTHTGEKPFACDICGRKFAHNSSLKDHTKIHLRSGTPHEVGVYTLSGTPHEVGVYTLAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN

[0326] ND1-Left-TALE-1397N-UGI ( Figure 62 : SEQ ID No: 415)

[0327] MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALLVGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML

[0328] ND1 - Left - TALE - 1397C - UGI ( Figure 62 : SEQ ID No: 416)

[0329] MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALLVGSAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML

[0330] ND1-right-TALE-1397N-UGI ( Figure 62 : SEQ ID No: 417)

[0331]

[0332] ND1-Right-TALE-1397C-UGI ( Figure 62 (SEQ ID No: 418)

[0333] MASVLTPLLLRGLTGSARRLPVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDKGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALLVGSAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML

[0334] ND4-left-TALE-AD ( Figure 62 Small figure b: SEQ ID No: 419)

[0335] MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAMDIADLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPQQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPQQVVAIASNGGGRPALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN

[0336] ND4-Right-TALE-AD ( Figure 62 Small image b: SEQ ID No: 420

[0337] MASVLTPLLLRGLTGSARRLPVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDKMDIADLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPQQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNGGGRPALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN

[0338] ND1 - left - TALE - 1333N - AD ( Figure 63 : SEQ ID No: 421)

[0339] MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALLVGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN

[0340] ND1 - Left - TALE - 1333C - AD ( Figure 63:SEQ ID No: 422)

[0341] MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALLVGSPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN

[0342] ND1 - Right - TALE - 1333N - AD ( Figure 63 : SEQ ID No: 423)

[0343]

[0344] ND1-Right-TALE-1333C-AD ( Figure 63 : SEQ ID No: 424)

[0345]

[0346] ND1-Left-TALE-1397C-AD Figures 62-63 (SEQ ID No: 425)

[0347] MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLLVGSAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN

[0348] ND1 - Right - TALE - 1397C - AD ( Figures 62-63 :SEQ ID No: 426)

[0349]

[0350] ND1 - Left - TALE - 1333N ( Figure 63 : SEQ ID No: 427)

[0351] MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGG

[0352] ND1 - Left - TALE - 1333C ( Figure 63 : SEQ ID No: 428)

[0353] MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSGSPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC

[0354] ND1 - Left - TALE - 1397N ( Figures 62-63 : SEQ ID No: 429)

[0355] MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEG

[0356] ND1 - Right - TALE - 1333N ( Figure 63 : SEQ ID No: 430)

[0357] MASVLTPLLLRGLTGSARRLPVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDKGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALLVGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGG

[0358] ND1 - Right - TALE - 1333C ( Figure 63 : SEQ ID No: 431)

[0359] MASVLTPLLLRGLTGSARRLPVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDKGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALLVGSPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC

[0360] ND1 - right - TALE - 1397N ( Figures 62-63 : SEQ ID No: 432)

[0361] MASVLTPLLLRGLTGSARRLPVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDKGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALLVGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEG

[0362] ND4-left-TALE-1333N-AD ( Figure 63 : SEQ ID No: 433)

[0363]

[0364] ND4 - Left - TALE - 1333C - AD ( Figure 63 : SEQ ID No: 434)

[0365]

[0366] ND4-LEFT-TALE-1397C-AD ( Figure 63 (SEQ ID No: 435)

[0367]

[0368] ND4-right-TALE-1333N-AD ( Figure 63 : SEQ ID No: 436)

[0369]

[0370] ND4-right-TALE-1333C-AD ( Figure 63 : SEQ ID No: 437)

[0371]

[0372] ND4-right-TALE-1397C-AD( Figure 63 (SEQ ID No: 438)

[0373]

[0374] ND1-Left-TALE-AD-GSVG ( Figure 64 (SEQ ID No: 439)

[0375]

[0376] ND1 - Left - TALE - AD - E1347A ( Figure 64 : SEQ ID No: 440)

[0377]

[0378] ND1-Left-TALE-AD-AAAAA Figure 64 (SEQ ID No: 441)

[0379]

[0380] ND1-Right-TALE-AD ( Figure 64 (SEQ ID No: 442)

[0381] MASVLTPLLLRGLTGSARRLPVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDKGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALLVSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN

[0382] ND1-Left-TALE-GSVG ( Figure 64 : SEQ ID No: 443)

[0383] MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALLVGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLEGKVFSSGGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGVIPVKRGATGETKVFTGNSNGPKSPTKGGC

[0384] ND1-Left-TALE-E1347A ( Figure 64: SEQ ID No: 444)

[0385] MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALLVGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVAGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC

[0386] ND1 - Left - TALE - AAAAA( Figure 64 SEQ ID No: 445)

[0387] MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALLVGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVAAGATGETAVFTGNSNSPASPTAGGC

[0388] ND1-TALE - Left (SEQ ID No: 446)

[0389] DLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAAL

[0390] ND1-TALE-Right (SEQ ID No: 447)

[0391] DLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAAL

[0392] ND4-TALE - Left (SEQ ID No: 448)

[0393] MDIADLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPQQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPQQVVAIASNGGGRPALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLG

[0394] ND4-TALE - Right (SEQ ID No: 449)

[0395] MDIADLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPQQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNGGGRPALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLG

[0396] TRAC-ZFP-left (SEQ ID No: 450)

[0397] FQCRICMRKFATSGSLTRHTKIHTGEKPFQCRICMRNFSRSDHLSTHIRTHTGEKPFACDICGRKFATSSNRTKHTKIHTHPRAPIPKPFQCRICMRNFSRSDNLSEHIRTHTGEKPFACDICGRKFAWHSSLRVHTKIHLR

[0398] TRAC-ZFP-Right (SEQ ID No: 451)

[0399] FQCRICMRNFSRSDHLSTHIRTHTGEKPFACDICGRKFADRSHLARHTKIHTGSQKPFQCRICMRKFALKQHLNEHTKIHTGEKPFQCRICMRNFSQSGNLARHIRTHTGEKPFACDICGRKFAHNSSLKDHTKIHLR

[0400] ND1-ZFP-Left (SEQ ID No: 452)

[0401] FQCRICMRNFSDSGNLRVHIRTHTGEKPYKCPDCGKSFSQSSSLIRHQRTHTGEKPYECDHCGKSFSQSSHLNVHKRTHTGEKPYRCKYCDRSFSISSNLQRHVRNIHLR

[0402] ND1-ZFP-Right (SEQ ID No: 453)

[0403] YKCPECGKSFSTKNSLTEHQRTHTGEKPYKCPECGKSFSSKKALTEHQRTHTGEKPYECNYCGKTFSVSSTLIRHQRIHTGEKPYRCKYCDRSFSISSNLQRHVRNIHLR

[0404] G1333N-DddAtox (SEQ ID No: 454)

[0405] GSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGG

[0406] G1333C-DddAtox (SEQ ID No: 455)

[0407] GSPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC

[0408] G1397N-DddAtox (SEQ ID No: 456)

[0409] GSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEG

[0410] G1397C-DddAtox (SEQ ID No: 457)

[0411] GSAIPVKRGATGETKVFTGNSNSPKSPTKGGC

[0412] Adenine deaminase (AD: ABE 8e) (SEQ ID No: 458)

[0413] SEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN

[0414] Full-length DddAtox variant GSVG (SEQ ID No: 459)

[0415] GSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLEGKVFSSGGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGVIPVKRGATGETKVFTGNSNGPKSPTKGGC

[0416] Full-length DddAtox variant E1347A (SEQ ID No: 460)

[0417] GSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVAGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC

[0418] Full-length DddAtox variant AAAAA (SEQ ID No: 461)

[0419] GSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVAAGATGETAVFTGNSNSPASPTAGGC

[0420] UGI (SEQ ID No: 462)

[0421] TNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML

[0422] Single module

[0423] (TALE)-Linker-AD-GSVG (SEQ ID No: 463)

[0424] SGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSINLVGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLEGKVFSSGGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGVIPVKRGATGETKVFTGNSNGPKSPTKGGC

[0425] (TALE)-Linker-AD-E1347A (SEQ ID No: 464)

[0426] SGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSINLVGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVAGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC

[0427] (TALE)-Linker-AD-AAAAA (SEQ ID No: 465)

[0428] SGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSINLVGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVAAGATGETAVFTGNSNSPASPTAGGC

[0429] Dual module

[0430] (TALE)-Linker-AD (SEQ ID No: 466)

[0431] SGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN

[0432] (TALE)-GSVG(SEQ ID No: 467)

[0433] GSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLEGKVFSSGGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGVIPVKRGATGETKVFTGNSNGPKSPTKGGC

[0434] (TALE)-E1347A(SEQ ID No: 468)

[0435] GSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVAGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC

[0436] (TALE)-AAAAA(SEQ ID No: 469)

[0437] GSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVAAGATGETAVFTGNSNSPASPTAGGC

[0438] (TALE)-1397C-AD(SEQ ID No: 470)

[0439] GSAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN

[0440] (TALE)-1333N-AD (SEQ ID No: 471)

[0441] GSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN

[0442] (TALE)-1333C-AD (SEQ ID No: 472)

[0443] GSPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN

[0444] Linker (AA represents amino acid)

[0445] 8AA: SGGGLGST (SEQ ID No: 473)

[0446] 16AA:SGSETPGTSESATPES(SEQ ID No: 474)

[0447] 32AA:SGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID No: 475)

[0448] SOD2 MTS-3xHA:

[0449] MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYA (SEQ IDNo: 476)

[0450] COX8A MTS-3xFLAG:

[0451] MASVLTPLLLRGLTGSARRLPVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDK (SEQ ID No: 477)

[0452] 6. Delivery

[0453] The fusion protein according to the invention can be delivered to cells by various methods known in the art, such as microinjection, electroporation, DEAE-glucan treatment, lipid transfection, nanoparticle-mediated transfection, protein transduction domain-mediated introduction, and PEG-mediated transfection, but the invention is not limited thereto.

[0454] Another aspect of the present invention relates to a nucleic acid encoding a fusion protein.

[0455] Nucleic acids are used interchangeably with "polynucleotides," "nucleotides," "nucleotide sequences," and "oligonucleotides." They can include nucleotides of any length in polymeric form, deoxyribonucleotides, or ribonucleotides or their analogues. Polynucleotides can have any three-dimensional structure and can perform any function, known or unknown. Polynucleotides can include at least one modified nucleotide, such as a methylated nucleotide or a nucleotide analogue. Modifications to the nucleotide structure can be made before or after polymer assembly.

[0456] Polynucleotides can have RNA sequences, DNA sequences, or combinations thereof (RNA-DNA combined sequences).

[0457] To express fusion proteins, known expression vectors can be used, such as plasmid vectors, granular vectors, phage vectors, etc., and those skilled in the art can easily construct vectors using DNA recombination technology according to any known method.

[0458] The vector can be a plasmid vector or a viral vector. In particular, examples of viral vectors include, but are not limited to, adenovirus, adeno-associated virus, lentivirus, and retroviral vectors.

[0459] The recombinant expression vector may contain a form of nucleic acid suitable for expression in a host cell and may include at least one regulatory element selected on a host cell basis to enable the recombinant expression vector to be used for expression, i.e., the regulatory element is operatively linked to the nucleic acid sequence to be expressed.

[0460] In recombinant expression vectors, "operably ligated" means that the target nucleotide sequence is ligated to a regulatory element in a manner that allows the expression of the nucleotide sequence (e.g., in an in vitro transcription / translation system or in a host cell when the vector is introduced into the host cell).

[0461] Recombinant expression vectors can be provided in a form suitable for messenger RNA synthesis, including a T7 promoter, which means that at least one regulatory element is included to enable in vitro mRNA synthesis, i.e. messenger RNA synthesis via T7 polymerase.

[0462] "Regulatory elements" can include promoters, enhancers, internal ribosome entry sites (IRES), and other expression control elements (such as transcription termination signals, like polyadenylation signals and poly-U sequences). Regulatory elements include those that direct the induced or constitutive expression of nucleotide sequences in many types of host cells, and those that direct the expression of nucleotide sequences only in certain host cells (e.g., tissue-specific regulatory sequences). Tissue-specific promoters can direct expression primarily in desired tissues of interest such as muscle, neurons, bone, skin, blood, specific organs (e.g., liver, pancreas), or specific cell types (e.g., lymphocytes). Regulatory elements can also direct expression in a transiently dependent manner, such as in a cell cycle or developmental stage-dependent manner, and may or may not be tissue- or cell-specific.

[0463] In some cases, the vector includes at least one pol III promoter, at least one pol II promoter, at least one pol I promoter, or a combination thereof. Examples of pol III promoters include, but are not limited to, the U6 and H1 promoters. Examples of pol II promoters include, but are not limited to, the retroviral Rous sarcoma virus (RSV) LTR promoter (optionally having an RSV enhancer), the cytomegalovirus (CMV) promoter (optionally having a CMV enhancer) (e.g., Boshart et al. (1985) Cell41:521-530), the SV40 promoter, the dihydrofolate reductase promoter, the β-actin promoter, the glycerol phosphokinase (PGK) promoter, and the EF1α promoter.

[0464] "Regulatory elements" may include enhancers such as WPRE; CMV enhancers; the R-U5' fragment in the LTR of HTLV-I; the SV40 enhancer; and intron sequences between exons 2 and 3 of rabbit β-globin. Those skilled in the art will understand that the design of expression vectors can depend on a variety of factors, such as the selection of host cells to be transformed, the desired expression level, etc. Vectors can be introduced into host cells to form transcripts, proteins, or peptides (e.g., regularly spaced clustered short palindromic repeats (CRISPR) transcripts, proteins, enzymes, their mutants, their fusion proteins, etc.) that encode fusion proteins or peptides as described herein. Useful vectors may include lentiviral and adeno-associated virus vectors, and these types of vectors may also be selected to target certain cell types.

[0465] Vectors can be delivered in vivo or into cells via microinjection (e.g., direct injection into lesions or target sites), electroporation, lipid transfection, viral vectors, nanoparticles, PTD (protein translocation domain) fusion protein methods, etc.

[0466] Nucleic acids can be injected in the form of ribonucleic acid, such as messenger ribonucleic acid (mRNA), allowing for unrestricted editing of gene bases in cells such as animal or plant cells.

[0467] The nucleic acid according to the invention can be in the form of mRNA, and when delivered in the form of mRNA, the transcription to mRNA process is unnecessary compared to delivery using a vector with DNA, and therefore gene editing can be initiated rapidly. The possibility of transient protein expression is high.

[0468] The inventors of this application have discovered that when a cytosine base editor is injected into plant cells in the form of ribonucleic acid (e.g., messenger RNA) for plant organelle gene editing, off-target effects are reduced compared to delivery with plasmids. In plant organelle gene editing, the advantage of off-target effects compared to plasmids is demonstrated for the first time when the cytosine base editor is transformed into plant cells in the form of mRNA.

[0469] mRNA can be delivered directly or via a vector. In some cases, the mRNA of nucleases and / or cleavage factors can be chemically modified or delivered directly as synthetic self-replicating RNA.

[0470] Methods for delivering mRNA molecules into cells in vitro or in vivo are considered, including methods for delivering mRNA into cells in vivo or into cells of organisms such as humans or animals. For example, mRNA molecules can be delivered into cells using lipids (e.g., liposomes, micelles, etc.), nanoparticles or nanotubes, or cationic compounds (e.g., polyethyleneimine or PEI). In some cases, bioemission methods such as gene guns or bioemission particle delivery systems can be used to deliver mRNA into cells.

[0471] Examples of carriers may include, but are not limited to, cell-penetrating peptides (CPPs), nanoparticles, and polymers.

[0472] CPP is a short peptide that promotes the uptake of various molecular cargoes by cells, from nanoparticles to small chemical molecules and large DNA fragments.

[0473] Regarding nanoparticles, the compositions according to the invention can be delivered via polymer nanoparticles, metal nanoparticles, metal / inorganic nanoparticles, or lipid nanoparticles. Polymer nanoparticles can be, for example, DNA nanowires or linear DNA nanoparticles synthesized via rolling circle amplification. DNA nanowires or linear DNA nanoparticles can be loaded with mRNA and coated with PEI to improve endosome escape. These complexes bind to the cell membrane, are internalized, and then delivered to the cell nucleus via endosome escape.

[0474] Regarding metal nanoparticles, gold particles can be linked and complexed with cationic endosome-disrupting polymers, and thus delivered to cells. Cationic endosome-disrupting polymers may include, for example, polyethyleneimine, poly(arginine), poly(lysine), poly(histidine), poly-[2-{(2-aminoethyl)amino}-ethyl-asparagine (pAsp(DET)), block copolymers of poly(ethylene glycol) (PEG) and poly(arginine), block copolymers of PEG and poly(lysine), or block copolymers of PEG and poly{N-[N-(2-aminoethyl)-2-aminoethyl]asparagine} (PEG-pAsp(DET)).

[0475] Regarding metal / inorganic nanoparticles, mRNA can be encapsulated using, for example, zeolite imidazole ester backbone-8 (ZIF-8).

[0476] In some cases, negatively charged mRNA can be coupled with cationic materials to form nanoparticles, which can penetrate cells through receptor-mediated endocytosis or phagocytosis.

[0477] Examples of cationic polymers may include polyallylamine (PAH); polyethyleneimine (PEI); poly(L-lysine) (PLL); poly(L-arginine) (PLA); polyvinylamine homopolymers or copolymers; poly(vinylbenzyl-tri-C1-C4-alkylammonium salt); polymers of aliphatic or alicyclic dihalides and aliphatic N,N,N',N'-tetra-C1-C4-alkyl-alkylene diamines; poly(vinylpyridine) or poly(vinylpyridinium salt); poly(N,N-diallyl-N,N-di-C1-C4-alkyl-ammonium halide); homopolymers or copolymers of quaternized acrylic acid or di-C1-C4-alkyl-aminoethyl methacrylate; POLYQUAD™; polyaminoamides, etc.

[0478] Cationic lipids can include cationic liposome formulations. The lipid bilayer of the liposome protects the encapsulated nucleic acid from degradation and prevents specific neutralization by antibodies capable of binding to the nucleic acid. During endosome maturation, the endosome membrane and liposome fuse, allowing cationic lipid-nucleases to efficiently escape from the endosome. Examples of cationic lipids include polyethyleneimine, star-radial polyamide amine (PAMAM) dendritic macromolecules, Lipofectin (a combination of DOTMA and DOPE), liposomal enzymes, LIPOFECTAMINE® (e.g., Lipofectamine® 2000, Lipofectamine® 3000, Lipofectamine® RNAiMAX, Lipofectamine® LTX), SAINT-RED (Synvolux Therapeutics, Groningen, Netherlands), DOPE, Cytofectin (Gilead Sciences, Foster City, California), and Eufectin (JBL, San Luis Obispo, California). Representative cationic liposomes can be prepared from N-[1-(2,3-dioleoyloxy)-propyl]-N,N,N-trimethylammonium chloride (DOTMA), N-[1-(2,3-dioleoyloxy)-propyl]-N,N,N-trimethylmethylammonium sulfate (DOTAP), 3β-[N-(N',N'-dimethylaminoethane)carbamoyl]cholesterol (DC-Chol), 2,3-dioleoyloxy-N-[2(speramidoamino)ethyl]-N,N-dimethyl-1-trifluoroacetate propylamine ester (DOSPA), 1,2-dimyristoyloxypropyl-3-dimethyl-hydroxyethylammonium bromide, or dioctadecyl dimethylammonium bromide (DDAB).

[0479] Regarding lipid nanoparticles, they can be delivered using liposomes as carriers. Liposomes are spherical vesicle structures consisting of a single or multiple lipid bilayer surrounding an internal aqueous compartment and a relatively impermeable outer lipophilic phospholipid bilayer. Liposome formulations may primarily contain natural phospholipids and lipids such as 1,2-distearate-sn-glycerol-3-phosphatidylcholine (DSPC), sphingomyelin, phosphatidylcholine, or monosialogangliosides. In some cases, cholesterol or 1,2-dioleoyl-sn-glycerol-3-phosphate ethanolamine (DOPE) can be added to the lipid membrane to address instability in plasma. The addition of cholesterol reduces the rapid release of the encapsulated bioactive compound into the plasma, or 1,2-dioleoyl-sn-glycerol-3-phosphate ethanolamine (DOPE) increases stability.

[0480] 7. Base Editing

[0481] Another aspect of the invention relates to a composition for base editing comprising a fusion protein or nucleic acid.

[0482] Another aspect of the invention relates to a base editing method comprising treating cells with the composition.

[0483] After DNA-binding proteins such as TALE or ZFP (zinc finger proteins) bind to target DNA, the fusion protein's cytosine deaminase hydrolyzes the amino group of cytosine, converting it to uracil. Since uracil can form base pairs with adenine, during intracellular DNA replication, the cytosine-guanine base pair can ultimately be edited into a thymine-adenine base pair via the uracil-adenine base pair. Furthermore, the fusion protein's adenine deaminase hydrolyzes the amino group of adenine to convert it to hypoxanthine. Similarly, since hypoxanthine can form base pairs with cytosine, during intracellular DNA replication, the adenine-thymine base pair can be edited into a guanine-cytosine base pair via the hypoxanthine-cytosine base pair.

[0484] The cells can be eukaryotic cells (e.g., fungi such as yeast, eukaryotic animal and / or eukaryotic plant-derived cells (e.g., animal embryonic cells, stem cells, somatic cells, germ cells, etc.), eukaryotic animals (e.g., primates such as humans, monkeys, dogs, pigs, cattle, sheep, goats, mice, rats, etc.) or eukaryotic plants (e.g., algae such as green algae, corn, soybeans, wheat, rice, etc.), but are not limited to these.

[0485] (1) Base editing of plant cell DNA

[0486] This invention relates to a composition or method for base editing of plant cell DNA. The composition for base editing in plant cells comprises a fusion protein or nucleic acid encoding thereof; and a nuclear localization signal (NLS) peptide, a chloroplast transport peptide, a mitochondrial targeting signal (MTS), a nuclear export signal, or nucleic acid encoding thereof.

[0487] The present invention also provides a composition for base editing in plant cells, the composition comprising a fusion protein or nucleic acid; and a nuclear localization signal (NLS) peptide or nucleic acid encoding thereon.

[0488] The present invention also provides a composition for base editing in plant cells, the composition comprising a fusion protein or nucleic acid; and a chloroplast transport peptide or a nucleic acid encoding the same.

[0489] The present invention also provides a composition for base editing in plant cells, the composition comprising a fusion protein or nucleic acid; and a mitochondrial targeting signal (MTS) or a nucleic acid encoding thereon.

[0490] In some cases, the present invention also provides a composition for base editing in plant cells, said composition further comprising a nuclear output signal or nucleic acid encoding thereon.

[0491] Specifically, the present invention relates to a composition or method for base editing of nuclear DNA, mitochondrial DNA or chloroplast DNA in plant cells.

[0492] Specifically, the fusion protein can be delivered to plant cells through the following methods:

[0493] Gene gun injection (bombardment);

[0494] PEG-mediated protoplast transfection;

[0495] Protoplast transfection via electroporation; or

[0496] Protoplast injection via microinjection.

[0497] The polynucleotide sequence encoding the fusion protein according to the present invention can be an RNA sequence, a DNA sequence, or a combination thereof (RNA-DNA combined sequence).

[0498] The polynucleotide encoding the fusion protein can be delivered to plant cells via the following methods:

[0499] Transformation was performed using Agrobacterium species such as Agrobacterium tumefaciens and Agrobacterium rhizogenes.

[0500] - Binary carrier,

[0501] - Viral vectors: Dual virus, tobacco brittle virus (TRV), tomato mosaic virus (ToMV), foxtail mosaic virus (FoMV), barley yellow-striped mosaic virus (BYSMV), sow thistle yellow-webbed bullet virus (SYNV), etc.;

[0502] Using the virus for transfection;

[0503] Injection (bombardment, gene gun);

[0504] PEG-mediated protoplast transfection;

[0505] Protoplast transfection via electroporation; or

[0506] Protoplast injection via microinjection.

[0507] Examples of viruses may include dual viruses in viral vectors, such as tobacco brittle virus (TRV), tomato mosaic virus (ToMV), foxtail mosaic virus (FoMV), barley yellow-striped mosaic virus (BYSMV), and sow thistle yellow-webbed rash virus (SYNV).

[0508] Vectors can be delivered into cells via microinjection (e.g., direct injection into lesions or target sites), electroporation, lipid transfection, viral vectors, nanoparticles, PTD (protein translocation domain) fusion protein methods, etc.

[0509] Regarding the proteins or nucleic acids that are transported to plant organelles, these organelles can be mitochondria, chloroplasts, or plastids (leucoplasts, chromoplasts).

[0510] Proteins transported to plant organelles can be, for example, chloroplast transport peptides or mitochondrial targeting signals (MTS).

[0511] For example, chloroplast transport peptides (CTPs) or mitochondrial targeting signals (MTS) bind and are then delivered to chloroplasts or mitochondria in plant cells. Upon delivery to the chloroplasts or mitochondria, the remainder, except for the N-terminal CTP or MTS, is delivered as a precursor protein. During entry into the chloroplasts or mitochondria, the delivered protein portion is dissociated and targets the chloroplasts or mitochondria to induce site-specific base editing.

[0512] In addition to fusion proteins or the nucleic acids encoding them, chloroplast transport peptides (CTPs) or the nucleic acids encoding them, or mitochondrial targeting signals (MTS) or the nucleic acids encoding them, can be fused and delivered to plant cells, enabling base editing of plant mitochondrial, chloroplast, chromoplast, or leucoplast DNA.

[0513] When the nuclear output signal is linked to a base-editing protein during mitochondrial gene editing, base editing can be achieved with greater efficiency. The nuclear output signal can be derived from, for example, MVM (mouse parvovirus), but the invention is not limited thereto. The nuclear output signal may include, for example, the amino acid sequence of SEQ ID NO: 31, but is not limited thereto.

[0514] VDEMTKKFGTLTIHDTEK (SEQ ID NO: 31)

[0515] The present invention also includes a TAL (transcription activator-like) effector (TALE)-FokI nuclease or a nucleic acid encoding thereof, said nuclease cutting wild-type DNA sequences but not the edited base sequences; or a ZFN (zinc finger nuclease) or a nucleic acid encoding thereof, particularly the mitochondrial nuclease mitoTALEN (mitochondrial TALE nuclease) or a nucleic acid encoding thereof; or a ZFN (zinc finger nuclease) or a nucleic acid encoding thereof, thereby anticipating more efficient mitochondrial base editing even when mitochondrial sequences are used to cut proteins simultaneously.

[0516] (2) Base editing of animal cell DNA

[0517] This invention relates to a composition or method for base editing of DNA in animal cells. The composition for base editing in animal cells comprises a fusion protein or nucleic acid encoding thereof; and a nuclear localization signal (NLS) peptide, a mitochondrial targeting signal (MTS), a nuclear export signal, or nucleic acid encoding thereof.

[0518] The present invention also provides a composition for base editing in animal cells, the composition comprising a fusion protein or nucleic acid; and a nuclear localization signal (NLS) peptide or nucleic acid encoding thereon.

[0519] The present invention also provides a composition for base editing in animal cells, the composition comprising a fusion protein or nucleic acid; and a mitochondrial targeting signal (MTS) or a nucleic acid encoding the MTS.

[0520] In some cases, the present invention also provides a composition for base editing in animal cells, said composition further comprising a nuclear output signal or nucleic acid encoding thereon.

[0521] Animal cells are non-human animal cells, and treatment with nuclear output signals or nucleic acids encoding them and / or mitochondrial targeting signals (MTS) or nucleic acids encoding them enables base editing of mitochondrial DNA in non-human animal cells.

[0522] For example, in addition to fusion proteins or the nucleic acids encoding them, mitochondrial targeting signals (MTS) bind and are delivered to the mitochondria. Upon delivery to the mitochondria, the remainder, except for the N-terminal MTS, is delivered as a precursor protein. During entry into the mitochondria, the delivered protein portion is dissociated and targets mitochondrial DNA to induce site-specific base editing.

[0523] This invention relates to a composition or method for base editing of mitochondrial DNA in non-human animal cells, wherein a nuclear export signal (NES) or a nucleic acid encoding it is coupled with a mitochondrial targeting signal, a TAL effector, and a cytosine deaminase (DddA). tox The TALE-DdCBE (TALE DddA-derived cytosine base editor) or the nucleic acid fusion encoding it. Adding a nuclear output signal to the fusion protein can reduce nuclear DNA base editing at sites with similar DNA sequences.

[0524] According to the present invention, more efficient base editing of animal mitochondrial DNA can be achieved by including nuclear export signals or nucleic acids encoding them. Furthermore, nuclear DNA base editing of mitochondrial nucleus-like sequences can be reduced due to nuclear export signals, thereby allowing editing only mitochondrial DNA.

[0525] The nuclear output signal may be derived from, for example, MVM (mouse parvovirus), but the invention is not limited thereto. The nuclear output signal may include, for example, the amino acid sequence VDEMTKKFGTLTIHDTEK (SEQ ID NO: 31), but is not limited thereto.

[0526] Before simultaneously or sequentially editing (1) the nuclear output signal or the nucleic acid encoding it and (2) the DNA-binding protein, deaminase or its variant or the nucleic acid encoding it, the wild-type DNA base sequence can be cut but not the edited base sequence or the nucleic acid encoding it, or the ZFN (zinc finger nuclease) or the TAL (transcription activator-like) effector (TALE)-FokI nuclease encoding the nucleic acid, or the ZFN (zinc finger nuclease) or the nucleic acid encoding it, the wild-type DNA base sequence can be injected into eukaryotic cells.

[0527] In particular, base editing of mitochondrial genes in eukaryotic cells may include nuclear export signals or the nucleic acids encoding them and / or mitochondrial targeting signals (MTS) or the nucleic acids encoding them.

[0528] According to the present invention, when nuclear output signals are linked to base editing proteins during animal mitochondrial gene editing, base editing may be more efficient, and non-specific base editing of homologous sequences in the cell nucleus is also suppressed in animal embryos.

[0529] In this invention, by further including a mitochondrial nuclease, namely mitoTALEN or the nucleic acid encoding it, mitochondrial base editing can be achieved with higher efficiency, even when mitochondrial nucleases are used simultaneously. Mitochondrial DNA can be cut using the mitochondrial DNA nuclease mitoTALEN, and the wild-type mitochondrial genome can be cut to obtain a base-edited genome in animals with high efficiency.

[0530] In this invention, by further including a mitochondrial nuclease mitoTALEN or a nucleic acid encoding it, mitochondrial base editing can be achieved with higher efficiency, even when mitochondrial nucleases are used simultaneously. Specifically, it may include a fusion protein (mitoTALEN) comprising a TAL effector domain or ZFN connecting a mitochondrial targeting signal and a FokI nuclease, or a nucleic acid encoding it.

[0531] Mitotemic DNA can be cut using the mitochondrial DNA nuclease mitoTALEN (mitochondrial TALE nuclease), and wild-type mitochondrial genomes can be cut to obtain base-edited genomes in animals with high efficiency.

[0532] In some cases, UGI (uracil DNA glycosyltransferase inhibitor) can be added to increase base editing efficiency. UGI increases base editing efficiency by inhibiting the activity of UDG (uracil DNA glycosyltransferase), an enzyme that repairs mutated DNA by catalyzing the removal of U from DNA.

[0533] Specifically, a DddA-derived cytosine base editor (DdCBE), consisting of the fissuring bacterial intertoxin DddAtox, a transcription activator-like effector (TALE) designed to bind to DNA, and a uracil glycosylase inhibitor (UGI), enables targeted cytosine-thymine base editing in mitochondrial DNA. According to the implementation scheme, efficient mitochondrial DNA editing is possible in mouse embryos. Targeting MT-ND5 (ND5), the subunit encoding the NADH dehydrogenase that catalyzes NADH dehydration and electron transfer to ubiquinone in mitochondrial genes, induces mutations associated with human mitochondrial diseases, such as m.G12918A, and mutations producing early stop codons, such as m.C12336T. Therefore, mitochondrial disease models can be constructed in mice, suggesting the potential for treating mitochondrial diseases.

[0534] (2) A DNA-binding protein, a deaminase or a variant thereof or a nucleic acid encoding thereof may be linked to (1) a nuclear output signal or a nucleic acid encoding thereof, and (3) a nuclease mitoTALEN or a nucleic acid encoding thereof may be linked to (1). For delivery of (1) to (3), a single delivery vector or multiple delivery vectors may be used in combination with the same or different conformations.

[0535] (1) It may be included in a first delivery carrier, (2) It may be included in a second delivery carrier, and (3) It may be included in a third delivery carrier. These individual delivery systems may be viral delivery carriers, may be viral and non-viral delivery carriers, or may be non-viral delivery carriers at the same time.

[0536] The nuclear output signals (1) to (3), DdCBE and mitoTALEN can be mixed and delivered.

[0537] At least one of (1) to (3) can be delivered to the nuclear output signal, DdCBE or mitoTALEN, and some can be delivered by positioning the DNA sequence encoding (1) to (3) on a vector.

[0538] The DNA sequences encoding (1) to (3) above may be located on the same vector and delivered simultaneously through a single vector, or may be located on different vectors and delivered.

[0539] Animals according to the present invention may include humans or non-human animals. Examples of non-human transgenic animals may be insects, annelids, mollusks, brachiopods, nematodes, coelenterates, sponges, chordates, and vertebrates. Vertebrates may be fish, amphibians, reptiles, birds, or mammals. Insects may be fruit flies, nematodes may be Caenorhabditis elegans, fish may be zebrafish, and mammals may be primates, carnivores, insectivores, rodents, even-toed ungulates, perissodactyls, or proboscis animals. Rodents may include rats or mice.

[0540] Base-edited animals can be produced by introducing the composition according to the invention into the embryos of non-human animals, transferring the embryos to a surrogate mother, and resulting in pregnancy. The composition according to the invention can be introduced into fertilized eggs of animals and cultured.

[0541] The resulting fertilized eggs can be transferred to a surrogate mother and delivered. Further steps may include verifying whether the non-human transgenic animal is transgenic after delivery. The non-human transgenic animal can be mated to produce offspring transgenic animals.

[0542] "Offspring" refers to all viable offspring of transgenic animals produced by mating non-human transgenic animals. More specifically, it can be the F1 generation produced by mating transgenic animals with each other or with normal animals, the F2 generation produced by mating F1 animals with normal animals, and subsequent generations, but the invention is not limited thereto.

[0543] The characteristics of mating can be either mating of transgenic animals or normal animals. This invention may include cells, tissues, and byproducts isolated from transgenic animals or their offspring. Byproducts may include any material derived from transgenic rabbits, and are preferably selected from blood, serum, urine, feces, saliva, organs, and skin, but are not limited thereto.

[0544] Example

[0545] A better understanding of the present invention can be obtained through the following embodiments. These embodiments are for illustrative purposes only and should not be construed as limiting the scope of the invention, as will be apparent to those skilled in the art.

[0546] Example 1. Zinc finger deaminase (ZFD)

[0547] Base editing of nuclear or mitochondrial DNA is widely used in biomedical research, medicine, and biotechnology. The ZFD platform comprises a DNA-binding protein, the cleaving bacterial intertoxin deaminase DDDAtox, and a uracil glycosylase inhibitor (UGI). Here, ZFD catalyzes targeted C-to-T base conversion without inducing unwanted small insertions and deletions (insertion-deletion) in human cells. Using publicly available zinc finger resources, plasmids encoding ZFD are constructed, enabling base editing at frequencies up to 60% in nuclear DNA and 30% in mitochondrial DNA. Unlike CRISPR-based base editing, ZFD does not produce single-strand or double-strand breaks through DNA cleavage, thus avoiding unwanted insertions and deletions (insertion-deletion) at the target site caused by error-prone non-homologous end joining. Furthermore, recombinant ZFD proteins purified from *E. coli* spontaneously penetrate human cells to induce targeted base conversion. This demonstrates a proof-of-concept for gene therapy without genetic factors.

[0548] Technologies used for genome editing in eukaryotic cells and organisms include, but are not limited to, zinc finger nucleases (ZFNs), transcription activator-like effector (TALE) nucleases (TALENs), TALE-linked cytosine base editors derived from the intermittent bacterial deaminase toxin DddA (also known as DdCBEs), CRISPR-Cas9, and Cas9-linked deaminases (also known as base editors) without cleavage activity. These tools are, in principle, composed of two functional units: a DNA-binding part and a catalytic part. Thus, the zinc finger array or TALE array serves as the DNA-binding part, while the nuclease (FokI in ZFNs and TALENs) or deaminase (the splitting DddAtox in DdCBEs and APOBEC1 in CBEs) serves as the catalytic unit. CRISPR-Cas9 possesses both nuclease and RNA-directed DNA-binding protein functions. Custom-designed programmable nucleases such as ZFNs, TALENs, and Cas9 cleave DNA, creating double-strand breaks, whose repair induces gene knockout and knock-in in a targeted manner. However, programmable nuclease-induced double-strand breaks can cause undesirable large gene deletions at the target site, p53 activation, and chromosomal rearrangements during two parallel DSB repairs at both the target and off-target sites. In contrast, programmable base editors, including cytosine and adenine base editors (CBE and ABE), do not generate DSBs, thus avoiding these undesirable events in the cell, and efficiently catalyze single nucleotide conversions without requiring repair templates or donor DNA. However, CBEs or ABEs containing Cas9 nickase variants cleave the target DNA strand to create nicks or single-strand breaks, resulting in undesirable insertions or deletions at the gene target site.

[0549] CBEs catalyze C-to-T base conversion in nuclear and mitochondrial DNA. We used a custom-designed DdCBE to demonstrate mitochondrial DNA editing in mice and chloroplast DNA editing in plants. We also generated zinc finger deaminases (ZFDs) by linking split DddAtox hemispheres to custom-designed zinc finger proteins for precise, insertion- and deletion-free base editing in human and other eukaryotic cells. Because zinc finger arrays (2 × 0.3–0.6 kbase pairs) are smaller than TALE arrays (2 × 1.7–2 kbase pairs) or Streptococcus pyogenes Cas9 (4.1 kbase pairs), genes encoding ZFDs can be readily packaged into viral vectors with limited cargo space, such as AAVs for in vivo studies and gene therapy applications. Unlike TALE arrays, zinc finger arrays lack large domains at the C- or N-terminus, making them engineer-friendly. Split DddAtox hemispheres can be fused to the C- or N-terminus of zinc finger proteins. Furthermore, zinc finger proteins, with their intrinsic ability to penetrate cells, enable nucleic acid-free gene editing in human cells. These properties make zinc finger proteins an ideal platform for DNA-binding modules used for base editing in the nucleus or other organelles.

[0550] 1-1. Materials and Methods

[0551] plasmid construction

[0552] p3s-ZFD plasmids for mammalian expression were generated by modifying the p3s-ABE7.10 plasmid (Addgene, #113128) after digestion with HindIII and XhoI (NEB) enzymes. The digested p3s plasmids and synthetic insert DNA were assembled using the HiFi DNA Assembly Kit (NEB). All insert DNA encoding MTS, ZFP (Toolgen, Sangamo, and Barbas modules), split DddA, or UGI was synthesized by IDT. pTarget plasmids were designed to determine the optimal length of the spacer sequence for ZFD activity. Each pTarget plasmid was constructed by inserting the ZFP binding sequence and spacer sequence into the pRGS-CCR5-NHEJ reporter plasmid digested with two enzymes (EcoRI and BamHI, NEB), with spacers of varying lengths between the two ZFP binding sites. The pET-ZFD plasmid for protein production in *E. coli* was generated by modifying the pET-Hisx6-rAPOBEC1-XTEN-nCas9-UGI-NLS plasmid (Addgene, #89508) after digestion with NcoI and XhoI (NEB) enzymes. The ZFD sequence was amplified from the p3s-ZFD plasmid using PCR, and Hisx6 and GST tag sequences were synthesized as oligonucleotides (Macrogen). All plasmids were generated using the HiFi DNA Assembly Kit (NEB) to insert the sequences encoding the ZFD and the tags for protein purification into the digested pET plasmid. The plasmids were transformed using chemically competent DH5α *E. coli* cells, and purified using the AccuPrep Plasmid Microextraction Kit (Bioneer) according to the manufacturer's protocol. After identifying the entire sequence using Sanger sequencing, the desired plasmids were selected.

[0553] HEK293T cell culture and transfection

[0554] HEK 293T cells (ATCC CRL-11268) were cultured in Durbeco modified Eagle medium (Welgene) supplemented with 10% fetal bovine serum (Welgene) and 1% antibiotic-antifungal solution (Welgene). HEK 293T cells (7.5 × 10⁻⁶) were cultured at a density of 10% fetal bovine serum (Welgene) and 1% antibiotic-antifungal solution (Welgene). 4Cells were seeded into 48-well plates. After 18–24 hours, cells were transfected with plasmids encoding left and right ZFDs (500 ng each) or along with the pTarget plasmid (10 ng) at 70–80% confluence using Lipofectamine 2000 (1.5 μl, Invitrogen). Cells were harvested 96 hours post-transfection and then lysed by incubating in 100 μl of cell lysis buffer (50 mM Tris-HCl (pH 8.0) (Sigma-Aldrich), 1 mM EDTA (Sigma-Aldrich), 0.005% sodium dodecyl sulfate (Sigma-Aldrich)) supplemented with 5 μl proteinase K (Qiagen) at 55°C for 1 hour, followed by incubation at 95°C for 10 minutes. For whole-mtDNA sequencing, HEK 293T cells were transfected with serially diluted concentrations of the mitoZFD pair targeting ND1 or ND2. It shows 7.5 × 10 4 The amount of construct delivered per cell (ng). mtDNA was isolated from cells 96 hours post-transfection.

[0555] K562 cell culture and transfection

[0556] K562 cells were cultured in RPMI 1640 supplemented with 10% fetal bovine serum (Welgene) and 1% antibiotic-antifungal solution (Welgene). To deliver ZFD into K562 cells via electroporation, an Amaxa 4D-Nucleofector™ X-cell system with programmed FF-120 (Lonza) was used. When using a 16-well Nucleocuvette™ Strip, the maximum volume of substrate solution added to each sample was 2 μl. K562 cells (1 × 10⁻⁶) were cultured in RPMI 1640 supplemented with 10% fetal bovine serum (Welgene) and 1% antibiotic-antifungal solution (Welgene). 5Transfect each left and right ZFD protein with 220 pmol (for maximum capacity) or 110 pmol (for half of maximum capacity) or 500 ng of plasmids encoding left and right ZFD. 96 hours post-treatment, cells were harvested by centrifugation at 100 g for 5 min and lysed by incubation at 55ºC for 1 h followed by incubation at 95ºC for 10 min in 100 μl cell lysis buffer supplemented with 5 μl proteinase K (Qiagen) (50 mM Tris-HCl (pH 8.0) (Sigma-Aldrich), 1 M EDTA (Sigma-Aldrich), 0.005% sodium dodecyl sulfate (Sigma-Aldrich)) supplemented with 5 μl proteinase K (Qiagen). To deliver ZFD or plasmids encoding ZFD directly into K562 cells, refer to previous methods used for direct ZFN delivery. Dilute a mixture of left and right ZFD proteins (final concentration 50 μM) or a mixture of plasmids encoding left and right ZFD (500 ng each) to a final volume of 20 μl using serum-free medium containing 100 mM L-arginine and 90 μM ZnCl2 at pH 7.4. In K562 cells (1 × 10⁻⁶), dilute to a final volume of 20 μl. 5 Centrifuge at 100 g for 5 minutes and discard the supernatant. Resuspend the cells in diluted ZFD solution and incubate at 37ºC for 1 hour. After incubation, centrifuge the cells at 100 g for 5 minutes and resuspend them in fresh culture medium. Incubate the cells at 30ºC (for transient cryogenic conditions) or 37ºC for 18 hours, then allow them to regrow at 37ºC for two days. Perform this procedure twice on some cells. Analyze the cells 96 hours after treatment.

[0557] ZFD protein expression and purification

[0558] Plasmids encoding each pair of ZFDs (each with a C-terminal GST tag) were transformed into Rosetta (DE3) competent cells and cultured in LB agar plates containing kanamycin. After overnight culture, single colonies were picked and cultured overnight at 37ºC in liquid medium containing 50 μg / ml kanamycin and 100 μM ZnCl2 (pre-culture). The next day, a portion of the pre-culture was transferred to a large volume of liquid medium and then cultured at 37ºC until the absorbance at A600 nm was approximately 0.5–0.70. The culture was placed on ice for approximately 1 hour, and then ZFD protein expression was induced by adding 0.5 mM isopropyl β-D-1-thiogalactopyranoside (IPTG; GoldBio), and the culture was incubated at 18ºC for 14 hours.

[0559] During protein purification, cells were resuspended in lysis buffer (50 mM Tris-HCl (Sigma-Aldrich), 500 mM NaCl (Sigma-Aldrich), 1 mM MgCl2 (Sigma-Aldrich), 10 mM 1,4-dithiothreitol (DTT; GoldBio), 1% Triton X-10 (Sigma-Aldrich), 10% glycerol, 1 mM benzyl sulfonyl fluoride (Sigma-Aldrich), 1 mg / ml lysozyme from egg white (Sigma-Aldrich), 100 µM ZnCl2 (Sigma-Aldrich), 100 mM arginine (Sigma-Aldrich, pH 8.0) and then sonicated (3 min total, 5 s on, 10 s off) for further lysis. The solution was then centrifuged (13,000 rpm) to extract only the supernatant. The supernatant was incubated for 1 hour with the addition of glutathione agarose 4B (GE Healthcare). After incubation, the resin-lysate mixture was placed in a column and then washed three times with wash buffer (50 mM Tris-HCl (Sigma-Aldrich), 500 mM NaCl (Sigma-Aldrich), 10 mM DTT (GoldBio), 1 mM MgCl2 (Sigma-Aldrich), 100 µM ZnCl2 (Sigma-Aldrich), 10% glycerol, 100 mM arginine (Sigma-Aldrich, pH 8.0). Proteins attached to the resin were eluted from the resin using elution buffer (50 mM Tris-HCl (Sigma-Aldrich), 500 mM NaCl (Sigma-Aldrich), 1 mM MgCl2 (Sigma-Aldrich), 40 mM glutathione (Sigma-Aldrich), 10% glycerol, 1 mM DTT (GoldBio), 100 µM ZnCl2 (Sigma-Aldrich), 100 mM arginine (Sigma-Aldrich, pH 8.0). Finally, the eluted protein was concentrated to a concentration of approximately 15 ng / μl (200-240 pmol / μl, depending on protein size).

[0560] In vitro deamination of PCR amplicon via ZFD

[0561] Amplicons containing the TRAC site were prepared using PCR. 8 µg of amplicons were incubated with 2 µg of each ZFD protein (left-G1397N and right-G1397C) in NEB3.1 buffer containing 100 µM ZnCl2 at 37ºC for 1–2 h. After the reaction, the ZFD protein was removed by incubation with 4 µl of proteinase K solution (Qiagen) at 55ºC for 30 min, and the amplicons were purified using a PCR purification kit (MGmed). 1 µg of purified amplicons were incubated with 2 units of USER enzyme (NEB) at 37ºC for 1 h. The amplicons were then incubated with 4 µl of proteinase K solution (Qiagen) and purified again using a PCR purification kit (MGmed). The purified PCR products were electrophoresed on an agarose gel and imaged.

[0562] Targeted deep sequencing

[0563] To analyze the base editing ratios at the target and off-target sites, overlapping primary, secondary, and tertiary PCR amplifications of the target sites were performed using PrimeSTAR® GXL DNA polymerase (TAKARA) and primers containing TruSeq HT double indexes to generate a deep sequencing library. Paired-end sequencing of the library was then performed using Illumina MiniSeq.

[0564] mRNA preparation

[0565] DNA templates containing the T7 RNA polymerase promoter upstream of the ZFD sequence were generated by PCR using forward and reverse primers (forward: 5'-CATCAATGGGCGTGGATAG-3' SEQ ID No: 116, reverse: 5'-CATCAATGGGCGTGGATAG-3' SEQ ID No: 117, reverse: 5'-GACACCTACTCAGACAATGC-3' SEQ ID No: 118). mRNA was then synthesized in vitro using the mMESSAGE mMACHINE™ T7 ULTRA Transcription Kit (Thermo Fisher). The in vitro transcribed mRNA was purified using the MEGAclear™ Transcription Clearance Kit (Thermo Fisher) according to the manufacturer's protocol.

[0566] Whole mitochondrial genome sequencing

[0567] Whole mitochondrial genome sequencing requires three steps. 1. Extract mtDNA from isolated mitochondria: 96 hours after transfection with mitoZFD targeting ND1 or ND2, 3 × 10⁻⁶ mtDNA samples are extracted. 5HEK293T cells were trypsinized and collected by centrifugation (500 g, 4 min, 4ºC). Cells were then washed with phosphate-buffered saline (Welgene) and collected again by centrifugation. The supernatant was removed, and mitochondria were isolated from the cultured cells using a reagent-based method with a Mitochondrial Isolation Kit for Cultured Cells (ThermoFisher), according to the manufacturer's protocol. mtDNA was then extracted from the isolated mitochondria using the DNeasy Blood and Tissue Kit (Qiagen). 2. NGS Library Generation: An NGS library was generated from the extracted mtDNA using the Illumina DNA Preparation Kit with Nextera™ DNA CD Index (Illumina). 3. NGS: The library was pooled and loaded onto a MiniSeq sequencer (Illumina). Average sequencing depth > 50.

[0568] Mitochondrial whole-genome DNA editing analysis

[0569] To analyze NGS data from whole-mitochondrial genome sequencing, Fastq files were aligned to the GRCh38.p13 (v102) reference genome using BWA, and BAM files with SAMtools (v.1.9) were generated by fixing the readings and pairing information and tags. Then, the REDItoolDenovo.py script from REDItools (v.1.2.1) was used to identify positions in the mitochondrial genome with a base editing rate of 1% or higher in all cytosine and guanine. Positions with a base editing rate of 50% or higher were considered single nucleotide variants in the cell line and excluded from all samples. For off-target analysis, target sites for each ZFD were excluded. Remaining positions with an editing frequency ≥ 1% were considered off-target sites, and the number of edited C / G nucleotides was counted. To calculate the average C / G to T / A base editing frequency for all C / G nucleotides in the mitochondrial genome, the editing rate at off-target sites was averaged. The specificity ratio was calculated by dividing the average on-target editing frequency by the average off-target editing frequency. A mitochondrial genome map is generated by plotting the base editing rates at the target and off-target sites.

[0570] 1-2. Optimization of ZFD Constructs

[0571] To develop a ZFD for base editing in human and other eukaryotic cells, the amino acid linker and spacer lengths of the zinc finger protein (ZFP) attached to the splitting DddAtox hemisphere were optimized. C-to-T base conversion was induced in the spacer between the left and right ZFP binding sites. Well-characterized ZFN pairs targeting the human CCR5 gene were selected. Using the same method, ZFDs with various linkers of 2, 5, 10, 16, 24, and 32 amino acids were prepared, and a series of target plasmids with various spacers ranging in length from 1 to 24 base pairs, along with the left and right ZFP binding sites of the ZFD and repeating TC sequences, were constructed. Figure 1 Small picture a and Figure 1 (See Figure b and Table 1).

[0572] [Table 1] pTarget library sequence

[0573]

[0574] DddAtox can split at two sites (G1333 and G1397), and each half can fuse to either the left or right ZFP. For each of the 24 target plasmids with spacers, the base editing efficiency of the resulting 24 (= 6 adapters × 2 splitting sites × ZFP fusion sites (left or right)) ZFD constructs was measured. Measurements were performed using deep sequencing on day 4 after transfection in Hek293T cells.

[0575] [Constructor]

[0576] *Left-ZFD: SV40 NLS-ZFP (S162-Left)-Connector-DddAtox Half-Connector-UGI

[0577] *Right-ZFD: SV40 NLS-ZFP (S162-Right)-Connector-DddAtox Half-Connector-UGI

[0578] - SV40 NLS:

[0579] PKKKRKV (SEQ ID No: 478)

[0580] - ZFP (S162-left)

[0581] GIHGVPAAMAERPFQCRICMRNFSDRSNLSRHIRTHTGEKPFACDICGRKFAISSNLNSHTKIHTGSQKPFQCRICMRNFSRSDNLARHIRTHTGEKPFACDICGRKFATSGNLTRHTKIHLR (SEQ ID NO: 2)

[0582] - ZFP (S162-right)

[0583] GIHGVPAAMAERPFQCRICMRNFSRSDNLSVHIRTHTGEKPFACDICGRKFAQKINLQVHTKIHTGEKPFQCRICMRNFSRSDVLSEHIRTHTGEKPFACDICGRKFAQRNHRTTHTKIHLR (SEQ ID NO: 3)

[0584] - The linker between zinc finger proteins and the DddAtox hemisphere:

[0585] - 2aa: GS

[0586] - 5aa: TGEKP (SEQ ID No: 479)

[0587] - 10aa: SGAQGSTLDF (SEQ ID No: 9)

[0588] - 16aa: SGSETPGTSESATPES (SEQ ID No: 10)

[0589] - 24aa: SGTPHEVGVYTLSGTPHEVGVYTL (SEQ ID No: 115)

[0590] - 32aa:GSGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID No: 11)

[0591] -Split DddAtox G1333-N

[0592] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGG (SEQ ID No: 27)

[0593] -Split DddAtox G1333-C

[0594] PTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC (SEQ ID No: 272)

[0595] -Split DddAtox G1397-N

[0596] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEG (SEQ ID No: 273)

[0597] -Split DddAtox G1397-C

[0598] AIPVKRGATGETKVFTGNSNSPKSPTKGGC (SEQ ID No: 26)

[0599] -4aa connector

[0600] SGGS (SEQ ID NO: 480)

[0601] - UGI

[0602] TNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML (SEQ ID NO: 481)

[0603] ZFDs with short linkers (2 and 5 amino acid (AA) linkers) exhibit low or no efficiency. In contrast, ZFDs with 10 AA or more linkers show C-to-T base editing efficiencies in spacers with 4 or more base pairs ranging from 1% to 24%. Figure 1 Small image c Figure 2 Small picture a and Figure 2 (Figure c). Among the ZFD pairs, the ZFD pair with 24AA connectors showed the highest editing efficiency. To determine the optimal connector combination, ZFD combinations were constructed in which the left ZFD of the ZFD pair was secured with a 24AA connector and connectors of various lengths were used for the right ZFD, or vice versa, and their editing efficiency was measured (Figure c). Figure 1 Small image d and Figure 3 Using the same 24AA connector at both sites was found to be most effective. It was also found that the DddAtox splitting compound at G1397 was more effective than the DddAtox splitting compound at G1333. Figure 1 Small image c Figure 2 Small picture a and Figure 2 (Figure b). Therefore, these are the most efficient ZFD pairs for editing cytosine, with the highest efficiency of >6.7% in spacer subregions of 7–21 bp in length (Figure b). Figure 1 Small image c Figure 2 Small picture a and Figure 2Small image c).

[0604] 1-3. Base Editing in Nuclear DNA Targets

[0605] This study investigated whether ZFDs with a 24AA linker in human cells could catalyze C-to-T base editing at target sites on chromosomes in vivo. Twenty-two pairs of ZFDs targeting 11 sites across eight genes were constructed (one pair of two ZFDs per site). Figure 4 Fourteen pairs of ZFDs were assembled using publicly available zinc finger resources. An additional eight pairs of ZFDs were generated by modifying previously characterized ZFNs (specific to CCR5 and TRAC). This is because ZFNs cleave target DNA in spacers of 5–7 bp in length, while ZFDs function in spacers of at least 7 bp, resulting in ZFDs that function by attaching one or both zinc fingers to or separating them from a ZFN pair. Since four different conformations of ZFNs can be constructed by fusing a FokI nuclease to the N-terminus or C-terminus of a ZFP, two pairs of ZFDs with different conformations were constructed. Figure 4 Trac-NC in b), and tested whether the split DddAtox half could be fused with the N-terminus of the ZFP and the C-terminus of the existing ZFP ( Figure 4 Small picture a and Figure 5 (NC configuration shown).

[0606] [Constructor]

[0607] *Type C: SV40 NLS-Zinc Finger Protein-24aa Linker-DddA tox Half-part - 4aa connector - UGI

[0608] *N type: SV40 NLS-DddA tox Half-part - 24aa connector - zinc finger protein - 4aa connector - UGI

[0609] - SV40 NLS

[0610] PKKKRKV (SEQ ID No: 478)

[0611] - 24AA connector

[0612] SGTPHEVGVYTLSGTPHEVGVYTL

[0613] -SplitDddA tox G1397-N

[0614] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEG

[0615] - Split DddA tox G1397-C

[0616] AIPVKRGATGETKVFTGNSNSPKSPTKGGC

[0617] - 4aa linker

[0618] SGGS

[0619] - UGI

[0620] TNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML

[0621] - ZFP

[0622] CCR5-1 left (C type) [S162 ZFN-left]

[0623] GIHGVPAAMAERPFQCRICMRNFS DRSNLSR HIRTHTGEKPFACDICGRKFA ISSNLNS HTKIHTGSQKPFQCRICMRNFS RSDNLAR HIRTHTGEKPFACDICGRKFA TSGNLTR HTKIHLR

[0624] CCR5-1 right (C type) [S162 ZFN-right]

[0625] GIHGVPAAMAERPFQCRICMRNFS RSDNLSV HIRTHTGEKPFACDICGRKFA QKINLQV HTKIHTGEKPFQCRICMRNFS RSDVLSE HIRTHTGEKPFACDICGRKFA QRNHRTT HTKIHLR

[0626] CCR5-2 left (C type) [S162 ZFN-left]

[0627] GIHGVPAAMAERPFQCRICMRNFS DRSNLSR HIRTHTGEKPFACDICGRKFA ISSNLNS HTKIHTGSQKPFQCRICMRNFS RSDNLAR HIRTHTGEKPFACDICGRKFA TSGNLTR HTKIHLR (SEQ ID NO: 2)

[0628] CCR5-2 Right (Type C) [Using Barbas zinc finger module group with additional ZF modification S162 ZFN-Right]

[0629] GIHGVPAAMAERPFQCRICMRNFS QSGDLRR HIRTHTGEKPFACDICGRKFA RSDNLSV HTKIHTGSQKPFQCRICMRNFS QKINLQV HIRTHTGEKPFACDICGRKFA RSDVLSE HTKIHLR (SEQ ID NO: 482)

[0630] TRAC - Left (Type C) [Adapted from Paschon, DE et al., 2019]

[0631] GIHGVPAAMAERPFQCRICMRNFS DQSNLRA HIRTHTGEKPFACDICGRKFA TSSNRK THTKIHTGSQKPFQCRICMRNFS LQQTLAD HIRTHTGEKPFACDICGRKFA QSGNLAR HTKIHLR (SEQ ID NO: 483)

[0632] TRAC-Left (N-type) [Adapted from Paschon, DE et al., 2019]

[0633] FQCRICMRKFA TSGSLTR HTKIHTGEKPFQCRICMRNFS RSDHLST HIRTHTGEKPFACDICGRKFA T SSNRTK HTKIHTHPRAPIPKPFQCRICMRNFS RSDNLSE HIRTHTGEKPFACDICGRKFA WHSSLRV HTKIHLR (SEQ ID NO: 484)

[0634] TRAC-Right (Type C) [from Paschon, DE et al., 2019]

[0635] GIHGVPAAMAERPFQCRICMRNFS RSDHLST HIRTHTGEKPFACDICGRKFA DRSHLAR HTKIHTGSQKPFQCRICMRKFA LKQHLNE HTKIHTGEKPFQCRICMRNFS QSGNLAR HIRTHTGEKPFACDICGRKFA HNSSLKD HTKIHLR (SEQ ID NO: 485)

[0636] MFAP1 Left (Type C) [Redesigned using Toolgen Zinc Finger Module Group]

[0637] GIRIPGEKPYSCGICGKSFS DSSAKRR HCILHTGEKPYTCSDCGKAFR DKSCLNR HRRTHTGEKPYKCMECGKAFN RRSHLTR HQRIHTGEKPYECNYCGKTFS VSSTLIR HQRIHLR (SEQ ID NO: 486)

[0638] MFAP1 Right (Type C) [Using Barbas zinc finger module group with additional ZF modification S162 ZFN-Right]

[0639] GIRERPYACPVESCDRRFS TSGSLVR HIRIHTGQKPFQCRICMRNFS RSDELTR HIRTHTGEKPFACDICGRKFA RSDHLTT HTKIHTGEKPFQCRICMRKFA QSSNLVR HTKIHLR (SEQ ID NO: 487)

[0640] CCDC28B left (Type C) [Using Barbas zinc finger module group with additional ZF modification S162 ZFN-right]

[0641] GIRERPYACPVESCDRRFS DPGHLVR HIRIHTGQKPFQCRICMRNFS RSDELTR HIRTHTGEKPFACDICGRKFA RSDHLTT HTKIHTGEKPFQCRICMRKFA RSDKLVR HTKIHLR (SEQ ID NO: 488)

[0642] CCDC28B Right (Type C) [Redesigned using Toolgen Zinc Finger Module Group]

[0643] GIRIPGEKPYECNYCGKTFS VSSTLIR HQRIHTGEKPYHCDWDGCGWKFA RSDELTR HYRKHTGEKPFECKDCGKAFI QKSNLIR HQRTHTGEKPYHCDWDGCGWKFA RSDELTR HYRKHLR (SEQ ID NO: 489)

[0644] KDM4B Left (Type C) [Using Barbas zinc finger module group with additional ZF modification S162 ZFN-Right]

[0645] GIRERPYACPVESCDRRFS DCRDLAR HIRIHTGQKPFQCRICMRNFS RSDELTR HIRTHTGEKPFACDICGRKFA RSDHLTT HTKIHTGEKPFQCRICMRKFA RSDKLVR HTKIHLR (SEQ ID NO: 490)

[0646] KDM4B Right (Type C) [Redesigned using Toolgen Zinc Finger Module Group]

[0647] GIRIPGEKPFECKDCGKAFI QKSNLIR HQRTHTGEKPYRCEECGKAFR WPSNLTR HKRIHTGEKPYRCEECGKAFR WPSNLTR HKRIHTGEKPYSCGICGKSFS DSSAKRR HCILHLR (SEQ ID NO: 491)

[0648] NUMBL Left (Type C) [Using Barbas zinc finger module group with additional ZF modification S162 ZFN-Right]

[0649] GIRERPYACPVESCDRRFS DCRDLAR HIRIHTGQKPFQCRICMRNFS RSDELTR HIRTHTGEKPFACDICGRKFA RSDHLTT HTKIHTGEKPFQCRICMRKFA RSDKLVR HTKIHLR (SEQ ID NO: 492)

[0650] NUMBL Right (Type C) [Redesigned using Toolgen Zinc Finger Module Group]

[0651] GIRIPGEKPYTCSDCGKAFR DKSCLNR HRRTHTGEKPYKCGQCGKFYS QVSHLTR HQKIHTGEKPFECKDCGKAFI QKSNLIR HQRTHTGEKPYHCDWDGCGWKFA RSDELTR HYRKHLR (SEQ ID NO: 493)

[0652] INPP5D-1 Left (Type C) [Using Barbas zinc finger module group with additional ZF modification S162 ZFN-Right]

[0653] GIRERPYACPVESCDRRFS RSDKLVR HIRIHTGQKPFQCRICMRNFS RSDELTR HIRTHTGEKPFACDICGRKFA RSDHLTT HTKIHTGEKPFQCRICMRKFA RSDKLVR HTKIHLR (SEQ ID NO: 494)

[0654] INPP5D-1 Right (Type C) [Redesigned using Toolgen Zinc Finger Module Group]

[0655] GIRIPGEKPYTCSDCGKAFR DKSCLNR HRRTHTGEKPYECNYCGKTFS VSSTLIR HQRIHTGEKPYTCSDCGKAFR DKSCLNR HRRTHTGEKPYTCSDCGKAFR DKSCLNR HRRTHLR (SEQ ID NO: 495)

[0656] INPP5D-2 Left (Type C) [Using Barbas zinc finger module group with additional ZF modification S162 ZFN-Right]

[0657] GIRERPYACPVESCDRRFS RSDKLVR HIRIHTGQKPFQCRICMRNFS RSDELTR HIRTHTGEKPFACDICGRKFA RSDHLTT HTKIHTGEKPFQCRICMRKFA RSDKLVR HTKIHLR (SEQ ID NO: 496)

[0658] INPP5D-2 Right (Type C) [Redesigned using Toolgen Zinc Finger Module Group]

[0659] GIRIPGEKPYECNYCGKTFS VSSTLIR HQRIHTGEKPYTCSDCGKAFR DKSCLNR HRRTHTGEKPYTCSDCGKAFR DKSCLNR HRRTHTGEKPYHCDWDGCGWKFA RSDELTR HYRKH (SEQ ID NO: 497)

[0660] DVL3 Left (Type C) [Redesigned using Barbas Zinc Finger Module]

[0661] GIHGVPAAMAERPFQCRICMRNFS TSGHLVR HIRTHTGEKPFACDICGRKFA TSGHLVR HTKIHTGEKPFQCRICMRNFS TSGELVR HIRTHTGEKPFACDICGRKFA QSSNLVR HTKIHLR (SEQ ID NO: 498)

[0662] DVL3 Right (Type C) [S162 ZFN-Left]

[0663] GIHGVPAAMAERPFQCRICMRNFS DRSNLSR HIRTHTGEKPFACDICGRKFA ISSNLNS HTKIHTGSQKPFQCRICMRNFS RSDNLAR HIRTHTGEKPFACDICGRKFA TSGNLTR HTKIHLR (SEQ ID NO: 499)

[0664] In Hek 293T cells, the C-to-T base editing efficiency of ZFDs (including ZFDs with the NC conformation) ranged from 1.0% to 60%. On the other hand, insertion-deletion (insertion-deletion) was <0.4%, therefore it was rare. Figure 4 Small image b and Figure 6 As seen in the plasmid-based experiments described above, ZFD efficiency for targeting CCR5 with 5 bp spacers was very low. Among targets with at least 7 bp spacers, the other 20 ZFD pairs showed an average editing efficiency of 12.0% ± 3.4%, comparable to the average of 8.3% ± 2.2% for Cas9-derived base editor 2. Furthermore, C-to-T base editing occurred not only in the TC background but also in the A... Cor GC C middle( Figure 4 Small image c- Figure 4 Small image f). A at C6 of NUMBL. C GC at C7 of INPP5D-2 C The corresponding C-base editing efficiencies were 4.58% and 1.85%.

[0665]

[0666] 1-4. Direct delivery of purified ZFD protein into human cells

[0667] Delivering purified gene-editing proteins into cells, rather than delivering plasmid DNA encoding the gene-editing proteins, reduces off-target effects, avoids innate immune responses to foreign DNA, and prevents foreign plasmid DNA from inserting into the genome in vivo. Other groups have shown that ZFPs can spontaneously pass through mammalian cells both in vitro and in vivo. To confirm protein-mediated base editing by ZFDs, ZFD pairs that efficiently target the TRAC gene were selected, and recombinant ZFD proteins with one or four nuclear localization signals (NLS) were purified from *E. coli*. First, the base-editing efficiency of the ZFD proteins was tested in vitro using PCR amplicon with the TRAC site, confirming very high efficiency. Efficiency was confirmed based on gene cleavage using a uracil-specific excision reagent (USER), which is a mixture of uracil DNA glycosylase and DNA glycosylase-lyase endonuclease VIII. Figure 7 The TRAC-NC ZFD protein was delivered to difficult-to-transfect human leukemia K562 cells in two ways (electroporation or direct delivery without electroporation). The ZFD protein was highly effective. The C-to-T base editing efficiency was 26.5% (electroporation) and 17% (direct delivery). Figure 4 (See Figure g). Therefore, these results indicate that plasmids encoding ZFD or purified recombinant ZFD proteins can be used for base editing of nuclear DNA in human cells.

[0668] 1-5. Mitochondrial DNA base editing using ZFD

[0669] Unlike CRISPR-based systems, the splitting DddAtox system, fused with a custom-designed DNA-binding protein, can be used to edit organelle DNA, including mitochondrial DNA. This is a major advantage of the DddA system over CRISPR systems. To deliver ZFDs to mitochondria, mitoZFD is constructed by linking the mitochondrial targeting signal (MTS) and nuclear export signal (NES) to the N-terminal portions of nine ZFDs designed to target mitochondrial genes. Figure 8ZFD ZFP fragments were assembled using publicly available zinc finger resources. ZFDs were configured such that the spacer length was 7-15 bp and both the left and right DNA binding sites were 12 bp long.

[0670] [Constructor]

[0671] *Type C: MTS-FLAG tag-NES-zinc finger protein-24aa linker-DddAtox half-4aa linker-UGI

[0672] *N-type: MTS-HA tag-NES-DddAtox half-section-24aa linker-zinc finger protein-4aa linker-UGI

[0673] -MTS (mitochondrial targeting sequence of human mitochondrial ATP synthase F1β subunit)

[0674] MLGFVGRVAAAPASGALRRLTPSASLPPAQLLLRAAPTAVHPVRDYAAQ (SEQ ID No: 274)

[0675] -FLAG label (Type C)

[0676] DYKDDDDK (SEQ ID No:275)

[0677] -HA label (N type)

[0678] YPYDVPDYA (SEQ ID No:276)

[0679] -NES (Nuclear Output Signal)

[0680] VDEMTKKF (Mouse parvovirus; MVM NES)

[0681] -Split DddAtox G1397-N

[0682] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEG

[0683] -Split DddAtox G1397-C

[0684] AIPVKRGATGETKVFTGNSNSPKSPTKGGC

[0685] -4aa connector

[0686] SGGS

[0687] - UGI

[0688] TNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML

[0689] -ZFP

[0690] *The zinc finger is connected via a TGEKP connector. (ZF1-connector-ZF2-connector-ZF3-connector-ZF4).

[0691] *The mitoZFD targeting ND2 has a QQ variant. This variant is a Q-substituted R or K amino acid. It is marked with a red letter.

[0692]

[0693]

[0694] The mitochondrial DNA base editing efficiency of mitoZFD in HEK293T cells ranged from 2.6% to 30% (mean 14% ± 3%). Figure 9 (Figure a). Furthermore, mitoZFD with the NC configuration (16% ± 4.5%, n = 6) is equally effective as the CC configuration (8.3% ± 2.7%, n = 3). T in the spacer subregion C Most of the background cytosines are edited with varying efficiencies ( Figure 9 Small image b- Figure 9 Small figure g). Furthermore, A is present at the ND2 site. CC Cytosine in (C8 and C9) backgrounds also showed efficiencies of approximately 7.4% and 19.9%, respectively. Figure 9 b). This indicates that ZFD-mediated C-to-T base editing is not limited to the TC motif.

[0695] Furthermore, single-cell-derived clonal populations were isolated from mitochondrial DNA (mtDNA) mutant cells to demonstrate that mitoZFD is not cytotoxic and that the mtDNA mutation was maintained in the clonal populations. Of the 30 single-cell-derived clonal populations isolated from HEK293T cells treated with ND1-specific mitoZFD, five clonal populations showed ND1 gene base editing efficiency ranging from 35% to 98%. Figure 10 Figure a). Similarly, among the 36 single-cell-derived clonal populations isolated from HEK293T cells treated with ND2-specific mitoZFD, seven clonal populations showed ND2 gene base editing efficiencies ranging from 26% to 76%. Figure 10(Figure b). For all other clonal populations, a low efficiency of 0.4%–1.0% was observed, most likely due to sequencing errors. Similar efficiency was obtained in cells not treated with ZFD ( Figure 10 c). These results indicate that mitoZFD induces heterogeneous mutations unevenly across the cell population. Most ZFD-treated cells were wild-type, while cells with heterogeneous mutations exhibited mutation rates as high as 98%. These mutations persisted even after clonal expansion. Figure 11 (and Figure 12).

[0696] 1-6. mitoZFD and TALE-based DdCBE

[0697] The mutational pattern of the constructed ND1-specific mitoZFD was found to be different from that of the TALE-based DdCBE targeting the same gene. Figure 9 Small image f- Figure 9 Small image h). Two mitoZFDs catalyze C-to-T base editing at C5 or C8 ( Figure 9 Small image f), while DdCBE in C8, C9 and C 11 Position-induced base editing ( Figure 9 Small figure g). Therefore, the amino acid changes induced by mitoZFD are completely different from those induced by DdCBE ( Figure 9 (See figure h). The left and right sites linked to mizoZFD are separated by an 8 bp spacer, but in the case of DdCBE, these sites are separated by a 16 bp spacer, which may be responsible for differential mutation patterns. These results suggest that mizoZFD and DdCBE can generate a variety of mutations in mitochondrial DNA in a complementary manner.

[0698] To generate more mutation patterns, the possibility of mixing ZFD and DdCBE monomers to form hybrid pairs was tested. Ten hybrid pairs targeting the ND1 gene showed good activity in HEK293T cells, with an average base editing efficiency of 17% ± 3.4%. Figure 13 In fact, one hybrid pair (TALE-L / ZFD-R1) showed better efficiency than two DdCBE pairs and ten ZFD pairs targeting the same site, with the highest base editing efficiency reaching 41%. Figure 13 Figure b). Furthermore, the hybridization pairs produced mutation patterns different from those obtained using DdCBE and mitoZFD (…). Figure 13(Inset Figure c). A few hybrid pairs (e.g., ZFD-L1 / TALE-R and ZFD-L2 / TALE-R) induced C-to-T transformation at a single location without bystander editing. In contrast, most mitoZFD and DdCBE pairs induced C-to-T transformation at multiple locations in the spacer subregion. These results suggest that ZFD / DdCBE hybrid pairs can produce unique mutational patterns and generate certain mutations that cannot be obtained using ZFD or DdCBE pairs alone.

[0699] 1-7. Mitozonic genome-wide target specificity of mitoZFD

[0700] To confirm whether mitoZFD induced off-target editing, mitochondrial DNA was extracted from cells treated with mitoZFD targeting either ND1 or ND2 genes, and then whole-mitochondrial genome sequencing was performed. Different amounts (5–500 ng) of mRNA or plasmids encoding mitoZFD pairs were transfected into HEK293T cells. As expected, the mid-target editing efficiency was dose-dependent. High concentrations (100, 200, and 500 ng) of mRNA or plasmid produced >30% mid-target efficiency, but also resulted in >1% of hundreds of off-target edits. Figures 14-17 Low concentrations (5 and 10 ng) largely avoided these off-target edits but significantly reduced mid-target efficiency. Medium concentrations (50 ng) of mRNA were most suitable, maintaining high mid-target efficiency without inducing hundreds of off-target edits. To further eliminate remaining off-target edits, an R(-5)Q mutation was introduced into each zinc finger to eliminate non-specific DNA contacts. Compared to mtDNA from cells not treated with ZFD, the resulting ZFD variants (in...) Figure 18 (Displayed as QQ) maintained high target activity and showed precise specificity with minimal off-target editing. Figure 18 Small picture a and Figure 18 Small image b).

[0701] Base editing is a relatively new approach that allows for the editing of targeted bases without causing DNA double-strand breaks or DNA repair templates. Base editing enables C-to-T or A-to-G conversions in cells, animals, and plants, allowing for the study of the functional effects of single nucleotide polymorphisms (SNPs) and the correction of pathogenic point mutations for therapeutic applications. Two types of base editing technologies have been developed: CRISPR-based adenine and cytosine base editing and DddA-based base editing. CRISPR-based base editors consist of catalytically impaired Cas9 or Cas12a as DNA-binding units and single-stranded DNA-specific deaminases derived from rats or E. coli. DdCBE, on the other hand, consists of a TALE DNA-binding array and double-stranded DNA-specific DddAtox.

[0702] Compared to DdCBE, ZFD is smaller in size. This is because the zinc finger proteins in ZFD are compact, while the TALE arrays in DdCBE are bulky. Therefore, the genes encoded by ZFD pairs, rather than those encoded by DdCBE pairs, can be easily packaged into AAV vectors with small cargo space. Furthermore, the compact ZFP is engineer-friendly, allowing the fusion of split DddAtox halves into the N- or C-terminus of the ZFP to produce ZFDs that operate upstream or downstream of the ZFP binding site. In addition, recombinant ZFD proteins can spontaneously penetrate human cells without electroporation or lipid transfection, enabling gene-free gene therapy. ZFD pairs or ZFD / DdCBE hybrids can produce unique mutational patterns that cannot be obtained using DdCBE alone. These properties make ZFD a powerful new platform for modeling and treating mitochondrial diseases.

[0703] Example 2. Chloroplast and mitochondrial gene editing in plants using TALE-DdCBE

[0704] Plant organelles, including mitochondria and chloroplasts, each possess their own genome, encoding numerous genes essential for respiration and photosynthesis. Plant organelle gene editing—an unmet need in plant genetics and biotechnology—is limited by the lack of suitable tools to target DNA within these organelles. To assemble a DddA-derived cytosine base-editing plasmid (DdCBE), a Golden Gate cloning system consisting of 16 expression plasmids (8 for delivering the resulting protein to mitochondria and another 8 to chloroplasts) and a 424 TALE subarray plasmid was developed. The completed DdCBE plasmid was used to efficiently induce point mutations in mitochondria and chloroplasts. DdCBE base editing induced mutations in lettuce or rapeseed callus tissue with efficiencies of up to 25% (mitochondria) and 38% (chloroplasts). To avoid off-target mutations caused by DdCBE-encoding plasmids, DdCBE messenger RNA was transfected into lettuce protoplasts, demonstrating DNA-free base editing in chloroplasts. In addition, streptomycin or spectinomycin-resistant lettuce callus and buds with editing efficiency up to 99% were generated by introducing point mutations into the chloroplast 16S rRNA gene.

[0705] DdCBE is a heterodimer comprising an isolated, nontoxic domain derived from the bacterial cytosine deaminase toxin DddAtox, a site-specific TALE array, and a UGI, which induces cytosine-to-thymine substitution in spacers between TALE protein binding sites in target DNA. We used DdCBE results to demonstrate efficient organelle base editing in plants.

[0706] 2-1. Method

[0707] Construction of expression plasmids for plant protoplast experiments

[0708] The DdCBE Golden Gate target vector was constructed using the Gibson assembly method. Sequences encoding the TAL N-terminal domain, HA tag, FLAG tag, TAL C-terminal domain, split DddAtox, and UGI were codon-optimized for expression in the dicotyledonous plant *Arabidopsis thaliana* and synthesized via integrative DNA technology. Sequences encoding CTP from AtinfA and AtRbcS and MTS from the ATPase δ and ATPase γ subunits were amplified from *Arabidopsis thaliana* cDNA. For plant expression, the PcUbi promoter and pea3A terminator were used instead of the mammalian CMV promoter in the backbone plasmid. To construct a vector for in vitro DdCBE mRNA transcription, the T7 promoter cassette was cloned into the DdCBE Golden Gate target vector between the PcUbi promoter and the DdCBE coding region.

[0709] TALE array genes were constructed using single-factor Golden Gate assembly. Using 424 TALE array plasmids and the target vector, the DdCBE expression plasmid was constructed via BsaI digestion and T4 ligation of the Golden Gate assembly. Single-factor Golden Gate cloning was performed using the following steps: 37ºC and 50ºC for 5 minutes each, 20 cycles, followed by a final reaction at 50ºC for 15 minutes, and then at 80ºC for 5 minutes. All vectors used for plant protoplast transfection were purified using the Plasmid Plus Mini Preparation Kit (Qiagen). The DNA and amino acid sequences used in vector construction are as follows.

[0710] [Table 2]

[0711]

[0712] The specific amino acid sequences of DdCBE and TALE repeat sequences are as follows.

[0713] [Table 3]

[0714]

[0715]

[0716]

[0717]

[0718]

[0719]

[0720]

[0721] mRNA in vitro transcription

[0722] DdCBE DNA templates were prepared by PCR using Phusion DNA polymerase (Thermo Fisher). DdCBE mRNA was synthesized and purified using an in vitro mRNA synthesis kit (Enzynomics).

[0723] Protoplast isolation and transduction

[0724] Lettuce seeds were surface-sterilized in 70% ethanol for 30 seconds, then in 0.4% hypochlorite solution for 15 minutes, and then washed three times with sterile distilled water. The lettuce seeds were then germinated on 0.5xMS medium supplemented with 2% sucrose at 25ºC under 16 hours of light and 8 hours of darkness. Rapeseed seeds were surface-sterilized in 70% ethanol for 3 minutes, then in 1.0% hypochlorite solution for 30 minutes, and then washed three times with sterile distilled water. The rapeseed seeds were then germinated on 1xMS medium supplemented with 3% sucrose at 25ºC under 16 hours of light and 8 hours of darkness.

[0725] Protoplast isolation and transduction were performed as described previously. Cotyledons from 7-day-old lettuce and 14-day-old rapeseed plants were digested with enzyme solution in the dark with shaking (40 rpm) for 3 hours. The protoplast-enzyme mixture was washed with an equal volume of W5 solution, and then intact protoplasts were obtained from the sucrose solution by centrifugation at 80 g for 7 minutes. The protoplasts were treated with W5 solution at 4ºC for 1 hour, and then transfected with polyethylene glycol.

[0726] Lettuce and rapeseed protoplasts resuspended in MMG solution were transfected with plasmids or mRNA using PEG and then cultured at room temperature for 20 minutes. The PEG-protoplast mixture was washed three times with an equal volume of W5 solution under gentle inversion and then cultured for 10 minutes. The protoplasts were then precipitated by centrifugation at 100 g for 5 minutes.

[0727] protoplast culture

[0728] Lettuce protoplasts transfected with a plasmid encoding DdCBE were resuspended in lettuce protoplast medium (LPCM). The protoplasts in the medium were mixed 1:1 with medium containing 2.4% low-melting-point agarose and immediately placed in 6-well plates. After the mixture solidified, the embedded protoplasts were covered with 1 ml of liquid medium and cultured at 25ºC in the dark for 1 week. After the initial culture, the covering liquid medium was replaced weekly with fresh medium, and the embedded protoplasts were cultured for 1 week under 16 hours of light and 8 hours of darkness, followed by 2 weeks under 16 hours of light and 8 hours of darkness. Protoplast-induced microcalls were cultured in regeneration medium at 25ºC for 4 weeks under 16 hours of light and 8 hours of darkness. To prepare for base editing efficiency analysis, protoplasts were cultured in liquid medium at 25ºC in the dark for 1 week without embedding. To test antibiotic resistance, microcallus embedded for one month was cultured for 4 weeks at 25ºC under 16 hours of light and 8 hours of darkness in regeneration medium containing 50 mg / L streptomycin or 50 mg / L spectinomycin. After 4 weeks, antibiotic-resistant green callus or adventitious shoots were transferred to fresh regeneration medium containing 200 mg / L streptomycin or 50 mg / L spectinomycin.

[0729] Rapeseed protoplasts transfected with a plasmid encoding DdCBE were resuspended in rapeseed culture medium. The protoplast-medium mixture was transferred to 6-well plates and cultured at 25ºC in the dark for 2 weeks. After 2 weeks, the protoplasts were cultured for 3 weeks under 16 hours of light and 8 hours of darkness. The culture medium was then replaced with very fresh medium.

[0730] DNA and RNA extraction

[0731] Total DNA or RNA was extracted from cells or transgenic callus cultured in liquid medium using the DNeasy Plant Microkit or RNeasy Plant Microkit. Cultured cells or callus were harvested by centrifugation at 10,000 rpm for 1 minute. cDNA was then reverse transcribed from total RNA using the RNA to cDNA EcoDry Premix (TaKaRa).

[0732] Deep sequencing

[0733] The target region was amplified using a fusion enzyme and appropriate primers (Supplementary Table 1). To create the DNA sequencing library, three rounds of PCR were performed (round 1, nested PCR; round 2, PCR; round 3, indexing PCR). Equal volumes of DNA were collected and then sequenced using a MiniSeq system (Illumina). Paired-end sequencing files were analyzed using a Cas analyzer and the source code of a computer program.

[0734] 2-2. Results

[0735] The Golden Gate assembly system was developed to construct DdCBEs targeting chloroplasts (cp-DdCBE) and DdCBEs targeting mitochondria (mt-DdCBE). Figure 19 The expression plasmid encodes a fusion protein consisting of a chloroplast transport peptide or mitochondrial targeting sequence, the N- or C-terminal domain of TALE, split DddAtox hemispheres (G1333N, G1333C, G1397N, and G1397C), and UGI. The expression plasmid is optimized for expression in dicotyledonous plants under the control of the parsley ubiquitin (PcUbi) promoter and the pea3A terminator. DdCBE plasmids with custom-designed TALE DNA-binding sequences are constructed in a single subcloning step by mixing the expression vector and six TALE subarray plasmids in an E-tube. A total of 424 modular TALE subarray plasmids (6 × 64 tripartites + 2 × 16 bipartites + 2 × 4 monopartites) are available for preparing cp-DdCBE and mt-DdCBE sequences recognizing 16–20 bp long sequences, including the conserved T at the 5' end. Therefore, DdCBE heterodimers functionally recognize 32-40 bp DNA sequences.

[0736] To determine whether DdCBE can promote base editing in chloroplasts, four pairs of cp-DdCBE plasmids were constructed to encode RNA components of chloroplast 16S rRNA genes encoding the 30S ribosomal subunit, and each pair was co-transfected into lettuce and rapeseed protoplasts. Seven days later, base editing efficiency was measured by deep sequencing. Figure 20 Small picture a and Figure 20 (Inset b). The most efficient cp-DdCBE pair (left-G1397-N + right-G1397-C) induced C*G to T*A conversion in a 15 bp spacer between the two TALE binding sites, with an efficiency of 30% in lettuce protoplasts and 15% in rapeseed protoplasts. Figure 20 (Figure b). As with previous results in mammalian cells and mice, cytosine (C9 and C13) in the 5'-TC motif was preferentially converted to thymine via cp-DdCBE. Interestingly, in lettuce protoplasts, another cp-DdCBE (left-G1333-N + right-G1333-C) converted cytosine (C7) in the 5'-AC background to thymine with an efficiency of 4.2%. Furthermore, persistence of base editing by cp-DdCBE in lettuce protoplasts was observed during 14 days of culture. Figure 24 Editing efficiency continued to increase for up to 10 days and remained so throughout the entire culture period.

[0737] Base editing was tested in two additional chloroplast genes, psbA and psbB, which encode the D1 and CP-47 photosynthetic proteins of photosystem II, respectively. Figure 20 Small image c, Figure 20 (Figures d and 25). Among the cp-DdCBEs targeting the Psb gene, the most active one (left-G1397-C + right-G1397-N) can induce C*G to T*A conversion in lettuce protoplasts with an efficiency of up to 25%. Figure 20 (Inset d). The base editor efficiently converts only two cytosines (C11 and C12) in the 5'-TCC background to thymine. 5'-TCC can be converted to 5'-TTC first, then to 5'-TTT. In rapeseed protoplasts, other combinations (left-G1333-N + right-G1333-C) showed the highest efficiency at four cytosine positions (C3, C4, C11, and C12), up to 3.5% (C3). C3 and C4 are in the 5'-TCC background in the rapeseed gene, while they are in the 5'-ACC background in the lettuce counterpart due to single nucleotide polymorphisms, resulting in efficient editing of two cytosines (C3 and C4) by DdCBE in the rapeseed gene but not in the lettuce gene. Similarly, the cp-DdCBE combination targeting the psbB gene catalyzed the conversion of two cytosines in rapeseed protoplasts with efficiencies ranging from 0.36% to 4.1% in a TCC background. Figure 25 In summary, these results indicate that editing efficiency is determined by the position and sequence of cytosine, including the position (G1333 vs. G1397) and orientation (left-G1333-N vs. left-G1333-C) of the DddAtox cleavage fragment, and that cp-DdCBE can perform efficient base editing in plant chloroplast genomes.

[0738] Furthermore, an attempt was made to achieve base editing in plant mitochondrial DNA using a custom-designed mt-DdCBE. To this end, plasmids encoding mt-DdCBE targeting the atp6 gene in lettuce and the rps14 gene in rapeseed were constructed (using the Golden Gate cloning system) and introduced into lettuce and rapeseed protoplasts. Seven days after introduction, base editing efficiency was measured by deep sequencing. Figure 20 Small image e, 20 small images f and Figure 26 The most efficient mt-DdCBE combination (left-G1397-N + right-G1397-C in lettuce and left-G1397-C + right-G1397-N in rapeseed) catalyzed C*G to T*A conversion at the atp6 gene target site with an efficiency of 23% in lettuce protoplasts and 23% in rapeseed protoplasts. Figure 20Furthermore, the mt-DdCBE combination induced C*G to T*A conversion at the rps14 target site in rapeseed protoplasts with an efficiency of 11%. These results indicate that mitochondrial DNA in plants is readily edited using mt-DdCBE.

[0739] To investigate whether the editing of cpDNA and mtDNA by DdCBE was maintained during regeneration, regenerated lettuce and rapeseed callus tissues were collected from DdCBE-treated protoplasts 4 weeks after introduction. Figure 21 Figure a), and the base editing efficiency of each callus tissue was measured using deep sequencing and Sanger sequencing. Figure 21 Small image b and Figure 27 ). Base editing of chloroplast or mitochondrial genes induced by DdCBE showed efficiencies of up to 38% and 25% in 22 out of 26 lettuce calluses and 7 out of 14 rapeseed calluses, respectively. Figure 21 (Inset c). Furthermore, base editing of the chloroplast psbA gene showed an efficiency of up to 3.9% in lettuce callus ( Figure 27 Similarly, measurements of mitochondrial base editing in rapeseed callus tissue showed efficiencies as high as 25% and 1.9% at p6 and rps14, respectively. Figure 27 These results indicate that DdCBE expression in plant protoplasts is tolerable, and that DdCBE induces organelle base editing during protoplast regeneration.

[0740] Furthermore, we attempted to demonstrate the absence of DNA base editing in organelles by using in vitro transcribed cp-DdCBE mRNA instead of plasmids. After introducing an in vitro transcript encoding cp-DdCBE targeting the 16S rRNA gene into lettuce protoplasts, we analyzed the base editing efficiency at the target site. Figure 21 (Figure a). Measurement of C-to-T mutations in protoplasts with an efficiency of up to 25% ( Figure 21 Small image d and Figure 28 As expected, 7 days after protoplast introduction, no DdCBE mRNA or DNA sequence was found. Figure 29 This method can prevent plasmid DNA fragments from potentially integrating into the host genome.

[0741] Resistance to streptomycin and spectinomycin antibiotics was measured by the stable maintenance of organelle editing in callus regenerated from protoplasts. These antibiotics inhibit protein synthesis through irreversible binding to the 16S rRNA gene in chloroplast DNA. Multiple single nucleotide polymorphisms in the 16S rRNA gene are commonly observed in streptomycin-resistant prokaryotes and eukaryotes; in particular, the 16S rRNA C860T (E. coli coordinate C912) mutation leads to streptomycin resistance in tobacco. The C860T point mutation in tobacco is equivalent to the C9 position in lettuce (…). Figure 20 Small image a, Figure 20 Small image b, Figure 21 Small image b and Figure 21 (Inset d). Lettuce callus regenerated from DdCBE-treated protoplasts was transferred to a medium supplemented with streptomycin and spectinomycin. When exposed to the antibiotics, the simulated treatment group turned white, indicating protoplast dysfunction in the callus. In contrast, DdCBE-treated callus remained green, showing resistance to these antibiotics. The editing efficiency of DdCBE was analyzed in drug-resistant lettuce callus and plantlets. Similar to the C860T mutation, C-to-T conversion at position C9 was observed in callus and shoots obtained after drug treatment, with an efficiency as high as 98.6% (…). Figure 21 Small picture e and Figure 21 (Figure f). Interestingly, C-to-T editing near position C13 showed up to 20% efficiency in the absence of spectinomycin, but none in the presence of antibiotics, indicating selection for this mutation upon drug treatment. In summary, these results suggest that DdCBE-induced plant organelle mutations in protoplasts can be maintained even after cell division and plant development, and that allologous chloroplast editing can be achieved through drug selection.

[0742] In addition, off-target activity of TALE deaminase targeting the 16S rRNA site was analyzed in protoplasts, calluses, and shoots. In antibiotic-resistant calluses or shoots derived from single cells, off-target activity was observed near the target site (within 50 base pairs flanking it). Figure 31 ) or the top five candidate off-target sites in the chloroplast genome selected based on sequence homology ( Figure 32 No off-target mutations were detected. In contrast, when a plasmid encoding DdCBE was introduced into protoplasts, off-target TC to TT mutations were induced at three of the five candidate off-target sites with low efficiencies ranging from 1.2% to 4.1%. The off-target efficiency in protoplasts was significantly reduced when an in vitro transcript (mRNA) was used instead of a plasmid encoding TALE deaminase. Figure 22These results indicate that overexpression or prolonged plasmid-based expression of DdCBE increases off-target mutations, and that transient mRNA-based expression using mRNA is preferred for avoiding off-target base editing.

[0743] In summary, the Golden Gate cloning system was developed using a 424-TALE subarray plasmid and 16 expression plasmids to assemble plasmids encoding DdCBE for organelle base editing in plants. Custom-designed DdCBE targeting three genes in chloroplast DNA and two genes in mitochondrial DNA achieved efficient C-to-T conversion in lettuce and rapeseed protoplasts. Specifically, editing in plant organelles was maintained during cell division and plant development. Furthermore, antibiotic-resistant lettuce callus and plantlets with near-homogeneous (99%) heterology were obtained through mutations in the chloroplast 16S rRNA gene. Editing efficiency was 25% in mitochondria and 38% in chloroplasts without antibiotic selection. The Golden Gate cloning system is expected to be a valuable resource for organelle DNA editing in plants.

[0744] Example 3. Animal DNA Editing with TALE-DdCBE

[0745] A DddA-derived cytosine base editor (DdCBE), consisting of the bacterial intertoxin DddAtox, a transcription activator-like effector (TALE) designed to bind to DNA, and a uracil glycosylase inhibitor (UGI), enables desired cytosine-to-thymine base editing in mitochondrial DNA. Similarly, efficient mitochondrial DNA editing is possible in mouse embryos. Within mitochondrial genes, MT-ND5 (ND5), which targets the subunit of the NADH dehydrogenase that catalyzes NADH dehydration and electron transfer to ubiquinone, includes mutations associated with human mitochondrial diseases, such as m.G12918A, and mutations producing early stop codons, such as m.C12336T. Therefore, it is possible to generate mitochondrial disease models in mice, demonstrating the potential for treating mitochondrial diseases.

[0746] 3-1. Method

[0747] Plasmid Assembly. Expression plasmids containing the DddA half and the final TALE-DddAtox construct were constructed using the TALEN (transcription activator-like effector nuclease) system. In the TALEN system expression plasmids, the nuclear localization signal and monomer of the FokI dimer were replaced by the mitochondrial targeting signal (MTS), the DddA deaminase half, and the uracil glycosylation inhibitor (UGI). Sequences encoding MTS, DddA, and UGI were synthesized via IDT. To construct the expression vector, the DNA fragments required for Gibson assembly were amplified using Q5 DNA polymerase (NEB) and then purified. The purified gene fragments were assembled using the HiFi DNA Assembly Kit (NEB), chemically transformed into *E. coli* DH5α (Enzynomics), and then identified by Sanger sequencing. Thus, eight different expression plasmids were obtained, with the BsaI restriction site for the Golden Gate cloning located between the sequences encoding the N-terminal and C-terminal domains. To assemble the DdCBE plasmid, the expression plasmid and module vector (each encoding the TALE sequence), BsaI-HFv2 (10 U), T4 DNA ligase (200 U), and reaction buffer were loaded into a test tube. Subsequently, restriction enzyme and ligase reactions were performed in a thermal cycler for 20 cycles at 37ºC for 5 min and 50ºC for 5 min, followed by further reactions at 50ºC for 15 min and 80ºC for 5 min. The conjugated plasmid was introduced into *E. coli* DH5α by chemical transformation, and the final construct was identified by Sanger sequencing. For cell line introduction, the plasmid was prepared in medium quantities.

[0748] Mammalian cell line culture and transfection. NIH3T3 (CRL-1658, American Type Culture Collection (ATCC)) cell lines were cultured at 37ºC in a 5% CO2 environment. Cell lines were grown in DMEM (Gibco) supplemented with 10% (v / v) fetal bovine serum without antibiotics and without mycoplasma testing. For lipid transfection, cells were cultured at 1.5 × 10⁻⁶ cells / day for 18–24 hours prior to transfection. 4 Cells were seeded at the specified density in 12-well cell culture plates (SPL, Seoul, South Korea). A total of 1,000 ng of plasmid DNA was introduced using Lipofectamine 3000 (Invitrogen) with 500 ng per DdCBE cell divider. Cells were harvested 4 days post-transfection.

[0749] mRNA preparation. The mRNA template was amplified by PCR using Q5 DNA polymerase (NEB) with the following primers (F: 5'-CATCAA TGGGCGTGGATAG-3' SEQ ID No: 268, R: 5'-GACACCTACTCAGACAATGC-3 SEQ ID No: 269). DdCBE mRNA was synthesized using an in vitro RNA transcription kit (mMESSAGE mMACHINE T7 Ultra kit, Ambion) and then purified using a MEGAclear kit (Ambion).

[0750] Animals. All experiments involving mice were conducted with the approval of the Animal Care and Use Committee of the Institute for Basic Science. Superovulatory C57BL / 6J females were mated with C57BL / 6J males, and ICR strain females were used as surrogate mothers. Mice were housed in specific pathogen-free facilities under 12-hour diurnal and constant temperature and humidity conditions (20–26ºC, 40%–60%).

[0751] Microinjection was performed into mouse zygotes. Superovulation, embryo collection, and microinjection were performed shortly before microinjection, as previously described. For microinjection, a mixture of left DdCBE mRNA (300 ng / µl) and right DdCBE mRNA (300 ng / µl) was diluted with DEPC-treated injection buffer (0.25 mM EDTA, 10 mM Tris, pH 7.4) and injected into the cytoplasm of the zygotes using a Nikon ECLIPSE Ti micromanipulator and a FemtoJet 4i microinjector (Eppendorf). After microinjection, the embryos were placed in KSOM + AA (Millipore) microdroplets and cultured at 37ºC, 5% CO2 for 4 days. The 2-cell stage embryos were then transferred to the fallopian tubes of 0.5-dpc pseudopregnant surrogate mothers.

[0752] Genotyping. Embryos and tissues at the blastocyst stage were placed in digestion buffer (25 mM NaOH, 0.2 mM EDTA, pH 10) and incubated at 95ºC for 20 min. The pH was then adjusted to 7.4 using HEPES (free acid, unadjusted) to achieve a final concentration of 50 mM. Genomic DNA was isolated from mouse progeny using the DNeasy Blood and Tissue Kit (Qiagen) and analyzed by Sanger sequencing and targeted deep sequencing.

[0753] Mitochondrial DNA Isolation for High-Throughput Sequencing. To isolate mitochondria from cultured NIH3T3 cells in a 12-well plate, the cell culture medium was removed, and 200 μl of mitochondrial isolation buffer A (ScienCell) was added to the plate. Cells were scraped using a cell scraper and placed in microtubes, then homogenized using a disposable pestle. After homogenization 15 times, the well-homogenized slurry was centrifuged at 1,000 xg and 4ºC for 5 minutes. The supernatant was transferred to a new microtube and centrifuged at 10,000 xg and 4ºC for 20 minutes. The precipitate was resuspended in 20 μl of lysis buffer (25 mM NaOH, 0.2 mM EDTA, pH 10) and then boiled at 95ºC for 20 minutes. To lower the pH, 2 μl of 1 M HEPES (free acid, unadjusted) was added to the mitochondrial lysate. 1 µl of the solution was used for high-throughput sequencing in the PCR template strand.

[0754] High-throughput sequencing. To prepare deep sequencing libraries, nested primary and secondary PCR were performed using Q5 DNA polymerase, followed by the addition of final index sequences. The libraries were then used for paired-end read sequencing using MiniSeq (Illumina). For whole mitochondrial genome analysis, isolated mitochondrial DNA was prepared using a labeled DNA preparation kit (Illumina) according to the manufacturer's protocol. Paired-end sequencing results from all analyses were merged into a single fastq-join file and analyzed using the CRISPR RGEN tool (http: / / www.rgenome.net / ).

[0755] Data analysis and display. Graphs, charts, and tables were created using Microsoft Excel (2019) and PowerPoint (2019). Genome sequence alignment, primer construction, and cloning were performed using Geneious (version 2021.0.1) and Snapgene 5.2.3, with NC_005089 as the reference sequence.

[0756] 3-2. Results

[0757] DdCBE plasmid assembly. To facilitate the assembly of custom-designed TALE sequences in DdCBE, an expression plasmid encoding the split DddAtox hemisphere was constructed, and the Golden Gate cloning system was used with a total of 424 plasmids (6 × 64 tripartites + 2 × 16 bipartites + 2 × 4 monopartites). Figure 33 Figure a). As shown in Table 4 below, six TALE module plasmids and expression plasmids were mixed in the same test tube to construct a ready-to-use DdCBE plasmid with 15.5–18.5 repeating variable double residue sequences (…). Figure 38).

[0758] [Table 4]

[0759]

[0760] The sequences of the DdCBE constructs are shown in Table 5 below. Therefore, DdCBE recognizes 17-20 DNA sequences, including the conserved thymine sequence at the 5' end. Thus, functional DdCBE pairs recognize a total of 32-40 DNA sequences.

[0761] [Table 5]

[0762]

[0763]

[0764]

[0765]

[0766]

[0767] In vitro mitochondrial base editing. To attempt in vivo editing of mitochondrial DNA using the Golden Gate cloning system, the ND5 gene, encoding the NADH ubiquinone oxidoreductase chain 5 protein in mouse mitochondria, was selected. The ND5 protein is a key subunit of NADH dehydrogenase (ubiquinone) and catalyzes the transfer of electrons from NADH to the respiratory chain. In humans, mutations in the ND5 gene are known to be associated with MELAS (mitochondrial encephalomyopathy, lactic acidosis, and stroke-like attacks) and some symptoms of Leigh syndrome or LHON (Leber hereditary optic neuropathy). An attempt was made to establish a mouse model with genetically altered mitochondrial genes to mimic human functional impairment.

[0768] First, several DdCBE plasmids were assembled, designed to generate two silent mutations, m.C12539T and m.G12542A. These plasmids were transfected into the NIH3T3 mouse cell line, and the base editing frequency was measured after 3 days. As expected, cytosine bases within the target range were edited to thymine (C12539T) at an efficiency of up to 19%. Figure 34 Figure a). Previously, it was reported that DddAtox only deaminated cytosine in the “TC” sequence, but based on experimental results, only two cytosines in the TC background were edited. No significant insertions, deletions, or other types of point mutations were generated within the editing target area.

[0769] In vivo mitochondrial base editing. The most effective DdCBE pairs (left-G1397-N and right-G1397-C) were used in in vivo experiments. Four days after microinjection of the in vitro transcript encoding the DdCBE pairs into 1-cell stage C57BL6 / J embryos, 9 out of 32 embryos were successfully edited (28%, Table 6).

[0770] [Table 6]

[0771]

[0772] TALE-DddAtox deaminase efficiently produces C*G to T*A base conversions, with efficiencies of 2.2%-25% at m.C12539 and 0.63%-5.8% at m.G12542. Subsequently, embryos injected with DdCBE are transferred to surrogate mothers to obtain offspring with m.C12539T and m.G12542T. Figure 39 Three of the four cubs (F0) show C*G to T*A edits, with an efficiency of 1% to 27%. Figure 34 (Inset c). Both pups showed similar mutation levels in their toes and tails, and maintained this level for 14 days after birth. Furthermore, these mitochondrial DNA mutations were detected in various tissues of adult F0 mice at 50 days after birth. Figure 34 (Inset d). These results indicate that the heterogeneity of mitochondrial DNA in DdCBE-induced 1-cell stage fertilized eggs is maintained during development and differentiation.

[0773] To determine whether DdCBE-induced mutations are passed on to the next generation, female F0 mice were mated with wild-type C57BL6 / J males to produce F1 offspring. The m.C12539T and m.G12542T mutations were observed in both pups, with efficiencies ranging from 6% to 26%. Furthermore, similar mitochondrial editing was observed in 11 different tissues. Figure 35 Small image b).

[0774] DdCBE-mediated MT-ND5 G12918A mutation. Attempts were made to generate the m.G12918A mutation, which also causes mitochondrial diseases in humans. This mutation causes various mitochondrial diseases such as Leigh syndrome, MELAS syndrome, and LHON syndrome. Base editing using DdCBE is possible because the cytosine base at this position is adjacent to a thymine base. Figure 36 Figure a). Assembling four pairs of DdCBEs confirmed that editing in NIH 3T3 is possible with an efficiency of up to 6.4% ( Figure 36(Figure b). The most effective DdCBE combination was then microinjected into mouse zygotes, and its efficiency was observed in the blastocysts. Eleven (25%) of the 44 embryos carried the m.G12918A mutation, with efficiencies ranging from 0.25% to 23%. Figure 36 Figure c). Furthermore, DdCBE microinjected embryos were transferred to surrogate mothers to obtain offspring with the G12918A mutation (…). Figure 39 (Figure b). It has been confirmed that 4 out of 11 newborn mice have this mutation in approximately 3.9%–31.6% of them. Figure 36 (Figure d). Although the phenotype did not appear immediately after birth, presumably because the offspring were very young and wild-type mitochondrial DNA and mutant DNA coexisted in a heterozygous state, these results suggest that DdCBE may produce animal models of mitochondrial disease.

[0775] MT-ND5 nonsense mutation. Finally, the possibility of maintaining the loss-of-function ND5 mutation in mice was confirmed by generating nonsense mutations in the gene. Using m.C12336 as the target cytosine, an early stop codon (Q199*) was introduced at position 199 of the ND5 protein. Figure 37 Figure a). Specifically, four DdCBE combinations were transfected into the NIH3T3 cell line to confirm base editing efficiency, showing that the most effective DdCBE for inducing nonsense mutations was approximately 5.7% (…). Figure 37 Figure b). It has been confirmed that this DdCBE leads to cytosine-to-thymine editing and produces a silent mutation in m.G12341A (Q200Q), which, although slightly less efficient, is edited within the target range. In 19 out of 37 mouse embryos (=51%), the m.C12336T and m.G12341A mutations were confirmed, with efficiencies of 32% and 23%, respectively. Figure 37 Small image c).

[0776] Based on these results, mouse embryos were transferred to surrogate mothers to obtain offspring with m.C12336T and m.G12341A mutations. Figure 39 Inset c). Of the 27 F0 mice, 9 (23%) showed C*G to T*A editing, with efficiencies ranging from 0.22% to 57%. Figure 37 Small image d and Figure 37 Figure e shows that nonsense mutations in ND5 do not cause embryo knockout.

[0777] Example 4. Mitochondrial DNA Editing in Animals

[0778] 4-1. Constructing an expression vector with nuclear output signal for base editing in animal mitochondria.

[0779] A vector was constructed to express a protein with a nuclear output signal fused with TALE-DdCBE in animal cells. Figure 40 The vector uses a cytomegalovirus promoter (CMV promoter). The vector includes a mitochondrial targeting signal, a protein purification / detection tag, a TALE array N-terminal domain, a repeat sequence region, a C-terminal domain, a half-part of DddA cytosine deaminase cleavage, a uracil glycosylation inhibitor, and a nuclear export signal. Figure 40 (Figure a). For nuclear output signals, for example, the NS2 protein derived from MVM (mouse parvovirus) can be used, but other sequences can also be used. The expressed protein is released outside the nucleus and then translocated to the mitochondria, where base editing occurs. Here, the target DNA sites are selected from the mitochondrial ND5 gene-chromosome 4 ND5-like gene, mitochondrial TrnA-chromosome 5, and mitochondrial Rnr2-chromosome 6 (…). Figure 40 Small image b).

[0780] 4-2. DdCBE-NES in animal cell lines

[0781] The afternoon before transfection, the NIH3T3 cell line (ATCC CRL-1658) was injected with 1.5 × 10⁻⁶ cells. 4 / wells were allocated to 12-well plates containing 1 ml of cell growth medium (DMEM + 10% fetal bovine serum). The following morning, cells were transfected using Lipofectamine 3000 with the experimental groups containing DNA-untreated mimics, DdCBE, and DdCBE-MVM NES plasmids, according to the manufacturer's protocol. After three days of culture in an incubator (37ºC, 5% CO2), cells were harvested and total DNA was purified using a Qiagen blood and tissue kit, followed by amplification using mitochondrial gene-specific PCR primers, and then next-generation sequencing using an Illumina MiniSeq system. Base editing efficiency was then determined using a Cas analyzer (www.rgenome.net).

[0782] Figure 40 Small image c Figure 40 Small image d and Figure 40 Inset e shows the mutations in the mouse mitochondrial genes ND5, TrnA, and Rnr2 generated by transfecting the NIH3T3 mouse cell line with DdCBE-NES, indicating that the efficiency varies depending on the DdCBE combination.

[0783] 4-3. Mitotalen in animal cell lines

[0784] Building recognition Figure 40The sequence of TALEN is shown in small figure f, and MTS is linked to this construct to introduce it into mitochondria together with DdCBE. The experimental method is the same as in Examples 4-2. Mismatches were intentionally made at the TALE recognition site, resulting in one mismatch with wild-type mtDNA and two mismatches with mutant mtDNA. This is because TALE cannot distinguish a single nucleotide mismatch. The results confirmed that the +2 mismatch experimental group treated with DdCBE and TALEN was more efficient than that treated with DdCBE alone.

[0785] 4-4. DdCBE-NES in animal embryos

[0786] Using DdCBE or DdCBE-NES expression vectors as templates, PCR amplicons containing the T7 promoter and DdCBE or DdCBE-NES expression sites were obtained. These PCR amplicons were then used as templates to synthesize mRNA using T7 polymerase.

[0787] DdCBE mRNA pairs or DdCBE-NES mRNA pairs in microinjection solution were microinjected into mouse zygotes.

[0788] After fertilized eggs were cultured for four days to develop into blastocysts, the blastocysts were lysed. Using the same template, a portion of the target site in mitochondrial DNA that differed from the nuclear DNA was amplified by PCR, followed by further PCR amplification of the index and sequencing of the adaptor. High-throughput sequencing was performed using an Illumina MiniSeq system, and the base editing efficiency was analyzed using a Cas analyzer (www.rgenome.net). Additionally, DNA with sequences resembling the mitochondrial target region in the cell nucleus was amplified by PCR and sequenced.

[0789] Therefore, DdCBE induces mutations not only in mitochondrial DNA but also in similar DNA sequences in the nucleus (mitochondria: 13.1%, nucleus: 3.2%). Using DdCBE-NES, the mitochondrial target mutation efficiency increased to 18.2%, while the nuclear DNA mutation efficiency decreased to 0.2%. Figure 41 Figure a). For other targets TrnA or Rnr2, no nuclear DNA mutations occurred, but in mitochondrial DNA, the base editing efficiency was statistically significantly increased (*p < 0.05, **p < 0.01, ns: not significant).

[0790] 4-5: DdCBE and mitoTALEN in animal embryos

[0791] In addition to the ND5 gene-specific DdCBE, TALEN, which cleaves the unedited mitochondrial DNA sequence, was also injected to increase the proportion of edited mitochondrial DNA in C-to-T transformed cells. The microinjection and sequencing identification methods were the same as in Examples 4-2 and 4-4. The group treated with DdCBE alone showed an editing efficiency of 11%, and when treated with DdCBE and mitoTALEN, the efficiency increased to 33.3%, resulting in a statistically significant increase in editing efficiency. Furthermore, the group treated with DdCBE-NES alone showed an editing efficiency of 20.5%, and when treated with DdCBE-NES and mitoTALEN, an efficiency of 36.8% was observed, which was also statistically significant. Figure 41 Small image b).

[0792] Similarly, by transferring microinjected fertilized eggs to surrogate mothers, DdCBE showed an editing efficiency of 10.9% in newborn mouse pups, but achieved 23.4% efficiency when using DdCBE-NES and mitoTALEN. Figure 41 Small image c).

[0793] When nuclear export signals bind to base-editing proteins during mitochondrial gene editing in animals, base editing is achieved with greater efficiency, and in animal embryos, non-specific base editing of similar sequences in the cell nucleus is also inhibited. Furthermore, even more efficient mitochondrial base editing can be expected when mitochondrial sequence-cutting proteins are co-injected.

[0794] Example 5. Splitting DddA tox Deaminase variants

[0795] A high-precision DddA-derived cytosine base editor capable of reducing the off-target effects of DdCBE was developed. This off-target base editing effect is a phenomenon caused by the spontaneous assembly of DddAtox deaminase cleavage fragments, independent of the interaction between TALE and DNA. Therefore, HF-DdCBE was constructed by replacing amino acid residues located on the surface between DddAtox cleavage fragments with alanine. HF-DdCBE prevents a pair of TALE-linked deaminases from functioning normally when not bound to DNA. Whole-mitochondrial genome analysis confirmed that HF-DdCBE is highly efficient and precise, unlike conventional DdCBE, which induces many undesirable off-target C-to-T conversions in human mitochondrial DNA.

[0796] 5-1. Method

[0797] Plasmid construction. Point mutations were introduced into the DdCBE expression plasmid. The plasmid was amplified using mutagenic primers for Q5 site-directed mutagenesis (NEB) (Table 7), and the results were confirmed by Sanger sequencing.

[0798] [Table 7]

[0799]

[0800]

[0801]

[0802] To assemble the interface mutant, small-batch prepared mutant expression plasmids and module vectors (each encoding a TALE sequence), BsaI-HFv2 (10 U), T4 DNA ligase (200 U), and reaction buffer were mixed in a test tube. Restriction enzyme and ligase reactions were then performed in a thermal cycler for 20 cycles at 37ºC for 5 min and 50ºC for 20 min, followed by further reactions at 50ºC for 15 min and 80ºC for 5 min. The ligated plasmids were introduced into *E. coli* DH5α via chemical transformation, and the final constructs were identified by Sanger sequencing. Medium-batch plasmids were prepared for introduction into cell lines.

[0803] Mammalian cell line culture and transfection. HEK 293T / 17 (CRL-11268, American Type Culture Collection (ATCC)) cell lines were cultured at 37ºC in a 5% CO2 environment. Cell lines were grown in DMEM (Gibco) supplemented with 10% (v / v) fetal bovine serum without antibiotics and without mycoplasma testing. For lipid transfection, 18–24 hours prior to transfection, [the following was performed] at 1 × 10 [units of] [amount ... 5 Cell growth was initiated at a cell density of 1000 g in 24-well cell culture plates (SPL, Seoul, South Korea). A total of 1,000 ng of plasmid DNA was introduced using Lipofectamine 2000 (Invitrogen) with 500 ng per DdCBE cell divider. Cells were harvested 4 days post-transfection.

[0804] Genomic and mitochondrial DNA isolation for high-throughput sequencing. After removing cell culture medium to isolate genomic DNA, lysis buffer containing proteinase K from the DNeasy Blood and Tissue Kit (Qiagen) was added to cell culture plates to isolate cells from the bottom of the plate. Genomic DNA was then isolated according to the manufacturer's protocol. For whole mitochondrial genome sequencing, 200 μl of mitochondrial isolation buffer A (ScienCell) was added to a culture plate after removing cell culture medium. Cells were scraped using a cell scraper and placed in microtubes, followed by homogenization using a disposable pestle. After homogenization 20 times, the well-homogenized homogenate was centrifuged at 1,000 xg and 4ºC for 5 minutes. The supernatant was transferred to a new microtube and centrifuged at 10,000 xg and 4ºC for 20 minutes. The pellet was resuspended in 10 μl of lysis buffer (25 mM NaOH, 0.2 mM EDTA, pH 10) and then boiled at 95ºC for 20 minutes. To lower the pH, 1 μl of 1 M HEPES (free acid, unadjusted) was added to the mitochondrial lysate. 1 µl of this solution was then used in the PCR template strand for high-throughput sequencing.

[0805] High-throughput sequencing. Nested primary and secondary PCR were performed using Q5 DNA polymerase to construct deep sequencing libraries, and final index sequences were added. The libraries were then used for paired-end read sequencing using MiniSeq (Illumina). For whole mitochondrial genome analysis, isolated mitochondrial DNA was prepared using a labeled DNA preparation kit (Illumina) according to the manufacturer's protocol. Paired-end sequencing results from all analyses were merged using a single fastq-join file and analyzed using the CRISPR RGEN tool (http: / / www.rgenome.net / ).

[0806] 5-2. Results

[0807] When attempting to edit chloroplasts in plants, off-target base mutations occur in the chloroplast genome, raising questions about the accuracy of DdCBE. There are two reasons for off-target base editing in DdCBE. The first is the non-specific binding between the TALE protein and DNA, and the second is the unintentional spontaneous interaction between the DddAtox hemimeres. Figure 42 (Figure a). This study focuses on splitting the DddAtox hemisphere and designs the interface between the two protein splitters to prevent unwanted assembly of the DddAtox hemisphere.

[0808] Specifically, we examined whether each subunit (left-TALE or right-TALE) of the mitochondrial ND1 (mtND1) gene targets DNA and interacts with the other half of the DddAtox without a TALE to induce cytosine-to-thymine base editing. The DdCBE pair (left-TALE: G1397N (N-terminal G1397 DddAtox half fused with the C-terminus of the left-TALE array recognizing the left half) + right-TALE: G1397C (C-terminal G1397 DddAtox half fused with the C-terminus of the right-TALE array recognizing the right half) of the human mitochondrial ND1 (mtND1) gene in the human kidney embryonic cell line (HEK293T) effectively edited the C11 of the target sequence, converting cytosine to thymine with an efficiency of 60.7%. Figure 43 Small figure a). Furthermore, with another DddA without TALE. tox Each individual subunit of the half-pair (left-TALE: G1397N or right-TALE: G1397C) also induces base editing, although less efficiently than the original pair. Therefore, each left-TALE and right-TALE that binds to the target ND1 sequence when paired with another TALE-free DddAtox half induces base editing, with efficiencies of 31% or 8.1%. Figure 43 In other words, DdCBE with two TALE fusions is only 2.0 times (60.7% / 31%) or 7.5 times (60.7% / 8.1%) more effective than mismatch pairs with only one TALE fusion. Clearly, DddA with TALE arrays fused to binding half-sites... tox The N-terminal portion can recruit DddA without a TALE array. tox The C-terminal portion, and vice versa, is used to reconstruct functional deaminase.

[0809] Since DddAtox can split at two positions (G1333 and G1397), DdCBE pairs (left-TALE: G1333-N and right-TALE: G1333-C) targeting the mtND1 gene at position G1333 were also constructed to test whether the left-TALE: G1333-N and right-TALE: G1333-C constructs could recruit the TALE-free DddAtox hemisphere and induce C-to-T editing. As expected, each TALE fusion paired with the other TALE-free DddAtox hemisphere showed a base editing efficiency of 32.7% (left TALE conjugate) or 18.1% (right TALE conjugate) at position C8, compared to 56.1% for the original DdCBE pair. Therefore, the original contrast pair with two TALE fusions was only 1.7-fold (56.1% / 32.7%) or 3.1-fold (56.1% / 18.1%) more effective than the mismatch pair with only one TALE fusion. In summary, these results indicate that DdCBE can induce unwanted off-target mutations at sites where only one TALE array can bind. Because TALE proteins can bind to sites even with some mismatches, DdCBE pairs may induce numerous off-target mutations in organelles or the nuclear genome.

[0810] We are attempting to develop a high-fidelity DdCBE that does not exhibit the characteristics of split DddA. tox This off-target editing is caused by the spontaneous assembly of the half-components. We hypothesize that the interface of the split dimer can be engineered to suppress or prevent self-assembly. To this end, we used a Python script (InterfaceResidues.py) in PyMOL software to identify amino acid residues at the interface of two split DddAtox halves (split at G1333 and G1397) within a 1 square angstrom range. As a result, we found 9 amino acid residues in G1397-N (the N-terminal DddAtox half-split at G1397), 4 residues in G1397-C (the C-terminal DddAtox half-split at G1397), 14 amino acid residues in G1333-N (the N-terminal DddAtox half-split at G1333), and 15 amino acid residues in G1333-C (the C-terminal DddAtox half-split at G1333). Figure 42 Small image b and Figure 42 Small image c).

[0811] Subsequently, we generated various mutant DddAtox hemimorphs by substituting alanine for each of these amino acid residues. We then measured the editing frequencies of these interface mutant DdCBEs combined with wild-type DdCBE mates or TALE-free DddAtox hemimorphs in HEK293T cells. Many DddA cleavages of G1397-fragments containing interface mutations such as C1376A, M1390A, and F1412A were observed. tox The variants could not induce C-to-T conversion in the spacer region between the two TALE binding sites, even when combined with wild-type mates, indicating that these mutants could not interact with other wild-type DddA variants near the target site. tox Half-part interaction. Other DddA tox Variants (such as those containing V1377A and E1381A) together with the TALE-free half induce C-to-T editing at a high frequency, which is comparable to the wild-type DdCBE pair, thus indicating that these mutations are neutral and do not prevent fission dimer interactions.

[0812] Importantly, some mutations, such as K1389A, K1410A, and T1413A, exhibit high activity when paired with wild-type DdCBE pairs, but low activity when paired with the TALE-free half. For example, the K1410A mutation shows an efficiency of 53.2%, similar to its efficiency (60.7%) when paired with wild-type DdCBE pairs, but shows an efficiency of only 0.9% when paired with the TALE-free half, resulting in a 59.1-fold difference (= 53.2% / 0.9%). As mentioned above, the wild-type pair shows a 7.5-fold difference (= 60.7% / 8.1%). Furthermore, these variants are more selective in editing bases than the wild-type DdCBE pair. Therefore, these variants preferentially edit C8, C9, and C4 bases within the editing window. 13 Editor C 11 Wild-type DdCBE showed much less differentiation, editing all four cytosines at a high frequency of >6.7%. Figure 43 Small image b).

[0813] Furthermore, screening for 29 mutations in G1333 (14 mutations in G1333N and 15 mutations in G1333C) yielded several desired interface mutations. Figure 44Variants containing most of these mutations (e.g., I1299A, Y1316A, Y1317A, and F1329A) exhibit poor activity even when paired with wild-type mates, or, in the case of other mutations (e.g., S1300A and T1314A), undesirable activity when paired with TALE-free mates. Notably, variants containing several mutations (including K1389A, T1391A, and V1393A) show high activity when paired with wild-type mates but low efficiency when paired with TALE-free mates. For example, K1389A shows a 38-fold difference (= 45.4% / 1.2%), while the wild-type DdCBE pair shows only a 3.1-fold difference (= 56.1% / 18.1%). Furthermore, the K1389A variant is more selective than the wild-type pair. Therefore, this variant is preferred over C9 and C... 11 and C 13 The wild-type pair edited C8, while the wild-type pair was promiscuous, editing all four cytosines at a high frequency of >19%. It is also noteworthy that K1410A in G1397-C preferentially edited C8, while K1389A in G1333-C selectively edited C11 with higher efficiency. In contrast, the wild-type DdCBE pair (G1333 or G1397 DddA) edited C8. tox The splittings showed poor selectivity. These results indicate that the aforementioned interface mutants have the potential to reduce unwanted editing of multiple bases within the target site, which is typically observed with DdCBE.

[0814] Example 6. Full-length deaminase

[0815] The DddA-derived cytosine base editor (DdCBE) consists of the cleaving bacterial intertoxin DddAtox, a TALE array, and a uracil glycosylase inhibitor (UGI), enabling the conversion of target cytosine in eukaryotic nuclear DNA, mitochondrial DNA (mtDNA), and plant chloroplast DNA into thymine. DddAtox, a bacterial cytotoxic enzyme derived from Burkholderia cepacia, deaminates cytosine in double-stranded DNA. To avoid host cell toxicity, DddAtox is cleaved into two inactive halves, each fused with a TALE DNA-binding protein to form a DdCBE pair. The functional deaminase can only be reconstructed when the two inactive halves are bound together to the target DNA via two adjacent TALE binding proteins. A C-to-T base conversion is induced in a 14–18 bp spacer region between the two TALE binding sites.

[0816] Unlike CRISPR-derived base editors that cannot edit organelle DNA, DdCBE enables targeted base editing in both nuclear and organelle DNA. However, a drawback is the need for two TALE constructs instead of a single construct to induce this editing. The first drawback is that TALEs must bind to target DNA sites with thymine at the 5' and 3' ends, thus limiting the targetable sites when using two TALE arrays. Second, delivery of two TALE constructs instead of one is generally inefficient and challenging. Viral vectors with limited capacity, such as the adeno-associated virus (AAV) vector widely used in gene therapy (capacity: approximately 4.7 kbp), cannot accommodate the split DdCBE coding sequence because the dimeric DdCBE combination is too large (2 × 4.1 kbp, including the promoter and poly-A signal). Furthermore, due to the high similarity of the two TALE array sequences, cloning the DNA fragments encoding the two TALE arrays into a single large-capacity vector can become difficult. Finally, using two TALE arrays instead of one may exacerbate off-target effects. To overcome these limitations of the dimer DdCBE containing DddAtox fragments, we offer a non-toxic, full-length DddA tox Fusion-modified DdCBEs, known as mDdCBEs (monomer DdCBEs), are used for targeted C-to-T conversion in nuclear and organelle DNA.

[0817] 6-1. Method

[0818] Plasmid construction. Using synthetic full-length DddAtox (gBlock, IDT) as a template, the DddA variant was amplified by PCR using primers and Q5 DNA polymerase (NEB) as shown in Table 8. These PCR products were cloned at the p3s-BE3 site using Gibson assembly (NEB), where Apobec1 was digested with BamHI and Sma I (NEB). TALE-DddAtox (Addgene #158093, #158095, #157842, #157841) plasmids were digested with BamHI and Sma I, and the DddA variant was amplified by PCR using primers as shown in Table 8, followed by Gibson assembly cloning. The resulting plasmids were transformed into chemically prepared *E. coli* DH5α using a heat shock method, and the plasmid sequences of surviving colonies were analyzed by Sanger sequencing. The final plasmids were prepared in medium-volume (Macherey-Nagel) for cell transfection.

[0819] [Table 8]

[0820]

[0821] Random mutagenesis. Using synthetic full-length DddAtox (gBlock, IDT) as a template, error-prone PCR was performed using the GeneMorphII Random Mutagenesis Kit (Agilent) according to the manufacturer's protocol. In summary, 0–16 mutations / kb were introduced using 1 ng, 100 ng, and 700 ng of DddAtox DNA as templates. Full-length DddAtox gBlock was pre-amplified by PCR using primers listed in Table 8. All PCR products were pooled and cloned into p3s-UGI-Cas9 (H840A) digested with Sma1 and Xho1 using Gibson assembly (NEB). Chemically prepared *E. coli* DH5α was transformed with plasmids via heat shock, and the plasmid sequences of surviving colonies were analyzed by Sanger sequencing. In the analyzed plasmids, the p3s-UGI-nCas9(H840A)-DddAtox plasmid with the coding framework was transfected into HEK293T cells along with sgRNA, and editing activity was then determined by targeted deep sequencing.

[0822] Mammalian cell culture and transfection. HEK293T (ATCC, CRL-11268) and HeLa (ATCC, CCL-2) cells were cultured at 37ºC in 5% CO2. Cells were cultured in DMEM supplemented with 10% (v / v) fetal bovine serum (Welgene) and 1% penicillin / streptomycin (Welgene). Cells were cultured at a rate of 3 × 10⁶ cells / mL for 24 hours prior to transfection. 5 Cells (HEK293T) and 4×10 4 HeLa cells were seeded at a density in 48-well plates (Corning) and then transfected using Lipofectamine 2000 (Invitrogen) with a Cas9-fused DddA plasmid (750 ng) and sgRNA (250 ng). TALE-DddA was transfected into HEK293T cells using 200 ng of plasmid and Lipofectamine 2000. The sgRNA sequence is shown in Table 9 below.

[0823] [Table 9]

[0824]

[0825] Genomic and mitochondrial DNA preparation. Cells transfected with the Cas9 fusion-DddA variant were harvested 2 days post-transfection, and cells transfected with TALE-DddA were harvested 3 days post-transfection. Genomic and mitochondrial DNA were isolated using the DNeasy Blood and Tissue Kit (Qiagen). For large-scale analysis, DNA was extracted using 100 μl of cell lysis buffer (50 mM Tris-HCl (pH 8.0) (Sigma-Aldrich), 1 mM EDTA (Sigma-Aldrich), 0.005% sodium dodecyl sulfate (Sigma-Aldrich)) containing 5 μl proteinase K (Qiagen). The lysates were incubated at 55ºC for 1 hour, followed by incubation at 95ºC for 10 minutes.

[0826] 6-2. Results

[0827] Compare the amino acid sequences of the wild-type and the novel full-length DddA variant. The altered amino acids are... Figure 45 The text is represented by a gray box.

[0828] like Figure 46 As shown, a 16-amino acid linker is used to connect DddA upstream of the N-terminus of Cas9, while a 4-amino acid linker is used to connect UGI (uracil glycosylation inhibitor) and NLS (nuclear localization signal) to the C-terminus. Conversely, a 16-amino acid linker is used to connect DddA downstream of the C-terminus of Cas9, while a 4-amino acid linker is used to connect UGI and NLS to the N-terminus.

[0829] In this invention, we constructed and used DddA-Cas9 (D10A, D10A, and H840A)-UGI. A full-length single DddA module fused with a zinc finger protein or TALE module enables cytosine-to-thymine editing. Current fission systems require two modules, but full-length DddA requires only one. These two DNA-binding proteins can be linked to NLS (nuclear localization signal), MTS (mitochondrial targeting sequence), or CTP (chloroplast transport peptide), enabling thymine to replace cytosine not only in the nuclear genome but also in the mitochondrial and plant chloroplast genomes where Cas9 editing is not feasible. Figure 47As shown, thymine substitution of cytosine in the TC motif was demonstrated at the ROR1 (a), HEK3 (b), and TYRO3 (c) sites in the human cell genome background. Thymine substitution of cytosine at a distance of 25 bp from the target site in the TC motif was also demonstrated (a). For A1341D KRKKA, thymine substitution of the second cytosine in the CC motif was demonstrated (a, b). Thymine substitution of cytosine in the TC motif was also demonstrated for the catalytic mutant E1347A (a, b, c). Red underlines indicate Cas9 binding sites. Efficiency is expressed as the percentage of cytosine-to-thymine conversion in the insertion-deletion-free reads of the total sequencing reads. Furthermore, the insertion / deletion ratio in the total sequencing reads is expressed as a percentage.

[0830] like Figure 48 As shown, the red squares indicate the portions confirming DddAtox activity, and activity was measured by dividing the same portion into three target sites using full-length DddA. For the fragments, orthogonal Cas9s with different PAMs were used to convert cytosine between two Cas9s to thymine. Therefore, precisely replacing the desired cytosine with thymine is difficult. However, for full-length DddA, the portion where Cas9 binds to the same target site can be divided into three and targeted, making it possible to precisely replace the desired cytosine with thymine. Efficiency is expressed as the percentage of cytosine-to-thymine conversion in the insertion-deletion-free reads of the total sequencing reads. Furthermore, the insertion / deletion ratio in the total sequencing reads is expressed as a percentage.

[0831] like Figure 49 As shown, the activity of full-length DddA was measured in human cell genome background TRAC sites 1 (a), TRAC site 2 (b), FANCF (c), and HBB (d). Red underlines indicate Cas9 binding sites. Efficiency is expressed as the percentage of cytosine-to-thymine conversion in the insertion-deletion-free reads of the total sequencing reads. Furthermore, the insertion / deletion ratio in the total sequencing reads is expressed as a percentage.

[0832] like Figure 50 As shown, DddA activity was measured using DddA-dCas9(D10A, H840A)-UGI in human cell genome backgrounds at TYRO3 (a), ROR1 (b), HEK3 (c), EMX1 site 2 (d), TRAC site 1 (e), and HBB (f). Efficiency is expressed as the percentage of cytosine-to-thymine conversion in the total sequencing reads. No insertions or deletions were observed.

[0833] To obtain non-toxic, full-length DddAtox variants for base editing, two methods were used: structure-based site-specific mutagenesis and random mutagenesis. In the first method, DddAtox variants with reduced DNA binding or reduced catalytic activity were fused with inactive CRISPR-Cas9 (dCas9) or nickase (nCas9) variants to develop novel base editors in which target cytosine was replaced with thymine in cultured human cells. For this purpose, the positively charged amino acid of DddAtox was replaced with alanine, and it was subcloned into an expression vector (…). Figure 51 (Figure a). It is assumed that these variants may avoid toxicity by weakening their binding to negatively charged dsDNA. Most alanine-substituted variants cannot form E. coli transformants ( Figure 51 (Figure b). Based on sequencing analysis of the plasmid DNA isolated from the resulting transformants, various frameshift mutations were induced in the protein-coding region. This full-length DddAtox variant, despite being controlled by a mammalian promoter, was poorly expressed in *E. coli*, leading to cell death. Fortunately, several variants with triple, tetrad, or quintuplet (called "AAAAA") alanine substitutions without frameshift mutations were obtained. The active site mutation E1347A was also successfully cloned.

[0834] Furthermore, we investigated whether AAAAA variants fused with D10A nCas9 or dCas9 and UGI could induce base editing in human embryonic kidney 293T (HEK293T) cells. Figure 51 Small image c and Figure 51 (Inset d). A base editor 2 (or 3) composed of rat APBEC1 deaminase, uracil glycosylase inhibitor (UGI), and dCas9 (or D10A nCas9) is active in a narrow region within the prototype spacer, while the AAAAA variant induces cytosine-to-thymine conversion with an efficiency up to 43% upstream of the prototype spacer at the 5' position. Surprisingly, the E1347A mutation induces base editing at the same C-3 position at a frequency of 37% (nCas9 fusion) or 16% (dCas9 fusion). Figure 51 (See small figure c), which indicates that the E1347A mutation did not completely inactivate DddA. tox The E1347A mutant exhibits high residual deaminase activity, and its residual activity is sufficiently high to induce base editing in human cells. However, the E1347A variant bound to the quintuple AAAAAA mutation does not induce base editing. Furthermore, variants with E1347A, AAAAA, and other alanine substitutions were confirmed to have [the necessary deaminase activity]. Figure 53Without frameshift mutations involving fusion with dCas9 or nCas9 and UGI, editing was induced at positions up to 25 bases upstream of the prototype spacer, exhibiting editing efficiencies up to 26% at various sites (Fig. 54). Furthermore, editing was highly efficient in HeLa cells, reaching 60% efficiency (Fig. 55). The fusion protein-induced base editing persisted in cells for up to 21 days, indicating that this base editing was not cytotoxic. Figure 56 ).

[0835] To modify the editing window of the cytosine base editor, an attempt was made to fuse a variant with alanine substitution to the C-terminus of H840AnCas9. Unexpectedly, no complete construct without frameshift mutation was obtained. Therefore, error-prone PCR was performed to introduce random mutations into the DddAtox coding sequence, and a non-toxic full-length DddAtox variant (referred to as “GSVG”) with four point mutations S1326G, G1348S, A1398V, and S1418G in the amino acid sequence of SEQ ID NO: 269 was obtained, where S at position 37 was replaced by G; G at position 59 was replaced by S; A at position 109 was replaced by V; and S at position 129 was replaced by G, including the sequence of SEQ ID NO: 276. Figure 52 (Figure a). Furthermore, these variants were fused to the C-terminus of dCas9, D10A nCas9, and Cas9, as well as the N-terminus of dCas9, nCas9, and Cas9. In human cells, these fusion proteins induced cytosine-to-thymine conversion at various sites with an efficiency of up to 38%, in addition to wild-type Cas9. Figure 52 (Figures b, 57, and 58). Interestingly, fusion proteins containing GSVG variants fused to the C-terminus of dCas9, D10A nCas9, and H840A nCas9 induce cytosine base editing downstream of the 3' of the prototypical spacer adjacent motif (PAM), while fusion proteins containing the same variants fused to their N-terminus, dCas9 and nCas9, induce base editing upstream of the 5' of the prototypical spacer. Figure 52 (Figure c). As expected, fusion proteins containing Cas9 result in insertions or deletions rather than base substitutions.

[0836] To identify which mutations in the GSVG variants are important, four revertants—SSVG, GGVG, GSAG, and GSVS—were constructed via site-directed mutagenesis. SSVG, GSAG, and GSVS revertants were obtained, but the GGVG variant fused to the C-terminus of nCas9 was not. G1348 is adjacent to E1347, a key catalytic site. The G1348S mutation reduced catalytic activity, avoiding cytotoxicity to *E. coli*. Editing frequencies at two target sites in transfected cells were determined for up to 21 days in the three revertants and the GSVG variant. The frequency of cytosine-to-thymine editing induced by GSAG and GSVS gradually decreased to about half from day 3 to day 21 post-transfection, indicating that these two revertants are cytotoxic to some extent, while GSVG and SSVG were preserved. Figure 59 These results indicate that G1348S is required in the GSVG variants, S1326G is neutral, while A1398V and S1418G reduce cytotoxicity.

[0837] In summary, our results indicate that non-toxic full-length DddA exhibits reduced affinity for dsDNA (AAAAA), weakened deaminase activity (E1347A and possibly GSVG), or reduced cytotoxicity (GSVG). tox Variants can be fused with dCas9 or nCas9 to produce new base editors with altered editing windows. These base editors, known as dCas9-mDdBE (a DddA-derived base editor consisting of a full-length monomeric DddAtox variant fused to the C-terminus of dCas9), nCas9-mDdBE, mDdCE-dCas9, and mDdCE-nCas9, can be used for base editing upstream or downstream of the prototype spacer subregion beyond the BE2 or BE3 range.

[0838] We also investigated whether non-toxic, full-length DddAtox variants could be used for mitochondrial DNA editing. Of the various variants, only two, GSVG and E1347A, successfully fused to the C-terminus of TALE arrays designed to bind to mitochondrial genes ND4 and ND6. The monomeric DdCBE (mDdCBE), including the GSVG variant, achieved base editing at the target nucleotide position, with an efficiency as high as 31% (ND4). Figure 60 Small figure a) and 27% (ND6) Figure 60Figure b) is equivalent to the original split DdCBE pair. The mDdCBE containing E1347A also converts the target cytosine to thymine, but with reduced efficiency, with editing rates of up to 7.2% (ND4) and 8.9% (ND6). Interestingly, the ND4 gene-specific original DdCBE pair (G1333 split) has an editing efficiency of 0.8% at position C4, while the two mDdCBEs containing GSVG show high editing efficiencies of 26% and 31%, respectively. These results suggest that the split-dimer DdCBE and mDdCBE exhibit different mutational patterns, indicating that mDdCBE may be complementary to the dimer DdCBE, inducing a variety of mutations at a given target site.

[0839] A potential advantage of mDdCBE over the split-dimer DdCBE is that off-target effects caused by non-specific TALE-DNA interactions are halved compared to the dimer DdCBE. The split-DddAtox dimer DdCBE can operate at only one subunit-binding half-site, leading to unwanted off-target mutations. The inactivated DddAtox half of the DdCBE pair can recruit the other inactivated half to form a functional deaminase. To confirm this hypothesis, HEK293T cells were co-transfected with a plasmid encoding one subunit of the dimer DdCBE and a plasmid encoding the TALE-free DddAtox half, and editing frequencies were measured at both mitochondrial target sites. As expected, cytosine-to-thymine editing was observed at the target sites at frequencies ranging from 0.7% to 3.6%. Figure 60 Small image c- Figure 60 (Figure f). These results suggest that unwanted off-target mutations may be caused by the interaction between the split DddAtox halves of the DdCBE pair at the half-site, and that mDdCBE can avoid half of the off-target mutations caused by the dimer DdCBE.

[0840] Example 7. Efficient A-to-G base editing in human cells using DdABE

[0841] Mitochondrial DNA base editing via a DddA-derived cytosine base editor (DdCBE) has enabled the establishment of disease models in various cell lines and animals, opening new avenues for the treatment of mitochondrial genetic diseases. However, since DdCBE almost exclusively induces TC to TT base editing, it can only cover about 1 / 8 of all cases. Therefore, TALE-linked deaminases (TALED) were developed by linking two types of deaminases with TALE (transcription activator-like effector). Here, TALE is custom-designed to bind the desired DNA moiety and fused to a non-catalytically active DddAtox cytosine deaminase variant and the TadA protein, which is a DNA adenine deaminase derived from E. coli. TALED is capable of base editing of A to G conversions, unlike conventional base editing techniques where cytosine base editing is only applicable to the TC background in human mitochondria. In fact, the custom-designed TALED can efficiently (up to about 50%) induce adenine base editing on a variety of targets in human cells.

[0842] To develop new base editing technologies, the TadA variants of ABE8e (TadA*) are selected from various TadA variants. This is because such variants can induce adenine editing with high efficiency and have been modified to be compatible with various DNA-binding proteins, thus being highly compatible with TALE or ZFP (zinc finger proteins) in practical applications.

[0843] TadA* and MTS (mitochondrial targeting sequence) were fused into TALEs tailored to ND1 or ND4 target sites, and the feasibility of actual base editing in mitochondrial DNA was tested. Based on the results of targeted deep sequencing, the adenosine base editing efficiency of the fusion protein was found to be very low but detectable. Adenine base editing was induced, with an efficiency as high as 1.2% at the ND1 site. Figure 67 Figure a shows an efficiency of up to 0.6% at the ND4 site. Figure 67 (Figure b). Although it is known that TadA* only specifically acts on single-stranded DNA, it has been found that base editing in double-stranded target DNA is also induced when TadA* is fused with TALE, although the efficiency is very low.

[0844] Since adenine base editing can occur in mitochondrial DNA, we sought to improve efficiency by fusing the DddAtox protein. DddAtox is a bacterial intertoxin derived from Burkholderia cepacia that deaminates cytosine. This protein acts on double-stranded DNA, thus helping TadA* adenine deaminase to better access the target DNA. For existing DdCBEs using DddAtox, the DddAtox protein was split into two halves, and these halves were then fused separately to a left-TALE (L-TALE) recognizing the left-half DNA site and a right-TALE (R-TALE) recognizing the right-half DNA site, as well as to a uracil glycosylase inhibitor (UGI) that increases cytosine base editing efficiency (TALE-split DddAtox-UGI). DddAtox is used in split form because using the full-length protein causes cytotoxicity. Specifically, TadA* was ligated to either side of the DdCBE targeting the ND1 site instead of UGI, and L-TALE-split DddAtox-TadA* and R-TALE-split DddAtox-UGI, or L-TALE-split DddAtox-UGI and R-TALE-1397C-TadA* forms were prepared and tested. Surprisingly, it was confirmed that when TadA* from one side paired with 1397C and UGI from the other side and was transferred to human cells, A-to-G and C-to-T conversions occurred. Figure 62 (Inset c). In conventional DdCBE, cytosine base editing occurs with an efficiency of approximately 20%, and adenine base editing does not occur at all. When UGI is replaced by TadA* on either side, cytosine base editing is reduced to about half, and adenine base editing occurs with an efficiency of approximately 10%. Figure 62 Figure c). In short, the TALE deaminase generated by fusing the TadA variant with the split DddAtox half of DdCBE can simultaneously induce A-to-G and C-to-T editing in human mtDNA. Figure 62 (Figure c). In conventional DdCBE, cytosine base editing occurs with an efficiency of approximately 20%, and adenine base editing does not occur at all. When TadA* is provided on either side, cytosine base editing is reduced to about half, and adenine base editing occurs with an efficiency of approximately 10%. Figure 62 (Small figure c). In short, adenine base editing and cytosine base editing have similar efficiencies. Figure 62 Small image c).

[0845] Simultaneous cytosine and adenine base editing can be used for random mutagenesis, but in the treatment of diseases, especially mitochondrial genetic diseases such as LOHN and MEALS caused by C-to-T mutations, it is desirable to induce only adenine base editing. Therefore, to eliminate this simultaneous cytosine base editing, UGI is removed. In DdCBE, when the cytosine deaminase DddAtox deaminates C to U, UGI is fused as an inhibitor of uracil glycosylase to prevent U from being repaired again by uracil glycosylase, a repair protein in the cell during DNA repair. Therefore, it was thought that removing this UGI would maintain adenine base editing efficiency and inhibit cytosine base editing. Surprisingly, it was confirmed that the TALE deaminase targeting ND1 without UGI induced almost no cytosine base editing (< 0.5%) and induced adenine base editing alone with high efficiency (approximately 50%). Figure 63 Small picture a and Figure 63 (See Figure c), which is significantly higher than with UGI. This is also confirmed in the TALE deaminase pair targeting ND4. Similar to targeting ND1, only adenine editing was detected with high efficiency (approximately 35%). Figure 63 Small image b and Figure 63 (Inset d). In this way, TALED, a novel adenine deaminase acting on double-stranded DNA, was developed, incorporating the DddAtox system and TadA*, and adenine base editing was enabled for the first time in human mitochondria. Furthermore, adenine base editing was ultimately induced with approximately 50-fold higher efficiency compared to TALE fused solely with TadA*.

[0846] Furthermore, attempts were made to induce adenine base editing using full-length E1347A DddAtox variants with eliminated catalytic activity or variants that maintained catalytic activity but eliminated cytotoxicity (AAAAA and GSVG). Since adenine base editing, rather than cytosine base editing, occurs in single-stranded DNA, the full-length E1347ADddAtox variants lacking cytosine base editing activity could still be used to enhance A-to-G editing efficiency by promoting accessibility of TadA to double-stranded DNA. Additionally, based on the result that cytosine base editing is ineffective in the absence of UGI, variants that only eliminated cytotoxicity were used. Two types of TALEDs containing full-length variants were prepared. Figure 64 Figure a). The first type is configured such that both TadA*(AD) and the full-length DddAtox variant are contained in a single TALE (mTALED), and the second type is configured such that TadA*(AD) and the full-length DddAtox variant are each fused into their respective TALE (dTALD). Both types are then tested. Figure 64Figure a). Surprisingly, it was confirmed that both types of TALED targeting ND1 efficiently induce adenine base editing (…). Figure 64 (Inset b). Here, mTALED exhibits an efficiency of up to approximately 45%, and dTALED also shows an adenine base editing efficiency of approximately 50%. Figure 64 (Figure b). Similar experiments were performed at the ND4 site in addition to the ND1 site, and adenine base editing was induced with similar high efficiency. Figure 64 Figure c). Furthermore, when using the full-length E1347A DddAtox variant, which lacks cytosine base editing activity, adenine base editing was induced with high efficiency. Figure 64 b and Figure 64 c). This is thought to be because the function of facilitating TadA*'s full access to double-stranded DNA is preserved even in the absence of cytosine deamination activity. When these results are reviewed in detail at the single nucleotide level ( Figure 65 and Figure 66 This induces adenine base editing in the region immediately adjacent to the DNA binding site of the TALE. Furthermore, when two TALEs are used, base editing is induced only between them (the spacer), and strangely, even in mTALEDs using only one TALE, the target length is found to be similar. Figure 65 and Figure 66 ).

[0847] Furthermore, we investigated whether this system functions in conjunction with the zinc finger protein (ZFP) system in nuclear DNA. Therefore, we generated NC-type ZFPs targeting nuclear DNA and fused them with splitting DddAtox and TadA*. Figure 61 (Small figure a). Here, TadA* is fused to different positions in the ZFP ( Figure 61 (Figure b). Among the various constructs, those capable of inducing adenine base editing in nuclear DNA with an efficiency of up to 10% were produced. Figure 61 (Small figure d). Because UGI exists on either side, the cytosine base editing efficiency is also very high. Figure 61 (Figure c). The ZFP-DddAtox-TadA* system functions in human cell nuclear DNA, and its function in mitochondria was also tested. Therefore, the construct that functions most efficiently in nuclear DNA was referenced. The experiment was conducted in such a manner that, instead of the nuclear localization signal (NLS) for nuclear DNA, a mitochondrial targeting sequence (MTS) was attached to it, and 1397N was fused to the right ZFP targeting the ND1 site, while TadA* and 1397C were fused to the left ZFP. As a result, adenine base editing was induced with an efficiency of approximately 3%. Figure 61(Figure f). Although the adenine base editing efficiency is lower than that of TALED, the ZFP system can also induce adenine base editing with good efficiency if various conditions such as the linker connecting the protein are optimized.

[0848] Significant progress has been made in gene editing technology to date. CRISPR-based gene scissors (CRISPRCas9, base editors, guide editors, etc.) have been developed in various ways to improve off-target editing and increase efficiency. However, despite these advances, the treatment of mitochondrial genetic diseases remains limited. This is because, unlike proteins, in CRISPR-based technologies that involve both catalytic proteins and guide gRNAs as targets, there is no method to transfer gRNA to mitochondria. Therefore, there are no technologies for processing mitochondrial genes other than eliminating mitochondrial DNA by cutting DNA. David R. Liu's team in the United States first introduced DdCBE, which can induce base editing in mitochondria. Since DdCBE contains the cytosine deaminase DddAtox, which acts on double-stranded DNA, it is fused with the DNA-binding protein TALE to induce base editing. However, because DdCBE only induces limited base editing in the TC background, there are many limitations when creating disease models or treating genetic diseases in real-world applications. Therefore, TALED, which can induce adenine base editing in mitochondria, was developed for the first time. TALED exhibits high efficiency (up to 50%) and induces base editing of various types of adenine at target sites. Furthermore, TALED can induce both cytosine and adenine base editing in the presence of UGI, thus enabling its use in random mutagenesis. It can also be used as a specific adenine base editing technique, as cytosine base editing is induced only in the absence of UGI. It is also applicable to the ZFP system, and adenine base editing can be performed on nuclear DNA. The development of TALED will provide solutions for many mitochondrial genetic diseases, making it possible to create corresponding disease models, and TALED will be used in many previously unexplored mitochondrial gene-related studies.

[0849] Although specific embodiments of the invention have been disclosed in detail above, it will be apparent to those skilled in the art that the descriptions are merely preferred exemplary embodiments and should not be construed as limiting the scope of the invention. Therefore, the essential scope of the invention will be defined by the appended claims and their equivalents.

[0850] Industrial applicability

[0851] According to the present invention, by substituting specific amino acid residues at the interface of cytosine deaminase cleavage products in DdCBE, the nonselectivity of unwanted cytosine deaminase can be reduced.

[0852] Regarding full-length cytosine deaminases, sections that are difficult to edit with traditional cytosine base editors can be edited. Apobec1, used as a deaminase in current cytosine base editors, is known to be an oncogene, and its use for therapeutic purposes is limited, but the full-length deaminase developed in this paper may not have this problem.

[0853] The present invention is as small as about 2.5 kb and includes DNA-binding proteins, thus it can be used in gene therapy using AAV vectors to facilitate the delivery of mRNA and RNPs and to enable the production of useful materials using prokaryotes.

[0854] Sequence List Free Text

[0855] Attached electronic files.

[0856] This disclosure relates to the following implementation plan.

[0857] 1. A fusion protein, said fusion protein comprising:

[0858] (i) DNA-binding proteins; and

[0859] (ii) First and second cleavage products derived from cytosine deaminase or its variants.

[0860] Each of the first and second fission products is fused with the DNA-binding protein.

[0861] 2. A fusion protein, said fusion protein comprising:

[0862] (i) DNA-binding proteins; and

[0863] (ii) Non-toxic full-length cytosine deaminases derived from cytosine deaminases or variants thereof.

[0864] 3. The fusion protein according to embodiment 1, wherein each of the first cleavage product and the second cleavage product does not have cytosine deaminase activity.

[0865] 4. The fusion protein according to embodiment 1, wherein the first cleavage comprises an amino acid sequence starting from the N-terminal residue of SEQ ID NO: 1 to at least one residue selected from G33, G44, A54, N68, G82, N98 and G108.

[0866] 5. The fusion protein according to embodiment 1 or 2, wherein the cytosine deaminase is derived from double-stranded DNA deaminase (DddA) or its ortholog.

[0867] 6. The fusion protein according to embodiment 1, wherein the second cleavage comprises an amino acid sequence from at least one residue selected from G34, P45, G55, N69, T83, A99 and A109 of SEQ ID NO:1 to a C-terminal residue.

[0868] 7. The fusion protein according to embodiment 1, wherein the variant of said cytosine deaminase has at least one amino acid substituted at residues 3, 5, 10, 11, 13, 14, 15, 16, 17, 18, 19, 28, 30 and 31 in the first cleavage comprising the amino acid sequence from the N-terminal residue of SEQ ID NO: 1 to G44.

[0869] 8. The fusion protein according to embodiment 1, wherein at least one amino acid at residues 13, 16, 17, 20, 21, 28, 29, 30, 31, 32, 33, 56, 57, 58 and 60 in the second cleavage comprising the amino acid sequence from P45 to the C-terminal residue of SEQ ID NO: 1 is substituted with a different amino acid.

[0870] 9. The fusion protein according to embodiment 1, wherein the variant of the cytosine deaminase has at least one amino acid substituted at residues 87, 88, 91, 92, 95, 100, 101, 102 and 103 in the first cleavage comprising the amino acid sequence from the N-terminal residue of SEQ ID NO: 1 to G108.

[0871] 10. The fusion protein according to embodiment 1, wherein the variant of said cytosine deaminase has at least one amino acid substituted at residues 13, 14, 15 and 16 in the second cleavage comprising the amino acid sequence from A109 of SEQ ID NO: 1 to the C-terminal residue.

[0872] 11. The fusion protein according to embodiment 2, wherein at least one amino acid at residues 37, 59, 109 and 129 of the nontoxic full-length cytosine deaminase in the wild-type cytosine deaminase of SEQ ID NO: 1 is replaced by a different amino acid.

[0873] 12. The fusion protein according to any one of embodiments 7 to 11, wherein the different amino acid is alanine.

[0874] 13. The fusion protein according to embodiment 2, wherein the nontoxic full-length cytosine deaminase is at least one selected from SEQ ID NO: 12 to 22.

[0875] 14. The fusion protein according to embodiment 1 or 2, wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-associated nuclease.

[0876] 15. The fusion protein according to embodiment 1 or 2, wherein the DNA-binding protein is fused to cytosine deaminase or a variant thereof via a peptide linker comprising 2 to 40 amino acid residues.

[0877] 16. The fusion protein according to embodiment 15, wherein the linker comprises:

[0878] 2a.a Connector: GS;

[0879] 5a.a Connector: TGEKQ (SEQ ID NO: 8);

[0880] 10a.a Connector: SGAQGSTLDF (SEQ ID NO: 9);

[0881] 16a.a Connector: SGSETPGTSESATPES (SEQ ID NO: 10);

[0882] 24a.a connector: SGTPHEVGVYTLSGTPHEVGVYTL (SEQ ID NO: 115); or

[0883] 32a.a Connector: GSGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID NO: 11).

[0884] 17. The fusion protein according to embodiment 1, wherein each of the first cleavage and the second cleavage is fused to the N-terminus or C-terminus of a zinc finger protein.

[0885] 18. The fusion protein according to embodiment 1 or 2, wherein a single TALE array or each of the first TALE array and the second TALE array is fused with the cytosine deaminase.

[0886] 19. The fusion protein according to embodiment 1 or 2, wherein the fusion protein further comprises (iii) adenine deaminase.

[0887] 20. The fusion protein according to embodiment 19, wherein the adenine deaminase is a deoxyadenine deaminase that is a variant of Escherichia coli TadA.

[0888] 21. The fusion protein according to embodiment 19, wherein the adenine deaminase is fused to the N-terminus or C-terminus of the DNA-binding protein or the cytosine deaminase or a variant thereof.

[0889] 22. A nucleic acid that encodes a fusion protein according to any one of embodiments 1 to 21.

[0890] 23. The nucleic acid according to embodiment 22, wherein the nucleic acid is ribonucleic acid or DNA.

[0891] 24. A composition for base editing, said composition comprising a fusion protein according to any one of embodiments 1 to 21 or a nucleic acid according to embodiment 22.

[0892] 25. The composition according to embodiment 24, wherein the composition further comprises a uracil glycosylation inhibitor (UGI).

[0893] 26. A composition for base editing in eukaryotic cells, said composition comprising a fusion protein according to any one of embodiments 1 to 21 or a nucleic acid according to embodiment 22.

[0894] 27. A composition for base editing in plant cells, the composition comprising a fusion protein according to any one of embodiments 1 to 21 or a nucleic acid according to embodiment 22 and a nuclear localization signal (NLS) peptide or a nucleic acid encoding thereon.

[0895] 28. A composition for base editing in plant cells, said composition comprising a fusion protein according to any one of embodiments 1 to 21 or a nucleic acid according to embodiment 22, and a chloroplast transport peptide or a nucleic acid encoding thereon.

[0896] 29. A composition for base editing in plant cells, said composition comprising a fusion protein according to any one of embodiments 1 to 21 or a nucleic acid according to embodiment 22 and a mitochondrial targeting signal (MTS) or a nucleic acid encoding the MTS.

[0897] 30. The composition according to embodiment 29, wherein the composition further comprises a nuclear output signal (NES) or a nucleic acid encoding thereon.

[0898] 31. The composition according to any one of embodiments 27 to 30, wherein the fusion protein is delivered to plant cells by means of gene gun injection (bombardment), PEG-mediated protoplast transfection, protoplast transfection by electroporation, or protoplast injection by microinjection.

[0899] 32. The composition according to embodiment 31, wherein the nucleic acid is delivered to plant cells by means of transformation with Agrobacterium (e.g., Agrobacterium tumefaciens or Agrobacterium rhizogenes), viral transfection, injection (bombardment) using a gene gun, PEG-mediated protoplast transfection, protoplast transfection by electroporation, or protoplast injection by microinjection.

[0900] 33. The composition according to any one of embodiments 27 to 30, wherein the composition is used for base editing in mitochondria, chloroplasts or plastids (leucoplasts, chromoplasts) of plants.

[0901] 34. The composition according to any one of embodiments 27 to 30, wherein the composition further comprises a transcription activator-like effector (TALE)-FokI nuclease or zinc finger nuclease (ZFN) or nucleic acid encoding the thereof that cleaves the wild-type DNA sequence but not the edited base sequence.

[0902] 35. A method for base editing in the nucleus, mitochondria, or plastid DNA of a eukaryotic cell, the method comprising treatment with a composition according to any one of embodiments 27 to 30.

[0903] 36. The method according to embodiment 35, wherein the base editing efficiency is improved by further including a TALEN or ZFN or a nucleic acid encoding the wild-type DNA sequence but not the edited base sequence.

[0904] 37. A method for performing base editing in plant cells, the method comprising treating the plant cells with a composition according to any one of embodiments 27 to 30.

[0905] 38. A method for base editing in plant cells, the method comprising treating the plant cells with a fusion protein according to any one of embodiments 1 to 21 or a nucleic acid according to embodiment 22 and a nuclear localization signal (NLS) peptide or a nucleic acid encoding the same.

[0906] 39. A method for base editing in plant cells, the method comprising treating the plant cells with a fusion protein according to any one of embodiments 1 to 21 or a nucleic acid according to embodiment 22, and a chloroplast transport peptide or a nucleic acid encoding the same.

[0907] 40. A method for base editing in plant cells, the method comprising treating the plant cells with a fusion protein according to any one of embodiments 1 to 21 or a nucleic acid according to embodiment 22 and a mitochondrial targeting signal (MTS) or a nucleic acid encoding the same.

[0908] 41. A composition for base editing in animal cells, the composition comprising a fusion protein according to any one of embodiments 1 to 21 or a nucleic acid according to embodiment 22 and a nuclear localization signal (NLS) peptide or a nucleic acid encoding thereon.

[0909] 42. A composition for base editing in animal cells, the composition comprising a fusion protein according to any one of embodiments 1 to 21 or a nucleic acid according to embodiment 22 and a mitochondrial targeting signal (MTS) or a nucleic acid encoding the MTS.

[0910] 43. The composition according to embodiment 42, wherein the composition further comprises a nuclear output signal or a nucleic acid encoding thereon.

[0911] 44. The composition according to embodiment 42, wherein the composition further comprises a transcription activator-like effector (TALE)-FokI nuclease or ZFN or nucleic acid encoding the thereof that cleaves the wild-type DNA sequence but not the edited base sequence.

[0912] 45. A method for performing base editing in animal cells, the method comprising treating the animal cells with a composition according to embodiment 41 or 42.

[0913] 46. ​​A method for performing base editing in animal cells, the method comprising treating the animal cells with a fusion protein according to any one of embodiments 1 to 21 or a nucleic acid according to embodiment 22 and a nuclear localization signal (NLS) peptide or a nucleic acid encoding the same.

[0914] 47. A method for performing base editing in animal cells, the method comprising treating the animal cells with a fusion protein according to any one of embodiments 1 to 21 or a nucleic acid according to embodiment 22 and a mitochondrial targeting signal (MTS) or a nucleic acid encoding the same.

[0915] 48. The method according to embodiment 46 or 47, wherein the base editing efficiency is improved by further including a TALEN or ZFN or a nucleic acid encoding the wild-type DNA sequence but not the edited base sequence.

[0916] 49. A composition for performing A-to-G base editing in prokaryotic or eukaryotic cells, the composition comprising a fusion protein or nucleic acid encoding thereon according to embodiment 19, wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-associated nuclease, and the fusion protein is a cytosine deaminase or a variant thereof derived from bacteria and specific for double-stranded DNA.

[0917] 50. A composition for A-to-G base editing in prokaryotic or eukaryotic cells, the composition comprising a fusion protein or nucleic acid encoding thereon according to embodiment 19, wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-associated nuclease, the fusion protein having a cytosine deaminase or a variant thereof derived from bacteria and specific for double-stranded DNA, the DNA-binding protein being fused to the N-terminus of the cytosine deaminase or a variant thereof, and the DNA-binding protein being fused to the C-terminus of the adenine deaminase of the fusion protein.

[0918] 51. A composition for C-to-T base editing in prokaryotic or eukaryotic cells, the composition comprising a fusion protein or a nucleic acid encoding thereon according to embodiment 19 and a uracil glycosylase inhibitor (UGI), wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-associated nuclease, and the fusion protein is a cytosine deaminase or a variant thereof derived from bacteria and specific for double-stranded DNA.

[0919] 52. A method for performing A-to-G base editing in prokaryotic or eukaryotic cells, the method comprising treating the prokaryotic or eukaryotic cells with a fusion protein or nucleic acid encoding thereon according to embodiment 19, wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-associated nuclease, and the fusion protein is a cytosine deaminase or a variant thereof derived from bacteria and specific for double-stranded DNA.

[0920] 53. A method for performing A-to-G base editing in prokaryotic or eukaryotic cells, the method comprising treating the prokaryotic or eukaryotic cells with a fusion protein or nucleic acid encoding thereon according to embodiment 19, wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-associated nuclease, the fusion protein having a cytosine deaminase or a variant thereof derived from bacteria and specific for double-stranded DNA, the DNA-binding protein being fused to the N-terminus of the cytosine deaminase or a variant thereof, and the DNA-binding protein being fused to the C-terminus of the adenine deaminase of the fusion protein.

[0921] 54. A method for performing C-to-T base editing in prokaryotic or eukaryotic cells, the method comprising treating the prokaryotic or eukaryotic cells with a fusion protein according to embodiment 19 or a nucleic acid encoding thereon and a uracil glycosylase inhibitor (UGI), wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-associated nuclease, and the fusion protein is a cytosine deaminase or a variant thereof derived from bacteria and specific for double-stranded DNA.

Claims

1. A composition for base editing in plant cells, said composition comprising: (i) (a) a fusion protein comprising a DNA-binding protein and a first cleavage product derived from a cytosine deaminase or a variant thereof; and (b) a fusion protein comprising a DNA-binding protein and a second cleavage product derived from a cytosine deaminase or a variant thereof; or (ii) The nucleic acid encoding the fusion protein.

2. The composition for base editing in plant cells according to claim 1, wherein the cytosine deaminase is derived from double-stranded DNA deaminase (DddA) or its ortholog.

3. The composition for base editing in plant cells according to claim 1, wherein the first splitting compound comprises an amino acid sequence from an N-terminal residue of SEQ ID NO: 1 to at least one residue selected from G33, G44, A54, N68, G82, N98 and G108, and the second splitting compound comprises an amino acid sequence from at least one residue selected from G34, P45, G55, N69, T83, A99 and A109 of SEQ ID NO: 1 to a C-terminal residue.

4. The composition for base editing in plant cells according to claim 1, wherein each of the DNA-binding proteins is independently selected from the group consisting of zinc finger proteins, TALE proteins, or CRISPR-associated nucleases.

5. The composition for base editing in plant cells according to claim 1, wherein the DNA-binding protein is a zinc finger protein.

6. The composition for base editing in plant cells according to claim 1, wherein the DNA-binding protein is a TALE protein.

7. The composition for base editing in plant cells according to claim 1, wherein the DNA-binding protein is a TALE protein. The cytosine deaminase described therein is a cytosine deaminase or its homolog containing the amino acid sequence of SEQ ID NO: 1, and The composition is used for cytosine (C) to thymine (T) base editing.

8. The composition for base editing in plant cells according to claim 1, wherein the DNA-binding protein is a zinc finger protein. The cytosine deaminase described therein is a cytosine deaminase or its homolog containing the amino acid sequence of SEQ ID NO: 1, and The composition is used for cytosine (C) to thymine (T) base editing.

9. The composition for base editing in plant cells according to claim 1, wherein the fusion protein comprising the first cleavage and the fusion protein comprising the second cleavage further comprises adenine deaminase.

10. The composition for base editing in plant cells according to claim 9, wherein the adenine deaminase is a variant of Escherichia coli TadA (tRNA-specific adenine deaminase).

11. The composition for base editing in plant cells according to claim 9, wherein the DNA-binding protein is a TALE protein. The cytosine deaminase described therein is a cytosine deaminase or a homolog thereof containing the amino acid sequence of SEQ ID NO:

1. The adenine deaminase described therein comprises the amino acid sequence of SEQ ID NO: 458 or a variant thereof, and The composition is used for adenine (A) to guanine (G) base editing.

12. The composition for base editing in plant cells according to claim 1, wherein the composition is for nuclear DNA base editing and optionally further comprises a nuclear localization signal (NLS) peptide or a nucleic acid encoding thereon.

13. The composition for base editing in plant cells according to claim 1, wherein the composition is for mitochondrial DNA base editing, and optionally further comprises (1) a mitochondrial targeting signal (MTS) or a nucleic acid encoding thereof and (2) a nuclear export signal (NES) or a nucleic acid encoding thereof.

14. The composition for base editing in plant cells according to claim 1, wherein the composition is for base editing in chloroplasts, leucoplasts or chromoplasts, and optionally further comprises (1) a chloroplast transport peptide (CTP) or a nucleic acid encoding thereof and (2) a nuclear export signal (NES) or a nucleic acid encoding thereof.

15. The composition for base editing in plant cells according to claim 1, wherein the composition further comprises a uracil glycosylase inhibitor (UGI) or a nucleic acid encoding thereon.

16. A method for base editing in plant cells, the method comprising the step of treating the plant cells with a composition according to any one of claims 1 to 15.

17. A plant or seed obtained from the method of claim 16.