Targeted deaminase and base editing using it
Fusion proteins with DNA-binding proteins and non-toxic deaminases enable targeted base editing in plant organelles, addressing delivery and expression challenges, enhancing genetic research and crop improvement.
Patent Information
- Application Number
- JP2023517882
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Priority Date
- 2021-08-30
- Filing Date
- 2021-09-17
- Publication Date
- 2026-01-09
- Estimated Expiration
- 2041-09-17
Smart Images

Figure 0007796729000032 
Figure 0007796729000033 
Figure 0007796729000034
Abstract
Description
[Technical Field]
[0001] The present invention relates to an isolated form of cytosine or adenine deaminase or a mutant thereof, a non-toxic full-length cytosine deaminase or a mutant thereof, a fusion protein containing the same, a base editing composition, and a method for editing bases using the same. [Background technology]
[0002] Fusion proteins linking DNA-binding proteins and deaminase enzymes enable targeted nucleotide substitution or base editing in genes without generating DNA double-strand breaks (DSBs), edit point mutations that induce genetic disorders, or perform single nucleotide transversions in a targeted manner to introduce desired single nucleotide mutations in prokaryotic, human, and other eukaryotic cells.
[0003] Unlike nucleases like CRISPR-Cas9, which induce small insertions or deletions (indels) at a target site, this method converts a single base within a window of several nucleotides at the target site, thus allowing for the editing of point mutations or the generation of single nucleotide polymorphisms (SNPs) that cause genetic diseases in cultured cells, animals, and plants.
[0004] Fusion proteins linking DNA-binding proteins and deaminase enzymes include: 1) base editors (BEs) containing catalytically-deficient Cas9 (dCas9) or D10A Cas9 nickase (nCas9) from S. pyogenes and the rat cytosine deaminase rAPOBEC1; 2) Target-AID, which contains dCas9 or nCas9 and the sea lamprey AID (activation-induced cytidine deaminase) ortholog PmCDA1 or human AID; and 3) CRISPR-X, which contains dCas9 and sgRNAs linked to an MS2 RNA hairpin to recruit a hyperactivated AID mutant fused to an MS2-binding protein.
[0005] Thus, programmed gene editing tools, such as zinc finger nuclease (ZFN), transcription activator-like effector nuclease (TALEN), clustered regularly interspaced short palindromic repeat (CRISPR) systems, CRISPR-associated protein 9 (Cas9) mutants, and base editing technologies for base deaminase proteins, have been developed for plant genetic research and crop trait improvement through alteration of base sequences. However, these tools are not suitable for editing DNA sequences in plant organelles, including mitochondria and chloroplasts, primarily because of the difficulty of delivering guide RNAs to organelles or simultaneously expressing two compounds in organelles. Plant organelles encode numerous essential genes required for photosynthesis and respiration. Methods and tools for editing genes in these organelles are critically needed for studying the functions of these genes and improving crop productivity and traits. For example, targeted mutations in the mitochondrial atp6 gene can cause male sterility, a trait useful for reproduction, and specific point mutations in the 16S rRNA gene of the chloroplast genome can cause antibiotic resistance.
[0006] Bacterial toxin DddA tox is the enzymatic part of a bacterial toxin derived from Burkholderia cenocepacia, which is capable of deaminating cytosine in double-stranded DNA. An example of a deaminase is DddA. tox is toxic to cells, so to avoid toxicity in host cells, DddA tox is split into two inactivated split halves, each of which can be linked to a DNA-binding protein designed to bind to DNA to form a functional DdCBE pair.
[0007] In principle, this deamination enzymatic reaction can only be activated when the two inactive halves are present near the target DNA by DNA-binding proteins. Therefore, base editing of cytosine to thymine occurs in the middle of the binding sites of the two DNA-binding proteins. The two inactive forms bound to a TALE (transcription activator-like effector) array using DNA-binding proteins become functional when bound near the target DNA by the TALE. Base conversion from cytosine to thymine is induced in the region of 14 to 18 bases between the two TALE-binding sites. This split form of DddA tox However, there are many limitations to the experiment.
[0008] Full length DddA tox Because of its toxicity, it cannot be cloned using ordinary E. coli, so it is cloned using E. coli that also expresses an immunity gene that prevents toxicity.
[0009] Mitochondrial DNA plays a crucial role in cellular respiration, which is achieved through the mitochondrial oxidative phosphorylation (OXPHOS) mechanism. Because the OXPHOS mechanism is essential for survival, mitochondrial DNA mutations can cause severe dysfunction in various energy-hungry organs and muscles. In human mitochondrial diseases, normal mitochondrial DNA and mitochondrial DNA with single-base mutations coexist, resulting in mitochondrial DNA heterogeneity (heteroplasmy). The balance between mutations and normal mitochondrial DNA determines the development of clinically symptomatic mitochondrial diseases. Programmable nucleases have been used in vitro and in vivo to cleave mutant mitochondrial DNA without removing normal mitochondrial DNA. However, these nucleases cannot introduce or reverse specific mutations into mitochondria because double-stranded DNA breaks in mitochondria are not repaired as efficiently in mitochondria through nonhomologous end joining or homologous recombination as they are in the nucleus.
[0010] Mitochondrial base editing can help create models for various diseases that have not been previously possible or develop therapeutic agents for treating them. For this reason, there is an increasing need to develop highly efficient mitochondrial base editing enzymes.
[0011] Against this technical background, the present inventors have completed the present invention by substituting the residues of deaminase enzymes to reduce non-selective base editing or by creating novel full-length deaminase enzymes to eliminate toxicity, and by confirming that these enzymes can be used as the desired CBEs (cytosine base editors) or ABEs (adenine base editors) and can edit DNA. Summary of the Invention [Problem to be solved by the invention]
[0012] It is an object of the present invention to provide a DNA binding protein, a fusion protein comprising an isolated form of cytosine or adenine deaminase or a mutant thereof, a non-toxic full-length cytosine deaminase, or a mutant thereof.
[0013] An object of the present invention is to provide a nucleic acid encoding the fusion protein.
[0014] An object of the present invention is to provide a base editing composition comprising the fusion protein or nucleic acid.
[0015] An object of the present invention is to provide a base editing method comprising a step of treating a cell with the composition.
[0016] To achieve the above-mentioned object, the present invention provides a fusion protein comprising (i) a DNA-binding protein; and (ii) a first and second split body derived from cytosine deaminase or a mutant thereof, wherein the first and second split bodies are each capable of binding to the DNA-binding protein.
[0017] The present invention provides a fusion protein comprising: (i) a DNA-binding protein; and (ii) a non-toxic full-length cytosine deaminase derived from cytosine deaminase or a mutant thereof.
[0018] The present invention provides a fusion protein comprising (i) a DNA-binding protein; and (ii) cytosine deaminase or a mutant thereof, and (iii) adenine deaminase, wherein the cytosine deaminase or mutant thereof comprises (a) a non-toxic full-length cytosine deaminase or (b) a first split body and a second split body derived from the cytosine deaminase or the mutant thereof, and the first split body and the second split body are each in a form capable of binding to the DNA-binding protein.
[0019] The present invention provides a nucleic acid encoding the fusion protein. The present invention provides a base editing composition comprising the fusion protein or nucleic acid. The present invention provides a composition for base editing in eukaryotic cells, comprising the fusion protein or nucleic acid.
[0020] The present invention also provides a composition for base editing in plant cells, comprising the fusion protein or nucleic acid; and an NLS (nuclear localization signal) peptide or a nucleic acid encoding the same.
[0021] The present invention also provides a composition for base editing in plant cells, comprising the fusion protein or nucleic acid; and a chloroplast transit peptide or a nucleic acid encoding the same.
[0022] The present invention also provides a composition for base editing in plant cells, comprising the fusion protein or nucleic acid; and a mitochondrial targeting signal (MTS) or a nucleic acid encoding the same.
[0023] In some cases, the present invention also provides a composition for base editing in a plant cell, further comprising a nuclear export signal protein or a nucleic acid encoding the same.
[0024] The present invention also provides a method for base editing in plant cells, comprising treating the plant cells with the composition.
[0025] The present invention also provides a method for base editing in plant cells, comprising treating a plant cell with the fusion protein or nucleic acid; and an NLS (nuclear localization signal) peptide or a nucleic acid encoding the same.
[0026] The present invention also provides a method for base editing in plant cells, comprising treating a plant cell with the fusion protein or nucleic acid; and a chloroplast transit peptide or a nucleic acid encoding the same.
[0027] The present invention also provides a method for base editing in plant cells, comprising treating a plant cell with the fusion protein or nucleic acid; and a mitochondrial targeting signal (MTS) or a nucleic acid encoding the same.
[0028] The present invention also provides a composition for base editing in animal cells, comprising the fusion protein or nucleic acid; and an NLS (nuclear localization signal) peptide or a nucleic acid encoding the same.
[0029] The present invention also provides a composition for base editing in animal cells, comprising the fusion protein or nucleic acid; and a mitochondrial targeting signal (MTS) or a nucleic acid encoding the same.
[0030] The present invention also provides a composition for base editing in animal cells, which optionally further comprises a nuclear export signal protein or a nucleic acid encoding the same.
[0031] The present invention also provides a method for base editing in animal cells, comprising treating the animal cells with the composition.
[0032] The present invention also provides a method for base editing in animal cells, comprising treating the animal cells with the fusion protein or nucleic acid, and an NLS (nuclear localization signal) peptide or a nucleic acid encoding the same.
[0033] The present invention also provides a method for base editing in animal cells, comprising treating the animal cells with the fusion protein or nucleic acid, and a mitochondrial targeting signal (MTS) or a nucleic acid encoding the same.
[0034] The present invention also provides a composition for A-to-G base editing in prokaryotic or eukaryotic cells, comprising the fusion protein or a nucleic acid encoding it, wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-linked nuclease, and the cytosine deaminase or a mutant thereof of the fusion protein is derived from bacteria and is specific for double-stranded DNA.
[0035] The present invention also provides a composition for A-to-G base editing in prokaryotic or eukaryotic cells, comprising the fusion protein or a nucleic acid encoding it, wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-linked nuclease, the cytosine deaminase or a mutant thereof of the fusion protein is derived from a bacterium and is specific to double-stranded DNA, a DNA-binding protein is bound to the N-terminus of the cytosine deaminase or a mutant thereof, and a DNA-binding protein is bound to the C-terminus of the adenine deaminase of the fusion protein.
[0036] The present invention also provides a composition for C-to-T base editing in prokaryotic or eukaryotic cells, comprising the fusion protein or a nucleic acid encoding the same and UGI (uracil glycosylase inhibitor), wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-linked nuclease, the cytosine deaminase or its variant of the fusion protein is a full-length non-toxic cytosine deaminase, and the cytosine deaminase or its variant of the fusion protein is derived from a bacterium and is specific to double-stranded DNA.
[0037] The present invention also provides a composition for C-to-T base editing in prokaryotic or eukaryotic cells, comprising the fusion protein or a nucleic acid encoding the fusion protein and UGI (uracil glycosylase inhibitor), wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-linked nuclease, the cytosine deaminase or its variant of the fusion protein is a split cytosine deaminase comprising a first split body and a second split body, and the cytosine deaminase or its variant of the fusion protein is derived from a bacterium and is specific to double-stranded DNA.
[0038] The present invention also provides an A-to-G base editing method in a prokaryotic or eukaryotic cell, comprising a step of treating the fusion protein or a nucleic acid encoding the same in the prokaryotic or eukaryotic cell, wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-linked nuclease, and the cytosine deaminase or a mutant thereof of the fusion protein is derived from bacteria and is specific to double-stranded DNA.
[0039] The present invention also provides an A-to-G base editing method in a prokaryotic or eukaryotic cell, the method comprising the step of treating a prokaryotic or eukaryotic cell with the fusion protein or a nucleic acid encoding the same, wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-linked nuclease;
[0040] the cytosine deaminase or its variant of the fusion protein is derived from bacteria and is specific for double-stranded DNA;
[0041] wherein a DNA-binding protein is bound to the N-terminus of the cytosine deaminase or mutant thereof, and a DNA-binding protein is bound to the C-terminus of the adenine deaminase of the fusion protein.
[0042] The present invention also provides a C-to-T base editing method in a prokaryotic or eukaryotic cell, comprising the step of treating the prokaryotic or eukaryotic cell with the fusion protein or a nucleic acid encoding it and UGI (uracil glycosylase inhibitor), wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-linked nuclease, and the cytosine deaminase or a mutant thereof of the fusion protein is derived from bacteria and is specific to double-stranded DNA. [Brief explanation of the drawings]
[0043] [Figure 1]Figure 1 shows the results of ZFD optimization using the pTarget plasmid. a) ZFD structure. One half of the split-DddAtox is fused to the C-terminus of a ZFP (C-type). b) Optimization of the ZFD platform using the pTarget library. The pTarget plasmid contains a spacer ranging in size from 1 to 24 bp (shown in red) and a ZFP DNA binding site (shown in green). The ZFD structure includes AA linkers of various lengths (shown in yellow and orange) and alternative DddAtox split sites and orientations (shown in blue). c) ZFD activity was measured at the target site in the pTarget library to examine the effects of the variables described in (b). ZFD pairs with linkers of the same (c) or different (d) lengths were tested on the left and right ZFDs. Base editing frequencies were measured by targeted deep sequencing of the relevant regions of the pTarget plasmid. Data are shown as the mean ± standard error of the mean (s.e.m.) obtained from n = 2 biologically independent samples. [Figure 2] Figure 1 shows the results of ZFD efficiency testing with various linkers in the pTarget plasmid. a) Heatmap showing the frequency of C / G to non-C / G editing in various pTarget spacers with lengths ranging from 1 to 24 bp. Various ZFD configurations were tested, including various linkers between the ZFP and the split DddAtox and the site at which the DddAtox was split. b) Overall activity of each ZFD pair. The nomenclature used on the x-axis indicates the left ZFD at the bottom and the right ZFD at the top. c) Base editing efficiency by spacer length. "AA" indicates the number of amino acids in the linker. Data are shown as the mean ± standard error of the mean (sem) obtained from n = 2 biologically independent samples. [Figure 3]Figure 1 shows the results of confirming the efficiency of ZFDs with a 24 AA linker versus other linkers. The effect of ZFD linker length on C / G to non-C / G editing efficiency is shown in a heatmap. The left ZFD of a ZFD pair is fixed with a 24 AA linker, while the right ZFD contains a linker of variable length, and vice versa. Error bars represent the standard error of the mean (s.e.m.) of n = 2 biologically independent samples. [Figure 4] Figure 1 shows the results of confirming the activity of ZFDs targeting the nucleus in vivo. a) Structure of the nuclear DNA-targeting ZFD. Half of the Split-DddAtox is fused to the C-terminus (C-form) or N-terminus (N-form) of the ZFP. The ZFD pair was designed in CC or NC configuration, consisting of the left ZFD of the C-form and the right ZFD of the C-form, or the left ZFD of the N-form and the right ZFD of the C-form, respectively. b) Base editing frequency induced by ZFD at nuclear DNA target sites in HEK 293T cells. Data are shown as mean ± sem for n = 3 biologically independent samples. c) ZFD-induced base editing efficiency at each base position within the spacer of the NUMBL (c), INPP5D-2 (d), TRAC-CC (e), and TRAC-NC (f) target sites in HEK 293T cells. Data are shown as mean ± sem for n = 3 biologically independent samples. g, ZFD-induced base editing frequency in K562 cells after delivery of ZFD protein or ZFD-encoding plasmid by electroporation or direct delivery. ZFD proteins with one or four NLSs were tested, with left- and right-handed ZFDs used as isotypes. Electroporation was performed using an Amaxa 4D-Nucleofector. For direct delivery, K562 cells were incubated with cell culture medium containing left- and right-handed ZFD proteins. Cells were treated once (1x) or twice (2x) in the same manner. Data are shown as mean ± SEM for biologically independent samples (n = 2). [Figure 5] Schematic diagram showing the construction of ZFDs that target nuclear DNA. ad, Four possible ZFD constructions. The NC and CN constructions are structurally identical, but the constructional types of the left and right ZFDs are different. [Figure 6] Figure 1 shows the results of confirming the deletion rate of ZFDs targeted to the nucleus in vivo. All tested ZFDs generated deletions at frequencies less than 0.4%. Data are shown as mean ± sem for n = 3 biologically independent samples. [Figure 7] Figure 1 shows the results of an in vitro activity test on ZFD recombinant proteins. a) Purification steps of the ZFD pair targeting the TRAC site. GST-tagged proteins were purified in E. coli cell lysates using glutathione Sepharose beads. Polyacrylamide gel electrophoresis was used to monitor the purification steps. The gel was stained with Coomassie blue. Lane 1, molecular weight marker. Lane 2, sample from cells in which protein expression was not induced with IPTG. Lane 3, sample from cells in which protein expression was induced with IPTG. Lane 4, soluble fraction after sonication. Lane 5, insoluble fraction after sonication. Lane 6, column flow-through. Lane 6, wash fraction. Lane 7, elution fraction. The sizes of representative markers are indicated on the left. The red box indicates the ZFD protein. b) ZFD binding sites on the left and right. The red arrows indicate potential sites for ZFD-induced deamination. c) Overview of ZFD activity on PCR amplicons containing the TRAC site. The TRAC-NC ZFD pair first deaminates cytosine to generate uracil (shown in red). The USER enzyme then cleaves the uracil to generate a cleft (shown in red triangles). d, Untreated PCR amplicon (left) and ZFD pair-treated PCR amplicon (right) analyzed by agarose gel electrophoresis. [Figure 8] Schematic diagram of various configurations of mitochondrial DNA-targeting ZFDs. ad, Four possible mitoZFD configurations. The NLS of the conventional ZFD has been replaced with an MTS and NES. While the NC and CN configurations are structurally identical, the configurational typology of the left and right ZFDs differs. [Figure 9]This figure shows the results of verifying the mitochondrial gene base editing efficiency of mitoZFD. a) Frequency of mtDNA base editing induced by mitoZFD and TALE-DdCBE in HEK 293T cells. Data are shown as mean ± standard error of the mean (sem) obtained from two biologically independent samples. bg) MitoZFD-induced base editing efficiency at each base position within the spacer of the ND2 (b), ND4L (c), COX2 (d), ND6 (e), and ND1 (f) target sites in HEK 293T cells, and TALE-DdCBE-induced base editing efficiency at the ND1 (g) target site. Data are shown as mean ± standard error of the mean (sem) obtained from two biologically independent samples. Comparison of DNA and amino acid changes in the ND1 gene introduced by mitoZFD and TALE-DdCBE. The frequency (%) of sequencing reads for each mutant allele was measured by targeted deep sequencing. The spacer regions of the ZFD pair and the TALE-DdCBE pair are indicated by blue dotted lines. [Figure 10] Figure 1 shows the results of base editing efficiency of single-cell-derived clone populations isolated from MT-ZFD-treated HEK293T cells. Single-cell-derived clones were obtained for allelic analysis. The C / G-to-non-C / G editing frequency of each single-cell-derived clone was determined by targeted deep sequencing. a) Single-cell-derived clone from a HEK293T cell population treated with ND1-targeted mitoZFD. b) Single-cell-derived clone from a HEK293T cell population treated with ND2-targeted mitoZFD. c) Single-cell-derived clone from an untreated HEK293T cell population. ZFP binding sites are indicated in green. High editing frequencies of clones that underwent mitoZFD-induced editing are indicated in red. [Figure 11]This figure shows the results of confirming the base editing efficiency of a single-cell-derived clone population isolated from MT-ZFD-treated HEK293T cells. Allele analysis of single-cell-derived clones showing high frequency of base editing. The table shows the amino acids changed as a result of base editing in ND1. In the reference sequence at the top, red letters indicate spacers. In the alleles, red letters indicate changes in the amino acid sequence. (* indicates a stop codon.) [Figure 12a] and [Figure 12b] This figure shows the results of confirming the base editing efficiency of a single-cell-derived clone population isolated from MT-ZFD-treated HEK293T cells. Allele analysis of single-cell-derived clones showing high frequency of base editing. The table shows the amino acids changed as a result of base editing in ND2. In the reference sequence at the top, red letters indicate spacers. In the alleles, red letters indicate changes in the amino acid sequence. [Figure 13] Figure 1 shows the results of confirming the base editing efficiency of the combination of ZFD and TALE-DdCBE. a) DNA sequences of the binding regions of the mitoZFD and TALE-DdCBE pairs. The sites recognized by TALE-DdCBE are highlighted in green, while those for mitoZFD are highlighted in blue. The top sequence represents the mtDNA heavy strand, and the bottom sequence represents the mtDNA light strand. b) Frequency of cytosines edited by ZFD, TALE-DdCBE, and the ZFD / DdCBE mixed pair. Data were obtained using targeted deep sequencing, and the data are shown as mean ± standard error of the mean (s.e.m.) from two biologically independent samples. c) Heat map of base editing activity at each base position. The red box indicates the spacer region for each construct. The blue arrow indicates the mtDNA position. [Figure 14]This figure shows the results of confirming editing efficiency by mRNA vs. plasmid at different ZFD concentrations. The target specificity of ND1-targeted mitoZFD across the mitochondrial genome varies depending on the concentration of ZFD-encoding mRNA or plasmid. The on- and off-target base editing frequencies were determined by sequencing the entire mtDNA. The graph shows the results of HEK 293T cells transfected with the indicated concentrations of ND1-targeted mitoZFD-encoding plasmid or mRNA. The red arrow indicates the target site, the red dots indicate the base editing frequency at the target site, and the gray dots indicate SNPs also present in the control. Data are shown as the mean ± standard error of the mean (sem) obtained from n=2 biologically independent samples. [Figure 15] The results of confirming the editing efficiency of mRNA vs. plasmid at different ZFD concentrations are shown. a) The top panel shows ZFD binding at the ND1 site. The ZFD binding site is shown in green. The target cytosine between the spacers is shown in red. On-target activity determined from the entire mtDNA sequence data in Figure 14. Activity decreases as the amount of transfected mitoZFD-encoding plasmid or mRNA decreases. b) The number of C / G sites edited at a frequency of >1% for each plasmid or mRNA amount. c) The average C / G-to-T / A editing frequency of all C / Gs in the mitochondrial genome at each plasmid or mRNA concentration. Data are shown as the mean ± standard error of the mean (s.e.m.) obtained from n = 2 biologically independent samples. [Figure 16]The results show the editing efficiency of mRNA vs. plasmid at different ZFD concentrations. The on- and off-target base editing frequencies were determined by sequencing the entire mtDNA. The graph shows the results of HEK 293T cells transfected with the indicated concentrations of ND2-targeted mitoZFD-encoding plasmid or mRNA. The red arrow indicates the target site, the red dots indicate the base editing frequency at the target site, and the gray dots indicate SNPs present in the control. Data are shown as the mean ± standard error of the mean (sem) from n=2 biologically independent samples. [Figure 17] This figure shows the results of confirming editing efficiency by mRNA vs. plasmid at different ZFD concentrations. a) The top panel shows ZFD binding at the ND2 site. The ZFD binding site is shown in green. The target cytosine between the spacers is shown in red. On-target activity determined from the entire mtDNA sequence data in Figure 16. Activity decreases as the amount of transfected mitoZFD-encoding plasmid or mRNA decreases. b) The number of C / G sites edited at a frequency of >1% for each plasmid or mRNA amount. c) The average C / G-to-T / A editing frequency for all C / Gs in the mitochondrial genome at each plasmid or mRNA concentration. Data are shown as the mean ± standard error of the mean (sem) obtained from n = 2 biologically independent samples. [Figure 18]We generated whole-mitochondrial sequencing / QQ mutants and confirmed their editing efficiency. a) The QQ mitoZFD mutant contains an R(-5)Q mutation in each zinc finger of the ZFD to eliminate nonspecific DNA contacts. (If R is absent at position -5 of the zinc finger framework, a nearby K or R is converted to Q.) b) Whole-mtDNA sequencing of mitoZFD-treated cells. On-site and off-site editing frequencies are indicated by red and black dots, respectively. Data are shown as mean ± standard error of the mean (s.e.m.) from two biologically independent samples. All C / G-to-T / A base edits with >1% efficiency are shown. c) Editing efficiency and specificity depend on the dose of delivered ZFD-encoding mRNA. c) Average C / G-to-T / A editing frequency relative to all C / Gs in the mitochondrial genome. d) Number of edited C / Gs with base editing frequencies >1%. [Figure 19] Golden Gate assembly system for plant-based base editors. Schematic of Golden Gate assembly for the cp-DdCBE and mt-DdCBE constructs. TALE subarray plasmids were selected from a total of 424 sets (= 6 x 64 3 copies + 2 x 16 2 copies + 2 x 4 1 copy) for each position in the target sequence and mixed with the desired vector to generate plasmids encoding DdCBEs targeting specific sequences. [Figure 20]Plant chloroplast and mitochondrial base editing. a, b, c, d: Frequency and pattern of chloroplast base editing induced by cp-DdCBE in 16s rDNA (a, b) and psbA (c, d). The split DdCBE pair G1333 and G1397 was transfected into lettuce and rapeseed protoplasts. e: Mitochondrial base editing efficiency and pattern induced by mt-DdCBE in the fATP6 gene. The split DdCBE pair G1333 and G1397 was transfected into lettuce and rapeseed protoplasts. a, c, e: TALE binding regions are shown in blue, and spacer cytosines are shown in orange. In all graphs, error bars represent the mean ± standard deviation (SD) of three independent biological replicates. b, d, f: Converted nucleotides are shown in red. % edited alleles (mean ± standard deviation) were obtained from three independent experiments. [Figure 21] DNA editing in plant organs by DdCBE. a) Schematic of plant organ mutagenesis. b) Representative Sanger sequencing chromatogram showing the efficiency of C·G to T·A conversion in cp-DdCBE-transfected calli cultured in the absence of spectinomycin. The converted nucleotide is highlighted in red to the left. The arrow indicates the substituted nucleotide in the chromatogram. c) Summary of DdCBE-driven plant organ mutagenesis. Mutant calli are shown to have editing frequencies much higher than those in simulated calli. d) C to T conversion frequencies induced after transfection of lettuce protoplasts with mRNA encoding cp-DdCBE targeting 16s rDNA. Error bars represent the mean ± SD. n = 3 independent biological replicates. e) Editing frequency and pattern in spectinomycin-resistant calli at 2.5 months. f) Representative Sanger sequencing chromatogram showing the efficiency of C·G to T·A conversion in streptomycin-resistant plants transfected with DdCBE mRNA. Arrows indicate the substituted nucleotides in the chromatogram. Scale bar: 1 mm. [Figure 22]Chloroplast and mitochondrial base editing strategies. The cp-DdCBE and mt-DdCBE preproteins contain a chloroplast transit peptide (CTP) or a mitochondrial targeting signal (MTS), respectively, which allows them to be transported into chloroplasts and mitochondria after translation in plant cells. The preproteins cross the outer and inner membranes of the organelles, where the CTP and MTS are cleaved by stromal processing peptidases and mitochondrial processing peptidases, respectively, resulting in the formation of the mature cp-DdCBE and mt-DdCBE proteins. [Figure 23] Time course for editing via the DdCBE plasmid in lettuce protoplasts. Transfected protoplasts were collected at each time point and analyzed for editing efficiency by targeted deep sequencing. Frequencies (mean ± standard deviation) are from three independent experiments. [Figure 24] Base editing frequency of the psbB gene. After transfection of rapeseed protoplasts with a plasmid encoding cp-DdCBE versus Left-G1333-N + Right-G1333-C targeting the chloroplast psbB gene, the base editing efficiency of the spacer region was analyzed by targeted deep sequencing. The TALE binding region, target cytosine, and converted nucleotide are shown in blue, orange, and red, respectively. The frequency (mean ± standard deviation) was calculated from n = 3 independent experiments. [Figure 25] Base editing efficiency of the mitochondrial RPS14 gene. After transfection of rapeseed protoplasts with a plasmid encoding mt-DdCBE versus Left-G1333-N + Right-G1333-C targeting the RPS14 gene, the C to T conversion efficiency was analyzed by targeted deep sequencing. The TALE binding region, target cytosine, and converted nucleotide are shown in blue, orange, and red, respectively. Frequencies (mean ± standard deviation) were calculated from n = 3 independent experiments. [Figure 26] Base editing efficiency of the chloroplast genome in callus. Frequency and pattern of DdCBE-mediated base editing at target sites in 16srDNA and psbA in lettuce and rapeseed callus after 4 weeks of culture. Converted nucleotides in the spacer region are shown in red. [Figure 27] Efficiency of base editing in the mitochondrial genome in callus. The frequency and pattern of DdCBE-mediated base editing at target sites in the ATP6 and RPS14 genes in rapeseed callus were confirmed by targeted deep sequencing. Converted nucleotides in the target spacer regions are shown in red. [Figure 28] DNA-free base editing. Frequency and pattern of chloroplast base editing at target sites in 16s rDNA after transfection of lettuce protoplasts with DdCBE mRNA. Protoplasts were cultured for 7 days and then subjected to targeted dip sequencing. Converted nucleotides in the target spacer region are shown in red. [Figure 29] FIG. 1 shows the results of gel electrophoresis demonstrating the absence of DdCBE mRNA or DNA sequences in protoplasts and calli (M is a marker). [Figure 30] Selection of 16s rDNA mutations. The red arrow indicates the green streptomycin-resistant callus. [Figure 31] No off-target mutations were found near the DdCBE target site in antibiotic-resistant calli or seedlings. (a), (b) Off-target activity was analyzed by targeted dip sequencing. The TALE binding site and spacer region are underlined in green and red, respectively. (a) Spectinomycin-resistant calli cultured from lettuce protoplasts transfected with the DdCBE plasmid. (b) Shoots from streptomycin-resistant seedlings. [Figure 32]Off-target activity analysis of the five most homologous sites to the on-target site. Five potential off-target sites for the 16S rRNA gene-specific DdCBE in the lettuce chloroplast genome were selected, including up to nine mismatches in the TALE binding site. The TALE binding sequence and mismatched nucleotides are shown in blue and red, respectively. Off-target mutation frequencies were measured in protoplasts and drug-resistant calli or suit transfected with DdCBE plasmid or DdCBE mRNA using targeted dip-sequencing. Frequencies (mean ± standard deviation) were obtained from three independent experiments. [Figure 33] Schematic diagram of DdCBE assembly and mitochondrial DNA editing. a) Illustration of one-pot Golden-gate assembly for efficient DdCBE construction. A total of 424 sequences (64 triple-recognition sequences x 6 + 16 double-recognition sequences x 2 + 4 single-recognition sequences x 2) and expression vectors were mixed to generate left and right modules for the final plasmid construction. b) Schematic diagram of DdCBE interacting with the target gene ND5 in mouse mitochondrial DNA. TALE binding sites are shown in gray, and the base editing region is shown in black. Each repeat variable diresidue module is shown in orange, blue, green, and yellow, respectively, for adenine (NI), thymine (NG), guanine (NN), and cytosine (HD). [Figure 34]Mouse mitochondrial ND5 point mutations resulting from base editing with DdCBE. a) Target sequence and efficiency of DdCBE deaminase-mediated cytosine-to-thymine base editing in NIH3T3 cells. In the target sequence, the translation codon is underlined, and the editable site is highlighted in red. The combinations for DdCBE transfection are indicated as left or right, -G1333 or -G1397, and -N or -C. The P values for the C10 mutations of left-G1333-N + right-G1333-C, left-G1333-C + right-G1333-N, left-G1397-N + right-G1397-C, and left-G1397-C + right-G1397-N are 0.0012, 0.0003, 0.0014, and 0.0009, respectively, and for the C13 mutations are 0.0116, 0.0076, 0.0030, and 0.0003, respectively (*p<0.05 and **p<0.01, using Student's two-tailed t-test). b, Resulting base editing efficiency in mouse blastocysts. Sequence analysis data were obtained from blastocysts developed from conjugates microinjected with left-G1397-N and right-G1397-C-DdCBE mRNA. (c) Alignment of neonatal mutant sequences. Targeted dip sequencing was performed on genomic DNA extracted from the tail section obtained immediately after birth and from the toe section obtained at postnatal days 7 and 14. Edited bases are shown in red. The editing frequency of the mutant mitochondrial genome is shown. (d) Editing efficiency in various tissues of adult F0 mice (sipup-1). Sequencing data were obtained from each tissue at postnatal day 50. In all graphs, dark and light gray bars indicate the editing frequency of the m.C12539T (C10) and m.G12542A (C13) mutations, respectively. Error bars represent the standard error of the mean (s.e.m.) for n = 3 biologically independent samples. [Figure 35]Transmission of mutant mitochondrial DNA to germ cells. a) To observe gonadal transmission of mtDNA mutations, female F0 (sipup-3) mice were mated with wild-type C57BL6 / J males to obtain F1 offspring (101, 102), followed by targeted deep sequencing. Edited bases are shown in red. The editing frequency of the mutant mitochondrial genome is shown. b) Base editing efficiency in various tissues of F1 offspring (101) obtained using targeted deep sequencing of genomic DNA. Dark and light gray bars indicate the frequency of the m.C12539T (C10) and m.G12542A (C13) mutations, respectively. Error bars represent the standard error of the mean (s.e.m.) of n = 3 biologically independent samples. [Figure 36]Mouse mitochondrial ND5 G12918A mutation generated by DdCBE. a, DdCBE targets to create the m.G12918A point mutation, which causes a D393N change in the ND5 protein. The target codon is underlined, and the potential editing site is highlighted in red. b, Efficiency of cytosine-thymine base editing using DdCBE in NIH3T3 cells. Transfected DdCBE pair combinations are shown. Error bars are sem. n = 3 biologically independent samples (ns not significant, *p < 0.05, **p < 0.01 using Student's two-tailed t test). The P values for the C6 mutations left-G1333-N + right-G1333-C, left-G1333-C + right-G1333-N, left-G1397-N + right-G1397-C, and left-G1397-C + right-G1397-N are 0.0052, 0.0099, 0.0027, and 0.0040, respectively. The P value for ns is 0.4971. c, m. Base editing efficiency of the G12918A point mutation in mouse blastocysts. Sequencing data were obtained from blastocysts cultured after microinjection of mRNA encoding left-G1397-C and right-G1397-N-DdCBE into one-cell embryos. d, Mouse (F0) carrying the ND5 point mutation. F0 offspring carrying the ND5 point mutation developed after microinjection of DdCBE mRNA. Alignment of the mutant sequences identified in newborns. Edited bases are shown in red, and the right side shows the editing frequency of the mutant mitochondrial genes. [Figure 37]Mouse mitochondrial ND5 nonsense mutations generated by cytosine deaminase-mediated base editing. a) DdCBE target sequences for generating the m.C12336T nonsense mutation and the m.G12341A silent mutation. The m.C12336T (C9) mutation generates a Q199stop mutation in the ND5 protein, whereas the m.G12341A (C14) mutation causes a silent Q200Q mutation. The transcription triplet is underlined, and potential editing positions are highlighted in red. b) Efficiency of cytosine-thymine base editing to generate nonsense mutations in NIH3T3 cells. The transfected DdCBE pair combinations are shown. Dark and light gray bars indicate the frequency of the m.C12336T (C9) and m.G12341A (C14) mutations, respectively. Error bars indicate s.e.m. n = 3 biologically independent samples (ns not significant, *p<0.05, **p<0.01 using Student's two-tailed t test). P values for the C9 mutations of Left-G1333-N + Right-G1333-C, Left-G1333-C + Right-G1333-N, Left-G1397-N + Right-G1397-C, and Left-G1397-C + Right-G1397-N are 0.0065, 0.1143, 0.0266, and 0.0037, respectively, and for the C14 mutations are 0.0077, 0.0144, 0.0406, and 0.0214, respectively. c, Editing efficiency in mouse blastocysts. Sequence analysis data were obtained from blastocysts developed after microinjection of mRNA encoding the conjugates left-G1333-N and right-G1333-C-DdCBE. The dark and light gray bars indicate the frequencies of the C9 and C14 mutations, respectively. d, Mutation sequence alignment of newborns. Edited bases are shown in red, and the right side shows the editing frequency of the mutant mitochondrial genome. e, Sanger sequence analysis chromatograms of wild-type and edited mice. The red arrow indicates the substituted nucleotide. [Figure 38]This is a schematic diagram of the Golden Gate cloning process for generating DdCBE constructs. All reactions occur simultaneously in one tube. Arrows do not indicate sequential reaction processes. Using BsaI enzyme, the empty expression vector and module vector were cleaved to contain compatible cohesive ends and lack the linearized backbone and TALE module inserts. T4 DNA ligase was then used to combine the backbone and six module inserts to generate the final DdCBE construct. Eight DdCBE replicative backbone plasmids were used: Left-G1333-N, Left-G1333-C, Left-G1397-N, and Left-G1397-C for SOD2 MTS; and Right-G1333-N, Right-G1333-C, Right-G1397-N, and Right-G1397-C for COX8A MTS. [Figure 39] ND5 mutant mice (F0). (a) ND5 silent mutant mice, (b) ND5 G12918A mutant mice, (c) ND5 nonsense mutant mice generated after microinjection of DdCBE mRNA. [Figure 40] (a) Schematic diagram of vector containing DdCBE-NES and NES sequence. (b) Mouse m.G12918 ND5 gene and ND5-like sequence on nuclear chromosome 4. Mitochondrial TrnA and nuclear chromosome 5 sequence. Mitochondrial Rnr2 and nuclear chromosome 6 sequence. (c) Editing efficiency of DdCBE and DdCBE-NES in the ND5 gene using an NIH3T3 cell line. (d) Editing efficiency of DdCBE and DdCBE-NES in the TrnA gene using an NIH3T3 cell line. (e) Editing efficiency of DdCBE and DdCBE-NES in the Rnr2 gene using an NIH3T3 cell line. Orange and gray graphs show the editing efficiency of DdCBE and DdCBE-NES. (f) DNA recognition sequence of the mitoTALEN TALE sequence. (g) DdCBE base editing efficiency in experimental groups treated with or without mitoTALEN. All graphs are of n=2 and error bars are standard error of the mean. [Figure 41]Improved editing efficiency in mouse embryos and individuals using DdCBE-NES and mitoTALENs. (a) Base editing efficiency of multiple mitochondrial DNA targets (mtND5, mtTrnA, mtRNR2) in blastocysts using DdCBE-NES. (b) Comparison of m.G12918A base editing efficiency using DdCBE, DdCBE-NES, and mitoTALENs. (c) Comparison of m.G12918A base editing efficiency in individual mice. All graphs are n>=3, and error bars represent the standard error of the mean. (ns=statistically not significant; *p<0.05, **p<0.01, ***p<0.001 obtained using Student's two-tailed t-test.) [Figure 42] (a) Schematic diagram of DdCBE protein modification. (b, c) Crystal structure of DddAtox deaminase. Residues in the binding surface are indicated by bars. (b) G1397-N and G1397-C splits are shown in purple and light blue, respectively. (b) G1333-N and G1333-C splits are shown in orange and green, respectively. (d, e) Binding surface mutations G1397-N and G1397-C (d) and G1333-N and G1333-C (e) are shown in red. [Figure 43] (a) Graph of base editing efficiency of G1397-binding surface mutations. Editing range and target cytosine are indicated at the top. Mutations and wild-type / TALE-free DddAtox proteins were co-transfected as indicated on the side. For Left-DdCBE, the TALE-free DddAtox protein is G1397-N, and for Right-DdCBE, the TALE-free DddAtox protein is G1397-C. (b) Heatmap showing the target cytosine-thymine (guanine-adenine) base editing efficiency of DdCBE and mutations. [Figure 44](a) Graph of base editing efficiency of G1333-binding surface mutations. Editing range and target cytosine are indicated at the top. Mutations and wild-type / TALE-free DddAtox proteins were co-transfected as indicated on the side. For Left-DdCBE, the TALE-free DddAtox protein is G1333-N, and for Right-DdCBE, the TALE-free DddAtox protein is G1333-C. (b) Heatmap showing the target cytosine-thymine (guanine-adenine) base editing efficiency of DdCBE and mutations. [Figure 45] FIG. 1 shows the results of comparison of the amino acid sequences of wild-type and novel full-length DddA. [Figure 46] FIG. 1 shows the form in which full-length DddA is transferred to animal or plant cells. [Figure 47] FIG. 1 shows the results of confirming the activity of substituting cytosine in the TC motif with thymine in the human cell genome contexts ROR1 site (a), HEK3 site (b), and TYRO3 site (c). [Figure 48] FIG. 1 shows the advantages of full-length DddA. [Figure 49] FIG. 1 shows the results of measuring the activity of full-length DddA in the human cell genome contexts TRAC site 1 (a), TRAC site 2 (b), FANCF (c), and HBB (d). [Figure 50] Figure 1 shows the results of measuring DddA activity using DddA-dCas9(D10A, H840A)-UGI in the human cell genome contexts TYRO3 (a), ROR1 (b), HEK3 (c), EMX1 site 2 (d), TRAC site 1 (e), and HBB (f). [Figure 51]Base editing efficiency of full-length DddAtox in HEK293T cells. (a) Schematic diagram of screening full-length DddAtox using a structure-based approach. Red alanine indicates a positively charged amino acid residue replaced with alanine. (b) E. coli transformants of DddA mutants replaced with alanine are shown. E1347A was used as a control active site mutant. (c) Editing and indel frequencies of DddA AAAAA and CBE at the TYRO3 site. (d) Allele frequencies at the TYRO3 site, showing C to T substitutions in red. The protospacer is shown in blue, and the protospacer adjacent motif (PAM) is shown in orange. [Figure 52] Atoxic DddA GSVG. (a) Schematic diagram of error-prone PCR-based screening for atoxic full-length DddAtox mutants. (b) Editing frequencies of alleles (c) fused to the N- and C-termini of Cas9, nCas9(D10A), nCas9(H840A), and dCas9(D10A,H840A). The protospacer is shown in blue, and the protospacer adjacent motif (PAM) is shown in orange. [Figure 53] The frequency of editing of DddAtox mutants with alanine substitutions for positively charged amino acid residues at the TYRO3 site (a), ROR1 site 1 (b), and HEK3 site (c) into the N-terminus of nCas9(D10A). The protospacer is shown in blue, and the protospacer adjacent motif (PAM) is shown in orange. [Figure 54] ROR1 site 1 (a), ROR1 site 2 (b), ROR1 site 3 (c), FANCF site (d), HBB site (e), HEK3 site (f), TRAC5 site 1 (g), and EMX1 site (h, l, j), respectively. The protospacer is shown in blue, and the protospacer adjacent motif (PAM) is shown in orange. C to T substitutions are shown in red. The target window for DddA is counted 5' upstream of the protospacer and is indicated by negative numbers. [Figure 55]Editing frequencies of AAAAA and E1347A in HeLa cells at TYRO3 site (a), ROR1 site 1 (b), ROR1 site 2 (c), ROR1 site 3 (d), FANCF site (e), HB site (f), HEK3 site (g), TRAC5 site 1 (h), TRAC5 site 2 (i), and EMX1 site 2 (j). The protospacer is shown in blue, and the protospacer adjacent motif (PAM) is shown in orange. The target window for DddA is counted 5' upstream of the protospacer and is shown as a negative number. The target cytosine is shown in red. [Figure 56] Figure 1 shows the time-dependent base editing and indel ratios of AAAAA and E1347A at the TYRO3 site (a) and ROR1 site 1 (b). [Figure 57] Editing, indels, and allele frequencies of GSVG fused to the N-terminus of nCas9(D10A), nCas9(H840A), and dCas9 at EMX1 site 2 (a), FANCF site (b), TRAC5 site 1 (c), TRAC5 site 2 (d), ROR1 site 1 (e), ROR1 site 2 (f), ROR1 site 3 (g), and HBB site (h). The protospacer is shown in blue, and the protospacer adjacent motif (PAM) is shown in orange. C to T substitutions are shown in red. The target window of GSVG was counted 5' upstream of the protospacer and is shown as a negative number. [Figure 58] Editing, indels, and allele frequencies of GSVG fused to the C-terminus of nCas9(D10A), nCas9(H840A), and dCas9 at EMX1 site 2 (a), EMX1 site 4 (b), ROR1 site 2 (c), and HBB site (d). The protospacer is shown in blue, and the protospacer adjacent motif (PAM) is shown in orange. G to A substitutions are shown in red. The target window of GSVG is counted 3' downstream from the first position of the protospacer. [Figure 59]Time-dependent editing and indel frequencies of E1347A, GSVG, SSVG, GSAG, and GSVS fused to the C-terminus of nCas9(H840A) in TYRO3 (a) and EMX1 site 2 (b). [Figure 60] Mitochondrial base editing with mDdCBE in HEK293T cells. Editing efficiency in ND4 (a) and ND6 (b). Target cytosines and TALE binding sites are shown in red and gray, respectively. Editing efficiency in ND4 (c, d) and ND6 (e, f) when only half of the DddAtox is fused to the TALE array and the other half is TALE-less. The TALE sequences on the left and right are indicated by L and R, respectively. ND6 TALE array mismatches with the reference genome are underlined in purple. [Figure 61] (a) Schematic diagram of zinc finger cytosine (ZFD) using a conventional ZFP DNA-binding protein. (b) The position where adenine deaminase is inserted into the ZFD (the red arrow indicates the insertion position). (c) Base editing efficiency (C-to-T) of the engineered ZFDdABE at the nuclear DNA Trac site. (d) Base editing efficiency (A-to-G) of the engineered ZFDdABE at the nuclear DNA Trac site. (WT-ZFD is a C-to-T deaminase with only split DddAtox, without adenine deaminase.) (e) Efficiency (C-to-T) of the ZF-DdABE at the ND1 site targeting mitochondrial DNA. (f) Efficiency (A-to-G) of the ZF-DdABE at the ND1 site targeting mitochondrial DNA. [Figure 62](a) Schematic diagram of DdABE using TALE and split DddAtox (components include split DddAtox, adenine deaminase, and TALE array). (b) Base editing efficiency when only adenine deaminase is attached to a TALE targeting the mitochondrial ND4 site. (c) Base editing efficiency when adenine deaminase is attached to a TALE-split DddAtox targeting the mitochondrial ND1 site. (d) Base editing efficiency measured at a single nucleotide level when a DdCBE pair is attached on the left and adenine deaminase is attached to a TALE-split DddAtox on the right (green boxes indicate the sites where TALEs are attached). (e) Base editing efficiency measured at a single nucleotide level when adenine deaminase is attached to a TALE-split DddAtox on the left and a DdCBE pair is used on the right (green boxes indicate the sites where TALEs are attached). [Figure 63] (a) C-to-T and A-to-G base editing efficiencies of DdABE targeting mitochondrial ND1 site with and without UGI (red box indicates adenine deaminase). (b) C-to-T and A-to-G base editing efficiencies of DdABE targeting mitochondrial ND4 site with and without UGI (red box indicates adenine deaminase). (c) The most efficient DdABE construct targeting mitochondrial ND1 site, viewed at single nucleotide level (green box indicates TALE attachment site). (d) The most efficient DdABE construct targeting mitochondrial ND4 site, viewed at single nucleotide level (green box indicates TALE attachment site). [Figure 64](a) Top: Schematic diagram of a single TALE module with all constructs in one TALE module (components include full-length DddAtox, adenine deaminase, and a TALE array). Bottom: Dual TALE module using two TALE modules (components include full-length DddAtox and a TALE array on one side, and adenine deaminase and a TALE array on the other). (b) Base editing efficiency of single-module and dual-module DdABE targeting the mitochondrial ND1 site. (c) Base editing efficiency of single-module and dual-module DdABE targeting the mitochondrial ND4 site. [Figure 65] This figure shows the results of confirming the base editing efficiency of a single module targeting the ND1 site (components include a TALE array, adenine deaminase (AD), and full-length DddAtox (GSVG, AAAA, and E1347A are variants)). [Figure 66] This figure shows the results of confirming the base editing efficiency of a dual module targeting the ND1 site (components include a TALE array, adenine deaminase (AD), and full-length DddAtox (GSVG, AAAA, and E1347A are variants)). [Figure 67] (a) Base editing efficiency when only TadA (AD) adenine deaminase is attached to a TALE-binding protein targeting the ND1 site. (b) Base editing efficiency when only TadA (AD) adenine deaminase is attached to a TALE-binding protein targeting the ND4 site. [Figure 68]This figure shows the adenine and cytosine base editing efficiencies for dual modules, single modules, and split-DddA-AD in TALEs targeting the ND1 site (from bottom to top). When UGIs are present on both sides, only cytosine base editing occurs. When AD is attached to one side, both cytosine and adenine base editing occurs, and when UGI is absent, only adenine base editing occurs selectively. Similarly, it can be seen that only adenine base editing occurs selectively in dual modules and single modules. BEST MODE FOR CARRYING OUT THE INVENTION
[0044] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one skilled in the art to which this invention belongs. Generally, the nomenclature used herein is that which is well known and commonly used in the art.
[0045] As used herein, "editing" can be used interchangeably with "editing" and refers to a method of altering a nucleic acid sequence by selective deletion of a specific genomic target, including but not limited to a chromosomal region, a gene, a promoter, an open reading frame, or any nucleic acid sequence.
[0046] As used herein, "single base" refers to one and only one nucleotide in a nucleic acid sequence.When used in the context of single base editing, it means that the base at a specific position in a nucleic acid sequence is replaced with a different base.This replacement can occur through many mechanisms, including, without limitation, substitution or transformation.
[0047] As used herein, "target" or "target site" refers to a pre-defined nucleic acid sequence of any composition and / or length. Such target sites include, but are not limited to, chromosomal regions, genes, promoters, open reading frames, or any nucleic acid sequence.
[0048] As used herein, "on-target" refers to a subsequence of a specific genomic target that can be perfectly complementary to a programmable DNA-binding region and / or a single guide RNA sequence.
[0049] As used herein, "off-target" refers to a subsequence of a specific genomic target that may be partially complementary to the programmable DNA-binding region and / or single guide RNA sequence.
[0050] 1. Split cytosine deaminase The fusion protein according to the present invention is a fusion protein comprising cytosine deaminase or a mutant thereof, wherein the cytosine deaminase or mutant thereof comprises a first segment and a second segment derived from the cytosine deaminase or mutant thereof, and the first segment and the second segment are each bound to the DNA-binding protein.
[0051] The cytosine deaminase is an amino group-releasing enzyme that can convert cytosine (C) to uridine (U).
[0052] The cytosine deaminase may be a cytosine deaminase. Cytosine deamins include, for example, APOBEC1 (apolipoprotein B editing complex 1) and AID (activation-induced deaminase). However, most DNA deaminases act only on single-stranded DNA and may not be suitable for base editing when linked to a DNA-binding protein. Specifically, the cytosine deaminase may be derived from a deaminase (DddA) that acts on double-stranded DNA or an orthologue thereof. More specifically, the cytosine deaminase may be a double-stranded DNA-specific bacterial cytosine deaminase.
[0053] The cytosine deaminase is in a split form, and the cytosine deaminase comprises an isolated first split body and a second split body, and each of the first split body and the second split body does not have deaminase activity.
[0054] The full-length cytosine deaminase may include the sequence of SEQ ID NO: 1, which corresponds to the tox fragment. The cytosine deaminase includes a first fragment and a second fragment, and neither the first fragment nor the second fragment has deaminase activity.
[0055] [SEQ ID NO: 1] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC
[0056] In one embodiment, the first or second split body of cytosine deaminase may comprise, at the N-terminus, one or more sequences selected from the group consisting of G33, G44, A54, N68, G82, N98, and G108 in the sequence of SEQ ID NO: 1. The first or second split body of cytosine deaminase may comprise, at the C-terminus, one or more sequences selected from the group consisting of G34, P45, G55, N69, T83, A99, and A109 in the sequence of SEQ ID NO: 1.
[0057] Specifically, the cytosine deaminase may include a first split product (G1333-N) of SEQ ID NO: 23 and a second split product (G1333-C) of SEQ ID NO: 24, a first split product (G1397-N) of SEQ ID NO: 25 and a second split product (G1397-C) of SEQ ID NO: 26, a first split product (G1333-N) of SEQ ID NO: 23 and a second split product (G1397-C) of SEQ ID NO: 26, or a first split product (G1397-N) of SEQ ID NO: 25 and a second split product (G1333-C) of SEQ ID NO: 24.
[0058] (SEQ ID NO: 23) Wild type DddAtox G1333-N GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGG (SEQ ID NO: 27) GGCTCTGGTTCCTACGCCCTGGGTCCATATCAGATTAGTGCTCCCAACTCCCCGCCTACAACGGTCAGACAGTGGGGACCTTTTACTATGTCAACGACGCCGGGGGATTGGAATCCAAGGTTTTCTCTAGCGGTGGG
[0059] (SEQ ID NO: 24) Wild type G1333-C PTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC (SEQ ID NO: 28) CCAACACCTTATCCTAACTACGCTAACGCCGGGCACGTCGAGGGGCAGTCAGCTCTTTTTATGAGAGATAACGGCATTAGCGAAGGGCTTGTGTTCCATAATAATCCTGAGGGCACCTGTGGCTTCTGTGTAAATATGACC GAAACACTTCTGCTGAGAACGCTAAAATGACTGTCGTACCACCCGAAGGCGCAATCCCAGTTAAACGGGGCGCAACCGGCGAAACCAAAGTATTCACCGGAAACAGCAATAGTCCAAAGTCCCCCACCAAGGGAGGTTGC
[0060] (SEQ ID NO: 25) Wild type DddAtox G1397-N GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEG
[0061] (SEQ ID NO: 29) GGTAGCTACGCACTTGGTCCTTACCAGATTAGCGCACCCCAACTCCCCGCCTATAATGGTCAAACCGTCGGGACCTTTTACTACGTAAACGATGCTGGTGGGCTGGAATCCAAAGTATTCTCCTCAGGGGGCCCTACACCCTACCCCAACTACGCCAATGCT GGTCATGTAGAAGGGCAGTCAGCACTGTTTATGCGCGATAATGGTATAAGCGAGGGGTTGGTCTTCCATAACAACCCAGAGGGTACTTGTGGCTTCTGTGTGAATATGACTGAAACCCTTCTGCCCGAAAATGCCAAGATGACTGTCGTCCCACCTGAAGGC
[0062] (SEQ ID NO: 26) Wild type DddAtox G1397-C AIPVKRGATGETKVFTGNSNSPKSPTKGGC
[0063] (SEQ ID NO: 30) GCCATACCTGTGAAGCGGGGAGCAACAGGGGAGACAAAGGTGTTCACAGGCAACTCTAACAGTCCAAAGAGCCCCACCAAAGGCGGGTGT G1333N, G1333C, G1397N, and G1397C can be combined to form split deaminase enzymes, such as Left-G1333-N+Right-G133-C, Left-G1397-N+Right-G1397-C, Left-G1397-N+Right-G1333-C, or Left-G1333-N+Right-G1397-C.
[0064] 2. Mutants The inventors of this application attempted to suppress the undesired base editing caused by the DdCBE mutation through the substitution of amino acid residues. We present a highly accurate DddA-derived cytosine base editor that can reduce the non-targeted effects of DdCBE. Such non-targeted base editing effects are independent of the interaction between TALE and DNA, and are mediated by DddA. tox This phenomenon is caused by the spontaneous binding of deaminase fragments.
[0065] The amino acid residues to be mutated are located at the contact sites on the surface where the DddAtox dimers bind to each other. tox We created HF-DdCBE by substituting alanine for amino acid residues located on the surface between the cleavage sites, so that HF-DdCBE cannot function properly unless the two deaminase pairs linked to the TALE can bind to DNA. Whole-middle mitochondrial genome analysis confirmed that HF-DdCBE is highly efficient and precise, unlike conventional DdCBEs, which cause numerous undesired, untargeted C-to-T transversions in human mitochondrial DNA.
[0066] In the case of DddAtox, theoretically, base editing can only occur if all split dimers are recruited to the target site in DNA, but actual experiments have shown that targeted base editing can also occur when one DdCBE binds to DNA while the other does not. To solve this, the residues on the protein surface where the split dimers bind to each other are replaced to prevent the DdCBE pair from binding to undesired positions.
[0067] Based on this, the present invention relates to novel mutants of the cytosine deaminase enzyme DddAtox that substitute amino acid residues to reduce non-selective base editing.
[0068] The cytosine deaminase comprises a first split body (G1333-N) of SEQ ID NO: 23 and a second split body (G1333-C) of SEQ ID NO: 24, or a first split body (G1397-N) of SEQ ID NO: 25 and a second split body (G1397-C) of SEQ ID NO: 26.
[0069] The cytosine deaminase mutant may be a first fragment of SEQ ID NO: 23 in which one or more amino acids selected from the group consisting of positions 3, 5, 10, 11, 13, 14, 15, 16, 17, 18, 19, 28, 30, and 31 have been substituted with other amino acids, or a second fragment of SEQ ID NO: 24 in which one or more amino acids selected from the group consisting of positions 13, 16, 17, 20, 21, 28, 29, 30, 31, 32, 33, 56, 57, 58, and 60 have been substituted with other amino acids.
[0070] The mutant according to the present invention is a first fragment of SEQ ID NO: 25 in which one or more amino acids selected from the group consisting of positions 87, 88, 91, 92, 95, 10.0, 101, 102 and 103, or a second fragment of SEQ ID NO: 26 in which one or more amino acids selected from the group consisting of positions 13, 14, 15 and 16, have been substituted with other amino acids.
[0071] The "other amino acid" refers to an amino acid selected from alanine, isoleucine, leucine, methionine, phenylalanine, proline, tryptophan, valine, aspartic acid, cysteine, glutamine, glycine, serine, threonine, tyrosine, glutamic acid, arginine, histidine, lysine, and all known variants of the amino acids, excluding the amino acid present in the wild-type protein at the original variant position. In one example, the "other amino acid" may be alanine.
[0072] Specifically, among the first fragments of SEQ ID NO: 23, Y3A, L5A, I10A, S11A, V 23 A, G 24 A, T 25 A, F 26 A, Y 27 A, Y 28A, V 29 A, K 38 A, F 40 A, S 41 A (corresponding to Y1292A, L1294A, I1299A, S1300A, V1312A, G1313A, T1314A, F1315A, Y1316A, Y1317A, V1318A, K1327A, F1329A, and S1330A, respectively).
[0073] The second fragment of SEQ ID NO: 24 may contain one or more amino acid substitutions selected from the group consisting of V13A, Q16A, S17A, F20A, M21A, E28A, G29A, L30A, V31A, F32A, H33A, K56A, M57A, T58A, and V60A (corresponding to V1346A, Q1349A, S1350A, F1353A, M1354A, E1361A, G1362A, L1363A, V1364A, F1365A, H1366A, K1389A, M1390A, T1391A, and V1393A, respectively).
[0074] Specifically, the first fragment of SEQ ID NO: 25 may contain one or more amino acid substitutions selected from the group consisting of C87A, V88A, T91A, E92A, L95A, K100A, M101A, T102A, and V103A (corresponding to C1376A, V1377A, T1380A, E1381A, L1384A, K1389A, M1390A, T1391A, and V1392A, respectively).
[0075] The second fragment of SEQ ID NO: 26 may contain one or more amino acid substitutions selected from the group consisting of K13A, V14A, F15A, and T16A (corresponding to K1410A, V1411A, F1412A, and T1413A, respectively).
[0076] JPEG0007796729000001.jpg230170JPEG0007796729000002.jpg86170
[0077] JPEG0007796729000003.jpg234170JPEG0007796729000004.jpg251170JPEG0007796729000005.jpg96170
[0078] JPEG0007796729000006.jpg234170JPEG0007796729000007.jpg166170
[0079] JPEG0007796729000008.jpg78170
[0080] The cytosine deaminase mutant according to the present invention may comprise one or more sequences selected from the group consisting of the amino acid sequences shown in the table above. The cytosine deaminase mutant according to the present invention suggests the possibility of reducing the effect of causing undesired editing of multiple bases in a non-specific target site.
[0081] 3. Full-length cytosine deaminase The inventors of the present application have developed a method for the production of wild-type cytosine deaminase DddA, which causes toxicity in cells and is used in a split form. tox We have developed a new programmable cytosine deaminase using full-length DddA, which was engineered by modifying the positively charged amino acids.
[0082] The present invention relates to a fusion protein comprising: (i) a DNA-binding protein; and (ii) a cytosine deaminase or a mutant thereof, wherein the cytosine deaminase or mutant thereof is a non-toxic full-length cytosine deaminase.
[0083] DddA tox The C-terminus of DddA is specifically clustered with positively charged amino acids (KRKKK). Because DNA has a negative charge, it binds to the positively charged amino acids of proteins. By replacing the positively charged amino acids with amino acids that do not have an electrode, DddA toxThis weakens the ability of DddA to bind to DNA, reducing or eliminating its intracellular toxicity. In other words, replacing the positively charged amino acid with a non-polar amino acid to eliminate toxicity allows for cloning in E. coli and the full-length DddA to be isolated.
[0084] Wild-type DddA tox Due to intracellular toxicity, two separate forms are used, but this poses many experimental limitations. In particular, when using Cas9, an orthogonal Cas9 mutant that recognizes another PAM is used. In this case, the presence of a restrictive PAM often makes it difficult to precisely deaminate cytosine to thymine at the desired site. Furthermore, the target window in which the highest activity is observed is the 40-bp interval between the binding sites of the two Cas9 mutants, and deamination of the undesired cytosine at this site is possible. However, the full-length DddA is not separated and is therefore not restricted by the PAM. Furthermore, the highest activity is observed with TC motifs within 10 bp of the target site, resulting in high accuracy.
[0085] We confirmed that Cas9 deaminates cytosines in ACA, GC, and CC motifs in the R-loop formed by binding to the target site to thymine, an activity not observed in the isolated form.
[0086] Full-length DddA can be used with Cas9, TALE modules, and zinc finger proteins to convert cytosine to thymine at the desired site. While conventional DddAtox is in a separate form and must be delivered as a pair, full-length DddA can be used with only one module (TALE module or zinc finger protein), allowing for unrestricted targeting. Furthermore, it can target not only genomic sites but also DNA in mitochondria, plant chloroplasts, or plastids, converting cytosine to thymine in specific DNA.
[0087] Furthermore, its small size allows the entire construct to be inserted into AAV, a viral vector used in gene therapy. While conventional cytosine base editors (CBEs) replace cytosines in the R-loop formed by Cas9 binding to the target site with thymine, the newly invented full-length DddA deaminates cytosines outside the R-loop. Therefore, it can convert cytosines to thymines at positions where editing is restricted using conventional CBEs.
[0088] Based on this, the non-toxic full-length cytosine deaminase may have one or more, two or more, three or more, four or more, or five or more amino acids of the wild-type deaminase of SEQ ID NO: 1 substituted with other amino acids.
[0089] The "other amino acid" refers to an amino acid selected from alanine, isoleucine, leucine, methionine, phenylalanine, proline, tryptophan, valine, aspartic acid, cysteine, glutamine, glycine, serine, threonine, tyrosine, glutamic acid, arginine, histidine, lysine, and all known variants of the amino acids, excluding the amino acid that the wild-type protein has at the original variant position. In one example, the "other amino acid" can be alanine.
[0090] The non-toxic full-length DddA may comprise a sequence selected from the group consisting of SEQ ID NO: 12 to SEQ ID NO: 18, depending on the type.
[0091] Wild type (SEQ ID No: 1) GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC
[0092] A1341D KRKKA (SEQ ID No: 12) GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYDNAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTAGGC
[0093] AAAAA (SEQ ID No: 13) GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVAAGATGETAVFTGNSNSPASPTAGGC
[0094] AAAAK (SEQ ID No: 14) GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVAAGATGETAVFTGNSNSPASPTKGGC
[0095] AAKAA (SEQ ID No: 15) GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVAAGATGETKVFTGNSNSPASPTAGGC
[0096] AAKAK (SEQ ID No: 16) GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVAAGATGETKVFTGNSNSPASPTKGGC
[0097] KAAAA (SEQ ID No: 17) GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKAGATGETAVFTGNSNSPASPTAGGC
[0098] E1347A (SEQ ID No: 18) GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVAGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC
[0099] The full-length deaminase mutant may comprise one or more substitutions in the amino acid sequence of SEQ ID NO: 1 selected from the group consisting of: S at position 37 is replaced with G, G at position 59 is replaced with S, A at position 109 is replaced by V, and The S at position 129 is replaced with a G.
[0100] In one embodiment, the full-length mutant deaminase may comprise the sequence of SEQ ID NO: 19, which includes, in the amino acid sequence of SEQ ID NO: 1, a substitution of S at position 37 with G, a substitution of G at position 59 with S, a substitution of A at position 109 with V, and a substitution of S at position 129 with G.
[0101] GSVG (SEQ ID No: 19) GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLEGKVFSSGGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGVIPVKRGATGETKVFTGNSNGPKSPTKGGC
[0102] The full-length DddA GSVG can be cloned using standard E. coli. It has been confirmed in the human cell genomic context that full-length DddA GSVG deaminates the cytosine of the TC motif at the target site to thymine. Full-length DddA GSVG can be cloned into the N-terminus and C-terminus of Cas9. DddA GSVG ligated to the N-terminus of Cas9 at the same target site can replace cytosine with thymine. It has been confirmed in the human cell genomic context that DddA GSVG ligated to the C-terminus of Cas9 replaces the cytosine of the TC motif at the complementary sequence of the desired guanine with thymine, replaces the cytosine of the TC motif at the complementary sequence of the desired guanine with thymine, and replaces the desired guanine with adenine.
[0103] In one embodiment, the full-length deaminase mutant may comprise a sequence selected from the group consisting of SEQ ID NOs: 20-22.
[0104] SSVG (SEQ ID No: 20) GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGVIPVKRGATGETKVFTGNSNGPKSPTKGGC
[0105] GSAG (SEQ ID No: 21) GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLEGKVFSSGGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNGPKSPTKGGC
[0106] GSVS (SEQ ID No: 22) GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLEGKVFSSGGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGVIPVKRGATGETKVFTGNSNSPKSPTKGGC
[0107] 4. DNA-binding proteins The DNA binding protein may be, for example, a zinc finger protein, a TALE (Transcription activator-like effector) protein, a CRISPR-associated nuclease, or a combination of two or more thereof.
[0108] The zinc finger motif of the zinc finger protein has a DNA-binding domain, and the C-terminal portion of the finger specifically recognizes a DNA sequence. DNA-binding proteins containing 3 to 6 zinc finger motifs recognize a DNA sequence.
[0109] In one embodiment, each of the first split body of cytosine deaminase and the second split body of cytosine deaminase can be bound to the N-terminus or C-terminus of a zinc finger protein.
[0110] The C-terminus of a zinc finger protein (ZF-Left) binds to the N-terminus of the first cytosine deaminase fragment, and the C-terminus of a zinc finger protein (ZF-Right) binds to the N-terminus of the second cytosine deaminase fragment (CC configuration). The N-terminus of a zinc finger protein (ZF-Left) binds to the C-terminus of the first cytosine deaminase fragment, and the C-terminus of a zinc finger protein (ZF-Right) binds to the N-terminus of the second cytosine deaminase fragment (NC configuration). The C-terminus of a zinc finger protein (ZF-Left) is bound to the N-terminus of the first cytosine deaminase fragment, and the N-terminus of a zinc finger protein (ZF-Right) is bound to the C-terminus of the second cytosine deaminase fragment (CN configuration), or The N-terminus of a zinc finger protein (ZF-Left) can bind to the C-terminus of the first cytosine deaminase fragment, and the N-terminus of a zinc finger protein (ZF-Right) can bind to the C-terminus of the second cytosine deaminase fragment (NN configuration).
[0111] The ZF-Left may comprise the sequence of SEQ ID NO:2 below.
[0112] [SEQ ID NO: 2] GIHGVPAAMAERPFQCRICMRNFSDRSNLSRHIRTHTGEKPFACDICGRKFAISSNLNSHTKIHTGSQKPFQCRICMRNFSRSDNLARHIRTHTGEKPFACDICGRKFATSGNLTRHTKIHLR.
[0113] The ZF-Right may comprise the sequence of SEQ ID NO:3 below.
[0114] [SEQ ID NO: 3] GIHGVPAAMAERPFQCRICMRNFSRSDNLSVHIRTHTGEKPFACDICGRKFAQKINLQVHTKIHTGEKPFQCRICMRNFSRSDVLSEHIRTHTGEKPFACDICGRKFAQRNHRTTHTKIHLR.
[0115] The ZF sequence may vary depending on the DNA target. ZFs can be customized according to the DNA target sequence. Because ZFs recognize 3-bp DNA, typically 3-6 ZFs can be combined to create a ZF combination that recognizes 9-18 bp DNA. For example, ZFs can be created using a library containing modules such as GNN, TNN, CNN, or ANN.
[0116] In some cases, the zinc finger protein may be linked to the deaminase via a linker. The linker may be a peptide linker containing 2 to 40 amino acid residues. The linker may be, for example, 2 aa, 5 aa, 10 aa, 16 aa, 24 aa, or 32 aa in length, but is not limited thereto.
[0117] In one embodiment, the linker may comprise: 2a.a linker: GS, 5a.a linker: TGEKQ (SEQ ID NO: 8), 10a.a linker: SGAQGSTLDF (SEQ ID NO: 9), 16a.a linker: SGSETPGTSESATPES (SEQ ID NO: 10), 24a.a linker: SGTPHEVGVYTLSGTPHEVGVYTL (SEQ ID NO: 115), or 32a.a linker: GSGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID NO: 11)
[0118] In a specific embodiment of the present invention, the split deaminase and zinc finger protein can be linked via a linker, such that the zinc finger protein binds to the N-terminus of the first split-body half deaminase and the zinc finger protein binds to the N-terminus of the second split-body half deaminase. A C-to-T transversion can occur in the spacer between the attachment positions of the left and right ZFPs. It was confirmed that high editing efficiency was observed when both the left and right ZFPs were linked to the first split-body half deaminase and the second split-body half deaminase via a 24a linker.
[0119] The TAL effector (TALE) consists of a repeat of 33 to 34 amino acid sequences, with approximately nine repeated domains (RVD, Repeated Variant Domain). Each domain can recognize one nucleotide and bind to a specific DNA sequence according to the 12th to 13th amino acid sequence (HD → Cytosine, NI → Adenine, NG → Thymine, NN → Guanine). The TAL effector (TALE) recognizes a single DNA strand within a target site. The distance between target sites can be 12 to 14 nucleotides.
[0120] The TALE domain refers to a protein domain that binds to nucleotides in a sequence-specific manner through the combination of one or more TALE-Repeat. It includes, but is not limited to, at least one TALE-Repeat, specifically 1 to 30 TALE-Repeat. A TALE-Repeat is a site that recognizes a specific nucleotide sequence within a TALE domain.
[0121] The TALE domain comprises an N-terminal encompassing region of a TALE as a backbone structure and a C-terminal encompassing region of a TALE. The first TALE encompassing the N-terminal of the TALE can be encoded by SEQ ID NO: 4 or 5. The second TALE encompassing the C-terminal of the TALE can be encoded by SEQ ID NO: 6 or 7.
[0122] [Table 1]
[0123] Based on the cleavage site, depending on the position where the TALE domain binds, a single TALE array or a first TALE array and a second TALE array can each be bound.
[0124] A first TALE (left TALE) can be bound to the first cleavage body of the cytosine deaminase, and a second TALE (right TALE) can be bound to the second cleavage body of the cytosine deaminase, with the structures N'-TALE-first cleavage body-C' and N'-TALE-second cleavage body-C', respectively.
[0125] When the cytosine deaminase is full-length, a single module TALE can be bound to the N-terminus of the cytosine deaminase. A single TALE module and cytosine deaminase are included in the NC direction. A dual module TALE may be included, in which a first TALE is bound to the N-terminus of the full-length cytosine deaminase and a second TALE is included separately. The first TALE module and cytosine deaminase are included in the NC direction, and have the structures N'-TALE-cytosine deaminase-C' and N'-TALE-C'.
[0126] TALE arrays can be customized to match target DNA sequences. They consist of a repeating array of modules consisting of 33–35 amino acid residues. These modules are derived from the plant pathogen Xanthmonas. Each module recognizes one A, one C, one G, and one T base to bind to DNA. The base specificity of each module is determined by the 12th and 13th amino acid residues, called the repeat variable di-residue (RVD). For example, a module with an RVD of NN recognizes G, NI recognizes A, HD recognizes C, and NG recognizes T. TALE arrays can be designed to consist of at least 14 modules and up to 18 or more modules, recognizing target DNA sequences of 15–20 bp.
[0127] Regarding the CRISPR-associated nuclease, two RNAs are encoded in the CRISPR array: one is crRNA (CRISPR RNA) and the other is tracRNA (trans-activating CRPSPR RNA). crRNA is transcribed in the protospacer and binds to tracRNA to form a tertiary structure. These two types of RNA help recognize and cut external DNA.
[0128] The Cas proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9, Cas10, Cas12a, Cas12b, Cas12c, Cas12d, Cas12e, Cas12g, Cas12h, Cas12i, Cas12j, Cas13a, Cas13b, Cas13c, Cas13d, Cas14, Csy1, Csy2, Csy3, Cse1, and Cs The endonuclease may be, but is not limited to, e2, Csc1, Csc2, Csa5, Csn2, CsMT2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3 or Csf4.
[0129] The Cas protein is active against Corynebacter, Sutterella, Legionella, Treponema, Filifactor, Eubacterium, Streptococcus (Streptococcus pyogenes), Lactobacillus, Mycoplasma, Bacteroides, Flaviivola, Flavobacterium, Azospirillum, Gluconacetobacter, Neisseria, Roseburia, Parvibaculum, Staphylococcus (Staphylococcus aureus), The Cas protein may be derived from a microbial genera containing an orthologue of a Cas protein selected from the group consisting of Nitratifractor, Corynebacterium, and Campylobacter, and may be simply isolated or recombinantly isolated therefrom.
[0130] It may also include mutated forms of Cas proteins, which may mean mutated to lose endonuclease activity that cleaves DNA double strands, for example, one or more of a mutant target-specific nuclease that has lost endonuclease activity but mutated to have nickase activity, and a form that has lost both endonuclease activity and nickase activity.
[0131] If the enzyme has nickase activity, a nick may be introduced in the strand where the base conversion occurred or in the reverse strand (e.g., the reverse strand of the strand where the base conversion occurred) simultaneously with or sequentially, regardless of order, the base conversion (e.g., conversion of cytosine to uridine) by the cytosine deaminase (e.g., a nick is introduced in the reverse strand of the strand where the base conversion occurred, at a position between the third and fourth nucleotides toward the 5' end of the PAM sequence in the reverse strand of the strand where the PAM is located). Such a mutation (e.g., amino acid substitution, etc.) may occur in the catalytically active domain (e.g., the RuvC catalytic domain in the case of Cas9). In the case of Streptococcus pyogenes Cas9, the mutation may include a substitution of one or more of the catalytic aspartate residues (e.g., aspartic acid at position 10 (D10)), glutamic acid at position 762 (E762), histidine at position 840 (H840), asparagine at position 854 (N854), asparagine at position 863 (N863), and aspartic acid at position 986 (D986) with any other amino acid. In this case, the substituted amino acid may be, but is not limited to, alanine.
[0132] In some cases, the Cas9 protein derived from Streptococcus pyogenes may be mutated such that one or more of the aspartic acid at position 1135 (D1135), the arginine at position 1335 (R1335), and the threonine at position 1337 (T1337) of the Cas9 protein, for example, all three of these are substituted with other amino acids to recognize NGA (where N is any base selected from A, T, G, and C) that is different from the PAM sequence (NGG) of wild-type Cas9.
[0133] For example, the amino acid sequence of the Cas9 protein derived from Streptococcus pyogenes is (1) D10, H840, or D10+H840, (2) D1135, R1335, T1337, or D1135+R1335+T1337, or (3) Amino acid substitutions can occur at both residues (1) and (2).
[0134] The "other amino acid" refers to an amino acid selected from alanine, isoleucine, leucine, methionine, phenylalanine, proline, tryptophan, valine, aspartic acid, cysteine, glutamine, glycine, serine, threonine, tyrosine, glutamic acid, arginine, histidine, lysine, and all known variants of the amino acids, excluding the amino acid that the wild-type protein has at the original variant position. In one example, the "other amino acid" may be alanine, valine, glutamine, or arginine.
[0135] In some cases, the gene may further include a guide RNA, which may be, for example, one or more of the following: CRISPR RNA (crRNA), transactivating crRNA (tracrRNA), and single guide RNA (single guide RNA; sgRNA), specifically, a double-stranded crRNA:tracrRNA complex in which crRNA and tracrRNA are bound to each other, or a single-stranded guide RNA (sgRNA) in which crRNA or a portion thereof and tracrRNA or a portion thereof are linked via an oligonucleotide linker.
[0136] 5. Addition of adenine deaminase The present invention relates to a fusion protein comprising (i) a DNA-binding protein, (ii) cytosine deaminase or a mutant thereof, and (iii) adenine deaminase, wherein the cytosine deaminase or mutant thereof comprises first and second split bodies derived from (a) a non-toxic full-length cytosine deaminase or (b) cytosine deaminase or a mutant thereof, and the first and second split bodies are each in a form capable of binding to the DNA-binding protein.
[0137] The inventors of the present application have created a base editor capable of editing the A base by linking a DddAtox cytosine deaminase and an adenine deaminase capable of A-to-G conversion to a TALE or ZFP protein capable of binding to DNA.
[0138] The conventional DddAtox-based deaminase (DdCBE) is a cytosine deaminase that uses a TALE repeat as a DNA binding module. Unlike DdCBE, which only induces C-to-T, DdABE can induce A-to-G, thus generating different mutation patterns.
[0139] DdABE recognizes double-stranded DNA by itself and causes deamination, without requiring additional components such as RNA. The Crispr system could not be applied to mitochondria or chloroplasts because the RNA delivery mechanism is unknown. However, DdABE without such components can not only target genomic DNA within cells, but also target the DNA of organelles such as mitochondria and chloroplasts, causing specific A-to-G conversion of DNA.
[0140] Currently, DdCBE is the only gene editing technology that targets mitochondria and organelles. While all previous techniques combined could only introduce C-to-T mutations, DdABE can induce A-to-G mutations, significantly broadening the spectrum of mutations that can be introduced. This will enable the creation of mitochondrial disease models or treatments that were previously impossible.
[0141] Conventional DdCBEs require two TALE modules (attached to the left and right sides, respectively), making them unsuitable for gene therapy in AAV, a viral vector with low gene capacity. However, DdABE can be used as a "single module" that can use only one TALE module, making it suitable for incorporation into AAV viruses, giving it advantages in gene therapy.
[0142] DdABE is highly compatible because it can be used with either split DddAtox or full-length DddAtox variants as needed.
[0143] The adenine deaminase may be selected from the group consisting of apolipoprotein B editing complex 1 (APOBEC1), activation-induced deaminase (AID), and tRNA-specific adenosine deaminase (tadA), and may be tRNA-specific adenosine deaminase (tadA). The adenine deaminase may be a deoxyadenine deaminase, which is a mutant of Escherichia coli TadA.
[0144] The cytosine deaminase is contained in a split form, and the DNA-binding protein is a zinc finger protein, in which the N-terminus of the zinc finger protein (ZF-Left) is bound to the C-terminus of the first split body of cytosine deaminase, and the C-terminus of the zinc finger protein (ZF-Right) is bound to the N-terminus of the second split body of cytosine deaminase (NC configuration). Adenine deaminase can be bound to the C-terminus of the zinc finger protein (ZF-Left), the N-terminus or C-terminus of the first split body of cytosine deaminase, the N-terminus of the zinc finger protein (ZF-Right), or the N-terminus or C-terminus of the second split body of cytosine deaminase.
[0145] Adenine deaminase is bound to the N-terminus of the first cytosine deaminase segment by the C-terminus of a zinc finger protein (ZF-Left) and to the N-terminus of the second cytosine deaminase segment by the C-terminus of a zinc finger protein (ZF-Right) (CC configuration); to the N-terminus of the first cytosine deaminase segment by the C-terminus of a zinc finger protein (ZF-Left) and to the C-terminus of the second cytosine deaminase segment by the N-terminus of a zinc finger protein (ZF-Right) (CN configuration); or to the C-terminus of the first cytosine deaminase segment by the C-terminus of a zinc finger protein (ZF-Left) and to the C-terminus of the second cytosine deaminase segment by the N-terminus of a zinc finger protein (ZF-Right) (NN configuration). In the same configuration, the cytosine deaminase fragment can bind to the C-terminus of the zinc finger protein (ZF-Left), the N-terminus or C-terminus of the first cytosine deaminase fragment, the N-terminus of the zinc finger protein (ZF-Right), or the N-terminus or C-terminus of the second cytosine deaminase fragment.
[0146] When the cytosine deaminase is included in a split form and the DNA-binding protein is a TALE, a first TALE binds to the first split body of the cytosine deaminase, and a second TALE binds to the second split body of the cytosine deaminase, having the structures N'-TALE-first split body DDDA-C' and N'-TALE-second split body DDDA-C', respectively. Adenine deaminase can bind to the N-terminus or C-terminus of the first split body of the cytosine deaminase or the N-terminus or C-terminus of the second split body of the cytosine deaminase.
[0147] When the cytosine deaminase is included in its full-length form and the DNA-binding protein is a TALE, it may include a single TALE module, and the single TALE module and cytosine deaminase may be included in the N-terminal direction, and in this case, adenine deaminase may be bound to the C-terminal direction of the single TALE module, or to the N-terminus or C-terminus of the cytosine deaminase.
[0148] When the cytosine deaminase is included in its full-length form and the DNA-binding protein is a TALE, it may include a dual TALE module, including a first TALE module and cytosine deaminase in the N-C direction, and may further include a second segment including adenine deaminase and a second TALE, in which the adenine deaminase can be bound to the N-terminus or C-terminus of the TALE in the structures N'-TALE-cytosine deaminase-C' and N'-TALE-adenine deaminase-C'.
[0149] In some cases, the composition may further contain UGI (uracil DNA glycosylase inhibitor), which can enhance base editing efficiency by inhibiting the activity of UDG (uracil DNA glycosylase), an enzyme that repairs mutated DNA and catalyzes the removal of U from DNA.
[0150] The present invention relates to a composition for A-to-G base editing in prokaryotic or eukaryotic cells, comprising the fusion protein or a nucleic acid encoding it, wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-associated nuclease, and the cytosine deaminase or a mutant thereof of the fusion protein is derived from bacteria and is specific for double-stranded DNA.
[0151] The present invention relates to a composition for A-to-G base editing in prokaryotic or eukaryotic cells, comprising the fusion protein or a nucleic acid encoding the same, wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-associated nuclease; The composition relates to a fusion protein in which the cytosine deaminase or a mutant thereof is derived from bacteria and is specific to double-stranded DNA, a DNA-binding protein is bound to the N-terminus of the cytosine deaminase or a mutant thereof, and a DNA-binding protein is bound to the C-terminus of the adenine deaminase of the fusion protein.
[0152] The present invention relates to a composition for C-to-T base editing in prokaryotic or eukaryotic cells, comprising the fusion protein or a nucleic acid encoding the same and UGI (uracil glycosylase inhibitor), wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-associated nuclease, and the cytosine deaminase or a mutant thereof of the fusion protein is derived from bacteria and is specific to double-stranded DNA.
[0153] Specifically, the present invention provides a composition for A-to-G base editing in prokaryotic and eukaryotic cells (without UGI), comprising: 1) a DNA-binding protein; 2) a full-length double-stranded DNA-specific bacterial cytosine deaminase or a mutant thereof; and 3) deoxyadenine deaminase derived from E. coli TadA, wherein the DNA-binding protein is a zinc finger protein, a transcription activator-like effector (TALE) array, or a catalytically deficient CRISPR-Cas9 (nCas9 or dCas9) or Cas12a, and the full-length double-stranded DNA-specific bacterial cytosine deaminase is DddAtox derived from Burkholderia cenocepacia.
[0154] The present invention provides a composition (without UGI) for A-to-G base editing in prokaryotic and eukaryotic cells, comprising: 1) a left-hand DNA-binding protein operably linked to a full-length double-stranded DNA-specific bacterial cytosine deaminase or a mutant thereof; and 2) a right-hand DNA-binding protein operably linked to deoxyadenine deaminase derived from E. coli TadA, wherein the left-hand or right-hand DNA-binding protein is a zinc finger protein, a transcription activator-like effector (TALE) array, or a catalytically-deficient CRISPR-Cas9 (nCas9 or dCas9), or Cas12a; and the full-length double-stranded DNA-specific bacterial cytosine deaminase, DddAtox, derived from Burkholderia cenocepacia.
[0155] The present invention also provides a composition for A-to-G and C-to-T base editing in prokaryotic and eukaryotic cells, comprising: 1) a DNA-binding protein; 2) a full-length double-stranded DNA-specific bacterial cytosine deaminase or a mutant thereof; 3) E. coli TadA-derived deoxyadenine deaminase; and 4) UGI (uracil glycosylase inhibitor), wherein the DNA-binding protein is a zinc finger protein, a transcription activator-like effector (TALE) array, a catalytically deficient CRISPR-Cas9 (nCas9 or dCas9), or Cas12a; and the full-length double-stranded DNA-specific bacterial cytosine deaminase is Burkholderia cenocepacia-derived DddAtox.
[0156] The present invention further provides a composition for A-to-G base editing in prokaryotic and eukaryotic cells (without UGI), comprising: 1) a DNA-binding protein; 2) a split double-stranded DNA-specific bacterial cytosine deaminase or a mutant thereof; and 3) deoxyadenine deaminase from E. coli TadA, wherein the DNA-binding protein is a zinc finger protein or a transcription activator-like effector (TALE) array or catalytically-deficient CRISPR-Cas9 (nCas9 or dCas9) or Cas12a, and the split double-stranded DNA-specific bacterial cytosine deaminase is DddAtox from Burkholderia cenocepacia.
[0157] The present invention further relates to a composition for A-to-G and C-to-T base editing in prokaryotic and eukaryotic cells, comprising: 1) a DNA-binding protein; 2) a split double-stranded DNA-specific bacterial cytosine deaminase or a mutant thereof; 3) deoxyadenine deaminase derived from E. coli TadA; and 4) UGI (uracil glycosylase inhibitor), wherein the DNA-binding protein is a zinc finger protein, a transcription activator-like effector (TALE) array, or a catalytically deficient CRISPR-Cas9 (nCas9 or dCas9) or Cas12a; and the split double-stranded DNA-specific bacterial cytosine deaminase, DddAtox, derived from Burkholderia cenocepacia.
[0158] The present invention relates to an A-to-G base editing method in a prokaryotic or eukaryotic cell, comprising a step of treating the fusion protein or a nucleic acid encoding it in the prokaryotic or eukaryotic cell, wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-associated nuclease, and the cytosine deaminase or a mutant thereof of the fusion protein is derived from bacteria and is specific to double-stranded DNA.
[0159] The present invention provides an A-to-G base editing method in a prokaryotic or eukaryotic cell, comprising the step of treating the fusion protein or a nucleic acid encoding the same in the prokaryotic or eukaryotic cell, wherein the DNA-binding protein is a zinc pinker protein, a TALE protein, or a CRISPR-associated nuclease, the cytosine deaminase or a mutant thereof of the fusion protein is derived from a bacterium and is specific to double-stranded DNA, a DNA-binding protein is bound to the N-terminus of the cytosine deaminase or mutant thereof, and a DNA-binding protein is bound to the C-terminus of the adenine deaminase of the fusion protein.
[0160] The present invention provides a C-to-T base editing method in prokaryotic or eukaryotic cells, comprising a step of treating the prokaryotic or eukaryotic cells with the fusion protein or a nucleic acid encoding it and UGI (uracil glycosylase inhibitor), wherein the DNA-binding protein is a zinc finger protein, a TALE protein, or a CRISPR-associated nuclease, and the cytosine deaminase or a mutant thereof of the fusion protein is derived from bacteria and is specific to double-stranded DNA.
[0161] Specific arrangements of components included in the compositions or methods according to the present invention are as follows: ND1-ZFP-Right-1397C-AD (Figure 61f-g: SEQ ID No: 410) MLGFVGRVAAAPASGALRRLTPSASLPPAQLLLRAAPTAVHPVRDYAAQDYKDDDDKVDEMTKKFGTLTIHDTEKAAEFGIRIPGEKPFQCRICMRNFSDSGNLRVHIRTHTGEKPYKCPDCGKSFSQSSSLIRHQRTHTGEKPYECDHCGKSFSQSSHLNVHKRTHTGEKPYRCKYCDRSFSISSNLQRHVRNIHLRSGTPHEVGVYTLSGTPHEVGVYTLAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN
[0162] ND1-ZFP-Left-1397C-UGI (Figure 61f-g: SEQ ID No: 411) MLGFVGRVAAAPASGALRRLTPSASLPPAQLLLRAAPTAVHPVRDYAAQDYKDDDDKVDEMTKKFGTLTIHDTEKAAEFGIRIPGEKPFQCRICMRNFSDSGNLRVHIRTHTGEKPYKCPDCGKSFSQSSSLIRHQRTHTGEKPYECDHCGKSFSQSSHLNVHKRTHTGEKPYRCKYCDRSFSISSNLQRHVRNIHLRSGTPHEVGVYTLSGTPHEVGVYTLAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML
[0163] ND1-ZFP-Right-1397N-UGI (Figure 61f-g: SEQ ID No: 412) MLGFVGRVAAAPASGALRRLTPSASLPPAQLLLRAAPTAVHPVRDYAAQYPYDVPDYAVDEMTKKFGTLTIHDTEKAAEFGIHGVPAAMGGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGSGTPHEVGVYTLSGTPHEVGVYTLYKCPECGKSFSTKNSLTEHQRTHTGEKPYKCPECGKSFSSKKALTEHQRTHTGEKPYECNYCGKTFSVSSTLIRHQRIHTGEKPYRCKYCDRSFSISSNLQRHVRNIHLRSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML
[0164] Trac-ZFP-LEFT-1397N-UGI (Figure 61d: SEQ ID No: 413) MAPKKKRKVGIHGVPAAMGGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGSGTPHEVGVYTLSGTPHEVGVYTLFQCRICMRKFATSGSLTRHTKIHTGEKPFQCRICMRNFSRSDHLSTHIRTHTGEKPFACDICGRKFATSSNRTKHTKIHTHPRAPIPKPFQCRICMRNFSRSDNLSEHIRTHTGEKPFACDICGRKFAWHSSLRVHTKIHLRSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML
[0165] Trac-ZFP-right-1397C-AD (Figure 61d: SEQ ID No: 414) MAPKKKRKVGIHGVPAAMAERPFQCRICMRNFSRSDHLSTHIRTHTGEKPFACDICGRKFADRSHLARHTKIHTGSQKPFQCRICMRKFALKQHLNEHTKIHTGEKPFQCRICMRNFSQSGNLARHIRTHTGEKPFACDICGRKFAHNSSLKDHTKIHLRSGTPHEVGVYTLSGTPHEVGVYTLAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN
[0166] ND1-Left-TALE-1397N-UGI (Figure 62: SEQ ID No: 415) MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALLVGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML
[0167] ND1-Left-TALE-1397C-UGI (Figure 62: SEQ ID No: 416) MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALLVGSAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML
[0168] ND1-Right-TALE-1397N-UGI (Figure 62: SEQ ID No: 417)
[0169] ND1-Right-TALE-1397C-UGI (Figure 62: SEQ ID No: 418) MASVLTPLLLRGLTGSARRLPVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDKGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALLVGSAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML
[0170] ND4-Left-TALE-AD (Fig. 62b: SEQ ID No: 419) MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAMDIADLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPQQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPQQVVAIASNGGGRPALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN
[0171] ND4-Right-TALE-AD (Fig. 62b: SEQ ID No: 420) MASVLTPLLLRGLTGSARRLPVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDKMDIADLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPQQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNGGGRPALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN
[0172] ND1-Left-TALE-1333N-AD (Figure 63: SEQ ID No: 421) MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALLVGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN
[0173] ND1-Left-TALE-1333C-AD (Figure 63: SEQ ID No: 422) MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALLVGSPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN
[0174] ND1-Right-TALE-1333N-AD (Figure 63: SEQ ID No: 423)
[0175] ND1-Right-TALE-1333C-AD (Figure 63: SEQ ID No: 424)
[0176] ND1-Left-TALE-1397C-AD (Figure 62-63: SEQ ID No: 425) MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLLVGSAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN
[0177] ND1-Right-TALE-1397C-AD (Figure 62-63: SEQ ID No: 426)
[0178] ND1-Left-TALE-1333N (Figure 63: SEQ ID No: 427) MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGG
[0179] ND1-Left-TALE-1333C (Figure 63: SEQ ID No: 428) MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSGSPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC
[0180] ND1-Left-TALE-1397N (Figs. 62-63: SEQ ID No: 429) MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEG
[0181] ND1-Right-TALE-1333N (Figure 63: SEQ ID No: 430) MASVLTPLLLRGLTGSARRLPVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDKGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALLVGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGG
[0182] ND1-Right-TALE-1333C (Figure 63: SEQ ID No: 431) MASVLTPLLLRGLTGSARRLPVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDKGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALLVGSPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC
[0183] ND1-Right-TALE-1397N (Figures 62 - 63: SEQ ID No: 432) MASVLTPLLLRGLTGSARRLPVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDKGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALLVGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEG
[0184] ND4-LEFT-TALE-1333N-AD (Figure 63: SEQ ID No: 433)
[0185] ND4-LEFT-TALE-1333C-AD (Figure 63: SEQ ID No: 434)
[0186] ND4-LEFT-TALE-1397C-AD (Figure 63: SEQ ID No: 435)
[0187] ND4-Right-TALE-1333N-AD (Figure 63: SEQ ID No: 436)
[0188] ND4-Right-TALE-1333C-AD (Figure 63: SEQ ID No: 437)
[0189] ND4-Right-TALE-1397C-AD (Figure 63: SEQ ID No: 438)
[0190] ND1-Left-TALE-AD-GSVG (Figure 64: SEQ ID No: 439)
[0191] ND1-Left-TALE-AD-E1347A (Figure 64: SEQ ID No: 440)
[0192] ND1-Left-TALE-AD-AAAAA (Fig. 64: SEQ ID No: 441)
[0193] ND1-Right-TALE-AD (Fig.64: SEQ ID No: 442) MASVLTPLLLRGLTGSARRLPVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDKGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALLVSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN
[0194] ND1-Left-TALE-GSVG (Figure 64: SEQ ID No: 443) MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALLVGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLEGKVFSSGGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGVIPVKRGATGETKVFTGNSNGPKSPTKGGC
[0195] ND1-Left-TALE-E1347A (Figure 64: SEQ ID No: 444) MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALLVGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVAGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC
[0196] ND1-Left-TALE-AAAAA (Figure 64 SEQ ID No: 445) MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALLVGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVAAGATGETAVFTGNSNSPASPTAGGC
[0197] ND1-TALE-left (SEQ ID No: 446) DLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAAL
[0198] ND1-TALE-Right (SEQ ID No: 447) DLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAAL
[0199] ND4-TALE-left (SEQ ID No: 448) MDIADLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPQQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPQQVVAIASNGGGRPALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLG
[0200] ND4-TALE-Right (SEQ ID No: 449) MDIADLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPQQVVAIASHD GGKQALETVQRLLPVLCQAHGLTPQQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGKQALETVQALLPVLCQAHGLTPE QVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVL CQAHGLTPEQVVAIASNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALE TVQRLLPVLCQAHGLTPEQVVAIASNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGKQALETVQALLPVLCQAHGLTPEQVVAIAS NIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNGGGRPALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLG
[0201] TRAC-ZFP-LEFT (SEQ ID No: 450) FQCRICMRKFATSGSLTRHTKIHTGEKPFQCRICMRNFSRSDHLSTHIRTHTGEKPFACDICGRKFATSSNRTKHTKIHTHPRAPIPKPFQCRICMRNFSRSDNLSEHIRTHTGEKPFACDICGRKFAWHSSLRVHTKIHLR
[0202] TRAC-ZFP-Right (SEQ ID No: 451) FQCRICMRNFSRSDHLSTHIRTHTGEKPFACDICGRKFADRSHLARHTKIHTGSQKPFQCRICMRKFALKQHLNEHTKIHTGEKPFQCRICMRNFSQSGNLARHIRTHTGEKPFACDICGRKFAHNSSLKDHTKIHLR
[0203] ND1-ZFP-Left (SEQ ID No: 452) FQCRICMRNFSDSGNLRVHIRTHTGEKPYKCPDCGKSFSQSSSLIRHQRTHTGEKPYECDHCGKSFSQSSHLNVHKRTHTGEKPYRCKYCDRSFSISSNLQRHVRNIHLR
[0204] ND1-ZFP-Right (SEQ ID No: 453) YKCPECGKSFSTKNSLTEHQRTHTGEKPYKCPECGKSFSSKKALTEHQRTHTGEKPYECNYCGKTFSVSSTLIRHQRIHTGEKPYRCKYCDRSFSISSNLQRHVRNIHLR
[0205] G1333N-DddAtox (SEQ ID No: 454) GSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGG
[0206] G1333C-DddAtox (SEQ ID No: 455) GSPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC
[0207] G1397N-DddAtox (SEQ ID No: 456) GSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEG
[0208] G1397C-DddAtox (SEQ ID No: 457) GSAIPVKRGATGETKVFTGNSNSPKSPTKGGC
[0209] Adenine deaminase (AD: ABE 8e) (SEQ ID No: 458) SEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN
[0210] Full length DddAtox variant GSVG (SEQ ID No: 459) GSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLEGKVFSSGGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGVIPVKRGATGETKVFTGNSNGPKSPTKGGC
[0211] Full length DddAtox variant E1347A (SEQ ID No: 460) GSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVAGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC
[0212] Full length DddAtox variant AAAAA (SEQ ID No: 461) GSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVAAGATGETAVFTGNSNSPASPTAGGC
[0213] UGI (SEQ ID No: 462) TNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML
[0214] Single module (TALE)-linker-AD-GSVG (SEQ ID No: 463) SGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSINLVGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLEGKVFSSGGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGVIPVKRGATGETKVFTGNSNGPKSPTKGGC
[0215] (TALE)-linker-AD-E1347A (SEQ ID No: 464) SGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSINLVGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTYPNYANAGHVAGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC
[0216] (TALE)-linker-AD-AAAAA (SEQ ID No: 465) SGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGVAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSINLVGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVAAGATGETAVTFGNSNSASPTAGGC
[0217] Dual module (TALE) -linker-AD (SEQ ID No: 466) SGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN
[0218] (TALE)-GSVG (SEQ ID No: 467) GSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLEGKVFSSGGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGVIPVKRGATGETKVFTGNSNGPKSPTKGGC
[0219] (TALE)-E1347A (SEQ ID No: 468) GSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVAGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC
[0220] (TALE)-AAAAA (SEQ ID No: 469) GSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVAAGATGETAVFTGNSNSPASPTAGGC
[0221] (TALE)-1397C-AD (SEQ ID No: 470) GSAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN
[0222] (TALE)-1333N-AD (SEQ ID No: 471) GSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN
[0223] (TALE)-1333C-AD (SEQ ID No: 472) GSPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREV PVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN
[0224] Link(AA means amino acid) 8AA: SGGGLGST (SEQ ID No: 473) 16AA: SGSETPGTSESATPES (SEQ ID No: 474) 32AA: SGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID No: 475)
[0225] SOD2 MTS-3xHA: MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYA (SEQ ID No: 476) COX8A MTS-3xFLAG: MASVLTPLLLRGLTGSARRLPVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDK (SEQ ID No: 477)
[0226] 5. Transmission The fusion proteins of the present invention can be delivered to cells by various methods known in the art, including, but not limited to, microinjection, electroporation, DEAE-dextran treatment, lipofection, nanoparticle-mediated transfection, protein transduction domain-mediated introduction, and PEG-mediated transfection.
[0227] In another aspect, the present invention relates to a nucleic acid encoding the fusion protein. With respect to the nucleic acid, the terms "polynucleotide," "nucleotide," "nucleotide sequence," and "oligonucleotide" are used interchangeably. A polymeric form of nucleotides of any length may contain deoxyribonucleotides or ribonucleotides, or their analogs. A polynucleotide can have any three-dimensional structure and can perform any function, known or unknown. A polynucleotide can contain one or more nucleotides, such as methylated nucleotides and nucleotide analogs. Modifications to the nucleotide structure are possible before or after assembly of the polymer.
[0228] The polynucleotide may be an RNA sequence, a DNA sequence, or a combination thereof (combined RNA-DNA sequence).
[0229] As a means for expressing the fusion protein, known expression vectors such as plasmid vectors, cosmid vectors, and bacteriophage vectors can be used, and the vectors can be easily produced by those skilled in the art according to any known method using DNA recombination technology.
[0230] The vector may be a plasmid vector or a viral vector, and the viral vector may specifically be, but is not limited to, an adenovirus, an adeno-associated virus, a lentivirus, or a retrovirus vector.
[0231] A recombinant expression vector can contain a nucleic acid in a form suitable for expression of the nucleic acid in a host cell, meaning that the recombinant expression vector contains one or more regulatory elements, which can be selected based on the host cell, to be used for expression, i.e., operably linked to the nucleic acid sequence to be expressed.
[0232] Within a recombinant expression vector, "operably linked" means that the nucleotide sequence of interest is connected to regulatory elements in a manner that allows expression of the nucleotide sequence (e.g., in an in vitro transcription / translation system or in a host cell when the vector is introduced into a host cell).
[0233] The recombinant expression vector may comprise a form suitable for messenger RNA synthesis, including a T7 promoter, which means that it contains one or more regulatory elements that allow for in vitro mRNA synthesis, i.e., for messenger RNA synthesis by T7 polymerase.
[0234] "Regulatory elements" may include promoters, enhancers, internal ribosome entry sites (IRES), and other expression control elements (e.g., transcription termination signals, e.g., polyadenylation signals, and poly-U sequences). Regulatory elements include elements that direct inducible or constitutive expression of a nucleotide sequence in many types of host cells, and elements that direct expression of a nucleotide sequence only in specific host cells (e.g., tissue-specific regulatory sequences). Tissue-specific promoters can direct expression primarily in a desired tissue of interest, e.g., muscle, neurons, bone, skin, blood, a specific organ (e.g., liver, pancreas), or a specific cell type (e.g., lymphocytes). Regulatory elements can also direct expression in a temporally-dependent manner, such as a cell-cycle-dependent or developmental stage-dependent manner, which may or may not be tissue- or cell-type-specific.
[0235] In some cases, the vector contains one or more Pol III promoters, one or more Pol II promoters, one or more Pol I promoters, or a combination thereof. Examples of Pol III promoters include, but are not limited to, the U6 and H1 promoters. Examples of Pol II promoters include, but are not limited to, the retroviral Rous sarcoma virus (RSV) LTR promoter (optionally with the RSV enhancer), the cytomegalovirus (CMV) promoter (optionally with the CMV enhancer) (e.g., Boshart et al. (1985) Cell 41:521-530), the SV40 promoter, the dihydrofolate reductase promoter, the β-actin promoter, the phosphoglycerate kinase (PGK) promoter, and the EF1α promoter.
[0236] "Regulatory elements" may include enhancers, such as the WPRE; the CMV enhancer; the R-U5' segment of the HTLV-I LTR; the SV40 enhancer; and the intron sequence between exons 2 and 3 of rabbit β-globin. Those skilled in the art will recognize that the design of an expression vector can depend on factors such as the choice of host cell to be transformed, the desired expression level, and the like. The vector can be introduced into a host cell to produce transcripts, proteins, or peptides, including fusion proteins or peptides, encoded by the nucleic acids described herein (e.g., clustered regularly interspaced short palindromic repeats (CRISPR) transcripts, proteins, enzymes, mutants thereof, fusion proteins thereof, etc.). Useful vectors include lentiviruses and adeno-associated viruses, and the types of such vectors can also be selected to target specific cell types.
[0237] The vectors can be delivered in vivo or intracellularly by local injection (e.g., direct injection into a lesion or target site), electroporation, lipofection, viral vectors, nanoparticles, as well as PTD (protein translocation domain) fusion protein methods.
[0238] The nucleic acid can be injected in the form of ribonucleic acid, for example, messenger ribonucleic acid (mRNA), to enable gene base editing in cells, for example, gene base editing in animal cells or plant cells without limitation.
[0239] The nucleic acid according to the present invention may be in the form of mRNA. When delivered in the form of mRNA, gene editing can be initiated more quickly than when delivered in the form of a DNA vector because the transcription process into mRNA is not required, and there is a high possibility of transient protein expression.
[0240] The inventors of the present application have confirmed that for plant organelle gene editing, when a cytosine base editor is injected into plant cells in the form of ribonucleic acid, such as messenger ribonucleic acid, off-target effects are reduced compared to when it is delivered via a plasmid. This is the first time that the inventors have demonstrated that, in plant organelle gene editing, when a cytosine base editor is transformed into plant cells in the form of mRNA, it has an advantage in terms of off-target effects compared to when it is delivered via a plasmid.
[0241] The mRNA can be delivered directly or via a carrier. In some cases, the mRNA of the nucleic acid cleaving enzyme and / or cleavage factor can be chemically modified or directly delivered in the form of synthetic self-replicative RNA.
[0242] Methods for delivering mRNA molecules to cells in vitro or in vivo are contemplated, including methods for delivering mRNA to cells or methods for delivering mRNA to cells of organisms such as humans and animals in vivo. For example, mRNA molecules can be delivered to cells using lipids (e.g., liposomes, micelles, etc.), nanoparticles or nanotubes, or cationic compounds (e.g., polyethyleneimine or PEI). In some cases, bolistic methods such as gene guns or biolistic particle delivery systems can be used to deliver mRNA to cells.
[0243] The carrier may include, but is not limited to, for example, a cell penetrating peptide (CPP), a nanoparticle, or a polymer.
[0244] The CPPs are short peptides that facilitate the cellular uptake of a variety of molecular cargoes, from nano-sized particles to small chemical molecules and large fragments of DNA.
[0245] Regarding the nanoparticles, the compositions according to the present invention can be delivered via polymeric nanoparticles, metal nanoparticles, metal / inorganic nanoparticles, or lipid nanoparticles. The polymeric nanoparticles may be, for example, DNA nanoclues or filamentous DNA nanoparticles synthesized by rolling circle amplification. The DNA nanoclues and filamentous DNA nanoparticles can be loaded with mRNA and coated with PEI to improve their endosomal escape ability. These complexes can bind to the cell membrane, be internalized, and then transported to the nucleus via endosomal escape and delivered.
[0246] Regarding the metal nanoparticles, gold particles can be linked and delivered to cells by complexing with a cationic endosomal escape polymer, such as polyethyleneimine, poly(aladdin), poly(lysine), poly(histidine), poly-[2-{(2-aminoethyl)amino}-ethyl-aspartamide] (pAsp(DET)), a block copolymer of poly(ethylene glycol) (PEG) and poly(arginine), a block copolymer of PEG and poly(lysine), or a block copolymer of PEG and poly{N-[N-(2-aminoethyl)-2-aminoethyl]aspartamide} (PEG-pAspDET)).
[0247] With regard to the metal / inorganic nanoparticles, for example, mRNA can be encapsulated via ZIF-8 (zeolitic imidazolate framework-8).
[0248] In some cases, the mRNA may be anionic and may be bound to a cationic substance to form nanoparticles, which may penetrate into cells via receptor-mediated endocytosis or endocytosis.
[0249] Cationic polymers include polyallylamine (PAH), polyethyleneimine (PEI), poly(L-lysine) (PLL), poly(L-arginine) (PLA), polyvinylamine monopolymers or copolymers, poly(vinylbenzyl-tri-C1-C4-alkylammonium salts), polymers of aliphatic or araliphatic dihalides and aliphatic N,N,N',N'-tetra-C1-C4 alkyl alkylenediamines, poly(vinylpyridine) or poly(vinylpyridinium salts), poly(N,N-diallyl-N,N-di-C1-C4-alkyl-ammonium halides), monopolymers or copolymers of quaternized di-C1-C4-alkyl aminoethyl acrylate or methacrylate, and polyquads. TM , polyaminoamides, and the like.
[0250] The cationic lipid may include a cationic liposome formulation. The lipid bilayer of the liposome protects the encapsulated nucleic acid from degradation and prevents specific neutralization by antibodies capable of binding to nucleic acids. During endosome maturation, the endosomal membrane and liposome fuse, allowing the cationic lipid-nucleic acid cleaving enzyme to efficiently escape the endosome. Cationic lipids may include polyethyleneimine, polyamidoamine (PAMAM)-functionalized dendrimer, lipofectin (a combination of DOTMA and DOPE), lipofectase, LIPOFECTAMINE® (e.g., LIPOFECTAMINE® 2000, LIPOFECTAMINE® 3000, LIPOFECTAMINE® RNAiMAX, LIPOFECTAMINE® LTX), SAINT-RED (Synvolux Therapeutics, Groningen, The Netherlands), DOPE, cytopectin (Gilead Sciences, Poster City, California), and eupectin (JBL, San Luis Obispo, California). Representative cationic liposomes can be prepared from N-[1-(2,3-dioleooxy)-propyl]-N,N,N-trimethylammonium chloride (DOTMA), N-[1-(2,3-dioleooxy)-propyl]-N,N,N-trimethylammonium methylsulfate (DOTAP), 3β-[N-(N',N'-dimethylaminoethane)carbamoyl]cholesterol (DC-hol), 2,3-dioleyloxy-N-[2(sperminecarboxamido)ethyl]-N,N-dimethyl-1-propanamide trifluoroacetate (DOSPA), 1,2-dimyristyloxypropyl-3-dimethyl-hydroxyethylammonium bromide, or dimethyldioctadecylammonium bromide (DDAB).
[0251] Lipid nanoparticles can also be delivered to carriers using liposomes. Liposomes are spherical vesicular structures consisting of a single or multilamellar lipid bilayer surrounding an internal aqueous compartment and a relatively impermeable outer lipophilic phospholipid bilayer. Liposomal formulations may contain primarily natural phospholipids and lipids, such as 1,2-distearoyl-sn-glycero-3-phosphatidylcholine (DSPC), sphingomyelin, phosphatidylcholine, or monosialoganglioside. In some cases, cholesterol or 1,2-dioleoyl-sn-glycero-3-phosphoethanolamine (DOPE) can be added to the lipid membrane to eliminate instability in plasma. The addition of cholesterol reduces the rapid release of the encapsulated bioactive compound into plasma, while the addition of 1,2-dioleoyl-sn-glycero-3-phosphoethanolamine (DOPE) increases stability.
[0252] 7. Base editing In another aspect, the present invention relates to a base editing composition comprising the fusion protein or the nucleic acid. In another aspect, the present invention relates to a base editing method comprising a step of treating a cell with the composition. After a DNA-binding protein, such as a TALE or ZFP, binds to a target DNA, the cytosine deaminase of the fusion protein hydrolyzes the amino group of cytosine to convert it to uracil. Because uracil can form a base pair with adenine, during the intracellular DNA replication process, the cytosine-guanine base pair can be converted to a uracil-adenine base pair and finally to a thymine-adenine base pair. Furthermore, the adenine deaminase of the fusion protein hydrolyzes the amino group of adenine to convert it to hypoxanthine. Because hypoxanthine can form a base pair with cytosine, during the intracellular DNA replication process, the adenine-thymine base pair can be converted to a guanine-cytosine base pair via a hypoxanthine-cytosine base pair.
[0253] The cells may be, but are not limited to, eukaryotic cells (e.g., fungi such as yeast, cells derived from eukaryotic animals and / or eukaryotic plants (e.g., germ cells, stem cells, somatic cells, germ cells, etc.), eukaryotic animals (e.g., humans, primates such as monkeys, dogs, pigs, cows, sheep, goats, mice, rats, etc.), or eukaryotic plants (e.g., algae such as green algae, corn, soybeans, wheat, rice, etc.).
[0254] (1) Base editing of plant cell DNA The present invention relates to a composition for base editing of plant cell DNA or a base editing method, and the composition for base editing of plant cells comprises the fusion protein or a nucleic acid encoding the same, and an NLS peptide, a chloroplast transit peptide, a mitochondrial targeting signal (MTS), a nuclear export signal protein, or a nucleic acid encoding the same.
[0255] The present invention also provides a composition for base editing in plant cells, comprising the fusion protein or nucleic acid; and an NLS peptide or a nucleic acid encoding the same.
[0256] The present invention also provides a composition for base editing in plant cells, comprising the fusion protein or nucleic acid; and a chloroplast transit peptide or a nucleic acid encoding the same.
[0257] The present invention also provides a composition for base editing in plant cells, comprising the fusion protein or nucleic acid; and a mitochondrial transfer protein or a nucleic acid encoding the same.
[0258] In some cases, the present invention provides a composition for base editing in plant cells, further comprising a nuclear export signal protein or a nucleic acid encoding the same.
[0259] Specifically, the present invention relates to compositions and methods for editing plant nuclear DNA, mitochondrial DNA, or chloroplast DNA.
[0260] Specifically, the fusion protein can be delivered to plant cells via: injection using a Gene gun (Bombardment); PEG-mediated protoplast transfection; Protoplast transfection (electroporation); or Protoplast injection, microinjection.
[0261] The polynucleotide sequence encoding the fusion protein according to the invention may be an RNA sequence, a DNA sequence, or a combination thereof (a combined RNA-DNA sequence).
[0262] The polynucleotide encoding the fusion protein can be transferred to the plant cell via:
[0263] Transformation using Agrobacterium, e.g., Agrobacterium tumefaciens - binary vector, - viral vectors: geminivirus, Tobacco rattle virus (TRV), tomato mosaic virus (ToMV), foxtail mosaic virus (FoMV), barley yellow striate mosaic virus (BYSMV), Sonchus yellow net rhabdovirus (SYNV), etc. Transfection with viruses; injection (Bombardment, Gene gun); PEG-mediated protoplast transfection; Protoplast transfection (electroporation); or Protoplast injection, microinjection.
[0264] Examples of the virus include viral vectors such as geminivirus, tobacco rattle virus (TRV), tomato mosaic virus (ToMV), foxtail mosaic virus (FoMV), barley yellow striate mosaic virus (BYSMV), and Sonchus yellow net rhabdovirus (SYNV).
[0265] The vectors can be delivered into cells via local injection (e.g., direct injection into a lesion or target site), electroporation, lipofection, viral vectors, nanoparticles, as well as PTD (protein translocation domain) fusion protein methods.
[0266] With regard to the protein or the nucleic acid encoding it that is transported to the plant organ, the plant organ may be a mitochondria, a chloroplast, or a plastid (plastid: leucoplast, chromoplast).
[0267] The protein transported to the plant organ may be, for example, a chloroplast transport peptide or a mitochondrial transport protein.
[0268] For example, chloroplast transfer protein (CTP) or mitochondrial transfer signal protein (MTS) binds and is transferred to the chloroplasts and mitochondria in plant cells. When transferred to the chloroplasts and mitochondria, the remaining portion, excluding the N-terminal CTP or MTS protein, is transferred to the inside of the chloroplasts and mitochondria in preprotein form. During the process of entering the chloroplasts and mitochondria, the transfer protein portion separates, allowing for base editing of specific regions by targeting the chloroplasts and mitochondria.
[0269] In addition to the fusion protein or the nucleic acid encoding it, a chloroplast signaling protein (CTP) or a nucleic acid encoding it, or a mitochondrial signaling protein (MTS) or a nucleic acid encoding it can be treated or introduced into a plant cell to edit bases in plant mitochondrial, chloroplast, chromoplast, or leucoplast DNA.
[0270] When a nuclear export sequence is attached to a base editing protein during mitochondrial gene editing, base editing can be performed with higher efficiency. The nuclear export signal protein may be derived from, for example, but is not limited to, MVM (Mirute virus of mice). The nuclear export signal protein may include, for example, but is not limited to, the amino acid sequence of SEQ ID NO: 31.
[0271] VDEMTKKFGTLTIHDTEK (SEQ ID NO: 31) The present invention also includes a TAL (Transcription Activator-Like) effector (TALE) domain that cleaves wild-type DNA base sequences but cannot cleave edited base sequences, and FokI nuclease (FokI nuclease) or a nucleic acid encoding it, or ZFN or a nucleic acid encoding it, i.e., mitoTALEN (Mitochondrial TALE Nuclease), which is a mitochondrial nuclease, or a nucleic acid encoding it, or ZFN or a nucleic acid encoding it, thereby enabling more efficient mitochondrial base editing to be expected even when a mitochondrial sequence-cleaving protein is used simultaneously.
[0272] (2) Base editing of animal cell DNA The present invention relates to a composition for base editing of DNA in animal cells or a base editing method, the composition for base editing of animal cells comprising the fusion protein or a nucleic acid encoding the same, and an NLS peptide, a mitochondrial transport protein (MTS), a nuclear export signal protein, or a nucleic acid encoding any of them.
[0273] The present invention also provides a composition for base editing in animal cells, comprising the fusion protein or nucleic acid; and an NLS peptide or a nucleic acid encoding the same.
[0274] The present invention also provides a composition for base editing in animal cells, comprising the fusion protein or nucleic acid; and a mitochondrial transport protein (MTS) or a nucleic acid encoding the same. The present invention further provides a composition for base editing in animal cells, which optionally further comprises a nuclear export signal protein or a nucleic acid encoding the same.
[0275] The animal cells are non-human animal cells, and bases in the mitochondrial DNA of the non-human animal cells can be edited by processing a nuclear export signal protein or a nucleic acid encoding the same; and / or a mitochondrial transduction signal protein (MTS) or a nucleic acid encoding the same.
[0276] For example, the fusion protein or the nucleic acid encoding it is bound to a mitochondrial transfer protein (MTS) and delivered to the mitochondria. Upon delivery to the mitochondria, the remaining portion, excluding the N-terminal MTS protein, is delivered to the mitochondria in the form of a whole protein. During entry into the mitochondria, the transfer protein portion is detached, allowing for base editing of specific portions by targeting the mitochondria.
[0277] Mitochondrial signal transduction, TAL effectors, and cytosine deaminases (DddA) tox The present invention relates to a composition for base editing of mitochondrial DNA in non-human animal cells or a method for editing the bases of mitochondrial DNA, in which a nuclear export signal (NES) or a nucleic acid encoding the same is linked to a TALE-DdCBE (TALE DddA-derived cytosine base editor) containing the TALE-DdCBE or a nucleic acid encoding the same. The nuclear export signal can reduce nuclear DNA base editing of mitochondrial-nuclear analogous sequences.
[0278] According to the present invention, more efficient base editing of animal mitochondrial DNA can be achieved by including a nuclear export signal protein or a nucleic acid encoding the same. Furthermore, the nuclear export signal can reduce nuclear DNA base editing of the mitochondrial-nuclear mitotic sequence, allowing only mitochondrial DNA to be edited.
[0279] The nuclear export signal protein may be derived from, for example, but is not limited to, MVM (Mirute virus of mice) and may include, for example, but is not limited to, the amino acid sequence of VDEMTKKFGTLTIHDTEK (SEQ ID NO: 31).
[0280] A TAL effector (TALE) domain and FokI nuclease or a nucleic acid encoding the same, or a ZFN or a nucleic acid encoding the same, which cleaves wild-type DNA base sequences but cannot cleave edited base sequences, can be injected into eukaryotic cells simultaneously with, or before sequential editing of, the (1) nuclear export signal protein or a nucleic acid encoding the same and (2) a DNA binding protein, deaminase or a mutant thereof, or a nucleic acid encoding the same.
[0281] In particular, with regard to base editing of eukaryotic mitochondrial genes, this may include a nuclear export signal protein or a nucleic acid encoding the same and / or a mitochondrial transduction signal protein (MTS) or a nucleic acid encoding the same.
[0282] According to the present invention, when a nuclear export sequence is attached to a base editing protein during animal mitochondrial gene editing, bases are edited more efficiently, and in animal embryos, non-specific base editing of similar sequences within the nucleus is also suppressed.
[0283] Furthermore, by further including a mitochondrial nuclease, mitoTALEN (Mitochondrial TALE Nuclease), or a nucleic acid encoding it, the present invention is expected to achieve more efficient mitochondrial base editing, even when a mitochondrial sequence-cleaving protein is simultaneously used. Using a mitochondrial DNA nuclease, mitoTALEN, mitochondrial DNA can be cleaved, allowing for highly efficient base-edited genomes to be obtained in animals by cleaving the wild-type mitochondrial genome.
[0284] The present invention further includes a mitochondrial nuclease, mitoTALEN, or a nucleic acid encoding the same, which is expected to enable more efficient mitochondrial base editing even when a mitochondrial sequence-cleaving protein is simultaneously used. Specifically, the present invention may include a fusion protein (mitoTALEN) having a TAL effector domain bound to a mitochondrial signaling signal and FokI nuclease (FokI nuclease) or ZFN, or a nucleic acid encoding the same.
[0285] Using mitoTALEN, a mitochondrial DNA nuclease, mitochondrial DNA can be cleaved, and base-edited genomes can be obtained from animals with high efficiency by cleaving the wild-type mitochondrial genome.
[0286] In some cases, the composition may further contain UGI (uracil DNA glycosylase inhibitor), which can enhance base editing efficiency by inhibiting the activity of UDG (uracil DNA glycosylase), an enzyme that catalyzes the removal of U from DNA and repairs mutated DNA.
[0287] Specifically, DddA-derived cytosine base editors (DdCBEs), consisting of the split bacterial toxin DddAtox, a TALE designed to bind to DNA, and a uracil glycosylation inhibitor (UGI), enabled targeted cytosine-thymine base modifications in mitochondrial DNA. In our experiments, we demonstrated that highly efficient mitochondrial DNA editing is possible in mouse embryos. We targeted the mitochondrial gene MT-ND5 (ND5), which encodes a subunit of NADH dehydrogenase, which catalyzes the dehydration of NADH and electron transfer to ubiquinone. This gene contains mutations associated with human mitochondrial disease, such as m.G12918A, and mutations that create premature stop codons, such as m.C12336T. This enabled the generation of mitochondrial disease models in mice, suggesting the possibility of treating mitochondrial diseases.
[0288] (2) a DNA-binding protein, deaminase, or a mutant thereof, or a nucleic acid encoding the same, can be linked to the (1) nuclear export signal protein or a nucleic acid encoding the same, and (3) a mitoTALEN, a nuclease, or a nucleic acid encoding the same can be linked to the (2). To deliver (1) to (3), a single or multiple delivery means can be used in combination in the same or different configurations.
[0289] (1) may be contained in a first delivery vehicle, (2) may be contained in a second delivery vehicle, and (3) may be contained in a third delivery vehicle. Each delivery system may simultaneously be a viral delivery vehicle, one may be a viral delivery vehicle and the other a non-viral delivery vehicle, or both may be non-viral delivery vehicles.
[0290] The nuclear export signal proteins (1) to (3), DdCBE, and mitoTALEN can be mixed and delivered.
[0291] One or more of (1) to (3) above can be transferred to a nuclear export signal protein, DdCBE, or mitoTALEN, and some of them can be transferred by placing the DNA sequence encoding (1) to (3) on a vector.
[0292] The DNA sequences encoding (1) to (3) can be located on the same vector and transferred simultaneously via one vector, or they can be transferred by being located on separate, different vectors.
[0293] The animals of the present invention may include humans or non-human animals. The non-human transgenic animals may be insects, annelids, mollusks, brachiopods, nematodes, coelenterates, sponges, chordates, or vertebrates. The vertebrates may be fish, amphibians, reptiles, birds, or mammals. The insects may be Drosophila, the nematodes may be Caenorhabditis elegans, the fish may be zebrafish, the mammals may be Primates, Carnivora, Insectivora, Rodentia, Artiodactyla, Perissodactyla, or Proboscidea, and the rodents may be rats or mice.
[0294] A composition according to the present invention can be introduced into a human or non-human animal embryo, and the embryo can be implanted into a surrogate mother for gestation to produce a base-edited animal. A composition according to the present invention can be introduced into a fertilized animal egg and cultured.
[0295] The resulting fertilized egg can be implanted into a surrogate mother and delivered to term. The method for producing the non-human transgenic animal may further include a step of confirming whether the non-human transgenic animal is transgenic after delivery. The non-human transgenic animals can be mated to produce offspring transgenic animals.
[0296] The term "offspring" refers to all viable offspring transformed animals that can be crossed with the non-human transformed animal, and more specifically refers to the F1 generation produced by crossing the transformed animal with another transformed animal or the transformed animal with a normal animal, and the F2 generation and subsequent generations produced by crossing an F1 generation animal with a normal animal, but is not limited to this.
[0297] The mating may be characterized by mating the transgenic animal with a normal animal. The present invention may also include cells, tissues, and by-products isolated from the transgenic animal or progeny of the transgenic animal. The by-products refer to all substances derived from the transgenic rabbit, and are preferably selected from the group consisting of blood, serum, urine, feces, saliva, organs, and skin, but are not limited thereto. MODES FOR CARRYING OUT THE INVENTION
[0298] The present invention will be described in more detail below with reference to examples. It will be obvious to those skilled in the art that these examples are merely for the purpose of illustrating the present invention and should not be construed as limiting the scope of the present invention.
[0299] Example 1. Zinc Finger Deaminase (ZFD) Base editing of nuclear or mitochondrial DNA is widely useful in biomedical research, medicine, and biotechnology. The ZFD platform includes a DNA-binding protein, the interbacterial toxin deaminase DddAtox, and a uracil glycosylase inhibitor. UGI catalyzes targeted C-to-T transversions in human cells without inducing unwanted small insertions or deletions. Using publicly available zinc finger resources, we engineered plasmids encoding ZFD and achieved base editing frequencies of up to 60% in nuclear DNA and 30% in mitochondrial DNA. Unlike CRISPR-based base editing, ZFD does not cleave DNA to create single- or double-strand breaks, thereby avoiding unwanted insertions and deletions due to error-prone non-homologous end joining at the target site. Furthermore, recombinant ZFD protein purified in E. coli spontaneously translocates through human cells and converts the target base to another base (base conversion). This demonstrates the proof-of-principle of gene-free gene therapy. Technologies for genome editing in eukaryotic cells and organisms include, but are not limited to, zinc finger nucleases (ZFNs), transcription activator-like effector (TALE) nucleases (TALENs), TALE-linked split interbacterial deaminase toxin DddA-derived cytosine base editors (aka DdCBEs), CRISPR-Cas9, and deaminases (aka base editors) that have lost their cleavage activity. These tools essentially consist of two functional units: a DNA-binding part and a catalytic part. Thus, the zinc finger array or TALE array functions as the DNA-binding part, while the nuclease (FokI in the case of ZFNs and TALENs) or deaminase (split DddA in DdCBEs) functions as the catalytic part. toxand APOBEC1 in CBEs function as catalytic units. Crispr-Cas9 possesses both nuclease and RNA-guided DNA binding protein functions. Customized, programmable nucleases, such as ZFNs, TALENs, and Cas9, cleave DNA to create double-strand breaks, which then repair and induce targeted gene knockout and knockin. However, double-strand breaks induced by programmable nucleases can lead to large deletions of undesired genes at the target site, p53 activation, and chromosomal rearrangements during simultaneous DSB repair at both the target and non-target sites. In contrast, programmable base editors, including cytosine and adenine base editors (CBEs and ABEs), do not create DSBs, thereby avoiding these unwanted events in cells and efficiently catalyzing single nucleotide conversions without a repair template or donor DNA. However, because CBE and ABE contain Cas9 nickase mutants, they still cleave one strand of DNA, creating a nick or single-strand break, resulting in unwanted insertions or deletions at the gene target site.
[0300] It can catalyze C-to-T base substitution in nuclear and mitochondrial DNA within cells. Using a custom-designed DdCBE, we demonstrated mitochondrial DNA editing in mice and chloroplast DNA editing in plants. Zinc finger deaminase (ZFD), which performs deletion-free and precise base editing in human cells and other eukaryotic cells, was constructed by linking split DddAtox to a customized zinc finger protein. Zinc finger arrays (2 x 0.3-0.6 k base pairs) are smaller in size than TALE arrays (2 x 1.7-2 k base pairs) and S. pyogenes Cas9 (4.1 k base pairs). Therefore, for in vivo studies and gene therapy applications, ZFD-encoded genes can be easily loaded into viral vectors with limited viral cargo space, such as AAVs. Unlike TALE arrays, zinc finger arrays do not have bulky domains at the C- or N-terminus, making them more engineering-friendly. The split halves of DddAtox can be fused to the C- or N-terminal portions of zinc finger proteins. Furthermore, zinc finger proteins have unique cell-penetrating capabilities, enabling nucleic acid-free gene editing in human cells. These properties make zinc finger proteins an ideal platform for DNA-binding modules to edit bases in the nucleus or other organelles.
[0301] 1-1. Materials and Methods Plasmid production For mammalian expression, the p3s-ZFD plasmid was prepared by digesting the p3s-ABE7.10 plasmid (addgene, #113128) with HindIII and XhoI (NEB) enzymes and then modifying it. The digested p3s plasmid and the synthesized insert DNA were assembled using the HiFi DNA Assembly Kit (NEB). All insert DNAs encoding MTS-, ZFP (Toolgen, Sangamo and Barbas module-, split DddA-, or UGI-) were synthesized by IDT. pTarget plasmids were designed to determine the optimal length of the spacer sequence for ZFD activity. Each pTarget plasmid containing spacers of various lengths and two ZFP-binding sites was constructed by inserting the ZFP-binding sequence and spacer sequence into the pRGS-CCR5-NHEJ reporter plasmid digested with two enzymes (EcoRI and BamHI, NEB). pET-ZFD plasmids for protein production in E. coli were derived from the pET-Hisx6-rAPOBEC1-XTEN-nCas9-UGI-NLS plasmid (addgene, The ZFD sequence was amplified from the p3s-ZFD plasmid (#89508) using PCR, and Hisx6 tag and GST tag sequences were synthesized using oligonucleotides (Macrogen). All plasmids for protein purification were constructed by inserting the ZFD and tag-encoding sequences into enzyme-cleaved pET plasmids using the HiFi DNA Assembly Kit (NEB). Plasmids were transformed into chemically competent DH5a E. coli cells, and the plasmids were purified using the AccuPrep Plasmid Mini Extraction Kit (Bioneer) according to the manufacturer's protocol. The desired plasmids were selected after confirming the entire sequence by Sanger sequencing.
[0302] HEK293T cell culture and transfection HEK293T cells (ATCC CRL-11268) were cultured in Dulbecco's Modified Eagle Medium (Welgene) supplemented with 10% fetal bovine serum (Welgene) and 1% antibiotic-antimycotic solution (Welgene). HEK293T cells (7.5 × 10 4 ) were seeded into 48-well plates. After 18–24 hours, when cells reached 70–80% growth, they were transfected with plasmids encoding the left and right ZFDs (500 ng each) using Lipofectamine 2000 (1.5 μL, Invitrogen) or co-transfected with the pTarget plasmid (10 ng). After 96 hours of transfection, cells were harvested and then lysed in 10.0 μL of cell lysis buffer (50 mM Tris-HCl, pH 8.0 (Sigma-Aldrich), 1 mM EDTA (Sigma-Aldrich), 0.005% sodium dodecyl sulfate (Sigma-Aldrich)), 5 μL of proteinase K (Qiagen), and incubated at 55°C for 1 hour and 95°C for 10 minutes. To perform whole-mtDNA sequencing, HEK293T cells were transfected with serially diluted concentrations of plasmid or mRNA encoding the ND1 or ND2 targeting mitoZFD pair. The amount of construct (ng) was 7.5 x 10 4 mtDNA was extracted from cells 96 hours after transfection.
[0303] K562 cell culture and transfection K562 cells were cultured in RPMI 1640 medium supplemented with 10% fetal bovine serum (Welgene) and 1% antibiotic-antimycotic solution (Welgene). ZFD transduction into K562 cells by electroporation was performed using an Amaxa 4D-Nucleofector equipped with program FF-120 (Lonza). TM The X Unit system was used. 16-well Nucleocuvette TMWhen using strips, the maximum volume of substrate solution added to each sample was 2 μL. For proteins, 220 pmol (maximum dose) or 110 pmol (half the maximum dose) of the left and right ZFDs, respectively, were transfected into K562 cells (1 × 10 cells). For plasmids, 500 ng of plasmids encoding the left and right ZFDs were transfected. After 96 h, cells were harvested by centrifugation at 100 g for 5 min. Then, 100 μL of cell lysis buffer (50 mM Tris-HCl, pH 8.0 (Sigma-Aldrich), 1 M EDTA (Sigma-Aldrich), 0.005% sodium dodecyl sulfate (Sigma-Aldrich)) and 5 μL of proteinase K (Qiagen) were added and incubated at 55°C for 1 h and 95°C for 10 min. To directly transduce ZFDs or ZFD-encoding plasmids into K562 cells, we followed the method previously used for direct transduction of ZFNs. A mixture of left- and right-side ZFD proteins (final concentration 50 μM) or a mixture of plasmids encoding left- and right-side ZFDs (500 ng each) was diluted in serum-free medium, pH 7.4, containing 100 mM L-arginine and 90 μM ZnCl2, to a final volume of 20 μL. K562 cells (1 × 105) were centrifuged at 100 g for 5 minutes, and the supernatant was removed. The cells were then released into the diluted ZFD solution and incubated at 37°C for 1 hour. After incubation, the cells were centrifuged at 100 g for 5 minutes and replaced with fresh culture medium. The cells were maintained at 30°C (transient hypothermia) or 37°C for 18 hours, then grown at 37°C for 2 days. Some cells were treated twice as described above. Cells were analyzed 96 hours after treatment.
[0304] Expression and purification of ZFD proteins Plasmids encoding each pair of ZFDs, each with a C-terminal GST tag, were transformed into Rosetta (DE3) competent cells and then plated on LB agar plates containing kanamycin. After overnight incubation, single colonies were selected and cultured overnight (preculture) in liquid medium containing 50 μg / ml kanamycin and 10.0 μM ZnCl2 at 37°C. The next day, a portion of the preculture was transferred to a larger volume of liquid medium and cultured at 37°C until the absorbance at A600 nm reached ∼0.5–0.70. The culture was placed on ice for approximately 1 hour, and then 0.5 mM isopropyl β-D-1-thiogalactopyranoside (IPTG; GoldBio) was added to induce ZFD protein expression, and the culture was incubated at 18°C for 14 hours.
[0305] For protein purification, cells were lysed in a lysis bath containing 50 mM Tris-HCl (Sigma-Aldrich), 500 mM NaCl (Sigma-Aldrich), 1 mM MgCl2 (Sigma-Aldrich), 10 mM 1,4-dithiothreitol (DTT; GoldBio), 1% Triton X-10 (Sigma-Aldrich), 10% glycerol, 1 mM phenylmethylsulfonyl fluoride (Sigma-Aldrich), 1 mg / ml lysozyme extracted from hen's egg white (Sigma-Aldrich), 100 μM ZnCl2 (Sigma-Aldrich), 100 mM ZnCl2 (Sigma-Aldrich), pH 8.0. To further lyse the cells, the solution was sonicated (3 min total, 5 s on, 10 s off). The solution was then centrifuged (13,000 rpm) and the supernatant was extracted. The supernatant was incubated with Glutathione Sepharose 4B (GE Healthcare) for 1 hour. After incubation, the resin-lysate mixture was poured onto the column and washed three times with wash buffer (50 mM Tris-HCl (Sigma-Aldrich), 500 mM NaCl (Sigma-Aldrich), 10 mM DTT (GoldBio), 1 mM MgCl2 (Sigma-Aldrich), 100 μM ZnCl2 (Sigma-Aldrich), 10% glycerol, 100 mM arginine (Sigma-Aldrich), pH 8.0). Proteins adhering to the resin were separated from the resin with elution buffer (50 mM Tris-HCl (Sigma-Aldrich), 500 mM NaCl (Sigma-Aldrich), 1 mM MgCl2 (Sigma-Aldrich), 40 mM glutathione (Sigma-Aldrich), 10% glycerol, 1 mM DTT (GoldBio), 100 μM ZnCl2 (Sigma-Aldrich), 100 mM arginine (Sigma-Aldrich), pH 8.0).Finally, the extracted proteins were concentrated to a concentration of ~15 ng / µL (200–240 pmol / µL, depending on the size of the protein).
[0306] In vitro deamination of PCR amplicons by ZFD Amplicons containing TRAC sites were generated using PCR. 8 μg of amplicon was incubated with 2 μg of each ZFD protein (Left-G1397N and Right-G1397C) in NEB3.1 buffer containing 10.0 μM ZnCl2 for 1-2 hours at 37°C. After the reaction, 4 μL of proteinase K solution (Qiagen) was added and incubated at 55°C for 30 minutes to remove the ZFD. The amplicon was then purified using a PCR purification kit (MGmed). 1 μg of the purified amplicon was incubated with two USER enzymes (NEB) for 1 hour at 37°C. The amplicon was then incubated with 4 μL of proteinase K solution and purified again using the PCR purification kit. The purified PCR fragment was electrophoresed on an agarose gel and imaged.
[0307] Targeted Deep Sequencing For analysis of on- and off-target base editing ratios, target sites were amplified by overlapping primary PCR, secondary PCR, and tertiary PCR using TruSeq HT Dual Index-containing primers with PrimeSTAR® GXL DNA polymerase (TAKARA) to generate deep sequencing libraries. Libraries were then subjected to paired-end sequencing using an Illumina MiniSeq.
[0308] Preparation of mRNA A DNA template containing a T7 RNA polymerase promoter upstream of the ZFD sequence was generated by PCR using forward and reverse primers (forward: 5'-CATCAATGGGCGTGGATAG-3' SEQ ID No: 116, reverse: 5'-CATCAATGGGCGTGGATAG-3' SEQ ID No: 117, reverse: 5'-GACACCTACTCAGACAATGC-3' SEQ ID No: 118). The mRNA was expressed as 100 mM MESSAGE 100 mM ATP. TM In vitro transcribed mRNA was synthesized using the T7 ULTRA transcription kit (Thermo Fisher). The in vitro transcribed mRNA was purified using MEGAclear according to the manufacturer's protocol. TM Purification was performed using a Transcription Clean-Up Kit (Thermo Fisher).
[0309] Whole mitochondrial genome sequencing For whole mitochondrial genome sequencing, three steps are required. 1. mtDNA extraction from isolated mitochondria: First, 3 × 10 HEK293T cells were trypsinized and harvested by centrifugation (500 g, 4 min, 4°C) 96 hours after transfection with the ND1 or ND2-targeting mitoZFD pair. The cells were then washed with phosphate-buffered saline (Welgene) and harvested again by centrifugation. The supernatant was removed, and mitochondria were isolated from the cultured cells using the reagent-based method of the Mitochondria Isolation Kit for Cultured Cells (Thermo Fisher) according to the manufacturer's protocol. mtDNA was extracted from the isolated mitochondria using the DNeasy Blood & Tissue Kit (Qiagen). 2. NGS library generation: To generate an NGS library from the extracted mtDNA, Nextera TMWe used the Illumina DNA Prep Kit, which includes DNA CD Indexes (Illumina). 3. Next Generation Systems: Libraries were pooled and loaded onto a MiniSeq sequencer (Illumina). The average sequencing depth was greater than 50.
[0310] Analysis of global DNA editing of the mitochondrial genome To analyze the NGS data from the whole mitochondrial genome sequencing, Fastq files were aligned to the GRCh38.p13 (release v102) reference genome using BWA, and read pairing information and flags were corrected to generate BAM files using SAMtools (v.1.9). The REDItoolDenovo.py script in REDItools (v.1.2.1) was then used to identify all cytosine and guanine positions in the mitochondrial genome with base editing rates of 1% or higher. Positions with base editing rates of 50% or higher in the cell line were considered single nucleotide mutations and excluded from all samples. For off-target analysis, the target sites of each ZFD were removed. Remaining positions with an editing frequency of ≥1% were considered off-target, and the number of edited C / G nucleotides was calculated. To calculate the average C / G to T / A base editing frequency for off-targets in the mitochondrial genome, all C / G nucleotides were averaged. The specificity ratio was calculated by dividing the average on-target by the average off-target. A graph of the entire mitochondrial genome was generated displaying base editing rates at on-target and off-target positions.
[0311] 1-2. Optimization of ZFD structure To develop ZFNs for base editing in human and other eukaryotic cells, we first optimized the amino acid linker and spacer length of the Zing finger protein (ZFP) attached to the DddAtox half. A C-to-T transversion occurs in the spacer between the left and right ZFP attachment sites. We selected a well-characterized ZFN pair targeting the human CCR5 gene. Using this, we engineered ZFNs with various amino acid linkers (2, 5, 10, 16, 24, and 32) and created a series of targeting plasmids with various spacers (1 to 24 base pairs long) containing repeated T-T sequences and the left and right ZFP binding sites of the ZFN (Figure 1a, b, and Table 1).
[0312] [Table 1]
[0313] DddAtox can be split at two positions (G1333 and G1397), and each half can be fused to either the left or right ZFP. The base editing efficiency of the resulting 24 ZFP constructs (= 6 linkers x 2 split positions x ZFP fusion position (left or right)) was measured for each target plasmid containing 24 spacers. This was measured by deep sequencing in Hek293T cells on day 4 after transfection.
[0314] [structure] ·Left-ZFD: SV40 NLS-ZFP(S162-left)-linker-DddAtox half-4aa linker-UGI ·Right-ZFD: SV40 NLS-ZFP(S162-right)-linker-DddAtox half-4aa linker-UGI - SV40 NLS: PKKKRKV (SEQ ID No: 478) - ZFP(S162-left) GIHGVPAAMAERPFQCRICMRNFSDRSNLSRHIRTHTGEKPFACDICGRKFAISSNLNSHTKIHTGSQKPFQCRICMRNFSRSDNLARHIRTHTGEKPFACDICGRKFATSGNLTRHTKIHLR (SEQ ID NO: 2) - ZFP(S162-right) GIHGVPAAMAERPFQCRICMRNFSRSDNLSVHIRTHTGEKPFACDICGRKFAQKINLQVHTKIHTGEKPFQCRICMRNFSRSDVLSEHIRTHTGEKPFACDICGRKFAQRNHRTTHTKIHLR (SEQ ID NO: 3)
[0315] - Linker between the zinc finger protein and the DddAtox half: - 2aa: GS - 5aa: TGEKP (SEQ ID No: 479) - 10aa: SGAQGSTLDF (SEQ ID No: 9) - 16aa: SGSETPGTSESATPES (SEQ ID No: 10) - 24aa: SGTPHEVGVYTLSGTPHEVGVYTL (SEQ ID No: 115) - 32aa: GSGGSGGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID No: 11) - Split-DddAtox G1333-N GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGG (SEQ ID No: 27) - Split-DddAtox G1333-C PTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPV KRGATGETKVFTGNSNSPKSPTKGGC (SEQ ID No: 272)
[0316] - Split-DddAtox G1397-N GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDN GISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEG (SEQ ID No: 273)
[0317] - Split-DddAtox G1397-C AIPVKRGATGETKVFTGNSNSPKSPTKGGC (SEQ ID No: 26) - 4aa linker SGGS (SEQ ID NO: 480) - UGI TNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWA LVIQDSNGENKIKML (SEQ ID NO: 481)
[0318] ZFDs with short linkers (2 and 5 amino acid (AA) linkers) showed low or no efficiency. On the other hand, ZFDs with linkers of 10 AA or more showed C-to-T base editing efficiencies ranging from 1% to 24% with spacers of 4 or more base pairs (Figure 1c and Figure 2a, c). Among ZFD pairs, those with a 24 AA linker showed the highest editing efficiency. To find the optimal linker combination, we fixed a 24 AA linker on the left ZFD of the ZFD pair and combined ZFDs with various linker lengths on the right side to confirm base editing efficiency, and then reversed the process (Figure 1d and Figure 3). We found that both 24 AA linkers were most efficient. Furthermore, we confirmed that DddAtox had better efficiency when split at the G1397 moiety than when split at the G1333 moiety (Figure 1c and Figure 2a, b). These most efficient ZFD pairs edit cytosins with high efficiency of >6.7% in spacers of 7 to 21 bp (Figure 1c and Figure 2a, c).
[0319] 1-3. Base editing in nuclear DNA targets in vivo Next, we investigated whether ZFDs with 24-AA linkers could catalyze C-to-T base editing at chromosomal target sites in vivo in human cells. Twenty-two ZFD pairs were constructed, targeting 11 sites (each consisting of two ZFD pairs) within eight genes (Figure 4). Fourteen ZFD pairs were obtained from publicly available zinc finger resources. The other eight ZFD pairs were constructed by modifying previously characterized ZFNs (specific to CCR5 and TRAC). While ZFNs typically require a spacer length of 5–7 bp to cleave target DNA, ZFDs function with a minimum spacer length of 7 or more bp. Therefore, we constructed ZFDs with one or two detachable zinc fingers in these ZFN pairs. Because FokI nuclease can be used to fuse portions to either the N- or C-terminal ends of ZFPs, allowing for the creation of ZFNs with four different configurations, we also created two pairs of ZFNs with other configurations (Trac-NC in Figure 4b) to test whether the split DddAtox halves could also be fused to the N-terminal end of ZFPs, as well as to the conventional ZFP C-terminus (the NC configurations are shown in Figures 4a and 5).
[0320] [structure] ·C type: SV40 NLS-Zinc finger protein-24aa linker-DddA tox half-4aa linker-UGI N type: SV40 NLS-DddA tox half-24aa linker-Zinc finger protein-4aa linker-UGI - SV40 NLS PKKKRKV (SEQ ID No: 478) - 24aa linker SGTPHEVGVYTLSGTPHEVGVYTL - Split-DddA tox G1397-N GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDN GISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEG - Split-DddA tox G1397-C AIPVKRGATGETKVFTGNSNSPKSPTKGGC - 4aa linker SGGS - UGI TNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWA LVIQDSNGENKIKML - ZFP CCR5-1 Left (C type) [S162 ZFN-Left] GIHGVPAAMAERPFQCRICMRNFS DRSNLSR HIRTHTGEKPFACDICGRKFA ISSNLNS HTKIHTGSQKPFQCRICMRNFS RSDNLAR HIRTHTGEKPFACDICGRKFA TSGNLTR HTKIHLR
[0321] CCR5-1 Right (C type) [S162 ZFN-Right] GIHGVPAAMAERPFQCRICMRNFS RSDNLSV HIRTHTGEKPFACDICGRKFA QKINLQV HTKIHTGEKPFQCRICMRNFS RSDVLSE HIRTHTGEKPFACDICGRKFA QRNHRTT HTKIHLR
[0322] CCR5-2 Left (C type) [S162 ZFN-Left] GIHGVPAAMAERPFQCRICMRNFS DRSNLSR HIRTHTGEKPFACDICGRKFA ISSNLNS HTKIHTGSQKPFQCRICMRNFS RSDNLAR HIRTHTGEKPFACDICGRKFA TSGNLTR HTKIHLR (SEQ ID NO: 2)
[0323] CCR5-2 Right (C type) [Modifying S162 ZFN-Right with additional ZF using the Barbas set of zinc finger modules] GIHGVPAAMAERPFQCRICMRNFS QSGDLRR HIRTHTGEKPFACDICGRKFA RSDNLSV HTKIHTGSQKPFQCRICMRNFS QKINLQV HIRTHTGEKPFACDICGRKFA RSDVLSE HTKIHLR (SEQ ID NO: 482)
[0324] TRAC-Left (C type) [Adapted from Paschon, D.E. et al., 2019] GIHGVPAAMAERPFQCRICMRNFS DQSNLRA HIRTHTGEKPFACDICGRKFA TSSNRK THTKIHTGSQKPFQCRICMRNFS LQQTLAD HIRTHTGEKPFACDICGRKFA QSGNLAR HTKIHLR (SEQ ID NO: 483)
[0325] TRAC-Left (N type) [Adapted from Paschon, D.E. et al., 2019] FQCRICMRKFA TSGSLTR HTKIHTGEKPFQCRICMRNFS RSDHLST HIRTHTGEKPFACDICGRKFA TSSNRTK HTKIHTHPRAPIPKPFQCRICMRNFS RSDNLSE HIRTHTGEKPFACDICGRKFA WHSSLRV HTKIHLR (SEQ ID NO: 484)
[0326] TRAC-Right (C type) [From Paschon, D.E. et al., 2019] GIHGVPAAMAERPFQCRICMRNFS RSDHLST HIRTHTGEKPFACDICGRKFA DRSHLAR HTKIHTGSQKPFQCRICMRKFA LKQHLNE HTKIHTGEKPFQCRICMRNFS QSGNLAR HIRTHTGEKPFACDICGRKFA HNSSLKD HTKIHLR (SEQ ID NO: 485)
[0327] MFAP1 Left (C type) [De novo designed using the Toolgen set of zinc finger modules] GIRIPGEKPYSCGICGKSFS DSSAKRR HCILHTGEKPYTCSDCGKAFR DKSCLNR HRRTHTGEKPYKCMECGKAFN RRSHLTR HQRIHTGEKPYECNYCGKTFS VSSTLIR HQRIHLR (SEQ ID NO: 486)
[0328] MFAP1 Right (C type) [Modifying S162 ZFN-Right with additional ZF using the Barbas set of zinc finger modules] GIRERPYACPVESCDRRFS TSGSLVR HIRIHTGQKPFQCRICMRNFS RSDELTR HIRTHTGEKPFACDICGRKFA RSDHLTT HTKIHTGEKPFQCRICMRKFA QSSNLVR HTKIHLR (SEQ ID NO: 487)
[0329] CCDC28B Left (C type) [Modifying S162 ZFN-Right with additional ZF using the Barbas set of zinc finger modules] GIRERPYACPVESCDRRFS DPGHLVR HIRIHTGQKPFQCRICMRNFS RSDELTR HIRTHTGEKPFACDICGRKFA RSDHLTT HTKIHTGEKPFQCRICMRKFA RSDKLVR HTKIHLR (SEQ ID NO: 488)
[0330] CCDC28B Right (C type) [De novo designed using the Toolgen set of zinc finger modules] GIRIPGEKPYECNYCGKTFS VSSTLIR HQRIHTGEKPYHCDWDGCGWKFA RSDELTR HYRKHTGEKPFECKDCGKAFI QKSNLIR HQRTHTGEKPYHCDWDGCGWKFA RSDELTR HYRKHLR (SEQ ID NO: 489)
[0331] KDM4B Left (C type) [Modifying S162 ZFN-Right with additional ZF using the Barbas set of zinc finger modules] GIRERPYACPVESCDRRFS DCRDLAR HIRIHTGQKPFQCRICMRNFS RSDELTR HIRTHTGEKPFACDICGRKFA RSDHLTT HTKIHTGEKPFQCRICMRKFA RSDKLVR HTKIHLR (SEQ ID NO: 490)
[0332] KDM4B Right (C type) [De novo designed using the Toolgen set of zinc finger modules] GIRIPGEKPFECKDCGKAFI QKSNLIR HQRTHTGEKPYRCEECGKAFR WPSNLTR HKRIHTGEKPYRCEECGKAFR WPSNLTR HKRIHTGEKPYSCGICGKSFS DSSAKRR HCILHLR (SEQ ID NO: 491)
[0333] NUMBL Left (C type) [Modifying S162 ZFN-Right with additional ZF using the Barbas set of zinc finger modules] GIRERPYACPVESCDRRFS DCRDLAR HIRIHTGQKPFQCRICMRNFS RSDELTR HIRTHTGEKPFACDICGRKFA RSDHLTT HTKIHTGEKPFQCRICMRKFA RSDKLVR HTKIHLR (SEQ ID NO: 492)
[0334] NUMBL Right (C type) [De novo designed using the Toolgen set of zinc finger modules] GIRIPGEKPYTCSDCGKAFR DKSCLNR HRRTHTGEKPYKCGQCGKFYS QVSHLTR HQKIHTGEKPFECKDCGKAFI QKSNLIR HQRTHTGEKPYHCDWDGCGWKFA RSDELTR HYRKHLR (SEQ ID NO: 493)
[0335] INPP5D-1 Left (C type) [Modifying S162 ZFN-Right with additional ZF using the Barbas set of zinc finger modules] GIRERPYACPVESCDRRFS RSDKLVR HIRIHTGQKPFQCRICMRNFS RSDELTR HIRTHTGEKPFACDICGRKFA RSDHLTT HTKIHTGEKPFQCRICMRKFA RSDKLVR HTKIHLR (SEQ ID NO: 494)
[0336] INPP5D-1 Right (C type) [De novo designed using the Toolgen set of zinc finger modules] GIRIPGEKPYTCSDCGKAFR DKSCLNR HRRTHTGEKPYECNYCGKTFS VSSTLIR HQRIHTGEKPYTCSDCGKAFR DKSCLNR HRRTHTGEKPYTCSDCGKAFR DKSCLNR HRRTHLR (SEQ ID NO: 495)
[0337] INPP5D-2 Left (C type) [Modifying S162 ZFN-Right with additional ZF using the Barbas set of zinc finger modules] GIRERPYACPVESCDRRFS RSDKLVR HIRIHTGQKPFQCRICMRNFS RSDELTR HIRTHTGEKPFACDICGRKFA RSDHLTT HTKIHTGEKPFQCRICMRKFA RSDKLVR HTKIHLR (SEQ ID NO: 496)
[0338] INPP5D-2 Right (C type) [De novo designed using the Toolgen set of zinc finger modules] GIRIPGEKPYECNYCGKTFS VSSTLIR HQRIHTGEKPYTCSDCGKAFR DKSCLNR HRRTHTGEKPYTCSDCGKAFR DKSCLNR HRRTHTGEKPYHCDWDGCGWKFA RSDELTR HYRKH (SEQ ID NO: 497)
[0339] DVL3 Left (C type) [De novo designed using Barbas zinc finger modules] GIHGVPAAMAERPFQCRICMRNFS TSGHLVR HIRTHTGEKPFACDICGRKFA TSGHLVR HTKIHTGEKPFQCRICMRNFS TSGELVR HIRTHTGEKPFACDICGRKFA QSSNLVR HTKIHLR (SEQ ID NO: 498)
[0340] DVL3 Right (C type) [S162 ZFN-left] GIHGVPAAMAERPFQCRICMRNFS DRSNLSR HIRTHTGEKPFACDICGRKFA ISSNLNS HTKIHTGSQKPFQCRICMRNFS RSDNLAR HIRTHTGEKPFACDICGRKFA TSGNLTR HTKIHLR (SEQ ID NO: 499)
[0341] In HEK293T cells, the C-to-T base editing efficiency of ZFDs, including ZFDs with NC constructs, ranged from 1.0% to 60%. Meanwhile, insertion-deletion was rare, occurring at <0.4% (Figures 4b and 6). As seen in the plasmid-based experiments, ZFDs targeting CCR5 with a 5-bp spacer showed significantly lower efficiency. The other 20 ZFD pairs targeting targets with a spacer length of at least 7 showed an average editing efficiency of 12.0±3.4%, comparable to the average of 8.3±2.2% achieved by a Cas9-based base editor (Cas9-derived base editor 2). Furthermore, C-to-T base editing was observed not only in TC contexts but also in A contexts. C and GC C This also occurred (Fig. 4c-f). C is NUMBL's C6, GC C showed C base editing efficiencies of 4.58% and 1.85%, respectively, at C7 of INPP5D-2.
[0342] JPEG0007796729000011.jpg249170
[0343] 1-4. Direct delivery of purified ZFD proteins into human cells Instead of delivering plasmid DNA encoding the gene-editing protein, delivering purified gene-editing proteins directly into cells reduces off-target effects, avoids the innate immune response induced by foreign DNA, and prevents foreign plasmid DNA from integrating into the genome in vivo. Other groups have demonstrated that ZFPs can spontaneously translocate and enter mammalian cells in vitro and in vivo. To demonstrate ZFD protein-mediated base editing, we selected ZFD pairs targeting the TRAC gene, which showed high efficiency, and purified ZFD recombinant proteins with one or four NLSs from E. coli. We first performed in vitro experiments using PCR amplicons containing the TRAC site to confirm the highly efficient base editing efficiency of the ZFD proteins. The efficiency was confirmed by gene cleavage using a uracil-specific cleavage reagent (USER), a mixture of Uracil DNA glycosylase and DNA glycosylase-lyase EndoNuclease VIII (Figure 7). We transfected TRAC-NC ZFD protein into difficult-to-transfect human leukemia K562 cells by two methods: electroporation and direct delivery without electroporation. The ZFD protein was highly efficient, demonstrating C-to-T base editing rates of 26.5% (electroporation) and 17% (direct delivery) (Figure 4g). Collectively, these results demonstrate that ZFD-encoding plasmids or purified recombinant ZFD protein can be used to base edit nuclear DNA in human cells.
[0344] 1-5.ZFD mitochondrial DNA base editing Unlike CRISPR-based systems, a major advantage of the system that fuses split DddAtox to custom-designed DNA-binding proteins is that these programmable base-editing scissors can be used to edit organelle DNA, such as mitochondrial DNA. To deliver ZFDs to mitochondria, we constructed mitoZFDs by linking MTS and NES to the N-terminal portions of nine ZFDs designed to target mitochondrial genes (Figure 8). The ZFP portions of the ZFDs were obtained from publicly available Zing finger resources. The ZFDs were engineered with spacer lengths of 7–15 bp, with 12 bp on both the left and right sides of the DNA-binding site.
[0345] [structure] ·C type: MTS-FLAG tag-NES-Zinc finger protein-24aa linker-DddA tox half-4aa linker-UGI N type: MTS-HA tag-NES-DddA tox half-24aa linker-Zinc finger protein-4aa linker-UGI - MTS (Mitochondrial Targeting Sequence of human mitochondrial ATP synthase F1β subunit) MLGFVGRVAAAPASGALRRLTPSASLPPAQLLLRAAPTAVHPVRDYAAQ (SEQ ID No:274) - FLAG tag (C type) DYKDDDDK (SEQ ID No:275) - HA tag (N type) YPYDVPDYA (SEQ ID No:276)
[0346] - NES (Nuclear export signal) VDEMTKKF (Minute virus if mice; MVM NES) - Split-DddA tox G1397-N GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDN GISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEG - Split-DddA tox G1397-C AIPVKRGATGETKVFTGNSNSPKSPTKGGC - 4aa linker SGGS - UGI TNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWA LVIQDSNGENKIKML
[0347] - ZFP * The zinc fingers are linked by a TGEKP linker (ZF1-linker-ZF2-linker-ZF3-linker-ZF4). * The ND2-targeted mitoZFD has a QQ mutant, in which an R or K amino acid is replaced with a Q, shown in red. JPEG0007796729000012.jpg246170
[0348] JPEG0007796729000013.jpg209170
[0349] In HEK293T cells, the mitochondrial DNA base editing efficiency of mitoZFD ranged from 2.6% to 30% (average 14 ± 3%) (Figure 9a). The efficiency of mitoZFD with the NC configuration (16 ± 4.5%, n = 6) was higher than that with the CC configuration (8.3 ± 2.7%, n = 3). CMost of the context-containing cytosins were base-edited with varying efficiencies (Figure 9b-g). CC Cytosins with the (C8 and C9) contexts also showed efficiencies of 7.4% and 19.9%, respectively (Figure (Figure9b). 9b), indicating that ZFD-mediated C-to-T base editing is not limited to the TC motif.
[0350] Next, we isolated single-cell derived clonal populations from mitochondrial DNA (mtDNA) mutant cells to demonstrate that mitoZFD is nontoxic and that mutant mtDNA persists in these clonal populations. Of 30 single-cell derived clonal populations isolated from HEK293T cells treated with ND1-specific mitoZFD, five showed base editing efficiencies in the ND1 gene ranging from 35% to 98% (Fig. 10a). Similarly, of 36 single-cell derived clonal populations isolated from HEK293T cells treated with ND2-specific mitoZFD, seven showed base editing efficiencies in the ND2 gene ranging from 26% to 76% (Fig. 10b). All other clonal populations showed low efficiencies ranging from 0.4% to 1.0%, most of which appeared to be sequence errors. Similar efficiencies were observed in cells not treated with ZFD (Fig. 10c). These results indicate that mitoZFD does not uniformly induce heteroplasmic mutations in the cell population. ZFD-treated cells were mostly wild-type, whereas those with heteroplasmic mutations had mutation rates of up to 98%, a change that persisted even after clonal expansion (Figures 11 and 12).
[0351] 1-6. mitoZFDs and TALE-based DdCBEs The mutation patterns of the constructed ND1-specific mitoZFD were found to be different from those of the TALE-based DdCBE targeting the same gene (Fig. 9f-h). In the case of the two mitoZFDs, base editing was performed from C to T at the cytosine at C5 or C8 (Fig. 9f), whereas DdCBE only edited C at C8, C9, and C. 11 The site is then base-edited (Figure 9g). As a result, the amino acid changes induced by mitoZFD are completely different from those induced by DdCBE (Figure 9h). The left and right sides of the mitoZFD attachment site are separated by an 8-bp spacer, whereas in the case of DdCBE, the spacer length is separated by a 16-bp spacer. This difference can result in different mutation patterns. These results demonstrate that mitoZFD and DdCBE can complement each other to generate various mutations in mitochondrial DNA.
[0352] To obtain a wider range of mutation patterns, we tested whether single ZFD and single DdCBE hybrids could be used as hybrid pairs. Ten hybrid pairs targeting the ND1 gene exhibited good activity in HEK293T cells, with an average base editing efficiency of 17 ± 3.4% (Figure 13). Indeed, one hybrid pair (TALE-L / ZFD-R1) outperformed two DdCBE pairs and ten ZFD pairs targeting the same site, achieving a maximum base editing efficiency of 41% (Figure 13b). Furthermore, the hybrid pairs exhibited mutation patterns distinct from those of DdCBE and mitoZFD (Figure 13c). Some (e.g., ZFD-L1 / TALE-R and ZFD-L2 / TALE-R) induced a single C-to-T mutation without any other concomitant mutations. In contrast, most mitoZFD and DdCBE pairs induced C-to-T changes at various locations within the spacer. These results indicate that the ZFD / DdCBE mixed pair can generate specific mutation patterns and produce specific mutations that cannot be produced by the ZFD pair or the DdCBE pair.
[0353] 1-7. Mitochondrial genome-wide targeting specificity of mitoZFD To confirm that mitoZFD induces off-target activity in mitochondrial DNA, we extracted genes from cells treated with two pairs of mitoZFD targeting the ND1 or ND2 gene, and performed whole mitochondrial genome sequencing. HEK293T cells were transfected with various amounts (5–500 ng) of mRNA or plasmid encoding the mitoZFD pair. As expected, the on-target efficiency was dose-dependent. High concentrations (100, 200, and 500 ng) of mRNA or plasmid demonstrated on-target efficiency of >30%, but also generated hundreds of off-targets exceeding 1% (Figures 14–17). Low concentrations (5 and 10 ng) almost eliminated these off-targets, but the on-target efficiency was also significantly lower. An intermediate concentration of 50 ng of mRNA was most appropriate. High on-target efficiency was maintained without generating hundreds of off-targets. To eliminate remaining off-targets, we introduced the R(-5)Q mutation into each zinc finger to eliminate nonspecific DNA contacts. The resulting ZFD mutant (shown as QQ in Fig. 18) maintained high on-target activity while exhibiting high specificity with almost no off-target activity compared to mtDNA in cells without ZFD treatment (Fig. 18a and b).
[0354] Base editing is a relatively new technology that can edit target bases without DNA repair templates and without causing double-strand breaks in DNA. Base editing can result in C-to-T or A-to-G base edits in cells, animals, and plants, allowing for the study of the functional effects of single-nucleotide polymorphisms (SNPs) and the repair of disease-causing point mutations for therapeutic applications. Two base editing technologies have been developed: CRISPR-based adenine and cytosine base editing scissors and DddA-based base editing. CRISPR-based base editing scissors contain catalytically impaired Cas9 or Cas12a as the DNA-binding unit and single-strand DNA-specific deaminases derived from rat or E. coli. Meanwhile, DdCBE, like the TALE DNA-binding array, contains double-strand DNA-specific DddA.
[0355] Compared to DdCBEs, ZFDs are smaller in size because the zinc finger proteins contained in ZFDs are small, while the TALE array contained in DdCBEs is large. As a result, genes encoding ZFD pairs can be easily delivered to AAV vectors, which have a small loading space, whereas genes encoding DdCBE pairs are not. Furthermore, the simple characteristics of ZFPs make them easy to engineer. The DddAtox half can be fused to the N- or C-terminal portion of ZFPs, allowing them to function upstream or downstream of the ZFP binding site. Furthermore, ZFD recombinant proteins spontaneously penetrate human cells without electroporation or lipofection, enabling gene-free gene therapy. ZFD pairs or ZFD / DdCBE mixed pairs can generate unique mutation patterns, which is not possible using DdCBEs alone. These properties make ZFDs a powerful new platform for modeling and treating mitochondrial diseases.
[0356] Example 2: TALE-DdCBE plant chloroplast and mitochondrial gene editing Plant organelles, including mitochondria and chloroplasts, each have their own genomes encoding many genes essential for respiration and photosynthesis. Gene editing in plant organelles, an unmet need in plant genetics and biotechnology, has been limited by the lack of suitable tools for targeting DNA in these organelles. To assemble DddA-derived cytosine base editing plasmids (DdCBEs), we developed a Golden Gate cloning system consisting of 16 expression plasmids (eight for delivering proteins to mitochondria and eight for delivering proteins to chloroplasts) and a 424-TALE subarray plasmid. We then used the completed DdCBE plasmid to efficiently create enhanced point mutations in mitochondria and chloroplasts. DdCBE base editing induced efficiencies of up to 25% (mitochondria) and 38% (chloroplasts) in lettuce or rapeseed. To avoid non-targeted mutations that could occur in the DdCBE-encoding plasmid, we transferred DdCBE messenger RNA into lettuce protoplasts and demonstrated base editing in chloroplasts without DNA. Furthermore, point mutations in the chloroplast 16S rRNA gene were introduced to generate lettuce and seedlings with up to 99% editing efficiency that were resistant to streptomycin or spectinomycin.
[0357] DdCBEs are heterodimers composed of an isolated non-toxic domain derived from the bacterial cytosine deaminase toxicant DddAtox, a TALE array designed for specific sites, and a UGI. They function as a cytosine deamination reaction, replacing cytosine with thymine in the spacer space between the TALE protein binding sites in target DNA. We rapidly and conveniently demonstrated the integration of DdCBE plasmids expressed in mitochondria and chloroplasts, and used the results of DdCBEs to demonstrate highly efficient organelle base editing in plants.
[0358] 2-1. Method Construction of expression plasmids in plant protoplasts The DdCBE Golden Gate Destination Vector was constructed using the Gibson assembly method. Sequences encoding the TAL N-terminal domain, HA tag, FLAG tag, TAL C-terminal domain, split DddAtox, and UGI were codon-optimized for dicotyledonous (Arabidopsis) expression and synthesized by Integrated DNA Technology. Sequences encoding the CTP from AtinfA and AtRbcS, and the MTS from the ATPase delta subunit and ATPase gamma subunit were amplified from Arabidopsis cDNA. For plant expression, the mammalian CMV promoter was replaced with the PcUbi promoter and pea3A terminator in the backbone plasmid. To construct a vector for in vitro DdCBE mRNA transcription, a T7 promoter cassette was introduced between the PcUbi promoter and the DdCBE-encoding region in the DdCBE Golden Gate Destination Vector.
[0359] The TALE array genes were constructed using a single Golden Gate assembly. The DdCBE expression plasmid was constructed using the 424 TALE array plasmids and the destination vector via Golden Gate assembly by BsaI digestion and T4 ligation. Single Golden Gate cloning was performed as follows: 20 cycles of 5 min at 37°C and 50°C, followed by a final 15 min at 50°C and 5 min at 80°C. All plant protoplast transduction vectors were purified using the Plasmid Plasmid Flap Kit (Qiagen). The DNA and amino acid sequences of the vectors are as follows:
[0360] [Table 2]
[0361] The specific amino acid sequences for the DdCBE construct and the specific amino acid sequences for the TALE repeats are as follows: [Table 3] JPEG0007796729000016.jpg247170JPEG0007796729000017.jpg246170JPEG00077967290 00018.jpg241170JPEG0007796729000019.jpg244170JPEG0007796729000020.jpg200170
[0362] mRNA in vitro transcription The DdCBE DNA template was prepared by PCR using Fusion DNA polymerase (Thermo Fisher). DdCBE mRNA was synthesized and purified using an in vitro mRNA synthesis kit (Enzynomics).
[0363] Protoplast extraction and transduction Lettuce seeds were surface-sterilized in 70% ethanol for 30 seconds and 0.4% Lux solution for 15 minutes, then rinsed three times with sterile distilled water. Lettuce seeds were sown on 0.5x MS medium and 2% sucrose at 25°C under 16 hours of light and 8 hours of darkness. Rapeseed seeds were surface-sterilized in 70% ethanol for 3 minutes and 1.0% Lux solution for 30 minutes, then rinsed three times with sterile distilled water. Rapeseed seeds were sown on 1x MS medium and 3% sucrose at 25°C under 16 hours of light and 8 hours of darkness.
[0364] Protoplast extraction and transduction were performed according to previous studies. Waxy leaves from 7-day-old lettuce and 14-day-old rapeseed were treated with enzyme solution for 3 hours under dark conditions with shaking (40 rpm). The protoplast-enzyme mixture was washed with an equal volume of W5 solution, and intact protoplasts were isolated from the sucrose solution by centrifugation at 80 g for 7 minutes. The protoplasts were treated with W5 solution for 1 hour at 4°C before being transduced with polyethylene glycol.
[0365] Lettuce and rapeseed protoplasts were resuspended in MMG solution and then transduced with plasmid or mRNA using PEG. The mixture was then incubated at room temperature for 20 minutes. The PEG-protoplast mixture was rinsed three times with the same volume of W5 solution, gently inverting each time, and then incubated for 10 minutes. The protoplasts were then pelleted by centrifugation at 100 g for 5 minutes.
[0366] Protoplast culture Lettuce protoplasts were transfected with a plasmid encoding DdCBE and then resuspended in lettuce protoplast culture medium (LPCM). The protoplasts were mixed 1:1 with medium containing 2.4% low-melting agarose and immediately plated in a 6-well dish. Once the mixture solidified, 1 ml of liquid medium was placed on the embedded protoplasts and cultured at 25°C in the dark for one week. After the initial culture, the liquid medium was replaced with fresh medium weekly. The embedded protoplasts were cultured under 16 hours of low light and 8 hours of darkness for one week, followed by 16 hours of light and 8 hours of darkness for two weeks. Small calli derived from the protoplasts were cultured in regeneration medium for four weeks at 25°C under 16 hours of light and 8 hours of darkness. For base editing efficiency analysis, protoplasts were cultured in liquid medium without embedding for one week at 25°C in the dark. To test for antibiotic resistance, small calli embedded for one month were cultured for four weeks at 25°C under 16 hours of light and 8 hours of darkness on regeneration medium containing 50 mg / L streptomycin or 50 mg / L spectinomycin. After four weeks, antibiotic-resistant green calli or adventitious shoots were transferred to new regeneration medium containing 200 mg / L streptomycin or 50 mg / L spectinomycin.
[0367] Rapeseed protoplasts transfected with a plasmid encoding DdCBE were resuspended in rapeseed culture medium. The protoplast-medium mixture was transferred to a 6-well dish and cultured at 25°C in the dark for two weeks. After two weeks, the protoplasts were cultured under a 16-hour dim light / 8-hour dark condition for three weeks. The medium was then replaced with fresh medium.
[0368] DNA and RNA extraction Total DNA or RNA from cells cultured in liquid medium or transformed calli was extracted using the Jien-Iji Plant Mini Kit or the Al-Iji Plant Mini Kit. Cultured cells or calli were harvested by centrifugation at 10,000 rpm for 1 minute. cDNA from total RNA was reverse transcribed using RNA to cDNA EcoDry Premix (TAKARA).
[0369] Deep sequencing The target region was amplified using the fusion enzyme and appropriate primers (Table 1). Three rounds of PCR (first, nested PCR, second, PCR, and third, indexing PCR) were performed to generate a DNA sequence analysis library. Equal amounts of DNA were collected and sequenced using the MiniSeq (Illumina) instrument. Paired sequence files were analyzed using the Cas-analyzer software and the source code of the computer program.
[0370] 2-2.Results We developed the Golden Gate assembly system to generate chloroplast-targeted DdCBEs (cp-DdCBEs) and mitochondria-targeted DdCBEs (mt-DdCBEs) (Figure 19). The expression plasmids encode the chloroplast transfer peptide or mitochondrial targeting sequence, the N- or C-terminal domain of the TALE, the split DddAtox (G1333N, G1333C, G1397N, and G1397C), and UGI, which are linked together in a single protein. These plasmids are optimized for expression in dicotyledonous plants and are regulated by the parsley ubiquitin (PcUbi) promoter and pea3A terminator. The customized TALE DNA-binding sequences and DdCBE plasmids can be combined with the expression vector and six TALE subarray plasmids in a single subcloning step. A total of 424 modular TALE subarrays (6x64 triplet recognition + 2x16 triplet recognition + 2x41 triplet recognition) can be constructed using the modular TALE subarray plasmids, which can recognize sequences of 16-20 base pairs containing a conserved T at the 5' end. As a result, functional DdCBE heterodimers recognize 32- to 40-bp DNA sequences.
[0371] To determine whether DdCBEs could promote chloroplast base editing, we constructed four pairs of cp-DdCBE plasmids targeting the chloroplast 16S rRNA gene, which encodes the RNA component of the 30S ribosomal subunit. We co-transfected each pair into lettuce and rapeseed protoplasts, and then measured the base editing efficiency by deep sequencing after 7 days (Figure 20a, b). The most efficient cp-DdCBE pair induced a C·G to T·A substitution in the 15-bp spacer region between the two TALE binding sites (Left-G1397-N + Right-G1397-C) with 30% efficiency in lettuce protoplasts and 15% efficiency in rapeseed protoplasts (Figure 20b). Similar to previous results in mammalian cells and mice, cp-DdCBE preferentially replaced cytosines (C9 and C13) in the 5'-TC motif with thymine. Interestingly, the 5'-AC motif, cytosine (C7), was replaced with thymine by another cp-DdCBE (Left-G1333-N + Right-G1333-C) in lettuce protoplasts with an efficiency of 4.2%. Furthermore, we investigated whether base editing by cp-DdCBE persisted in lettuce protoplasts over a 14-day culture period (Figure 24). Editing efficiency increased continuously up to 10 days and was maintained throughout the culture period.
[0372] Two additional experiments were performed on the chloroplast genes psbA and psbB, which encode the D1 and CP-47 photosynthetic proteins, respectively, in photosystem II (Figures 20c, 20d and 25). The most efficient cp-DdCBE targeting the Ps gene (Left-G1397-C+Right-G1397-N) induced up to 25% C·G to T·A substitutions in lettuce protoplasts (Figure 20d). This base editing scissors efficiently replaced only two cytosines (C11, C12) of 5'-TCC with thymines. This likely occurred after 5'-TCC was first replaced by 5'-TTC, followed by 5'-TTT. In rapeseed protoplasts, another combination (Left-G1333-N + Right-G1333-C) showed the highest efficiency at four cytosine positions (C3, C4, C11, and C12), with an efficiency of up to 3.5% (C3). While C3 and C4 have 5'-TCC sequences in rapeseed genes, lettuce has a 5'-ACC sequence due to a relatively small single nucleotide polymorphism. This allows DdCBE to efficiently edit two cytosines (C3 and C4) in rapeseed genes, but not in lettuce genes. Furthermore, the cp-DdCBE combination targeting the psbB gene in rapeseed protoplasts replaced two cytosines with TCC sequences with efficiencies ranging from 0.36% to 4.1% (Figure 25). Collectively, these results demonstrate that editing efficiency is determined by the cytosine position and sequence, including the split position (G1333 vs. G1397) and orientation (left-G1333-N vs. left-G1333-C) of DddAtox, and that cp-DdCBE is capable of efficient base editing of plant chloroplast genes.
[0373] Next, we attempted to achieve base editing of plant mitochondrial DNA using a customized mt-DdCBE. To do so, we constructed plasmids (using the Golden Gate cloning system) encoding mt-DdCBEs targeting the atp6 gene in lettuce and rapeseed, and the rps14 gene in rapeseed. We then transformed the plasmids into lettuce and rapeseed protoplasts, and measured base editing efficiency by deep sequencing 7 days after transfection (Figures 20e, f, and 26). The most efficient mt-DdCBE combinations (Left-G1397-N + Right-G1397-C in lettuce and Left-G1397-C + Right-G1397-N in rapeseed) resulted in a C·G to T·A substitution in 23% of lettuce protoplasts and 23% of rapeseed protoplasts at the atp6 gene target position (Figure 20). Furthermore, the mt-DdCBE fusion induced 11% C-G to T-A substitutions at the rps14 target site in rapeseed protoplasts. These results demonstrate that mt-DdCBE is effective in base editing plant mitochondrial DNA.
[0374] To investigate whether DdCBE-induced cpDNA and mtDNA editing is maintained during regeneration, lettuce and rapeseed calli redifferentiated from DdCBE-treated protoplasts were collected 4 weeks after injection (Figure 21a). The base editing efficiency of each calli was measured using deep sequencing and Sanger sequencing (Figure 21b, Figure 27). DdCBE-induced base editing of chloroplast or mitochondrial genes was measured at efficiencies of up to 38% and 25% in 22 of 26 lettuce calli and 7 of 14 rapeseed calli, respectively (Figure 21c). Furthermore, base editing of the chloroplast psbA gene showed efficiencies of up to 3.9% in lettuce calli (Figure 27). Furthermore, mitochondrial base editing in rapeseed calli was measured at efficiencies of up to 25% and 1.9% in atp6 and rps14, respectively (Figure 27). These results indicate that plant protoplasts can tolerate DdCBE expression and induce base editing in organelles during protoplast regeneration.
[0375] Next, we attempted to demonstrate DNA-free base editing in organelles using in vitro transcribed cp-DdCBE mRNA instead of a plasmid. After injecting in vitro transcripts encoding cp-DdCBE targeting the 16S rRNA gene in lettuce, we analyzed the base editing efficiency at the target position (Figure 21a). C-to-T mutations in the protoplasts were measured at up to 25% (Figure 21d, Figure 28). As expected, these mutations disappeared 7 days after introducing DdCBE mRNA and DNA sequences into the protoplasts (Figure 29). This method avoids potential integration of portions of the plasmid DNA into the host genome.
[0376] Using the stable maintenance of organelle editing in calli regenerated from protoplasts, we investigated resistance to streptomycin and spectinomycin, antibiotics that irreversibly bind to 16S rRNA and inhibit protein synthesis, through editing of the 16S rRNA gene in chloroplast DNA. Single nucleotide polymorphisms at several sites in the 16S rRNA gene commonly confer streptomycin resistance in prokaryotes and eukaryotes. In particular, the 16S rRNA C860T mutation (position C912 in E. coli) confers streptomycin resistance in tobacco. The C860T point mutation in tobacco is identical to the C9 mutation in lettuce (Figures 20a, 21b, 22b, and 22d). Lettuce calli regenerated from DdCBE-treated protoplasts were transferred to medium supplemented with streptomycin and spectinomycin. Mock-treated calli turned white upon exposure to antibiotics, indicating impaired callus protoplast function. On the other hand, DdCBE-treated calli remained green, demonstrating resistance to these antibiotics. We analyzed the editing efficiency of DdCBE in resistant lettuce and seedlings. A T substitution at the C9 position, such as the C860T mutation, was observed in up to 98.6% of calli and shoots obtained from drug treatment (Figure 21e, f). Interestingly, T editing at the nearby C13 position showed an efficiency of up to 20% in the absence of specniomycin but not in the presence of antibiotics, confirming that this mutation was selected by drug treatment. Collectively, these results suggest that plant organelle mutations induced by DdCBE in protoplasts can be maintained after cell division and plant development, and that chloroplast editing can achieve homoplasy through drug selection.
[0377] Furthermore, we analyzed the off-targets of TALE diaminase targeting 16S rRNA regions in protoplasts, calli, and soats. Five potential off-target locations were selected based on sequence similarity in the target site (50 base pairs on either side) (Figure 31) or in chloroplast genes in single-cell-derived calli and soats (Figure 32). No off-targets were observed. On the other hand, when a DdCBE-encoding plasmid was introduced into protoplasts, off-targets resulting in TC to TT changes were induced at low efficiencies (1.2% to 4.1%) at three of the five potential off-target locations. When in vitro transcripts (mRNA) were used instead of the TALE diaminase-encoding plasmid, the off-target efficiency in protoplasts was significantly reduced (Figure 22). These results suggest that overexpression or long-term plasmid-derived expression of DdCBE may increase off-target mutations, while transient mRNA-derived expression using mRNA is preferred to avoid off-target base editing.
[0378] In summary, we developed the Golden Gate cloning system using a 424-TALE subarray plasmid and 16 expression plasmids to assemble plasmids encoding DdCBEs for organelle base editing in plants. Customized DdCBEs targeting three chloroplast DNA genes and two mitochondrial DNA genes achieved highly efficient C-to-T substitutions in lettuce and rapeseed protoplasts. Notably, plant organelle editing was maintained throughout cell division and plant development. Furthermore, mutation of the chloroplast 16S rRNA gene resulted in nearly homozygous (99%) antibiotic-resistant lettuce calli and plants. Without antibiotic selection, editing efficiencies of 25% in mitochondria and 38% in chloroplasts were observed. We anticipate that the Golden Gate cloning system will be a valuable resource for plant organelle DNA editing.
[0379] Example 3. TALE?DdCBE Animal DNA Editing DddA-derived cytosine base editors (DdCBEs), containing a split bacterial toxin, DddAtox, and a TALE and UGI designed to enable binding to DNA, enabled targeted cytosine-thymine base modifications in mitochondrial DNA. Highly efficient mitochondrial DNA editing was demonstrated in mouse embryos. The mitochondrial gene MT-ND5 (ND5), encoding a subunit of NADH dehydrogenase, which catalyzes NADH dehydration and electron transfer to ubiquinone, was targeted. This gene contained mutations associated with human mitochondrial disease, such as m.G12918A, and mutations that generate premature stop codons, such as m.C12336T. This enabled the generation of mitochondrial disease models in mice, highlighting the potential for the treatment of mitochondrial diseases.
[0380] 3-1. Method Plasmid assembly. The TALEN (transcription activator-like effector nuclease) system was used to generate an expression plasmid containing the DddA half and the final TALE-DddAtox construct. In the TALEN system expression plasmid, the monomers in the nuclear import signal and FokI dimer were replaced with the mitochondrial transduction signal (MTS), the DddA deaminase half, and the uracil glycosylation inhibitor (UGI). The sequences encoding MTS, DddA, and UGI were synthesized by IDT. To generate the expression vector, the DNA fragments required for Gibson assembly were amplified using Q5 DNA polymerase (NEB) and subsequently purified. The purified gene fragments were assembled using the HiFi DNA assembly kit (NEB), chemically transformed into E. coli DH5a (Enzynomics), and verified by Sanger sequencing. Thus, eight different expression plasmids were obtained, each containing a BsaI restriction enzyme site for Golden Gate cloning between the sequences encoding the N- and C-terminal domains. To assemble the DdCBE plasmids, the expression plasmids were combined with the module vectors (each encoding a TALE sequence), BsaI-HFv2 (10 U), T4 DNA ligase (200 U), and reaction buffer in a single tube. The restriction enzyme and ligation reactions were then cycled 20 times in a thermocycler at 37°C for 5 minutes and 50°C for 5 minutes, followed by 50°C for 15 minutes and 80°C for 5 minutes. The ligated plasmids were then introduced into E. coli DH5a via chemical transformation, and the final constructs were confirmed by Sanger sequencing. For cell line transfection, the plasmids were midiprepped.
[0381] Mammalian cell line culture and transfection. NIH3T3 (CRL-1658, American Type Culture Collection (ATCC)) cell line was cultured at 37°C in a 5% CO2 environment. Cell lines were grown in 10% (v / v) elegant serum-enriched DMEM (Gibco) medium without antibiotics and were not mycoplasma tested. For lipofection, cells were plated at 1.5 × 10 in 12-well cell culture plates (SPL, Seoul, Korea). 4 Cells were grown at a density of 18–24 h before transfection. A total of 1,000 ng of plasmid DNA was transfected using 500 ng of each DdCBE aliquot using Lipofectamine 3000 (Invitrogen). Cells were harvested 4 days after transfection.
[0382] Preparation of mRNA. The mRNA template was PCR amplified using Q5 DNA polymerase (NEB) with the following primers (F: 5′-CATCAA TGGGCGTGGATAG-3′ SEQ ID No. 268, R: 5′-GACACCTACTCAGACAATGC-3 SEQ ID No. 269). DdCBE mRNA was synthesized using an in vitro RNA transcription kit (mMESSAGE mMACHINE T7 Ultra kit, Ambion) and purified using the MEGAclear kit (Ambion).
[0383] Animals. All experiments involving rats were performed with the approval of the Animal Care and Use Committee of the Graduate School of Basic Science. Superovulated C57BL / 6J females were mated with C57BL / 6J males, and ICR strain females were used as surrogate mothers. Mice were housed in a specific pathogen-free facility under conditions maintaining a 12-h day / night cycle, constant temperature, and humidity (20–26°C, 40–60%).
[0384] Microinjection into mouse conjugates. The steps immediately preceding microinjection, including superovulation, embryo collection, and microinjection, were performed as described in our previous paper. For microinjection, a mixture containing left DdCBE mRNA (300 ng / μl) and right DdCBE mRNA (300 ng / μl) was diluted in DEPC-treated injection buffer (0.25 mM EDTA, 10 mM Tris, pH 7.4) and injected into the conjugate cytoplasm using a Nikon ECLIPSE Ti micromanipulator and a FemtoJet 4i microinjector (Eppendorf). After microinjection, embryos were placed in KSOM + AA (Millipore) microdroplets and cultured at 37°C and 5% CO2 for 4 days. Two-cell embryos were transferred into the oviducts of 0.5-dpc pseudo-pregnant surrogate mothers.
[0385] Genotyping. Blastocyst-stage embryos and tissues were placed in digestion buffer (25 mM NaOH, 0.2 mM EDTA, pH 10) and incubated at 95°C for 20 minutes. The pH was then adjusted to 7.4 with HEPES (free acids, no pH adjustment) to a final concentration of 50 mM. Genotypic DNA from mouse pups was isolated using DNeasy Blood & Tissue Kits (Qiagen) and analyzed by Sanger sequencing and targeted deep sequencing.
[0386] Mitochondrial DNA isolation for high-throughput sequencing analysis. To isolate mitochondria from NIH3T3 cells cultured in 12-well plates, the cell culture medium was removed and 200 μl of Mitochondrial Isolation Buffer A (ScienCell) was added to the culture plate. The cells were scraped using a cell lifter, placed in a microtube, and crushed using a disposable pestle. After 15 rounds of trituration, the homogenate was centrifuged at 1,000 × g for 5 minutes at 4°C. The supernatant was placed in a new microtube and centrifuged at 10,000 × g for 20 minutes at 4°C. The precipitate was released into 20 μl of lysis solution (25 mM NaOH, 0.2 mM EDTA, pH 10) and then boiled at 95°C for 20 minutes. To lower the pH, 2 μl of 1 M HEPES (free acids, without pH adjustment) was added to the mitochondrial lysate. 1 μl of the resulting solution was used as a PCR template strand for high-throughput sequencing analysis.
[0387] High-throughput sequencing analysis. To generate deep sequencing libraries, nested primary and secondary PCRs were performed using Q5 DNA polymerase to add final index sequences. Libraries were used for fair-end sequence analysis using MiniSeq (Illumina). For whole mitochondrial genome analysis, isolated mitochondrial DNA was prepared using the tagmentation DNA prep kit (Illumina) according to the manufacturer's protocol. For all analyses, fair-end sequence results were combined into a single fastqjoin file and analyzed using CRISPR RGEN Tools (http: / / www.rgenome.net / ).
[0388] Data analysis and presentation. Microsoft Excel (2019) and PowerPoint (2019) were used to create pictures, graphs, and tables. Geneious (version 2021.0.1) and Snapgene 5.2.3 were used to design genome sequence alignments, primers, and cloning, and NC_005089 was used as the reference sequence.
[0389] 3-2.Results Assembly of DdCBE Plasmids. To facilitate assembly of customized TALE sequences in DdCBE, expression plasmids encoding the split halves of DddAtox were constructed using the Golden Gate cloning system, with a total of 424 sequences (6 × 64 triple sequences + 2 × 16 double sequences + 2 × 4 single sequences) (Figure 33a). As shown in Table 4 below, the six TALE plasmids and the expression plasmid were mixed in the same tube to create ready-to-use DdCBE plasmids with 15.5 to 18.5 repeat variable residue sequences (Figure 38).
[0390] [Table 4]
[0391] The sequences of the DdCBE constructs are shown in Table 5 below. As a result, DdCBEs recognize 17 to 20 DNA sequences, including the conserved thymine sequence at the 5' end. Therefore, a functional DdCBE pair recognizes a total of 32 to 40 DNA bases.
[0392] [Table 5] JPEG0007796729000023.jpg235161JPEG0007796729000024.jpg80170
[0393] In vitro mitochondrial base editing. To attempt in vivo mitochondrial DNA editing using the Golden Gate cloning system, we selected the ND5 gene from Mus musculus, encoding the mitochondrial NADH-ubiquinone oxidoreductase chain 5 protein. The ND5 protein is a core subunit of NADH dehydrogenase (ubiquinone) and catalyzes the transfer of electrons from NADH to the respiratory chain. In humans, mutations in the ND5 gene are known to be associated with MELAS (mitochondrial encephalomyopathy, lactic acidosis, and stroke-like episodes) and some symptoms of Leigh syndrome or LHON (Leber's hereditary optic neuropathy). We attempted to create a mouse model with genetic mutations in mitochondrial genes to mimic human dysfunction.
[0394] First, we constructed several DdCBE plasmids designed to induce two silent mutations, m.C12539T and m.G12542A. We transfected these plasmids into NIH3T3 mouse cells and measured the base editing frequency after three days. As expected, cytosine bases in the target region were edited to thymine with an efficiency of up to 19% (Figure 34a). DddAtox has previously been reported to exclusively deaminate cytosines in "TC" sequences, and our experimental results also showed that only two cytosines in the TC context were edited. No deletions or other types of point mutations were generated meaningfully in the editing target region.
[0395] In vivo mitochondrial base editing. The most effective DdCBE pair (left-G1397-N and right-G1397-C) was used for in vitro experiments. Four days after microinjection of in vitro transcripts encoding this DdCBE pair into one-cell C57BL6 / J embryos, 9 of 32 embryos were successfully edited (28%, Table 6).
[0396] [Table 6]
[0397] TALE-DddAtox deaminase efficiently generated C·G to T·A transitions, with efficiencies ranging from 2.2 to 25% in m.C12539 and 0.63 to 5.8% in m.G12542. DdCBE-injected embryos were then implanted into surrogate mothers, yielding offspring carrying m.C12539T and m.G12542T (Figure 39). Three of the four pups (F0) were C·G to T·A edited, with efficiencies ranging from 1 to 27% (Figure 34c). Two pups showed similar mutation counts in the toes and tail, maintaining this efficiency even after 14 days of age. Furthermore, these mitochondrial DNA mutations were detected in various tissues of adult F0 mice 50 days after birth (Figure 34d). These results suggest that the heteroplasmy of mitochondrial DNA generated by DdCBE in the 1-cell base complex is maintained during development and differentiation.
[0398] To confirm that the mutations induced by DdCBE are transmitted to the next generation, we mated female F0 mice with wild-type C57BL6 / J males to generate F1 offspring. The m.C12539T and m.G12542T mutations were observed in two offspring with efficiencies ranging from 6% to 26%. These mitochondrial edits were also observed with similar efficiency in 11 different tissues (Figure 35b).
[0399] DdCBE-mediated MT-ND5 G12918A mutation. Next, we attempted to create the m.G12918A mutation, which can also cause mitochondrial diseases in humans. Notably, this mutation causes several mitochondrial diseases, such as Leigh syndrome, MELAS syndrome, and LHON syndrome. The cytosine base at this position is adjacent to a thymine, making it amenable to base editing using DdCBE (Figure 36a). We assembled four pairs of DdCBEs and confirmed that editing was possible up to 6.4% in NIH3T3 cells (Figure 36b). The most efficient DdCBE combination was then microinjected into mouse conjugates, and its efficiency was monitored in blastocysts. Eleven of 44 (25%) embryos carried the m.G12918A mutation, with efficiencies ranging from 0.25 to 23% (Figure 36c). Next, we implanted the DdCBE-microinjected embryos into surrogate mothers to obtain offspring carrying the G12918A mutation (Figure 39b). Four of the 11 newborn mice were confirmed to carry the mutation, with a 3.9-31.6% harboring rate (Figure 36d). Although no phenotypes were observed immediately after birth, likely due to the very young age of the offspring and the presence of both normal and mutant mitochondrial DNA in a heterozygous state, these results suggest that DdCBE can be used to create an animal model of mitochondrial disease.
[0400] MT-ND5 Nonsense Mutation. Finally, we confirmed whether the loss-of-function mutation in ND5 can be maintained in mice by creating a nonsense mutation in the gene. Using m.C12336 as the target cytosine, a premature stop codon was introduced at position 199 of the ND5 protein (Q199*; Figure 37a). First, we transfected four DdCBE combinations into NIH3T3 cell lines to confirm base editing efficiency. We confirmed that the most effective DdCBE pair produced nonsense mutations with approximately 5.7% efficiency (Figure 37b). This DdCBE induced cytosine-thymine editing, and also confirmed that the m.G12341A silent mutation (Q200Q) was edited at the target level, albeit with slightly lower efficiency. We also confirmed the m.C12336T and m.G12341A mutations in 19 of 37 (51%) mouse embryos, all with efficiencies of 32% and 23%, respectively (Figure 37c).
[0401] Based on these results, we implanted mouse embryos into surrogate mothers and obtained offspring carrying the m.C12336T and m.G12341A mutations (Figure 39c). Of 27 total, 9 F0 mice (23%) showed C·G to T·A editing, with an efficiency ranging from 0.22 to 57% (Figures 37d, e), demonstrating that nonsense mutations in ND5 do not cause embryonic selection.
[0402] Example 4. Animal mitochondrial DNA editing 4-1. Construction of an animal mitochondrial base editing expression vector with a nuclear export signal We constructed a vector (Figure 40) that expresses proteins containing TALE-DdCBE-linked nuclear export signals in animal cells. The vector uses a cytomegalovirus (CMV) promoter. From the N-terminus, it contains a mitochondrial transport signal, a protein purification / detection tag, a TAL array N-terminal domain, a repeat site, a C-terminal domain, a DddA cytosine deaminase split half, a uracil glycosylase inhibitor, and a nuclear export signal (Figure 40a). The nuclear export signal can be derived from the NS2 protein of the Mirute virus of mice (MVM), for example, but other sequences are also possible. This allows the expressed protein to be exported from the nucleus and delivered to mitochondria for base editing. The target DNA used in this experiment was the mitochondrial ND5 gene (ND5-like gene on chromosome 4), mitochondrial TrnA (chromosome 5), and mitochondrial Rnr2 (chromosome 6) (Figure 40b).
[0403] 4-2. DdCBE-NES in animal cell lines NIH3T3 cell line (ATCC CRL-1658) was transfected with 1.5 × 10 4 The cells were then dispensed into a 12-well plate containing 1 ml of cell growth medium (DMEM + 10% Bovine Calf Serum) per well. The following morning, untreated Mock, DdCBE, and DdCBE-MVM NES plasmid-injected experimental groups were transfected into cells using Lipofectamine 3000 according to the manufacturer's protocol. After incubation in a carbon dioxide incubator (37°C, 5% CO2) for 3 days, the cells were harvested and total DNA was purified using the Qiagen Blood & Tissue Kit. Mitochondrial gene-specific PCR primers were then amplified and subjected to next-generation sequencing using an Illumina Miniseq system. The results were used to confirm base editing efficiency using the Cas-analyzer (www.rgenome.net).
[0404] Figures 40c, d, and e show the results of transfection of DdCBE-NES into NIH3T3 mouse cells to induce mutations in the mouse mitochondrial genes ND5, TrnA, and Rnr2. The efficiency of the transfection varies depending on the DdCBE combination.
[0405] 4-3. mitoTALEN in animal cell lines A TALEN recognizing the sequence shown in Figure 40f was constructed, and this construct was linked to MTS and introduced into mitochondria together with DdCBE. The experimental method was the same as in Example 4-2. Mismatches were intentionally introduced into the TALE recognition site, resulting in one mismatch with wild-type mtDNA and two mismatches with mutant mtDNA. This is because TALEs may not be able to recognize a single sequence mismatch. As a result, we confirmed that the efficiency of the +2 mismatch experimental group was slightly increased when treated with TALEN compared to that treated with DdCBE alone.
[0406] 4-4. DdCBE-NES in animal embryos Using the DdCBE expression vector and DdCBE-NES expression vector as templates, the T7 promoter and DdCBE or DdCBE-NES expression region are amplified by PCR to obtain DNA. This DNA is then used as a template to synthesize mRNA using T7 polymerase.
[0407] A pair of DdCBE mRNAs or a pair of DdCBE-NES mRNAs is mixed into a microinjection solution and microinjected into mouse fertilized eggs. After 4 days of incubation, the fertilized eggs reached the blastocyst stage, which was then lysed. Using this as a template, the target region in mitochondrial DNA, distinct from nuclear DNA, was amplified by PCR, followed by an additional PCR amplification of the index and sequencing adapter. High-throughput sequencing was performed using an Illumina Miniseq system, and the results were used to confirm base editing efficiency using the Cas-analyzer (www.rgenome.net). Additionally, nuclear DNA with a similar sequence to the mitochondrial target region was amplified by PCR and sequenced.
[0408] As a result, DdCBE induced mutations not only in mitochondrial DNA but also in similar DNA sequences in the nucleus (mitochondria: 13.1%, nuclear: 3.2%). In the case of DdCBE-NES, the mutation efficiency of the mitochondrial target increased to 18.2%, while the mutation efficiency of nuclear DNA decreased to 0.2% (Figure 41a). For other targets, TrnA or Rnr2, no nuclear DNA mutations occurred, but the base editing efficiency of mitochondrial DNA was found to increase statistically significantly (*p<0.05, **p<0.01, ns not statistically significant).
[0409] 4-5: DdCBE and mitoTALEN in animal embryos To increase the proportion of edited mitochondrial DNA in cells after C-to-T conversion, a TALEN that cleaves unedited mitochondrial DNA sequences was co-injected into the ND5 gene site using DdCBE. The microinjection and sequencing methods were the same as in Examples 4-2 and 4-4. The group microinjected with DdCBE alone showed an editing efficiency of 11%, which increased to 33.3% when treated simultaneously with mitoTALEN, resulting in a statistically significant increase in editing efficiency. Furthermore, the group microinjected with DdCBE-NES alone showed an editing efficiency of 20.5%, which increased to 36.8% when treated simultaneously with mitoTALEN. This was also statistically significant (Figure 41b).
[0410] The microinjected fertilized eggs were implanted into surrogate mothers, and the resulting offspring showed similar editing efficiency of 10.9% with DdCBE, whereas the combined use of DdCBE-NES and mitoTALENs resulted in an efficiency of 23.4% (Figure 41c).
[0411] When a nuclear export sequence is attached to a base editing protein during editing of animal mitochondrial genes, base editing is more efficient, and in animal embryos, non-specific base editing of similar sequences in the nucleus is also suppressed. Furthermore, even when a mitochondrial sequence cleavage protein is used at the same time, more efficient mitochondrial base editing is expected.
[0412] Example 5. Split DddA tox Deaminase mutants We present a highly accurate DddA-derived cytosine base editor that can reduce the non-targeting effect of DdCBE. This non-targeting effect is caused by the spontaneous binding of DddAtox deaminase fragments, independent of the interaction between TALEs and DNA. Therefore, we created HF-DdCBE by substituting the amino acid residues located on the surface between the DddAtox fragments with alanine. HF-DdCBE does not function properly if the two deaminase pairs linked to the TALE cannot bind to DNA. Whole-mitochondrial genome analysis confirmed that HF-DdCBE is highly efficient and precise, unlike conventional DdCBEs, which cause numerous undesired non-targeting C-to-T transversions in human mitochondrial DNA.
[0413] 5-1. Method Plasmid construction. Point mutations were introduced into the DdCBE expression plasmid. The plasmid was amplified using Q5 Site-Directed Mutagenesis (NEB) mutagenesis primers (Table 7), and the results were confirmed by Sanger sequencing.
[0414] [Table 7] JPEG0007796729000027.jpg249170JPEG0007796729000028.jpg249170JPEG0007796729000029.jpg191170
[0415] For assembly of the ligation surface mutants, minipreps of the mutant expression plasmids were combined with the module vector (each encoding a TALE sequence), BsaI-HFv2 (10 U), T4 DNA ligase (200 U), and reaction buffer in a single tube. The restriction enzyme and ligation reactions were then cycled 20 times in a thermocycler at 37°C for 5 minutes and 50°C for 5 minutes, followed by 50°C for 15 minutes and 80°C for 5 minutes. The ligated plasmids were then introduced into E. coli DH5a by chemical transformation, and the final constructs were confirmed by Sanger sequencing. For cell line transfection, the plasmids were midiprepped.
[0416] Mammalian cell line culture and transfection. HEK 293T / 17 (CRL-11268, American Type Culture Collection (ATCC)) cell line was cultured at 37°C in a 5% CO2 environment. The cell line was grown in DMEM supplemented with 10% (v / v) fetal bovine serum (Gibco) without antibiotics and was not mycoplasma tested. For lipofection, cells were grown at 1 × 10 in 24-well cell culture plates (SPL, Seoul, Korea). 5 Growth was initiated 18–24 hours before transfection at a cell density of 1000 μg / ml. A total of 1,000 μg of plasmid DNA was transfected using 500 μg of each DdCBE aliquot using Lipofectamine 2000 (Invitrogen). Cells were harvested 4 days after transfection.
[0417] Genomic and mitochondrial DNA isolation for high-throughput sequencing analysis. To isolate genomic DNA, the cell culture medium was removed, and then the lysis buffer supplemented with proteinase K from the DNeasy Blood & Tissue Kit (Qiagen) was added to the cell culture plate to separate the cells from the bottom. Genomic DNA was then isolated according to the manufacturer's protocol. For whole mitochondrial genome sequencing, 200 μl of Mitochondrial Isolation Buffer A (ScienCell) was added to the culture plate from which the cell culture medium had been removed. The cells were scraped using a cell lifter, placed in a microtube, and crushed using a disposable pestle. After 20 triturations, the homogenate was centrifuged at 1,000 × g for 5 minutes at 4°C. The supernatant was transferred to a new microtube and centrifuged at 10,000 × g for 20 minutes at 4°C. The precipitate was released into 10 μl of lysis buffer (25 mM NaOH, 0.2 mM EDTA, pH 10) and then boiled at 95°C for 20 minutes. To lower the pH, 1 μl of 1 M HEPES (free acids, without pH adjustment) was added to the mitochondrial lysate. 1 μl of the solution thus prepared was used as a PCR template strand for high-throughput sequence analysis.
[0418] High-throughput sequencing analysis. To generate deep sequencing libraries, nested primary and secondary PCRs were performed using Q5 DNA polymerase to add final index sequences. Libraries were used for fair-end sequence analysis using MiniSeq (Illumina). For whole mitochondrial genome analysis, isolated mitochondrial DNA was prepared using a tagmentation DNA prep kit (Illumina) according to the manufacturer's protocol. For all analyses, the results of the fair-end sequence analysis were combined into a single fastqjoin file and analyzed using CRISPR RGEN Tools (http: / / www.rgenome.net / ).
[0419] 5-2.Results When chloroplast editing was attempted in plants, non-targeted base mutations appeared in the chloroplast genome, thus raising questions about the accuracy of DdCBE. We considered two possible reasons for non-targeted base editing by DdCBE. First, the first reason is non-specific binding between TALE proteins and DNA, and second, natural and unintended binding between DddAtox halves (Figure 42a). In this study, we focused on the DddAtox split halves and improved the binding surfaces of the two split proteins to prevent undesired binding of the DddAtox halves.
[0420] First, we investigated the mitochondrial ND1 (mtND1) gene to see whether each subunit (Left-TALE or Right-TALE) binds to DNA alone and interacts with the other half of TALE-free DddAtox or TALE-DddAtox without a binding sequence to induce cytosine-thymine base editing. In human embryonic kidney cell line (HEK293T), a DdCBE pair targeting the human mitochondrial ND1 (mtND1) gene (Left-TALE:G1397N (the N-terminal G1397 DddAtox half binds to the C-terminus of the TALE sequence, taking over the left side of the entire recognition sequence) + Right-TALE:G1397C (the C-terminal G1397 DddAtox half binds to the C-terminus of the TALE sequence, taking over the right side of the entire recognition sequence)) effectively edited C11 of the target sequence from cytosine to thymine with an efficiency of 60.7% (Figure 43a). Furthermore, when treated with TALE-free DddAtox half subunits that can pair with each other, base editing still occurred, although at a lower efficiency than the original DdCBE pair. Thus, Left-TALE and Right-TALE bind to the target ND1 sequence and edit bases with 31% or 8.1% efficiency, respectively (Figure 43). In other words, the DdCBE pair fused with both TALE proteins, respectively, is 2.0-fold (=60.7% / 31%) or 7.5-fold (=60.7% / 8.1%) more efficient than when bound to either TALE alone. Clearly, the N-terminal half of DddAtox bound to a TALE binds to the C-terminal half of DddAtox without a TALE sequence, enabling deamination, and vice versa.
[0421] Because DddAtox splits into two halves (G1333 and G1397), we constructed DdCBE pairs targeting the mtND1 gene at G1333 (Left-TALE:G1333-N and Right-TALE:G1333-C). We then independently combined the Left-TALE:G1333-N and Right-TALE:G1333-C constructs with the corresponding TALE-free DddAtox halves to confirm whether they could induce cytosine-thymine editing. As expected, each TALE-fused DdCBE exhibited a base editing efficiency of 32.7% (left-TALE conjugate) or 18.1% (right-TALE conjugate) at C8, compared with 56.1% for the original DdCBE pair. Therefore, the TALE-bound DdCBE pair acts 1.7-fold (=56.1% / 32.7%) or 3.1-fold (=56.1% / 18.1%) more efficiently than either pair acting separately. Collectively, these results suggest that DdCBEs can induce undesired non-targeted base editing even at sites where only one TALE sequence binds. Because TALE proteins can bind to DNA even with a small number of mismatches, DdCBEs may induce non-targeted base editing in organelles or nuclear genes.
[0422] To reduce non-targeted base editing caused by the binding of split DddAtox halves to each other, we sought to develop a high-precision DdCBE. We hypothesized that self-association could be prevented or inhibited by modifying the binding surface of the split dimer. Therefore, in PyMOL software, we used Python code (InterfaceResidues.py) to find the amino acid residues on the binding surface of the two split DddAtox halves (split positions G1333 and G1397) within a 1 Ų area. As a result, we found 9 amino acid residues in G1397-N (the N-terminal DddAtox half split at G1397), 4 residues in G1397-C (the C-terminal DddAtox half split at G1397), 14 amino acid residues in G1333-N (the N-terminal DddAtox half split at G1333), and 15 amino acid residues in G1333-C (the C-terminal DddAtox half split at G1333) (Figures 42b and 42c). We then created several mutant DddAtox halves by substituting the identified amino acid residues with alanine. We then co-transfected these DdCBE binding surface mutants with wild-type DdCBE pairs or TALE-free DddAtox pairs into HEK293T cells to observe base editing efficiency. Several binding surface mutations in the G1397 fragment, including C1376A, M1390A, and F1412A, failed to demonstrate cytosine-thymine editing efficiency at the target site when coupled with wild-type or TALE-free DddAtox, suggesting that these mutations are unable to interact with other DddAtox halves even when positioned nearby. Some mutations, such as E1381A and V1377A, demonstrated highly efficient base editing when coupled with the TALE-free half, indicating that these mutations are independent of dimer interactions.
[0423] Importantly, some mutations, such as K1389A, K1410A, and T1413A, showed high activity when combined with the wild-type DdCBE pair but lower activity when used with the TALE-free pair. As an example, the K1410A mutation showed an efficiency of 53.2%, similar to the wild-type DdCBE pair (60.7%), but when used with the TALE-free pair, it showed an efficiency of 0.9%, representing a 59.1-fold difference (= 53.2% / 0.9%). As previously mentioned, the wild-type DdCBE showed a 7.5-fold difference (= 60.7% / 8.1%). Furthermore, these mutations selectively edited bases compared to the wild-type DdCBE. The mutants did not edit C8, C9, or C within the editing range. 13 than C 11 This shows a strong preference for cytosines, which can be compared to the wild type, which edited all four cytosines evenly to over 6.7% (Figure 43b).
[0424] Furthermore, we screened 29 mutations (14 G1333N and 15 G1333C) in G1333 and obtained several suitable binding surface mutations (Figure 44). Many mutations either decreased efficiency when coupled with the wild-type (I1299A, Y1316A, Y1317A, and F1329A) or increased efficiency when coupled with the TALE-free (S1300A and T1314A). However, it is noteworthy that mutations such as K1389A, T1391A, and V1393A frequently caused base editing when coupled with the wild-type but showed low efficiency when coupled with the TALE-free. For example, K1389A showed a 38-fold difference (=45.4% / 1.2%), while the wild-type DdCBE pair showed only a 3.1-fold difference (=56.1% / 18.1%). In addition, K1389A was sometimes observed to be more efficient than the wild type. 11 or C 13The editing efficiency was observed to be concentrated at C8, which is not a cytosine position, compared to wild-type DdCBE, which showed an increase in editing efficiency at all cytosine positions of over 19%. It is also noteworthy that G1397-C K1410A preferentially edits C8, whereas G1333-C K1389A selectively edits C11 with higher efficiency. On the other hand, wild-type DdCBE pairs (G1333 or G1397 DddA) tox These results suggest that the binding surface mutations described above may reduce the effect of DdCBE in causing undesired editing of multiple bases within the target site.
[0425] Example 6. Full-length deaminase The DddA-derived cytosine base editing enzyme (DdCBE), which contains the split bacterial toxin DddAtox, a TALE array, and a uracil glycosylase inhibitor (UGI), converts target cytosines to thymine in eukaryotic nuclei, mitochondrial DNA (mtDNA), and plant chloroplast DNA. DddAtox, which induces bacterial toxicity, is derived from Burkholderia cenocepacia and deaminates cytosines within double-stranded DNA. To avoid toxicity to host cells, DddAtox splits into two inactive halves, each of which binds to a TALE DNA-binding protein to create DdCBE. The two inactive forms bound to the TALE array are functional when bound to target DNA in close proximity by a TALE. Cytosine-to-thymine base conversion is induced in a region of 14–18 bases between the two TALE binding sites. Unlike CRISPR-derived base editors, which cannot edit organelle DNA, DdCBEs enable targeted base editing in both nuclear and organelle DNA. However, they have the drawback of requiring two TALE constructs rather than one to induce it. The first drawback is that the use of two TALE arrays limits the targetable sites because the TALE must bind to thymines at both the 5' and 3' ends of the target DNA site. The second drawback is that delivering two TALE constructs rather than one is often inefficient and challenging. Dose-limited viral vectors, such as adeno-associated virus (AAV) vectors (dose, ~4.7 kbps), widely used in gene therapy, cannot accommodate two DdCBE-encoding sequences because the dimeric DdCBE combination is too large (2 × 4.1 kbps, including the promoter and polyA signal). Cloning two TALE array DNAs into a single vector with a larger dose can be difficult due to the high similarity of the two TALE array sequences. Finally, using two TALE arrays rather than one may exacerbate off-target effects. To overcome these limitations of dimeric DdCBE with split DddAtox, we present a non-toxic full-length DddAtox form, mDdCBE (monomer DdCBE), which induces cytosine-to-thymine conversion in target DNA in the nucleus and organelles.
[0426] 6-1. Method Plasmid construction. The DddA mutant was PCR-amplified using the synthesized full-length DddAtox (gBlock, IDT) as a template with the primers listed in Table 8 and Q5 DNA polymerase (NEB). These PCR products were cloned using Gibson assembly (NEB) into the p3s-BE3 site of Apobec1, which had been digested with BamHI and SmaI (NEB). For TALE-DddAtox (Addgene #158093, #158095, #157842, #157841), the plasmid was digested with BamHI and SmaI, and the DddA mutant was PCR-amplified using the primers listed in Table 8 and cloned using Gibson assembly. The isolated plasmid was transformed into chemically prepared E. coli DH5a by heat shock, and the plasmid sequences of surviving colonies were analyzed by Sanger sequencing. The final plasmid was midiprepped (Macherey-Nagel) for cell transfection.
[0427] [Table 8]
[0428] Random mutagenesis. Error-prone PCR was performed using the GeneMorph II Random Mutagenesis Kit (Agilent) with full-length DddAtox (gBlock, IDT) synthesized according to the manufacturer's protocol as the template. Briefly, 1 ng, 100 ng, and 700 ng of DddAtox DNA were used as templates to introduce random mutations of 0 to 16 mutations / kb, respectively. The full-length DddAtox gBlock was pre-amplified by PCR using the primers listed in Table 8. The PCR products were combined and cloned into p3s-UGI-Cas9(H840A) digested with Sma1 and Xho1 using Gibson assembly (NEB). The plasmid was transformed into E. coli DH5a cells prepared by chemical methods using heat shock, and the plasmid sequences of surviving colonies were analyzed by Sanger sequencing. Among the analyzed plasmids, the p3s-UGI-nCas9(H840A)-DddAtox plasmid containing the coding frame was transfected into HEK293T cells together with sgRNA, and the editing activity was confirmed by targeted deep sequencing.
[0429] Mammalian cell culture and transfection. HEK293T (ATCC, CRL-11268) cells and HeLa (ATCC, CCL-2) cells were cultured at 37°C with 5% CO2. Cells were cultured in DMEM supplemented with 10% (v / v) fetal bovine serum (Welgene) and 1% penicillin / streptomycin (Welgene). Cells were seeded into 48-well plates (Corning) at a density of 3 x 105 cells (HEK293T) and 4 x 104 cells (HeLa) 24 h before transfection and transfected with Lipofectamine 2000 (Invitrogen) and Cas9-fused DddA plasmid (750 ng) and sgRNA (250 ng). TALE-DddA was transfected into HEK293T cells using 200 ng of plasmid and Lipofectamine 2000. The sgRNA sequences are shown in Table 9.
[0430] [Table 9]
[0431] Preparation of genomic and mitochondrial DNA. Cells transfected with Cas9-fused DddA mutants were harvested 2 days after transfection, and cells transfected with TALE-DddA were harvested 3 days after transfection. Genomic and mitochondrial DNA were isolated using the DNeasy Blood and Tissue Kit (Qiagen). For large-scale analysis, DNA was extracted using 100 μL of cell lysis buffer (50 mM Tris-HCl, pH 8.0 (Sigma-Aldrich), 1 mM EDTA (Sigma-Aldrich), 0.005% sodium dodecyl sulfate (Sigma-Aldrich)) containing 5 μL of proteinase K (Qiagen). The lysate was incubated at 55°C for 1 hour and then at 95°C for 10 minutes.
[0432] 6-2.Results The amino acid sequences of the wild-type and novel full-length DddA are compared, and the modified amino acids are indicated by gray boxes in Figure 45.
[0433] As shown in Figure 46, DddA was linked to the N-terminus of Cas9 using a 16-amino acid linker, and UGI (uracil glycosylase inhibitor) and NLS (nuclear localization signal) were linked to the C-terminus using a 4-amino acid linker. Conversely, DddA was linked to the C-terminus of Cas9 using a 16-amino acid linker, and UGI and NLS were linked to the N-terminus using a 4-amino acid linker.
[0434] In this study, we used DddA-Cas9(D10A, D10A and H840A)-UGI. DddA binds to a DNA-binding zinc finger protein, the TALE module, and can replace cytosine with thymine using only one module. While conventional splitting requires two modules, full-length DddA can be achieved with only one module. These two DNA-binding proteins link a nuclear localization signal (NLS), a mitochondrial targeting sequence (MTS), and a chloroplast transit peptide (CTP) to replace cytosine with thymine not only in genomic regions but also in mitochondria and plant chloroplasts, which Cas9 cannot. As shown in Figure 47, we confirmed the activity of replacing cytosine with thymine in the TC motif in the human cell genomic contexts of the ROR1 site (a), HEK3 site (b), and TYRO3 site (c). The activity of the target site was confirmed by substituting a cytosine in the TC motif, located 25 bp away, with a thymine (a). A1341D KRKKA was confirmed to have the activity of substituting the second cytosine in the CC motif with a thymine (a, b). The catalytic mutant E1347A was also confirmed to have the activity of substituting a cytosine in the TC motif with a thymine (a, b, c). The red underline indicates the binding site of Cas9. The efficiency is expressed as the percentage of cytosines that were substituted with thymine in the deletion-free reads of the entire sequencing reads. The efficiency is also expressed as the percentage of deletions in the entire sequencing reads.
[0435] As shown in Figure 48, the red square box indicates the segment where DddAtox activity was confirmed. The same segment was then split into three target sites using full-length DddA and activity was measured. The split fragments used orthogonal Cas9s with different PAMs to replace the cytosine between the two Cas9s with thymine. However, in this case, precise replacement of the desired cytosine with thymine was difficult. However, full-length DddA can target the same target site by splitting the Cas9 binding site into three segments, allowing for precise replacement of the desired cytosine with thymine. Efficiency is expressed as the percentage of cytosines replaced with thymine among all sequencing reads without deletions. The percentage of all sequencing reads with deletions is also shown.
[0436] As shown in Figure 49, the activity of full-length DddA was measured in the human cell genome contexts TRAC site 1 (a), TRAC site 2 (b), FANCF (c), and HBB (d). The red underlines indicate the Cas9 binding sites. Efficiency is expressed as the percentage of cytosine to thymine substitutions in deletion-free reads among all sequencing reads. It is also expressed as the percentage of deletions among all sequencing reads.
[0437] As shown in Figure 50, DddA activity was measured in the human cell genomic contexts TYRO3 (a), ROR1 (b), HEK3 (c), EMX1 site 2 (d), TRAC site 1 (e), and HBB (f) using DddA-dCas9(D10A, H840A)-UGI. Efficiency is expressed as the percentage of cytosine-to-thymine substitutions across all sequencing reads. No deletions were observed.
[0438] To obtain nontoxic full-length DddAtox mutants useful for base editing, we employed two approaches: structure-directed site-directed mutagenesis and random mutagenesis. In the first approach, we fused DddAtox mutants with reduced DNA binding or catalytic activity to inactive CRISPR-Cas9 (dCas9) or nickase (nCas9) mutants to develop novel base editors that substituted target cytosines with thymines in human cells. To achieve this, we attempted to subcloning DddAtox into expression vectors by substituting positively charged amino acids with alanines (Figure 51a). We reasoned that such mutants would weaken binding to negatively charged dsDNA and potentially avoid toxicity. Most alanine-substituted mutants failed to form E. coli transformants (Figure 51b). Base analysis of plasmid DNA isolated from the resulting transformants revealed various frame-shift mutations in the protein-forming region. Although these full-length DddAtox mutants were under the control of a mammalian promoter, they were weakly expressed in E. coli and caused apoptosis. Fortunately, we were able to obtain several triple, quadruple, or quintuple (AAAAA) alanine substitution mutants without frameshift mutations. Furthermore, the active site mutation E1347A was also successfully cloned.
[0439] We next investigated whether the AAAAA mutant fused to D10A nCas9 or dCas9 and UGI could induce base editing in human embryonic kidney 293T (HEK293T) cells (Figure 51c, d). Base editor 2 (or 3), consisting of rat APBEC1 deaminase, uracil glycosylase inhibitor (UGI), and dCas9 (or D10A nCas9), is active in a narrow region within the protospacer region, whereas the AAAAA mutant induced up to 43% cytosine-to-thymine substitutions immediately 5' above the protospacer region. Unexpectedly, the E1347A mutation induced the same C mutation at a frequency of 37% (nCas9 fusion) or 16% (dCas9 fusion). -3We induced base editing at the E1347A position (Figure 51c), confirming that the E1347A mutation did not completely inactivate DddAtox by deamination and had sufficiently high deamination activity to achieve base editing in human-derived cells. However, the E1347A mutant combined with a quintuple AAAAA mutation failed to induce base editing. Furthermore, we confirmed that E1347A, AAAAA, and other mutants substituted with alanine without a translocation mutation (Figure 53) fused to dCas9 or nCas9 and UGI resulted in editing at positions up to 25 bases away above the protospacer region, with editing efficiencies of up to 26% at many other sites (Figure 54). Furthermore, the fusion protein was highly efficient in HeLa cells, with editing frequencies of up to 60% (Figure 55). Base edits induced by such fusion proteins were maintained in cells for up to 21 days, suggesting that such base editing is not cytotoxic (Figure 56).
[0440] Due to the altered editing window of the cytosine base editor, we attempted to fuse alanine-substituted mutants to the C-terminus of H840A nCas9. Unexpectedly, we failed to obtain an intact construct without frameshift mutations. Therefore, we performed error-prone PCR to introduce random mutations into the DddAtox coding sequence, resulting in a non-toxic full-length DddAtox mutant with four point mutations: S1326G, G1348S, A1398V, and S1418G (referred to as "GSVG"). (S1326G, G1348S, A1398V, and S1418G comprise the sequence of SEQ ID NO: 276, which contains a substitution of S at position 37 for G; a substitution of G at position 59; an substitution of A at position 109 for V; and a substitution of S at position 129 for G, respectively, of the amino acid sequence of SEQ ID NO: 269; Figure 52a). Furthermore, these mutants were fused to the C-terminus of dCas9, D10A nCas9, and Cas9, and to the N-terminus of dCas9, nCas9, and Cas9. These fusion proteins, except for wild-type Cas9, induced cytosine-to-thymine conversions at various sites with efficiencies up to 38% in human-derived cells (Figures 52b, 57, and 58). Interestingly, fusion proteins containing GSVG mutants fused to the C-terminus of dCas9, D10A nCas9, and H840A nCas9 exhibited cytosine base editing 3' downstream of the protospacer adjacent motif (PAM), whereas fusion proteins containing the same mutants fused to the N-terminus of dCas9 and nCas9 induced base editing 5' upstream of the protospacer (Figure 52c). As expected, Cas9-containing fusion proteins induced indels rather than base substitutions.
[0441] To find out which mutations are important for GSVG variants, S SVG, G G VG, GS A G, and GSV SWe attempted to generate four revertants of nCas9 by site-directed mutagenesis. SSVG, GSAG, and GSVS revertants were obtained, but not the GGVG mutant fused to the C-terminus of nCas9. G1348 is located immediately adjacent to E1347, the core site of the catalytic site. The G1348S mutation reduced catalytic activity and avoided cytotoxicity in E. coli. We measured the editing frequencies of the three revertants and the GSVG mutant at two target sites in transfected cells for up to 21 days. The frequency of cytosine-to-thymine editing induced by GSAG and GSVS gradually decreased ~2-fold from 3 to 21 days posttransfection. These two revertants were slightly cytotoxic, while GSVG and SSVG maintained their cytotoxicity (Figure 59). These results suggest that G1348S is essential in the GSVG mutant, S1326G is neutral, and A1398V and S1418G reduce cytotoxicity.
[0442] In summary, these results suggest that a non-toxic, full-length DddAtox mutant with reduced affinity for dsDNA (AAAAA), weakened deamination activity (E1347A and possibly GSVG), or reduced cytotoxicity (GSVG) can be fused to dCas9 or nCas9 to create a base editor with a novel, altered editing window. These base editors are hereafter referred to as dCas9-mDdBE (a DddA-derived base editor consisting of a full-length monomeric DddAtox mutant fused to the C-terminus of dCas9), nCas9-mDdBE, mDdCE-dCas9, and mDdCE-nCas9, and are used for base editing upstream or downstream of the protospacer, which is inaccessible to BE2 or BE3.
[0443] We further investigated whether nontoxic full-length DddAtox mutants could be used for mitochondrial DNA editing. Among the numerous mutants, only two, the GSVG and E1347A mutants, were successfully fused to the C-terminus of a TALE array designed to bind to mitochondrial genes, ND4 and ND6. Monomeric DdCBEs (mDdCBEs) containing the GSVG mutant achieved base editing at target nucleotide positions up to 31% (ND4) (Figure 60a) and 27% (ND6) (Figure 60b), comparable to the original split DdCBE pair. mDdCBEs containing E1347A also converted target cytosines to thymines, albeit with reduced efficiency, with editing ratios of up to 7.2% (ND4) and 8.9% (ND6). Interestingly, the original DdCBE pair (split G1333) dedicated to the ND4 gene had an editing efficiency of 0.8% at the C4 position, whereas the two mDdCBEs containing GSVG showed high editing efficiencies of 26% and 31%. This result indicates that split-dimeric DdCBE and mDdCBE may have different mutation patterns and suggests that mDdCBE can be complementary to dimeric DdCBE and thus induce diverse mutations at a given target site.
[0444] One potential advantage of mDdCBE over split-dimeric DdCBE is that off-target effects due to nonspecific TALE-DNA interactions are half as severe as those of dimeric DdCBE. Dimeric DdCBE with split DddAtox can function at half sites that can bind only one subunit, leading to undesired off-target mutations. The inactive DddAtox half of a DdCBE pair can recruit the other inactive half to form a functional deaminase. To confirm this hypothesis, we co-transfected HEK293T cells with a plasmid encoding one subunit of dimeric DdCBE and a plasmid encoding the TALE-less DddAtox half and measured the editing frequency at both mitochondrial target sites. As expected, cytosine-to-thymine editing was observed at the target site at frequencies of 0.7–3.6% (Figure 60c–f). This result suggests that the DddAtox halves split by the DdCBE pair may interact with each other at half sites to cause unwanted off-target mutations, and that mDdCBE can avoid half of the off-target mutations induced by dimeric DdCBE.
[0445] Example 7. Highly efficient A-to-G base editing in human cells using DdABE Mitochondrial DNA base editing using DddA-directed cytosine base editing (DdCBE) has enabled the generation of various cell lines and animal disease models, opening up new avenues for treating mitochondrial genetic diseases. However, DdCBE is limited to TC-to-TT base editing, covering only approximately one-eighth of all cases. Therefore, we developed a transcription activator-like effector (TALE) tethered to two deaminase enzymes. The TALE is customized to bind to the desired DNA segment and contains the catalytically inactive DddAtox cytosine deaminase mutant and the E. coli-derived DNA adenine deaminase TadA protein. Unlike previous base editing techniques, which only allowed cytosine base editing in the conventional TC context in human mitochondria, TALED enables all A-to-G base editing. Indeed, the customized TALED was able to induce adenine base editing at multiple targets in human cells with high efficiency (up to 50%).
[0446] To develop a novel base editing technology, we selected the TadA mutant of ABE8e (TadA*) from among numerous TadA mutants because it not only enables highly efficient adenine editing but also has been improved to be compatible with a variety of DNA-binding proteins, thereby improving compatibility with actual TALEs or ZFPs.
[0447] For the first time, we fused TadA* and MTS to a TALE customized to target ND1 or ND4, and tested whether they could actually induce base editing in mitochondrial DNA. Targeted deep sequencing showed that the fusion protein induced adenosine base editing with a very low efficiency. It induced adenine base editing at the ND1 site with a maximum efficiency of 1.2% (Figure 67a) and at the ND4 site with a maximum efficiency of 0.6% (Figure 67b). While TadA* was previously known to function specifically in single-stranded DNA, its fusion to a TALE demonstrated that it could also induce base editing in double-stranded target DNA, even though the efficiency was very low.
[0448] Based on the results demonstrating that adenine base editing can occur in mitochondrial DNA, we considered fusing the previously known DddAtox protein to further increase its efficiency. DddAtox is an interbacterial toxin derived from Burkholderia cenocepacia that deaminates cytosine. Because this protein functions on double-stranded DNA, its use allows TadA* adenine deaminase to better access the target DNA. In conventional DdCBEs using DddAtox, the DddAtox protein is split into two halves: a left-TALE (L-TALE) that attaches to the left side of the DNA and a right-TALE (R-TALE) that attaches to the right side. The TALE is then fused to a uracil glycosylase inhibitor (UGI) to increase cytosine base editing efficiency (TALE-Split DddAtox-UGI). The reason for using DddAtox in this split form is due to the cytotoxicity that occurs when the full-length protein is used. First, we experimented with constructing L-TALE-Split-DddAtox-TadA* and R-TALE-1397C-TadA* constructs by replacing UGI with TadA* on one side of the DdCBE targeting the ND1 site. Surprisingly, when we transfected these constructs, with TadA* on one side and 1397N and UGI on the other side, into human cells, we observed both A-to-G and C-to-T base editing (Figure 62c). Conventional DdCBEs exhibited approximately 20% cytosine base editing and no adenine base editing. However, with TadA* on one side, cytosine base editing was reduced to half its normal level, with adenine base editing occurring at approximately 10% (Figure 62c). Adenine and cytosine base editing efficiencies were similar (Figure 62c).
[0449] While both cytosine and adenine base editing may be useful for randomly inducing mutations, it is desirable to induce adenine base editing only when treating diseases, particularly mitochondrial genetic diseases such as LOHN and MEALS, which are caused by C-to-T mutations. Therefore, we eliminated UGI to eliminate the simultaneous cytosine base editing. In DdCBE, UGI is fused to the cytosine deaminase DddAtox, which deaminates C to U, to prevent the re-repair of U by uracil glycosylase, a cellular repair protein, during the DNA repair process. Therefore, we hypothesized that removing UGI would maintain adenine base editing efficiency while suppressing cytosine base editing. Surprisingly, we confirmed that ND1-targeting TALE deaminase pairs without UGIs rarely induced cytosine base editing (<0.5%) and only induced adenine base editing at a high efficiency (approximately 50%) (Figure 63a, c). This is significantly higher efficiency than those with UGIs. Furthermore, we confirmed this with ND4-targeting TALE deaminase pairs, which edited only adenine bases at a high efficiency (approximately 35%), similar to that of ND1-targeting TALE deaminase pairs (Figure 63b, d). Thus, we developed a new adenine deaminase, TALED, that functions on double-stranded DNA by fusing the DddAtox system with TadA*, and achieved adenine base editing in human mitochondria for the first time. Furthermore, we achieved a final adenine base editing efficiency approximately 50-fold higher than that of TALEs with TadA* alone.
[0450] Next, we attempted to induce adenine base editing using the full-length E1347A DddAtox mutant, which lacks catalytic activity, or the mutants we developed that retain catalytic activity but lack cytotoxicity (AAAAA and GSVG). Because our goal was not to induce cytosine base editing but to induce adenine base editing in DNA duplexes, we could use the full-length E1347A DddAtox mutant, which lacks cytosine base editing activity. Since we confirmed that cytosine base editing efficiency is lost without UGI, we could also use a mutant that lacks cytotoxicity. We created two types of TALEDs containing the full-length mutants (Figure 64a). The first type contained both TadA*(AD) and the full-length DddAtox mutant in a single TALE (mTALED). The second type fused TadA*(AD) and the full-length DddAtox mutant separately to each TALE (dTALED). We tested these two types (Figure 64a). Surprisingly, both types of ND1-targeting TALEDs induced adenine base editing with high efficiency (Figure 64b). mTALED showed a maximum efficiency of approximately 45%, while dTALED showed an adenine base editing efficiency of approximately 50% (Figure 64b). Experiments were performed at both the ND1 and ND4 sites, and similarly high adenine base editing efficiency was observed (Figure 64c). Furthermore, we confirmed that adenine base editing was highly efficient when using the full-length E1347A DddAtox mutant, which lacks cytosine base editing activity (Figures 64b, c). This indicates that despite the loss of cytosine deamination ability, TadA* still retains its role in facilitating TadA*'s access to the DNA duplex. Detailed analysis of these results at the single-base level (Figures 65, 66) reveals that adenine base editing occurs in the immediate vicinity of the TALE's binding to DNA. Furthermore, when two TALEs are used, base editing occurs only in the spacer between them, but strangely, mTALEDs used by a single TALE also have similar target lengths (Figures 65 and 66).
[0451] We were also curious to know whether the system would function in a zinc finger protein (ZFP) system or in nuclear DNA. Therefore, we constructed a nuclear DNA-targeting NC-type ZFP and fused it to the split DddAtox and TadA* (Figure 61a). TadA* was fused to the ZFP at various positions (Figure 61b). Among the various morphologies, we were able to create individuals that edited up to 10% of adenine bases in nuclear DNA (Figure 61d). Because the UGI is present on one side, cytosine base editing efficiency was also high (Figure 61c). Having confirmed that the ZFP-DddAtox-TadA* system functions in human nuclear DNA, we further tested whether it would function in mitochondria. Here, we used the structure that functioned most efficiently in nuclear DNA as a reference. Instead of a nuclear localization signal (NLS) that transmits to nuclear DNA, we attached a mitochondrial targeting sequence (MTS) that targets mitochondria. We then fused TadA* to a ZFP (NC type) that targets the ND1 site, with 1397N on the left ZFP and 1397C on the right ZFP. The results showed that adenine base editing occurred with approximately 3% efficiency (Figure 61.g). Although the adenine base editing efficiency was lower than that of TALED, optimizing various conditions, such as the linker connecting the proteins, may enable efficient adenine base editing using this ZFP system.
[0452] To date, gene editing technology has made remarkable progress. CRISPR-based genetic scissors (e.g., CRISPR-Cas9, base editor, prime editor) have improved off-targeting, increased efficiency, and diversified. However, despite these advances, methods for treating mitochondrial genetic diseases have remained limited. This is because CRISPR-based technology consists of a catalytic protein and a gRNA that guides it to the target. However, unlike proteins, there is no way to deliver gRNA to mitochondria. Therefore, the only way to manipulate mitochondrial genes is to cut the DNA and remove the mitochondrial DNA. However, a group led by David R. Liu in the United States demonstrated DdCBE, a method for inducing base editing in mitochondria. DdCBE contains the cytosine deaminase DddAtox, which functions in double-stranded DNA. By fusing DdCBE to the DNA-binding protein TALE, it was possible to deliver only the protein and cause base editing. However, DdCBE only edits TC-type cytosines, limiting its effectiveness in creating disease models and treating genetic diseases. Therefore, we developed the world's first TALED capable of adenine base editing in mitochondria. The TALED exhibited high efficiency (up to 50%), editing various adenine bases at the target site. Furthermore, in the presence of UGI, cytosine and adenine bases can be simultaneously edited, making it useful for random mutagenesis. In contrast, in the absence of UGI, only adenine base editing occurs, preventing cytosine base editing, making it a specific adenine base editing technique. Furthermore, we demonstrated that this TALED can be applied to the ZFP system, enabling adenine base editing in nuclear DNA. The development of this TALED will provide solutions to many mitochondrial genetic diseases, enable the creation of corresponding disease models, and serve as a valuable tool for many unexplored mitochondrial genetic research projects.
[0453] Although the specific details of the present invention have been described above, it is obvious to those skilled in the art that such specific descriptions are merely preferred embodiments and do not limit the scope of the present invention. Therefore, the true scope of the present invention is defined by the appended claims and their equivalents. [Industrial Applicability]
[0454] According to the present invention, the undesired non-selectivity of cytosine deaminase can be reduced by substituting specific amino acid residues on the cytosine deaminase split binding surface during gene editing.
[0455] The full-length cytosine deaminase makes it possible to edit regions that are difficult to edit using conventional cytosine base editors. Among conventional cytosine base editing methods, Apobec1, which is used as a deaminase, is known to be a cancer-inducing gene, which limits its use for therapeutic purposes. However, the newly developed full-length deaminase is believed to be free of such issues.
[0456] It is small, about 2.5 kb including the DNA-binding protein, and can be used in gene therapy using AAV vectors. It is easy to transfer mRNA and RNP, and can be used for the production of useful substances using prokaryotes.
Claims
1. (i) (a) a fusion protein comprising a first split body of a programmable DNA-binding protein and a cytosine deaminase, and (b) a fusion protein comprising a second split body of a programmable DNA-binding protein and a cytosine deaminase; or (ii) a nucleic acid encoding the fusion protein A base editing composition comprising: the cytosine deaminase is derived from double-stranded DNA deaminase (DddA) or an orthologue thereof, and has amino acid substitutions at residues located at the binding interface of the split dimer; the first fragment comprises an amino acid sequence from the N-terminus to a residue selected from the group consisting of G44 and G108 in the sequence of SEQ ID NO: 1, and the second fragment comprises an amino acid sequence from a residue selected from the group consisting of P45 and A109 in the sequence of SEQ ID NO: 1 to the C-terminus, A base editing composition, wherein the residues located on the binding surface of the split dimer are one or more residues corresponding to one or more residues selected from the group consisting of positions 3, 5, 10, 11, 23, 24, 25, 26, 27, 28, 29, 38, 40, and 41 of SEQ ID NO: 23 and positions 13, 16, 17, 20, 21, 28, 29, 30, 31, 32, 33, 56, 57, 58, and 60 of SEQ ID NO: 24, or one or more residues corresponding to one or more residues selected from the group consisting of positions 87, 88, 91, 92, 95, 100, 101, 102, and 103 of SEQ ID NO: 25 and positions 13, 14, 15, and 16 of SEQ ID NO:
26.
2. The base editing composition described in Claim 1, wherein the amino acid after amino acid substitution at the residue located on the binding surface of the split dimer is alanine.
3. The base editing composition of claim 1, wherein the DNA binding proteins are independently selected from the group consisting of zinc finger proteins, TALE proteins, and CRISPR-associated nucleases.
4. The base editing composition of claim 1, wherein the DNA binding protein is a zinc finger protein.
5. The base editing composition of claim 1, wherein the DNA binding protein is a TALE protein.
6. The base editing composition of claim 1, wherein at least one of the fusion protein comprising the first split body and the fusion protein comprising the second split body further comprises adenine deaminase.
7. The base editing composition according to claim 6, wherein the adenine deaminase is Escherichia coli TadA (tRNA-specific adenosine deaminase) or a mutant thereof.
8. the DNA binding protein is a TALE protein; the adenine deaminase comprises the amino acid sequence of SEQ ID NO: 458, or a variant thereof; and The composition is for base editing of adenine (A) to guanine (G). The base editing composition according to claim 6.
9. The base editing composition of claim 1, wherein the composition is for base editing in animal cells.
10. The base editing composition of claim 9, wherein the animal cell is a human cell.
11. The base editing composition of claim 1, wherein the composition is for base editing in plant cells.
12. The base editing composition of claim 1, wherein the composition is for nuclear DNA base editing and further comprises an NLS (nuclear localization signal) peptide or a nucleic acid encoding the same.
13. The composition is for base editing of mitochondrial DNA, and (1) a mitochondrial targeting signal (MTS) or a nucleic acid encoding the same; and (2) The base editing composition of claim 1, further comprising either a nuclear export signal (NES) or a nucleic acid encoding the same, or both.
14. the composition is for base editing in chloroplasts, leucoplasts, or chromoplasts; and (1) a chloroplast transit peptide (CTP) or a nucleic acid encoding the same; and (2) The base editing composition of claim 1, further comprising either a nuclear export signal (NES) or a nucleic acid encoding the same, or both.
15. The base editing composition of claim 1, further comprising a UGI (uracil glycosylase inhibitor) or a nucleic acid encoding the same.
16. The composition of claim 1, wherein the composition further comprises a TALEN (TALE nuclease) or ZFN (zinc finger nuclease), or a nucleic acid encoding the same, that cleaves wild-type DNA base sequences but does not cleave edited base sequences, and the TALEN comprises a TALE protein and a FokI nuclease.