Mitochondrial base mutation correction system for Leber's hereditary optic neuropathy

The base editor system addresses the inadequacy of current LHON treatments by correcting specific mitochondrial DNA mutations, providing a potential cure by restoring normal genotypes and preventing blindness.

JP2026501611APending Publication Date: 2026-01-16EDGENE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025538629
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-12-29
Filing Date
2023-12-28
Publication Date
2026-01-16

AI Technical Summary

Technical Problem

Current treatments for Leber's hereditary optic neuropathy (LHON), caused by mitochondrial DNA mutations, are inadequate, with idebenone being the only available treatment that does not provide a fundamental cure, leading to blindness in patients.

Method used

A base editor system using fusion proteins or polynucleotides that correct specific mitochondrial DNA mutations (G3460A, G11778A, or T14484C) to normal genotypes, employing adenine or cytosine deaminases and DNA-binding proteins like zinc finger proteins or TALE proteins to target and correct mutations in mitochondrial genes.

Benefits of technology

The system effectively corrects mitochondrial DNA mutations, potentially preventing or treating LHON by restoring normal genotypes, offering a more fundamental approach than existing treatments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026501611000005
    Figure 2026501611000005
  • Figure 2026501611000006
    Figure 2026501611000006
  • Figure 2026501611000007
    Figure 2026501611000007
Patent Text Reader

Abstract

The present invention relates to a base correction system that corrects mitochondrial DNA mutations G3460A, G11778A, or T14484C, which are present in patients with Leber's hereditary optic neuropathy (LHON), to a normal genotype. Specifically, the present invention provides a base editor capable of correcting a mutation site in a mitochondrial gene of an LHON patient to a normal genotype. The present invention also provides a method for correcting a mitochondrial gene mutation using a fusion protein or a polynucleotide encoding such a fusion protein that recognizes a specific site in the mitochondrial gene of an LHON patient and specifically corrects the adenine base at position 3460, the adenine base at position 11778, or the cytosine base at position 14484. The base editor or polynucleotide according to the present invention can be used in cells or in an extracellular test tube environment to correct DNA mutations specifically expressed in LHON, and more preferably, can be used as a gene therapy agent to prevent or treat the disease. Thus, the present invention also provides a use of the substance for preventing or treating Leber's hereditary optic neuropathy.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a base correction system that corrects mitochondrial DNA mutations G3460A, G11778A, or T14484C, which are present in patients with Leber's hereditary optic neuropathy (LHON), to a normal genotype. Specifically, the present invention provides a base editor capable of correcting a mutation site in a mitochondrial gene of an LHON patient to a normal genotype. The present invention also provides a method for correcting a mitochondrial gene mutation using a fusion protein or a polynucleotide encoding such a fusion protein that recognizes a specific site in the mitochondrial gene of an LHON patient and specifically corrects the adenine base at position 3460, the adenine base at position 11778, or the cytosine base at position 14484. The base editor or polynucleotide according to the present invention can be used in cells or in an extracellular test tube environment to correct DNA mutations specifically expressed in LHON, and more preferably, can be used as a gene therapy agent to prevent or treat the disease. Thus, the present invention also provides a use of the substance for preventing or treating Leber's hereditary optic neuropathy. [Background technology]

[0002] Mitochondria are organelles within eukaryotic cells that generate the energy (ATP) used for biological activity. They possess independent genes separate from nuclear genes. Mitochondria generate over 80% of the ATP, the energy source used by cells. Mutations in mitochondrial DNA (mtDNA) can cause fatal defects in the central nervous system, heart, muscles, eyesight, and hearing. mtDNA is inherited maternally, meaning that mutations in a mother's mitochondrial DNA are passed on to the next generation. While most patients diagnosed with mtDNA mutations inherit their family history from their mother, approximately 40% of cases have been reported to develop spontaneously. mtDNA mutations that cause mitochondrial disease occur in one in 5,000 people. Hereditary diseases caused by mitochondrial DNA mutations are highly diverse, but precise treatments and prevention methods are unavailable for most. Representative mitochondrial genetic substitutions include Leber hereditary optic neuropathy (LHON), Mitochondrial encephalopathy, lactic acidosis, and stroke-like episode (MELAS), and Leigh syndrome.

[0003] Among the various mitochondrial diseases, LHON was the first genetic disease caused by a mitochondrial gene mutation to be discovered, and was first described in 1871 by German ophthalmologist Theodore Leber. LHON is also known as Leber's optic neuropathy. Unlike other diseases, LHON develops without clear warning signs or pain, progresses rapidly, and can affect people of all ages. The average age of onset is known to be men in their 20s and 30s, and most patients will lose vision in both eyes simultaneously or sequentially after a few months.

[0004] LHON is a major disease that accounts for approximately 30-50% of optic neuropathies of unknown cause that result in loss of vision in both eyes. LHON patients have a G to A substitution at base 3460 of the ND1 gene in mtDNA, a G to A substitution at base 11778 of the ND4 gene, or a T to C substitution at base 14484 of the ND6 gene. LHON progresses by impairing the function of complex 1, which is made up of proteins encoded by these genes. These three point mutations account for more than 90% of LHON cases. Consequently, LHON is known to be a major cause of optic neuropathies of unknown cause in both eyes. Summary of the Invention [Problem to be solved by the invention]

[0005] Despite extensive ongoing research into the genetic disease LHON, there is currently no suitable treatment, leaving patients unable to access alternative treatments and ultimately leading to blindness. At the time of the present invention, idebenone, developed by Santhera Pharmaceuticals and approved under the trade name Raxone, was the only available treatment for LHON. However, while it has the effect of delaying blindness, it has the limitation of not being a fundamental cure. The objective of the present invention is to provide a base editor that corrects the mitochondrial DNA mutation that causes LHON, thereby providing a method for preventing and / or treating LHON. [Means for solving the problem]

[0006] The base correction system according to the present invention uses an adenine base editor having the activity of correcting A at base 3460 of the mitochondrial DNA of an LHON patient to G, an adenine base editor having the activity of correcting A at base 11778 to G, or a cytosine base editor having the activity of correcting C at base 14484 to T. The base editor uses a combination of one or more fusion proteins or polynucleotides encoding the same, where each of the one or more fusion proteins independently contains a DNA binding protein and further contains a deaminase (one or more of adenine deaminase and cytosine deaminase). The DNA-binding protein used in the base correction system of the present invention may be a zinc finger protein (also referred to as "ZF"), a transcriptional activator-like effector (TALE) protein, a CRISPR-linked nuclease, or a combination thereof. The deaminase may be apolipoprotein B editing complex (APOBEC), activation-induced deaminase (AID), tRNA-specific adenosine deaminase (TadA) or a mutant thereof, and DddAtox or a mutant thereof (either full-length or two fragments), or a combination thereof. The fusion protein used in the base correction system of the present invention may additionally include one or more of UGI (uracil glycosylase inhibitor), NES (nuclear export signal), and MTS (mitochondrial targeting sequence). The base editor of the present invention may be in the form of one or more of the above-mentioned fusion proteins or a polynucleotide composition encoding the same. Prior to the present invention, no method was known for correcting point mutations in mitochondrial genes present in LHON patients to normal genotypes through base correction to prevent or treat LHON. [Brief explanation of the drawings]

[0007] [Figure 1] The efficiency of m.C14484T correction using DdCBE (DddA-derived cytosine base editor) was screened. 1397N indicates the N-terminal G1397DddAtox fragment, and 1397C indicates the C-terminal G1397DddAtox fragment. The boxed portion on the left indicates the DNA sequence recognized by the first fusion protein containing 1397N or 1397C, and the boxed portion on the right indicates the DNA sequence recognized by the second fusion protein containing 1397N or 1397C. The boxed "C" indicates the target base, 14484 (T14484C), and the degree of correction indicates the efficiency of C to T correction. The bar graph on the right shows the frequency of 14484C→T correction. [Figure 2] The figure shows the efficiency of m.A3460G correction using TALED (TALE-linked deaminase). 1397N indicates the N-terminal G1397DddAtox fragment, and 1397C indicates the C-terminal G1397DddAtox fragment. The boxed portion on the left indicates the DNA sequence recognized by the first fusion protein containing TALE and 1397N, or TALE, 1397C, and TadA8e. The boxed portion on the right indicates the DNA sequence recognized by the second fusion protein containing TALE and 1397N, or TALE, 1397C, and TadA8e. The boxed "A" indicates the target base, 3460 (G3460A), and the degree of correction indicates the efficiency of A-to-G correction. The bar graph on the right shows the frequency of 3460A-to-G correction. [Figure 3]The efficiency of m.A11778G correction using TALED and ZFD (zinc finger deaminase) was screened. 1397N indicates the N-terminal G1397DddAtox fragment, 1397C indicates the C-terminal G1397DddAtox fragment, and ZF indicates zinc finger protein. The boxed portion on the left indicates the DNA sequence recognized by the first fusion protein containing a TALE or ZF protein and 1397N, or a TALE or ZF protein, 1397C, and TadA8e. The boxed portion on the right indicates the DNA sequence recognized by the second fusion protein containing a TALE and 1397N, or a TALE, 1397C, and TadA8e. The boxed "A" indicates the 11778th base (G11778A) to be corrected, and the degree of correction from A to G is indicated by the intensity. The bar graph on the right shows the frequency of correction from 11778A to G. DETAILED DESCRIPTION OF THE INVENTION

[0008] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one skilled in the art to which this invention belongs. Generally, the terms used herein are well known and commonly used in the art.

[0009] As used herein, the terms "correction," "editing," and "editing" are used interchangeably and refer to the process of altering a nucleic acid sequence by selective mutation of a specific game target, including but not limited to a gene, promoter, open reading frame, or any nucleic acid sequence.

[0010] As used herein, the term "base editor" or "mitochondrial DNA base editor" refers to a substance that has the activity of altering a nucleic acid sequence by selectively mutating a mitochondrial target, and includes a combination of one or more different base editors. As used herein, the term "base editor" or "mitochondrial DNA base editor" can refer to a composition comprising a polypeptide (which may be a fusion protein) or a polynucleotide, depending on the context. That is, as used herein, the term "base correction composition" refers to a combination of one or more different base editors, where the different base editors can be used simultaneously or separately.

[0011] As used herein, the term "target" or "target site" refers to a pre-defined nucleic acid sequence of any composition and / or length. Such target sites include, but are not limited to, a gene, promoter, open reading frame, or any nucleic acid sequence.

[0012] The present invention relates to a base editor that has the activity of correcting mitochondrial DNA mutations in patients with Leber's hereditary optic neuropathy (LHON), wherein the base editor comprises one or more fusion proteins, each of which independently comprises a DNA-binding protein that specifically binds to the mitochondrial DNA of LHON patients, and may be in the form of a composition further comprising either adenine deaminase or cytosine deaminase. The cytosine deaminase may exist in the form of a full-length or two-part fragment.

[0013] The base editor according to the present invention has the activity of correcting the adenine (A) base at position 3460 of the mitochondrial ND1 DNA of an LHON patient to guanine (G), and / or correcting the adenine (A) base at position 11778 of the mitochondrial ND4 DNA to guanine (G), and / or correcting the cytosine (C) base at position 14484 of the mitochondrial ND6 DNA to thymine (T).

[0014] As used herein, the 3460th base of ND1 DNA refers to the 3460th base of all bases constituting mitochondrial DNA, i.e., the base constituting the ND1 gene. The 11778th base of ND4 DNA refers to the 11778th base of all bases constituting mitochondrial DNA, i.e., the base constituting the ND4 gene. The 14484th base of ND6 DNA refers to the 14484th base of all bases constituting mitochondrial DNA, i.e., the base constituting the ND6 gene. This is obvious to those of ordinary skill in the art.

[0015] A cytosine deaminase that can be used in a base editor according to the present invention refers to an amino-deaminase having the activity of converting cytosine to uridine, and can be derived from and / or mutated (e.g., engineered and / or evolved) any organism (e.g., eukaryote or prokaryote), including, but not limited to, algae, bacteria, fungi, plants, invertebrates, and mammals. For example, the cytosine deaminase can be derived from and / or mutated by APOBEC (apolipoprotein B editing complex), AID (activation-induced deaminase), the bacterial adenine deaminase TadA (tRNA-specific adenosine deaminase), or an ortholog thereof, or a cytosine deaminase or fragment thereof derived from and / or mutated by the bacterial cytosine deaminase DddA or an ortholog thereof. The cytosine deaminase mutated from TadA mentioned above may be, for example, a polypeptide in which one or more of the amino acid residues at positions 6, 26, 27, 28, 46, 48, 49, 61, 74, 76, 77, 82, 96, 107, 108, 112, 114, 115, 119, 122, 127, 142, 143, 151, 154, and 158 of the amino acid sequence of SEQ ID NO: 1 are mutated to other amino acids. For example, the polypeptide may be a polypeptide in which the amino acid at position 27 of the amino acid sequence of SEQ ID NO: 1 is mutated to lysine, the amino acid at position 28 to alanine, the amino acid at position 61 to isoleucine, and the amino acid at position 96 to asparagine. With regard to the configuration of cytosine deaminase that can be used in the present invention, reference may be made to content that was already known prior to the filing of this application, including International Patent Application Publications WO2022 / 060185, WO2023 / 086953, etc., which are incorporated by reference in their entirety in this application.

[0016] When a base editor according to the present invention includes a cytosine deaminase, the cytosine deaminase may be in the form of a first segment and a second segment, or may be in the form of a full-length segment. When the cytosine deaminase has the form of a first segment and a second segment, the first segment and the second segment each have a form that allows them to be linked to a DNA-binding protein.

[0017] As used herein, when two proteins are "linked," they can be directly linked or indirectly linked via a linker or another protein (or proteins).

[0018] The cytosine deaminase used in the present invention may be DddAtox, which is a part of a bacterial toxin derived from Burkholderia cenocepacia that exhibits enzymatic function and can deaminate cytosine in double-stranded DNA. DddAtox may comprise the amino acid sequence of SEQ ID NO: 2.

[0019] SEQ ID NO: 2: wild-type DddAtox GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC

[0020] Since DddAtox is toxic to cells, it can be used in the form of two inactivated split bodies, i.e., the first split body and the second split body, to avoid toxicity in host cells. When the cytosine deaminase used in the present invention is used in the form of the first split body and the second split body, each of the first split body and the second split body has no deamination activity.

[0021] The first segment of the DddAtox cytosine deaminase can comprise the sequence from the N-terminus to G33, G44, A54, N68, G82, N98, or G108 in the amino acid sequence of SEQ ID NO: 2. The second segment can comprise the sequence from G34, P45, G55, N69, T83, A99, or A109 in the amino acid sequence of SEQ ID NO: 2 to the C-terminus.

[0022] Preferably, the first segment of DddAtox cytosine deaminase comprises the sequence from the N-terminus to G44 of the amino acid sequence of SEQ ID NO: 2 (SEQ ID NO: 3 below), and the second segment comprises the sequence from P45 to the C-terminus (SEQ ID NO: 4 below).

[0023] SEQ ID NO: 3: wild-type DddAtox G1333-N GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGG

[0024] SEQ ID NO: 4: wild-type DddAtox G1333-C PTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC

[0025] Preferably, the first segment of DddAtox cytosine deaminase comprises the sequence from the N-terminus to G108 of the amino acid sequence of SEQ ID NO: 2 (SEQ ID NO: 5 below), and the second segment may comprise the sequence from A109 to the C-terminus (SEQ ID NO: 6 below).

[0026] SEQ ID NO: 5: wild-type DddAtox G1397-N GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEG

[0027] SEQ ID NO: 6: wild-type DddAtox G1397-C AIPVKRGATGETKVFTGNSNSPKSPTKGGC

[0028] When the first and second segments of DddAtox are used as cytosine deaminase, one or more amino acids located on the surface where the first and second segments of cytosine deaminase bind to each other may be substituted with other amino acids. For example, the first and second segments of DddAtox may have the amino acid sequences of SEQ ID NO: 3 (G1333-N) and SEQ ID NO: 4 (G1333-C), respectively. In this case, one or more amino acids selected from the group consisting of positions 3, 5, 10, 11, 13, 14, 15, 16, 17, 18, 19, 28, 30, and 31 of SEQ ID NO: 3, or one or more amino acids selected from the group consisting of positions 13, 16, 17, 20, 21, 28, 29, 30, 31, 32, 33, 56, 57, 58, and 60 of SEQ ID NO: 4 may be substituted with other amino acids, but are not limited thereto. As another example, the first fragment of DddAtox may comprise the amino acid sequences of SEQ ID NO: 5 (G1397-N) and SEQ ID NO: 6 (G1397-C), in which case one or more amino acids selected from the group consisting of positions 87, 88, 91, 92, 95, 100, 101, 102, and 103 of SEQ ID NO: 5, or one or more amino acids selected from the group consisting of positions 13, 14, 15, and 16 of SEQ ID NO: 6 may be substituted with other amino acids, but is not limited thereto. The "other amino acids" refer to alanine, isoleucine, leucine, methionine, phenylalanine, proline, tryptophan, valine, aspartic acid, cysteine, glutamine, glycine, serine, threonine, tyrosine, aspartate, glutamic acid, arginine, histidine, lysine, and all known variants of these amino acids, excluding the amino acids present in the wild-type protein at the original mutated positions. By using such mutants, the pair of two DddAtox segments, each linked to a DNA-binding protein, does not function properly when not bound to DNA, preventing undesired non-targeted C-to-T correction, and thus enabling highly efficient and precise C-to-T correction.Examples of such mutants include the first fragment of DddAtox having the amino acid sequence of SEQ ID NO: 139 (which may be referred to as "G1397-N" or "G1397N"), and the second fragment of DddAtox having the amino acid sequence of SEQ ID NO: 140 (which may be referred to as "G1397-C" or "G1397C").

[0029] As used herein, the terms "G1333-N," "G1333N," or "1333N" can refer to the first fragment of wild-type DddAtox having the amino acid sequence of SEQ ID NO: 3 or an amino acid variant thereof, and the terms "G1333-C," "G1333C," or "1333C" can refer to the second fragment of wild-type DddAtox having the amino acid sequence of SEQ ID NO: 4 or an amino acid variant thereof.

[0030] As used herein, the terms "G1397-N," "G1397N," or "1397N" can refer to the first fragment of wild-type DddAtox having the amino acid sequence of SEQ ID NO: 5 or 1397 or an amino acid variant thereof, and the terms "G1397-C," "G1397C," or "1397C" can refer to the second fragment of wild-type DddAtox having the amino acid sequence of SEQ ID NO: 6 or 140 or an amino acid variant thereof.

[0031] The cytosine deaminase used in the present invention can be in its full-length form. In this case, the full-length cytosine deaminase (e.g., DddAtox) has its amino acid sequence modified to be non-toxic and / or have only low toxicity. Positively charged amino acids are specifically concentrated at the C-terminus of DddAtox. Because DNA is negatively charged, it binds to positively charged amino acids in proteins. Substitution of these positively charged amino acids weakens the DNA-binding ability of DddAtox, thereby reducing and / or eliminating its intracellular toxicity. In other words, if the positively charged amino acids are substituted to eliminate toxicity, cloning can be performed in E. coli, allowing full-length DddAtox to be obtained. Such a non-toxic full-length cytosine deaminase can be provided by substituting one or more, two or more, three or more, four or more, or five or more amino acids in the wild-type amino acid sequence of SEQ ID NO: 2 with other amino acids. The "other amino acid" refers to an amino acid selected from alanine, isoleucine, leucine, methionine, phenylalanine, proline, tryptophan, valine, aspartic acid, cysteine, glutamine, glycine, serine, threonine, tyrosine, aspartate, glutamic acid, arginine, histidine, lysine, and all known variants of the amino acids, excluding the amino acid that the wild-type protein has at the original mutation position. For example, the other amino acid may be alanine.

[0032] The non-toxic full-length DddAtox may comprise an amino acid sequence selected from the group consisting of the following amino acid sequences:

[0033] A1341D KRKKA mutation GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYDNAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTAGGC

[0034] AAAAA mutant GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVAAGATGETAVFTGNSNSPASPTAGGC

[0035] AAAAK mutant GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVAAGATGETAVFTGNSNSPASPTKGGC

[0036] AAKAA mutation GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVAAGATGETKVFTGNSNSPASPTAGGC

[0037] AAKAK mutant GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVAAGATGETKVFTGNSNSPASPTKGGC

[0038] KAAAA mutant GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKAGATGETAVFTGNSNSPASPTAGGC

[0039] E1347A mutant GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVAGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC

[0040] Preferably, the full-length cytosine deaminase mutant that can be used in the present invention may have one or more amino acid substitutions selected from the group consisting of S to G at position 37, G to S at position 59, A to V at position 109, and S to G at position 129 of the amino acid sequence of SEQ ID NO: 2.

[0041] More preferably, the full-length cytosine deaminase mutant that can be used in the present invention may have all of the following substitutions in the amino acid sequence of SEQ ID NO: 2: S to G at position 37, G to S at position 59, A to V at position 109, and S to G at position 129. In this case, the sequence is as follows:

[0042] GSVG mutants GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLEGKVFSSGGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGVIPVKRGATGETKVFTGNSNGPKSPTKGGC

[0043] As another example, a full-length cytosine deaminase mutant that can be used in the present invention can comprise the following sequence:

[0044] SSVG mutant GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSGSGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGVIPVKRGATGETKVFTGNSNGPKSPTKGGC

[0045] GSAG variant GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLEGKVFSGSGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNGPKSPTKGGC

[0046] GSVS variant GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLEGKVFSGSGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGVIPVKRGATGETKVFTGNSNSPKSPTKGGC

[0047] An adenine deaminase that can be used in a base editor according to the present invention refers to an amino-deaminase that converts adenine bases to inosine. The adenine deaminase can be derived from any organism (e.g., a eukaryote or a prokaryote), including, but not limited to, algae, bacteria, fungi, plants, invertebrates, and mammals, such as E. coli, S. aureus, S. typhi, S. putrefaciens, H. influenzae, or C. crescentus, and / or can be mutated (e.g., engineered and / or evolved). Such an adenine deaminase can be, for example, APOBEC, AID, or TadA, or a mutant thereof. The aforementioned TadA can be, for example, TadA8e (SEQ ID NO: 1) or a truncated or mutant thereof (e.g., a mutant that has been improved and evolved to be applicable to deoxynucleotides). The TadA8e mutants mentioned above may be those in which one or more of the amino acid residues at positions 23, 28, 30, 36, 46, 48, 49, 51, 76, 82, 84, 106, 108, 110, 111, 146, 147, 152, 154, 155, 156, and 157 of SEQ ID NO: 1 have been mutated to other amino acids. Regarding the construction of adenine deaminase that can be used in the present invention, reference may be made to International Patent Application Publications WO2022 / 060185, WO2023 / 086952, etc., which are incorporated herein by reference in their entirety. The adenine deaminase that can be used in the present invention may comprise the amino acid sequence of SEQ ID NO: 1 or a conservative amino acid substitution thereof.

[0048] The DNA-binding protein used in the base editor according to the present invention may be a zinc finger protein, a TALE protein, or a CRISPR-linked nuclease, or a combination thereof. Regarding the configuration of zinc finger proteins, TALE proteins, and CRISPR-linked nucleases, reference may be made to International Patent Application Publications WO2022 / 060185, WO2023 / 086952, etc., which are incorporated herein by reference in their entirety, as well as other prior art documents.

[0049] A zinc finger is a representative DNA-binding protein structure that forms a major DNA-binding protein motif, and the interaction between the α-helix of the zinc finger and the major groove of DNA enables strong and specific recognition of DNA sequences. When one or more zinc finger motifs are combined, the DNA-binding protein used in the base editor of the present invention may comprise an amino acid sequence of any of SEQ ID NOs: 7 to 9.

[0050] The DNA-binding protein used in the base editor according to the present invention may be a TALE protein. The TALE protein of the present invention refers to a protein that binds to nucleotides in a sequence-specific manner via one or more TALE-repeat modules. The TALE protein includes at least one TALE-repeat module, preferably 1 to 30 TALE-repeat modules, but is not limited to this. As used herein, a TALE-repeat module may also be referred to as a "TALE array," and the term "TALE protein" refers to a configuration including an N-terminal domain and a C-terminal domain (which may include half domains) on both sides of the TALE array. As used herein, the term "TALE" may refer only to a "TALE array" or to a "TALE protein" depending on the context. When the DNA-binding protein used in the base editor according to the present invention is a TALE protein, it may include an amino acid sequence of SEQ ID NOs: 10 to 65.

[0051] When a TALE protein is used as the DNA binding protein of the base editor used in the present invention, a single-module TALE array or a multi-module TALE array (e.g., a dual-module TALE array consisting of a first TALE (or left TALE) array and a second TALE (or right TALE) array) can also be used.

[0052] When the base editor used in the present invention comprises two fusion proteins each containing a TALE protein, the two fusion proteins each contain a first TALE protein (or left TALE) and a second TALE protein (or right TALE). The first TALE protein and the second TALE protein can each be independently linked, directly or indirectly (e.g., via a linker and / or other protein components), to at least one of cytosine deaminase and adenine deaminase. For example, a first fusion protein containing the first TALE protein (left TALE) (a fusion protein that binds to DNA 5' upstream from the base correction target site) can contain a first cleavage body of cytosine deaminase, and a second fusion protein containing the second TALE protein (right TALE) (a fusion protein that binds 3' downstream from the base correction target site) can contain a second cleavage body of cytosine deaminase, and either or both of the first and second fusion proteins can contain adenine deaminase. In another approach, a first fusion protein containing a first TALE protein (left TALE) may contain a second fragment of cytosine deaminase, and a second fusion protein containing a second TALE protein (right TALE) may contain a first fragment of cytosine deaminase, and either or both of the first and second fusion proteins may contain adenine deaminase.In another approach, either one of the first fusion protein containing a first TALE protein and the second fusion protein containing a second TALE protein may contain the full-length form of cytosine deaminase, and either or both of the first and second fusion proteins may contain adenine deaminase.

[0053] When the base editor used in the present invention comprises two fusion proteins, one of the two fusion proteins may comprise a TALE protein and the other may comprise a zinc finger protein. For example, the first fusion protein (the fusion protein that binds to DNA 5' upstream of the base correction target site) may comprise a TALE protein (left TALE), and the second fusion protein (the fusion protein that binds to DNA 3' downstream of the base correction target site) may comprise a zinc finger protein (right ZF). In another approach, the first fusion protein (the fusion protein that binds to DNA 5' upstream of the base correction target site) may comprise a zinc finger protein (left ZF), and the second fusion protein (the fusion protein that binds to DNA 3' downstream of the base correction target site) may comprise a TALE protein (right TALE). In either of these two approaches, adenine deaminase may be included in either or both of the first and second fusion proteins.

[0054] When cytosine deaminase is included in a fusion protein comprising a TALE protein or a zinc finger protein, the cytosine deaminase can be linked directly or indirectly (e.g., through a linker and / or other protein components) to the N-terminus or C-terminus (preferably the C-terminus) of the TALE protein, or to the N-terminus or C-terminus (preferably the N-terminus) of the zinc finger protein. When adenine deaminase is included in a fusion protein comprising a TALE protein or a zinc finger protein, the adenine deaminase can be linked directly or indirectly (e.g., through a linker and / or other protein components) to the N-terminus or C-terminus (preferably the C-terminus) of the TALE protein, or to the N-terminus or C-terminus (preferably the N-terminus) of the zinc finger protein. When a fusion protein containing a TALE protein or a zinc finger protein contains both cytosine deaminase and adenine deaminase, the adenine deaminase can be linked directly or indirectly (e.g., via a linker and / or other protein components) to the N-terminus or C-terminus (preferably the C-terminus) of the cytosine deaminase.

[0055] When a base editor according to the present invention comprises two fusion proteins, the fusion proteins may have different compositions and sequences of protein components (e.g., a DNA-binding protein, a cytosine deaminase, and an adenine deaminase). For example, the first fusion protein (the fusion protein that binds to DNA 5' upstream of the base correction target site) may contain an adenine deaminase, while the second fusion protein (the fusion protein that binds to DNA 3' downstream of the base correction target site) may not contain an adenine deaminase, or vice versa. For example, the first fusion protein (the fusion protein that binds to DNA 5' upstream of the base correction target site) may contain a TALE protein, while the second fusion protein (the fusion protein that binds to DNA 3' downstream of the base correction target site) may contain a zinc finger protein, or vice versa.

[0056] The protein components of the fusion protein can be arranged independently of each other in the first fusion protein (the fusion protein that binds to DNA 5' upstream of the base correction target site) and the second fusion protein (the fusion protein that binds to DNA 3' downstream of the base correction target site). For example, in the first fusion protein, adenine deaminase can be located after the C-terminus of cytosine deaminase, while in the second fusion protein, adenine deaminase can be located before the N-terminus of cytosine deaminase, or vice versa. For example, in the first fusion protein, a DNA-binding protein can be located after the C-terminus of cytosine deaminase, while in the second fusion protein, the DNA-binding protein can be located before the N-terminus of cytosine deaminase, or vice versa.

[0057] One or more fusion proteins included in the base editor according to the present invention may each independently further comprise a UGI (uracil glycosylase inhibitor). UGI can increase the efficiency of base correction by inhibiting the activity of UDG (uracil DNA glycosylase), an enzyme that catalyzes the removal of U from DNA and repairs mutated DNA. When UGI is used in the base editor according to the present invention, its position may vary, for example, it may be linked directly or indirectly (e.g., via a linker and / or other protein components) to the C-terminus of cytosine deaminase, but is not limited thereto. When UGI is used in the base editor according to the present invention, UGI may comprise the amino acid sequence of SEQ ID NO: 126.

[0058] One or more fusion proteins included in the base editor according to the present invention may additionally contain a nuclear export signal (NES). Attaching a nuclear export signal to a base correction protein can result in more efficient base correction. The NES sequence may be any signal sequence (e.g., SEQ ID NO: 127) that confers the ability to translocate out of the nucleus, and a natural NES or an artificially synthesized NES may be used. For example, it may be derived from, but is not limited to, MVM (mirute virus of mice). When an NES is used in a base editor according to the present invention, its position may vary. For example, it may be linked directly or indirectly (e.g., via a linker and / or other protein components) to the N-terminus of cytosine deaminase, but is not limited thereto. When an NES is used in a base editor according to the present invention, the NES may comprise the amino acid sequence of SEQ ID NO: 127.

[0059] The base editor (fusion protein) according to the present invention may further comprise an MTS (mitochondrial targeting sequence). The MTS may be any signal sequence capable of translocating into mitochondria, such as a naturally occurring MTS present at the N-terminus of various mitochondrial proteins, or an artificially synthesized MTS. When an MTS is used in a base editor according to the present invention, its position may vary, for example, it may be linked directly or indirectly (e.g., via a linker and / or other protein components) to the N-terminus of a DNA-binding protein or the NES. When an MTS is used in a base editor according to the present invention, the MTS may comprise any one of the amino acid sequences of SEQ ID NOs: 128 to 130.

[0060] The DNA-binding protein used in the base editor of the present invention may contain, in whole or in part, the mitochondrial ND1 DNA sequence 5'-TACGGGCTACTACAACCCTTCGCTGAC A The base editors of the present invention can recognize and bind to the nucleotide sequence 5'-TACGGGCTACTACAACCCTTCG-3' or a portion thereof, and / or 5'-TAAAACTCTTCACCAAAGAGCCCCTAAA-3' or a portion thereof, of the mitochondrial ND1 DNA sequence (the underlined and bolded A is the 3460th base). Preferably, they can recognize the nucleotide sequence 5'-TACGGGCTACTACAACCCTTCG-3' or a portion thereof, and / or 5'-TAAAACTCTTCACCAAAGAGCCCCTAAA-3' or a portion thereof, of the mitochondrial ND1 DNA sequence. When the base editors of the present invention are in the form of a composition of one or more different base editors, one base editor (first fusion protein) can recognize 5'-TACGGGCTACTACAACCCTTCG-3' or a portion thereof, of the mitochondrial DNA sequence, and another base editor (second fusion protein) can recognize 5'-TAAAACTCTTCACCAAAGAGCCCCTAAA-3' or a portion thereof, of the mitochondrial DNA sequence.

[0061] The DNA-binding protein used in the base editor of the present invention may contain, in whole or in part, the mitochondrial ND4 DNA sequence 5'-CAAACTCAAACTACGAACGCACTCACAGTC AThe base editor can recognize and bind to the nucleotide sequence 5'-CATCATAATCCTCTCTCAAGGACTTCAAAC-3' (the underlined and bolded A is the 11,778th base) or a portion thereof. Preferably, it can recognize 5'-CAAACTCAAACTACGAACGCACTCACAGTC-3' or a portion thereof, and / or 5'-CATAATCCTCTCTCAAGGACTTCAAAC-3' or a portion thereof, of the mitochondrial DNA sequence. When the base editor of the present invention is in the form of a composition of one or more different base editors, one base editor (first fusion protein) can recognize 5'-CAAACTCAAACTACGAACGCACTCACAGTC-3' or a portion thereof, of the mitochondrial DNA sequence, and another base editor (second fusion protein) can recognize 5'-CATAATCCTCTCTCAAGGACTTCAAAC-3' or a portion thereof, of the mitochondrial DNA sequence.

[0062] The DNA-binding protein used in the base editor of the present invention may contain, in whole or in part, the mitochondrial ND6 DNA sequence 5'-TAGCCATCGCTGTAGTATATCCAAAGACAACCA C CATTCCCCCTAAATAAATTAAAAAAACTA-3' (the underlined and bolded A is the 14484th base) or a part thereof, preferably 5'-TCGCTGTAGTATATCCAAAGACAACCA CATTCCCCCTAAATAAATTAAAAAAACTA-3' or a portion thereof. Desirably, it can recognize 5'-TCGCTGTAGTATATCCAAAGACA-3' or a portion thereof, and / or 5'-TCCCCCTAAATAAATTAAAAA-3' or a portion thereof, of the mitochondrial DNA sequence. When the base editor according to the present invention is in the form of a composition of one or more different base editors, one base editor (first fusion protein) can recognize 5'-TCGCTGTAGTATATCCAAAGACA-3' or a portion thereof, of the mitochondrial DNA sequence, and another base editor (second fusion protein) can recognize 5'-TCCCCCTAAATAAATTAAAAA-3' or a portion thereof, of the mitochondrial DNA sequence.

[0063] A DNA binding protein for use in a base editor according to the present invention can comprise an amino acid sequence from SEQ ID NO: 7 to SEQ ID NO: 65, or a conservative amino acid substitution thereof.

[0064] A "conservative amino acid substitution" is one in which an amino acid residue is replaced with another amino acid residue having a side chain (R group) with similar chemical properties (e.g., charge or hydrophobicity). Generally, conservative amino acid substitutions do not substantially alter the functional properties of a protein. When two or more amino acid sequences differ from each other by conservative substitutions, the percent sequence identity or degree of similarity may be adjusted upwards to correct for the conservative nature of the substitution. Means for making such adjustments are well known to those skilled in the art (see Pearson (1994) Methods Mol. Biol. 24:307-331). Examples of amino acid groups with chemically similar side chains include: (1) aliphatic side chains: glycine, alanine, valine, leucine, and isoleucine; (2) aliphatic-hydroxyl side chains: serine and threonine; (3) amide-containing side chains: asparagine and glutamine; (4) aromatic side chains: phenylalanine, tyrosine, and tryptophan; (5) basic side chains: lysine, arginine, and histidine; (6) acidic side chains: aspartate and glutamate; and (7) sulfur-containing side chains: cysteine ​​and methionine. Preferred conservative amino acid substitution groups are as follows: valine-leucine-isoleucine, phenylalanine-tyrosine, lysine-arginine, alanine-valine, glutamate-aspartate, and asparagine-glutamine. Alternatively, a conservative substitution can be any change that has a positive value in the PAM250 log-probability matrix described in Gonnet et al. (1992) Science 256:1443-45.

[0065] The base correction composition of the present invention has the activity of correcting the 3460th adenine (A) base of mitochondrial ND1 DNA to guanine (G) in LHON patients, and comprises two fusion proteins, each of which contains a TALE protein and a DddAtox fragment that specifically bind to mitochondrial ND1 DNA, and one of the two fusion proteins may additionally contain TadA8e.

[0066] One fusion protein of the base correction composition may contain a TALE protein (left TALE) comprising an amino acid sequence selected from the group consisting of SEQ ID NO: 10 to SEQ ID NO: 17 or a conservative amino acid substitute thereof, and the other fusion protein may contain a TALE protein (right TALE) comprising an amino acid sequence selected from the group consisting of SEQ ID NO: 18 to 31 or a conservative amino acid substitute thereof.

[0067] The base correction composition of the present invention has the activity of correcting the 11778th adenine (A) base of mitochondrial ND4 DNA to guanine (G) in LHON patients, and comprises two fusion proteins, each of which comprises a TALE or zinc finger protein that specifically binds to mitochondrial ND4 DNA and a DddAtox fragment, and one of the two fusion proteins may additionally contain TadA8e.

[0068] One fusion protein of the base correction composition may comprise a TALE protein (left TALE) comprising an amino acid sequence selected from the group consisting of SEQ ID NO: 32 to SEQ ID NO: 42 or a conservative amino acid substitute thereof, or a zinc finger protein (left ZF) comprising an amino acid sequence selected from the group consisting of SEQ ID NO: 7 to SEQ ID NO: 9 or a conservative amino acid substitute thereof, and the other fusion protein may comprise a TALE protein (right TALE) comprising an amino acid sequence selected from the group consisting of SEQ ID NO: 43 to SEQ ID NO: 53 or a conservative amino acid substitute thereof, or a zinc finger protein (right ZF) comprising an amino acid sequence selected from the group consisting of SEQ ID NO: 7 to SEQ ID NO: 9 or a conservative amino acid substitute thereof.

[0069] The base correction composition of the present invention has the activity of correcting the 14484th cytosine (C) base of mitochondrial ND6 DNA to thymine (T) in LHON patients, and comprises two fusion proteins, each of which may contain a TALE protein and a DddAtox fragment that specifically binds to mitochondrial ND6 DNA.

[0070] One fusion protein of the base correction composition may contain a TALE protein (left TALE) having an amino acid sequence selected from the group consisting of SEQ ID NO: 54 to SEQ ID NO: 9 or a conservative amino acid substitute thereof, and the other fusion protein may contain a TALE protein (right TALE) having an amino acid sequence selected from the group consisting of SEQ ID NO: 58 to 65 or a conservative amino acid substitute thereof.

[0071] The term "fusion protein" as used herein refers to a polypeptide formed by linking two or more different polypeptides via peptide bonds. The fusion protein of the present invention has the activity of correcting A at base 3460 to G, and / or A at base 11778 to G, and / or C at base 14484 to T in the mitochondrial DNA of an LHON patient. It contains a DNA-binding protein, a deaminase (one or more of adenine deaminase and cytosine deaminase), and a UGI, NES, and / or MTS. Methods for designing and constructing a fusion protein (or a polynucleotide encoding the fusion protein) can be any method known in the art. The polynucleotide can be inserted into a vector, and the vector can be introduced into a cell. The individual proteins constituting the fusion protein of the present invention are typically cloned into a single polynucleotide and expressed as a single polypeptide (fusion protein). However, one or more of the individual proteins can also be cloned into separate polynucleotides and expressed as two or more separate polynucleotides, and such cases are also within the scope of the present invention.

[0072] The linker that can be used in the fusion protein of the present invention can be a peptide linker containing 2 to 40 amino acid residues. The length can be, for example, 2, 5, 10, 16, 24, or 32 amino acids, but is not limited thereto. The linker used in the present invention can include, for example, the following:

[0073] GS SGSETPGTSESATPES (SEQ ID NO: 131) SGTPHEVGVYTLSGTPHEVGVYTL (SEQ ID NO: 132) AAEFGIHGVPAAMG (SEQ ID NO: 133) AAEFGIHGVPAAMGGS (SEQ ID NO: 134) SGGS (SEQ ID NO: 135)

[0074] The present invention provides a polynucleotide encoding any one of the one or more fusion proteins contained in the base correction composition. A base editor according to the present invention can include a polynucleotide encoding the fusion protein described above. A base editor according to the present invention can be in the form of a composition of different polynucleotides encoding different base editors.

[0075] Exemplary Compositions of Base Editors (or Base Correction Compositions) According to the Invention

[0076] A base editor (or base correction composition) according to the present invention can comprise a combination of a fusion protein of an amino acid sequence selected from the group consisting of SEQ ID NO:78 to SEQ ID NO:85 and a fusion protein of an amino acid sequence selected from the group consisting of SEQ ID NO:86 to SEQ ID NO:99.

[0077] For example, a base editor (or base corrector composition) according to the present invention can comprise a combination of fusion proteins selected from the group consisting of the following two fusion protein combinations:

[0078] a fusion protein of the amino acid sequence of SEQ ID NO: 78 and a fusion protein of the amino acid sequence of SEQ ID NO: 90; a fusion protein of the amino acid sequence of SEQ ID NO: 78 and a fusion protein of the amino acid sequence of SEQ ID NO: 91; a fusion protein of the amino acid sequence of SEQ ID NO: 79 and a fusion protein of the amino acid sequence of SEQ ID NO: 92; a fusion protein of the amino acid sequence of SEQ ID NO: 79 and a fusion protein of the amino acid sequence of SEQ ID NO: 93; a fusion protein of the amino acid sequence of SEQ ID NO: 79 and a fusion protein of the amino acid sequence of SEQ ID NO: 94; a fusion protein of the amino acid sequence of SEQ ID NO: 79 and a fusion protein of the amino acid sequence of SEQ ID NO: 90; a fusion protein of the amino acid sequence of SEQ ID NO: 79 and a fusion protein of the amino acid sequence of SEQ ID NO: 91; a fusion protein of the amino acid sequence of SEQ ID NO: 80 and a fusion protein of the amino acid sequence of SEQ ID NO: 95; a fusion protein of the amino acid sequence of SEQ ID NO: 80 and a fusion protein of the amino acid sequence of SEQ ID NO: 90; a fusion protein of the amino acid sequence of SEQ ID NO: 80 and a fusion protein of the amino acid sequence of SEQ ID NO: 91; a fusion protein of the amino acid sequence of SEQ ID NO: 81 and a fusion protein of the amino acid sequence of SEQ ID NO: 96; a fusion protein of the amino acid sequence of SEQ ID NO: 81 and a fusion protein of the amino acid sequence of SEQ ID NO: 91; a fusion protein of the amino acid sequence of SEQ ID NO: 82 and a fusion protein of the amino acid sequence of SEQ ID NO: 97; a fusion protein of the amino acid sequence of SEQ ID NO: 82 and a fusion protein of the amino acid sequence of SEQ ID NO: 90; a fusion protein of the amino acid sequence of SEQ ID NO: 83 and a fusion protein of the amino acid sequence of SEQ ID NO: 88; a fusion protein of the amino acid sequence of SEQ ID NO: 83 and a fusion protein of the amino acid sequence of SEQ ID NO: 87; a fusion protein of the amino acid sequence of SEQ ID NO: 83 and a fusion protein of the amino acid sequence of SEQ ID NO: 89; a fusion protein of the amino acid sequence of SEQ ID NO: 84 and a fusion protein of the amino acid sequence of SEQ ID NO: 88; a fusion protein of the amino acid sequence of SEQ ID NO: 84 and a fusion protein of the amino acid sequence of SEQ ID NO: 86; A fusion protein of the amino acid sequence of SEQ ID NO: 84 and a fusion protein of the amino acid sequence of SEQ ID NO: 87; and A fusion protein of the amino acid sequence of SEQ ID NO: 85 and a fusion protein of the amino acid sequence of SEQ ID NO: 88.

[0079] A base editor (or base corrector composition) according to the present invention can comprise a combination of a fusion protein of an amino acid sequence selected from the group consisting of SEQ ID NOs: 100 to 114 and a fusion protein of an amino acid sequence selected from the group consisting of SEQ ID NOs: 115 to 125:

[0080] For example, a base editor (or base corrector composition) according to the present invention can comprise a combination of fusion proteins selected from the group consisting of the following two fusion protein combinations:

[0081] a fusion protein of the amino acid sequence of SEQ ID NO: 100 and a fusion protein of the amino acid sequence of SEQ ID NO: 118; a fusion protein of the amino acid sequence of SEQ ID NO: 101 and a fusion protein of the amino acid sequence of SEQ ID NO: 119; a fusion protein of the amino acid sequence of SEQ ID NO: 101 and a fusion protein of the amino acid sequence of SEQ ID NO: 118; a fusion protein of the amino acid sequence of SEQ ID NO: 102 and a fusion protein of the amino acid sequence of SEQ ID NO: 118; a fusion protein of the amino acid sequence of SEQ ID NO: 103 and a fusion protein of the amino acid sequence of SEQ ID NO: 120; a fusion protein of the amino acid sequence of SEQ ID NO: 103 and a fusion protein of the amino acid sequence of SEQ ID NO: 119; a fusion protein of the amino acid sequence of SEQ ID NO: 103 and a fusion protein of the amino acid sequence of SEQ ID NO: 118; a fusion protein of the amino acid sequence of SEQ ID NO: 113 and a fusion protein of the amino acid sequence of SEQ ID NO: 118; a fusion protein of the amino acid sequence of SEQ ID NO: 104 and a fusion protein of the amino acid sequence of SEQ ID NO: 119; a fusion protein of the amino acid sequence of SEQ ID NO: 104 and a fusion protein of the amino acid sequence of SEQ ID NO: 118; a fusion protein of the amino acid sequence of SEQ ID NO: 104 and a fusion protein of the amino acid sequence of SEQ ID NO: 121; a fusion protein of the amino acid sequence of SEQ ID NO: 104 and a fusion protein of the amino acid sequence of SEQ ID NO: 122; a fusion protein of the amino acid sequence of SEQ ID NO: 105 and a fusion protein of the amino acid sequence of SEQ ID NO: 120; a fusion protein of the amino acid sequence of SEQ ID NO: 105 and a fusion protein of the amino acid sequence of SEQ ID NO: 123; a fusion protein of the amino acid sequence of SEQ ID NO: 105 and a fusion protein of the amino acid sequence of SEQ ID NO: 119; a fusion protein of the amino acid sequence of SEQ ID NO: 105 and a fusion protein of the amino acid sequence of SEQ ID NO: 118; a fusion protein of the amino acid sequence of SEQ ID NO: 105 and a fusion protein of the amino acid sequence of SEQ ID NO: 121; a fusion protein of the amino acid sequence of SEQ ID NO: 105 and a fusion protein of the amino acid sequence of SEQ ID NO: 122; a fusion protein of the amino acid sequence of SEQ ID NO: 106 and a fusion protein of the amino acid sequence of SEQ ID NO: 119; a fusion protein of the amino acid sequence of SEQ ID NO: 106 and a fusion protein of the amino acid sequence of SEQ ID NO: 118; a fusion protein of the amino acid sequence of SEQ ID NO: 106 and a fusion protein of the amino acid sequence of SEQ ID NO: 121; a fusion protein of the amino acid sequence of SEQ ID NO: 106 and a fusion protein of the amino acid sequence of SEQ ID NO: 122; a fusion protein of the amino acid sequence of SEQ ID NO: 106 and a fusion protein of the amino acid sequence of SEQ ID NO: 124; a fusion protein of the amino acid sequence of SEQ ID NO: 111 and a fusion protein of the amino acid sequence of SEQ ID NO: 120; a fusion protein of the amino acid sequence of SEQ ID NO: 111 and a fusion protein of the amino acid sequence of SEQ ID NO: 118; a fusion protein of the amino acid sequence of SEQ ID NO: 107 and a fusion protein of the amino acid sequence of SEQ ID NO: 118; a fusion protein of the amino acid sequence of SEQ ID NO: 108 and a fusion protein of the amino acid sequence of SEQ ID NO: 120; a fusion protein of the amino acid sequence of SEQ ID NO: 108 and a fusion protein of the amino acid sequence of SEQ ID NO: 119; a fusion protein of the amino acid sequence of SEQ ID NO: 108 and a fusion protein of the amino acid sequence of SEQ ID NO: 118; a fusion protein of the amino acid sequence of SEQ ID NO: 108 and a fusion protein of the amino acid sequence of SEQ ID NO: 121; a fusion protein of the amino acid sequence of SEQ ID NO: 108 and a fusion protein of the amino acid sequence of SEQ ID NO: 122; a fusion protein of the amino acid sequence of SEQ ID NO: 108 and a fusion protein of the amino acid sequence of SEQ ID NO: 124; a fusion protein of the amino acid sequence of SEQ ID NO: 112 and a fusion protein of the amino acid sequence of SEQ ID NO: 120; a fusion protein of the amino acid sequence of SEQ ID NO: 112 and a fusion protein of the amino acid sequence of SEQ ID NO: 118; a fusion protein of the amino acid sequence of SEQ ID NO: 112 and a fusion protein of the amino acid sequence of SEQ ID NO: 121; a fusion protein of the amino acid sequence of SEQ ID NO: 112 and a fusion protein of the amino acid sequence of SEQ ID NO: 124; a fusion protein of the amino acid sequence of SEQ ID NO: 114 and a fusion protein of the amino acid sequence of SEQ ID NO: 115; a fusion protein of the amino acid sequence of SEQ ID NO: 114 and a fusion protein of the amino acid sequence of SEQ ID NO: 116; a fusion protein of the amino acid sequence of SEQ ID NO: 109 and a fusion protein of the amino acid sequence of SEQ ID NO: 117; a fusion protein of the amino acid sequence of SEQ ID NO: 109 and a fusion protein of the amino acid sequence of SEQ ID NO: 115; A fusion protein of the amino acid sequence of SEQ ID NO: 109 and a fusion protein of the amino acid sequence of SEQ ID NO: 116; and A fusion protein of the amino acid sequence of SEQ ID NO: 110 and a fusion protein of the amino acid sequence of SEQ ID NO: 115.

[0082] A base editor (or base corrector composition) according to the present invention can comprise a combination of a fusion protein of an amino acid sequence selected from the group consisting of SEQ ID NO:66 to SEQ ID NO:69 and a fusion protein of an amino acid sequence selected from the group consisting of SEQ ID NO:70 to SEQ ID NO:77.

[0083] For example, a base editor (or base corrector composition) according to the invention can comprise a combination of fusion proteins selected from the group consisting of the following two fusion protein combinations:

[0084] a fusion protein of the amino acid sequence of SEQ ID NO: 66 and a fusion protein of the amino acid sequence of SEQ ID NO: 70; a fusion protein of the amino acid sequence of SEQ ID NO: 67 and a fusion protein of the amino acid sequence of SEQ ID NO: 70; a fusion protein of the amino acid sequence of SEQ ID NO: 66 and a fusion protein of the amino acid sequence of SEQ ID NO: 71; a fusion protein of the amino acid sequence of SEQ ID NO: 67 and a fusion protein of the amino acid sequence of SEQ ID NO: 71; a fusion protein of the amino acid sequence of SEQ ID NO: 66 and a fusion protein of the amino acid sequence of SEQ ID NO: 72; a fusion protein of the amino acid sequence of SEQ ID NO: 66 and a fusion protein of the amino acid sequence of SEQ ID NO: 73; a fusion protein of the amino acid sequence of SEQ ID NO: 67 and a fusion protein of the amino acid sequence of SEQ ID NO: 72; a fusion protein of the amino acid sequence of SEQ ID NO: 67 and a fusion protein of the amino acid sequence of SEQ ID NO: 73; a fusion protein of the amino acid sequence of SEQ ID NO: 68 and a fusion protein of the amino acid sequence of SEQ ID NO: 74; a fusion protein of the amino acid sequence of SEQ ID NO: 68 and a fusion protein of the amino acid sequence of SEQ ID NO: 75; a fusion protein of the amino acid sequence of SEQ ID NO: 68 and a fusion protein of the amino acid sequence of SEQ ID NO: 76; a fusion protein of the amino acid sequence of SEQ ID NO: 68 and a fusion protein of the amino acid sequence of SEQ ID NO: 77; a fusion protein of the amino acid sequence of SEQ ID NO: 69 and a fusion protein of the amino acid sequence of SEQ ID NO: 74; A fusion protein of the amino acid sequence of SEQ ID NO: 69 and a fusion protein of the amino acid sequence of SEQ ID NO: 75; and A fusion protein of the amino acid sequence of SEQ ID NO: 69 and a fusion protein of the amino acid sequence of SEQ ID NO: 76.

[0085] A base editor (or base correction composition) according to the present invention can comprise a combination of a polynucleotide that encodes an amino acid sequence selected from the group consisting of SEQ ID NO:78 to SEQ ID NO:85 and a polynucleotide that encodes an amino acid sequence selected from the group consisting of SEQ ID NO:86 to SEQ ID NO:99.

[0086] For example, a base editor (or base corrector composition) according to the present invention can comprise a polynucleotide combination selected from the group consisting of combinations of the following two polynucleotide sequences:

[0087] a polynucleotide encoding the amino acid sequence of SEQ ID NO: 78 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 90; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 78 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 91; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 78 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 92; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 78 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 93; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 78 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 94; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 78 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 90; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 78 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 91; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 78 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 95; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 78 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 90; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 78 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 91; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 78 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 96; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 78 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 91; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 78 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 97; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 78 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 90; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 78 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 88; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 78 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 87; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 78 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 89; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 84 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 88; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 84 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 86; a polynucleotide encoding the amino acid sequence of SEQ ID NO:84 and a polynucleotide encoding the amino acid sequence of SEQ ID NO:87; and A polynucleotide encoding the amino acid sequence of SEQ ID NO:85 and a polynucleotide encoding the amino acid sequence of SEQ ID NO:88.

[0088] A base editor (or base correction composition) according to the present invention can comprise a combination of a polynucleotide that encodes an amino acid sequence selected from the group consisting of SEQ ID NO:100 to SEQ ID NO:114 and a polynucleotide that encodes an amino acid sequence selected from the group consisting of SEQ ID NO:115 to SEQ ID NO:125.

[0089] For example, a base editor (or base corrector composition) according to the present invention can comprise a polynucleotide combination selected from the group consisting of combinations of the following two polynucleotide sequences:

[0090] a polynucleotide encoding the amino acid sequence of SEQ ID NO: 100 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 118; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 101 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 119; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 101 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 118; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 102 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 118; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 103 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 120; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 103 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 119; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 103 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 118; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 113 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 118; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 104 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 119; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 104 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 118; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 104 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 121; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 104 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 122; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 105 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 120; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 105 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 123; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 105 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 119; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 105 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 118; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 105 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 121; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 105 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 122; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 106 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 119; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 106 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 118; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 106 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 121; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 106 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 122; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 106 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 124; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 111 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 120; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 111 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 118; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 107 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 118; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 108 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 120; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 108 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 119; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 108 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 118; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 108 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 121; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 108 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 122; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 108 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 124; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 112 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 120; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 112 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 118; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 112 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 121; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 112 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 124; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 114 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 115; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 114 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 116; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 109 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 117; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 109 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 115; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 109 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 116; and A polynucleotide encoding the amino acid sequence of SEQ ID NO:110 and a polynucleotide encoding the amino acid sequence of SEQ ID NO:115.

[0091] A base editor (or base correction composition) according to the present invention can comprise a combination of a polynucleotide that encodes an amino acid sequence selected from the group consisting of SEQ ID NO:66 to SEQ ID NO:69 and a polynucleotide that encodes an amino acid sequence selected from the group consisting of SEQ ID NO:70 to SEQ ID NO:77.

[0092] For example, a base editor (or base correcting composition) according to the present invention can comprise a polypeptide combination selected from the group consisting of combinations of the following two polynucleotide sequences:

[0093] a polynucleotide encoding the amino acid sequence of SEQ ID NO: 66 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 70; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 67 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 70; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 66 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 71; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 67 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 71; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 66 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 72; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 66 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 73; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 67 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 72; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 67 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 73; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 68 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 74; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 68 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 75; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 68 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 76; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 68 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 77; a polynucleotide encoding the amino acid sequence of SEQ ID NO: 69 and a polynucleotide encoding the amino acid sequence of SEQ ID NO: 74; A polynucleotide encoding the amino acid sequence of SEQ ID NO:69 and a polynucleotide encoding the amino acid sequence of SEQ ID NO:75; and A polynucleotide encoding the amino acid sequence of SEQ ID NO:69 and a polynucleotide encoding the amino acid sequence of SEQ ID NO:76.

[0094] The present invention provides a method for correcting a base mutation in the mitochondrial DNA of an LHON patient, comprising contacting the mitochondrial DNA of an LHON patient with a mitochondrial DNA base editor described herein (or a base correction composition; for example, the combination of fusion proteins or the combination of polynucleotides described under the section "Example Compositions of Base Editors (or Base Correction Compositions) According to the Present Invention"), wherein the correction corrects the adenine (A) base at position 3460 of the mitochondrial ND1 DNA to guanine (G), and / or corrects the adenine (A) base at position 11778 of the mitochondrial ND4 DNA to guanine (G), and / or corrects the cytosine (C) base at position 14484 of the mitochondrial ND6 DNA to thymine (T).

[0095] A mitochondrial DNA base editor according to the present invention can correct the adenine (A) base at position 3460 of mitochondrial ND1 DNA to guanine (G), and / or the adenine (A) base at position 11778 of mitochondrial ND4 DNA to guanine (G), and / or the cytosine (C) base at position 14484 of mitochondrial ND6 DNA to thymine (T) at a frequency of 0.5% or more, 1% or more, 2% or more, 3% or more, 4% or more, 5% or more, 6% or more, 7% or more, 8% or more, 9% or more, 10% or more, 11% or more, 12% or more, 13% or more, 14% or more, 15% or more, 16% or more, 17% or more, 18% or more, 19% or more, or 20% or more. Correction of the adenine (A) base at position 3460 of mitochondrial ND1 DNA, the adenine (A) base at position 11778 of mitochondrial ND4 DNA, or the cytosine (C) base at position 14484 of mitochondrial ND6 DNA according to the present invention means that the corresponding base is changed when compared with the base sequence of the corresponding mitochondrial DNA that has not been contacted with the base correction composition according to the present invention. Whether or not the corresponding base has been changed can be confirmed by DNA sequencing.

[0096] The present invention provides a method for preventing or treating Leber's hereditary optic neuropathy, comprising administering to a patient in need thereof an effective amount of a mitochondrial DNA base editor (or base correcting composition; for example, a combination of fusion proteins or a combination of polynucleotides described under the section "Example Compositions of Base Editors (or Base Correcting Compositions) According to the Present Invention") described herein, i.e., a fusion protein (including a combination of one or more different fusion proteins) having the activity of correcting the adenine (A) base at position 3460 to guanine (G) in mitochondrial ND1 DNA, and / or correcting the adenine (A) base at position 11778 to guanine (G) in mitochondrial ND4 DNA, and / or correcting the cytosine (C) base at position 14484 to thymine (T) in mitochondrial ND6 DNA, or a polynucleotide comprising a gene encoding the fusion protein (including a combination of different polynucleotides each encoding one or more different fusion proteins).

[0097] As used herein, the term "effective amount" refers to an amount of a biologically active agent sufficient to induce a desired biological response. In some embodiments, an effective amount is the amount necessary to ameliorate disease symptoms in an untreated patient. The effective amount of an active ingredient used to practice disease therapeutic methods and treatments may vary depending on the mode of administration, the age, weight, and general health of the subject. In one embodiment, an effective amount is the amount of a base editor (fusion protein, or polynucleotide, or vector or nanoparticle containing the same) described herein sufficient to introduce an alteration in a gene of interest (e.g., mitochondrial DNA) in a cell (e.g., in vitro, in vivo, or ex vivo). In one embodiment, an effective amount is the amount of a base editor (fusion protein, or polynucleotide, or vector or nanoparticle containing the same) necessary to achieve a therapeutic effect (e.g., to reduce and / or control symptoms or pathology in a LHON patient). Such a therapeutic effect need not be sufficient to alter all mitochondrial DNA in all cells of a subject's body, tissue, or organ, but may be sufficient to alter mitochondrial DNA in about 1%, 5%, 10%, 25%, 50%, 75% or more of the cells present in the subject's body, tissue, or organ, or may be sufficient to alter mitochondrial DNA in about 1%, 5%, 10%, 25%, 50%, 75% or more of the total number of mitochondrial DNA copies present in those cells. In one embodiment, an effective amount is sufficient to ameliorate one or more symptoms of LHON.

[0098] The present invention provides a pharmaceutical composition for preventing or treating LHON, comprising an effective amount of the base correction composition or the polynucleotide.The present invention also provides a pharmaceutical composition for preventing or treating Leber's hereditary optic neuropathy, comprising a mitochondrial DNA base editor according to the present invention, i.e., a fusion protein (including a combination of one or more different fusion proteins) having the activity of correcting the adenine (A) base at position 3460 of mitochondrial ND1 DNA to guanine (G), and / or correcting the adenine (A) base at position 11778 of mitochondrial ND4 DNA to guanine (G), and / or correcting the cytosine (C) base at position 14484 of mitochondrial ND6 DNA to thymine (T) in a patient with Leber's hereditary optic neuropathy, or a polynucleotide (including a combination of different polynucleotides encoding one or more different fusion proteins) comprising a gene encoding the fusion protein, and a pharmaceutically acceptable excipient, carrier, or vehicle.

[0099] As used herein, the term "pharmaceutical composition" refers to a composition formulated for pharmaceutical use. In some embodiments, a pharmaceutical composition further comprises a pharmaceutically acceptable carrier. Those of ordinary skill in the pharmaceutical arts will be familiar with pharmaceutical carriers that can be commonly used to formulate the base correction compositions of the present invention for pharmaceutical use. In some embodiments, a pharmaceutical composition may include additional agents (e.g., agents for specific delivery, increased half-life, or other therapeutic compounds). The term "pharmaceutically acceptable carrier" refers to a pharmaceutically acceptable substance, composition, or vehicle, such as a liquid or solid filler, diluent, excipient, manufacturing aid, or solvent encapsulating compound that is accompanied by a compound to carry or transport the compound from one location in the body (e.g., a delivery site) to another location (e.g., an organ, tissue, or body part). A pharmaceutically acceptable carrier is "acceptable" in the sense of being compatible with the other ingredients of the formulation and not harmful to the subject's tissues (e.g., physiological compatibility, sterility, physiological pH, etc.). The terms "excipient," "carrier," "pharmaceutically acceptable carrier," "vehicle," and the like may be used interchangeably.

[0100] The present invention provides a gene vector comprising a polynucleotide encoding any one of one or more fusion proteins contained in the base correcting composition of the present invention. The gene vector may be a viral vector, preferably an adeno-associated viral vector. The gene vector may also be in the form of a lipid nanoparticle or a polymeric nanoparticle. For example, the polynucleotide or combination of polynucleotides of the present invention may be complexed with a lipid or a polymer.

[0101] The present invention provides a gene therapy agent for preventing or treating LHON, comprising the gene vector. The pharmaceutical composition described above may be in the form of a gene therapy agent containing, as an active ingredient, a polynucleotide (including a combination of different polynucleotides encoding one or more different fusion proteins) encoding a fusion protein (including a combination of one or more different fusion proteins) used as a mitochondrial DNA base editor according to the present invention.

[0102] Polynucleotides (including combinations of different polynucleotides encoding one or more different fusion proteins) containing genes encoding fusion proteins (including combinations of one or more different fusion proteins) used as mitochondrial DNA base editors according to the present invention can be delivered to patients. Examples of delivery methods include using viruses as vectors, non-viral methods using synthetic phospholipids or synthetic cationic polymers, and electroporation, which involves applying a transient electrical impulse to the cell membrane to introduce genes. When using viruses as carriers, viruses with a low gene carrying capacity (e.g., adeno-associated viruses (AAVs), which are approximately 4.7 kbp in size) are difficult to use due to the size of the DNA-correcting fusion proteins. However, if a zinc finger protein is used as the DNA-binding protein and a full-length deaminase is used as the deaminase, such viruses can also be used as vectors. The vector may be an adeno-associated virus.

[0103] Based on the above, the present invention may relate to the following (1) to (35), but is not limited thereto.

[0104] (1) A base correction composition having activity for correcting mitochondrial DNA mutations in patients with Leber's hereditary optic neuropathy (LHON), the base correction composition comprising one or more fusion proteins, each of which independently comprises a DNA-binding protein that specifically binds to the mitochondrial DNA of LHON patients, and further comprises one or more of adenine deaminase and cytosine deaminase, wherein the cytosine deaminase is present in a full-length form or in a two-part form.

[0105] (2) A base correcting composition according to (1), which has the activity of correcting the adenine (A) base at position 3460 of mitochondrial DNA to guanine (G), and / or correcting the adenine (A) base at position 11778 of mitochondrial DNA to guanine (G), and / or correcting the cytosine (C) base at position 14484 of mitochondrial DNA to thymine (T) in a patient with LHON.

[0106] (3) A base correcting composition according to (1) or (2), wherein the cytosine deaminase is APOBEC (apolipoprotein B editing complex), AID (activation-induced deaminase), TadA (tRNA-specific adenosine deaminase), or DddAtox, or a mutant thereof.

[0107] (4) In (1) or (2), the cytosine deaminase is DddAtox, and is included in the form of a first segment and a second segment, and one or more amino acids located on the surface where the first segment and the second segment are bonded to each other are substituted with other amino acids.

[0108] (5) In (4), a base correcting composition, wherein the first fragment of DddAtox comprises the amino acid sequence of SEQ ID NO: 5 or 139 or a variant thereof, wherein the variant has one or more amino acids selected from the group consisting of positions 87, 88, 91, 92, 95, 100, 101, 102 and 103 of the amino acid sequence of SEQ ID NO: 5 substituted with other amino acids, and the second fragment of DddAtox comprises the amino acid sequence of SEQ ID NO: 6 or 140 or a variant thereof, wherein the variant has one or more amino acids selected from the group consisting of positions 13, 14, 15 and 16 of the amino acid sequence of SEQ ID NO: 6 substituted with other amino acids.

[0109] (6) The base correcting composition according to (1) or (2), wherein the adenine deaminase is APOBEC, AID, TadA, or a mutant thereof.

[0110] (7) In (6), the base correcting composition, wherein the adenine deaminase comprises the amino acid sequence of SEQ ID NO: 1 or a conservative amino acid substitution thereof.

[0111] (8) In any one of (1) to (7), the base correction composition, wherein the DNA binding protein is selected from the group consisting of zinc finger proteins, TALE proteins, and CRISPR-associated nucleases.

[0112] (9) In any one of (1) to (8), the base correcting composition, wherein the one or more fusion proteins each independently contain a UGI (uracil glycosylase inhibitor).

[0113] (10) In any one of (1) to (9), the base correction composition, wherein the one or more fusion proteins each independently contain a nuclear export signal (NES).

[0114] (11) In any one of (1) to (10), the base correcting composition, wherein the one or more fusion proteins each independently contain a mitochondrial targeting sequence (MTS).

[0115] (12) In any one of (1) to (8), any one DNA binding protein is a mitochondrial ND1 DNA sequence A base corrector composition that binds to the nucleotide sequence TACGGGCTACTACAACCCTTCGCTGACACCATAAAACTCTTCACCAAAGAGCCCCTAAA (5' to 3') or a portion thereof.

[0116] (13) In any one of (1) to (8), any one DNA binding protein is a mitochondrial ND4 DNA sequence, A base corrector composition that binds to the nucleotide sequence CAAACTCAAACTACGAACGCACTCACAGTCACATCATAATCCTCTCTCAAGGACTTCAAAC (5' to 3') or a portion thereof.

[0117] (14) In any one of (1) to (8), any one DNA binding protein is a mitochondrial ND6 DNA sequence, A base corrector composition that binds to the nucleotide sequence TCGCTGTAGTATATCCAAAGACAACCACCATTCCCCCTAAATAAATTAAAAAAACT (5' to 3') or a portion thereof.

[0118] (15) A base correction composition according to any one of (1) to (8), wherein any one of the DNA binding proteins comprises an amino acid sequence selected from the group consisting of SEQ ID NO: 7 to SEQ ID NO: 65, or a conservative amino acid substitution thereof.

[0119] (16) A base correction composition according to (1), comprising two fusion proteins, each of which comprises a TALE protein and a DddAtox fragment that specifically bind to mitochondrial ND1 DNA, one of which additionally contains TadA8e, and which has the activity of correcting the adenine (A) base at position 3460 of mitochondrial DNA to guanine (G) in LHON patients.

[0120] (17) A base correction composition according to (1), comprising two fusion proteins, each of which comprises a TALE or zinc finger protein that specifically binds to mitochondrial ND4 DNA and a DddAtox fragment, one of which additionally contains TadA8e and has the activity of correcting the adenine (A) base at position 11778 of mitochondrial DNA to guanine (G) in LHON patients.

[0121] (18) A base correction composition in (1), comprising two fusion proteins, each of which comprises a TALE protein and a DddAtox fragment that specifically bind to mitochondrial ND6 DNA, and which has the activity of correcting the cytosine (C) base at position 14484 of mitochondrial DNA to thymine (T) in LHON patients.

[0122] (19) In (16), a base correction composition, wherein one fusion protein comprises a TALE protein having an amino acid sequence selected from the group consisting of SEQ ID NO: 10 to SEQ ID NO: 17 or a conservative amino acid substitution thereof, and the other fusion protein comprises a TALE protein having an amino acid sequence selected from the group consisting of SEQ ID NO: 18 to SEQ ID NO: 31 or a conservative amino acid substitution thereof.

[0123] (20) In (17), a base correction composition is provided, wherein one fusion protein comprises a TALE protein having an amino acid sequence selected from the group consisting of SEQ ID NO: 32 to SEQ ID NO: 42 or a conservative amino acid substitute thereof, or a zinc finger protein having an amino acid sequence selected from the group consisting of SEQ ID NO: 7 to SEQ ID NO: 9 or a conservative amino acid substitute thereof, and the other fusion protein comprises a TALE protein having an amino acid sequence selected from the group consisting of SEQ ID NO: 43 to SEQ ID NO: 53 or a conservative amino acid substitute thereof, or a zinc finger protein having an amino acid sequence selected from the group consisting of SEQ ID NO: 7 to SEQ ID NO: 9 or a conservative amino acid substitute thereof.

[0124] (21) In (18), a base correction composition, wherein one fusion protein comprises a TALE protein having an amino acid sequence selected from the group consisting of SEQ ID NO: 54 to SEQ ID NO: 57 or a conservative amino acid substitution thereof, and another fusion protein comprises a TALE protein having an amino acid sequence selected from the group consisting of SEQ ID NO: 58 to SEQ ID NO: 65 or a conservative amino acid substitution thereof.

[0125] (22) In (16) or (19), the base correction composition comprises a combination of a fusion protein of an amino acid sequence selected from the group consisting of SEQ ID NO: 78 to SEQ ID NO: 85 and a fusion protein of an amino acid sequence selected from the group consisting of SEQ ID NO: 86 to SEQ ID NO: 99.

[0126] (23) In (17) or (20), the base correction composition comprises a combination of a fusion protein of an amino acid sequence selected from the group consisting of SEQ ID NO: 100 to SEQ ID NO: 114 and a fusion protein of an amino acid sequence selected from the group consisting of SEQ ID NO: 115 to SEQ ID NO: 125.

[0127] (24) In (18) or (21), the base correction composition comprises a combination of a fusion protein of an amino acid sequence selected from the group consisting of SEQ ID NO: 66 to SEQ ID NO: 69 and a fusion protein of an amino acid sequence selected from the group consisting of SEQ ID NO: 70 to SEQ ID NO: 77.

[0128] (25) A polynucleotide encoding one or more fusion proteins contained in the base composition according to any one of (1) to (24), or a combination of two or more of the polynucleotides.

[0129] (26) In (25), a combination of a polynucleotide encoding an amino acid sequence selected from the group consisting of SEQ ID NO: 78 to SEQ ID NO: 85 and a polynucleotide encoding an amino acid sequence selected from the group consisting of SEQ ID NO: 86 to SEQ ID NO: 99.

[0130] (27) In (25), a combination of a polynucleotide encoding an amino acid sequence selected from the group consisting of SEQ ID NO: 100 to SEQ ID NO: 114 and a polynucleotide encoding an amino acid sequence selected from the group consisting of SEQ ID NO: 115 to SEQ ID NO: 125.

[0131] (28) In (25), a combination of a polynucleotide encoding an amino acid sequence selected from the group consisting of SEQ ID NO: 66 to SEQ ID NO: 69 and a polynucleotide encoding an amino acid sequence selected from the group consisting of SEQ ID NO: 70 to SEQ ID NO: 77.

[0132] (29) A method for correcting a base mutation in the mitochondrial DNA of an LHON patient, comprising contacting the mitochondrial DNA of the LHON patient with a base correction composition according to any one of (1) to (24), wherein the correction is to correct the adenine (A) base at position 3460 of the mitochondrial DNA of the LHON patient to guanine (G), and / or the adenine (A) base at position 11778 of the mitochondrial DNA to guanine (G), and / or the cytosine (C) base at position 14484 of the mitochondrial DNA to thymine (T).

[0133] (30) A method for preventing or treating LHON, comprising administering to a patient in need thereof a base-correcting composition according to any one of (1) to (24) or a composition comprising a polynucleotide or a combination of polynucleotides according to any one of (25) to (28).

[0134] (31) A pharmaceutical composition for preventing or treating LHON, comprising a base-correcting composition according to any one of (1) to (24) or a polynucleotide or a combination of polynucleotides according to any one of (25) to (28).

[0135] (32) A genetic vector comprising a polynucleotide or a combination of polynucleotides according to any one of (25) to (28).

[0136] (33) The gene vector according to (32), which is an adeno-associated virus vector.

[0137] (34) The gene vector according to (32), which is a lipid nanoparticle or a polymeric nanoparticle.

[0138] (35) A gene therapy agent for preventing or treating LHON, comprising a gene vector according to any one of (32) to (34).

[0139] Examples of the present invention will be described below. However, the following examples are for the purpose of illustrating the present invention and are not intended to limit the present invention.

[0140] Example 1: Amplification of template DNA carrying the m.T14484C mutation

[0141] A Japanese strand DNA sequence mimicking the mitochondrial genome of a patient with the m.T14484C mutation was synthesized as a gBlock DNA fragment by IDT (Integrated DNA Technologies). The resulting template DNA sequence is as follows:

[0142] Template DNA was amplified using the forward primer (GACTGGTTCCAATTGACAACG) and reverse primer (GCAAATGGCATTCTGACATCC) and then purified using a PCR purification kit (GeneAll). Immediately before the experiment, the amplified DNA was freshly diluted in distilled water to 10 ng / ul.

[0143] Example 2: Synthesis of a fusion protein to correct the m.T14484C mutation

[0144] The template DNA obtained in Example 1 was quantified at 10 ng / µl, and then 10 ng of template DNA, 0.5 µg of a plasmid containing DNA encoding the first fusion protein, 0.5 µg of a plasmid containing DNA encoding the second fusion protein, and 20 µL of an in vitro coupled transcription / translation (IVTT) kit mixture (containing distilled water at 25 µL) were mixed in a tube and reacted at 30°C for 6 hours, followed by 16 hours at 37°C.

[0145] The first and second fusion proteins are linked in the order of [MTS]-[tag]-[TALE protein]-[linker]-[DddAtox fragment]-[linker]-[UGI]. The tag is 3X HA (SEQ ID NO: 136) or 3X FLAG (SEQ ID NO: 137). The proteins used are located between the CMV promoter and the T7 promoter and termination sequence.

[0146] Example 3: Sequencing to confirm m.T14484C mutation correction and measurement of base correction efficiency

[0147] The reaction product obtained in Example 2 was used as a template without purification and sequenced using targeted deep sequencing. The constructed base editors were screened for the efficiency of m.C14484T correction, and the DNA binding sites and m.C14484T correction efficiencies of 15 base editors with high base correction efficiencies are shown in Figure 1.

[0148] The amino acid sequences of the first fusion protein (including Left TALE) and the second fusion protein (including Right TALE) used in the 15 base editors that were confirmed to have high base correction efficiency are as follows:

[0149] First fusion protein

[0150] SEQ ID NO: 66 MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYA GIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPE QVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPAQVVAIASN GGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALE TVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFISGGPTPYPNYVSAGHVEGQSALFMRDNGISEGLVFHNNPKGTCGFCVNMIETLLPENAAMTVVPPEGSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML (The underlined parts correspond to TALE proteins containing the N-terminal domain and the C-terminal domain (including the half domain).)

[0151] SEQ ID NO: 67 MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYA GIRIQDLRTLGYSQQQQEKIKKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPAQVVAIASNIGGKQALETVQRLLPVL CQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVL CQDHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVL CQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGS GSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFISGGPTPYPNYVSAGHVEGQSALFMRDNGISEGLVFHNNPKGTCGFCVNMIETLLPENAAMTVVPPEGSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML (The underlined parts correspond to TALE proteins containing the N-terminal domain and the C-terminal domain (including the half domain).)

[0152] SEQ ID NO: 68 MASVLTPLLLRGLTGSARRLVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDK GIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGS GSAIPVKRGATGETKVFIGNSNSPKSPTKGGCSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML (The underlined parts correspond to TALE proteins containing the N-terminal domain and the C-terminal domain (including the half domain).)

[0153] SEQ ID NO: 69 MASVLTPLLLRGLTGSARRLVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDK GIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGS GSAIPVKRGATGETKVFIGNSNSPKSPTKGGCSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML (The underlined parts correspond to TALE proteins containing the N-terminal domain and the C-terminal domain (including the half domain).)

[0154] The second fusion protein

[0155] SEQ ID NO: 70 MASVLTPLLLRGLTGSARRLVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDK GIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGS GSAIPVKRGATGETKVFIGNSNSPKSPTKGGCSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML (The underlined parts correspond to TALE proteins containing the N-terminal domain and the C-terminal domain (including the half domain).)

[0156] SEQ ID NO:71 MASVLTPLLLRGLTGSARRLVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDK GIRIQDLRTLGYSQQQQEKPKVRSTVAQHHEALVGHGFTFAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEAL LTVAGELRGPPLQLDTGQLLKIAKRGGVTAAVHAWRNALTGAPLNLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASN IGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTP AQVVAIASNGGGKQALETVQRLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVL CQAHGLTPDQVVAIASNGGGKQALETVQRLPVLCQAHGLTQVVAIASNGGGKQALETVQRLPVLCQAHGLTPAQVVAIASNGGGKQALE TVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASN NGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGKQALETVQRLLPVLCQAHGLTP AQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGS GSAIPVKRGATGETKVFIGNSNSPKSPTKGGCSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML (The underlined parts correspond to TALE proteins containing the N-terminal domain and the C-terminal domain (including the half domain).)

[0157] SEQ ID NO:72 MASVLTPLLLRGLTGSARRLVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDK GIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGS GSAIPVKRGATGETKVFIGNSNSPKSPTKGGCSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML (The underlined parts correspond to TALE proteins containing the N-terminal domain and the C-terminal domain (including the half domain).)

[0158] SEQ ID NO: 73 MASVLTPLLLRGLTGSARRLVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDK GIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGS GSAIPVKRGATGETKVFIGNSNSPKSPTKGGCSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML (The underlined parts correspond to TALE proteins containing the N-terminal domain and the C-terminal domain (including the half domain).)

[0159] SEQ ID NO:74 MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYA GIRIQDLRTLGYSQQQQEKIKKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLVLLP CQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVL CQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVL CQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGS GSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFISGGPTPYPNYVSAGHVEGQSALFMRDNGISEGLVFHNNPKGTCGFCVNMIETLLPENAAMTVVPPEGSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML (The underlined parts correspond to TALE proteins containing the N-terminal domain and the C-terminal domain (including the half domain).)

[0160] SEQ ID NO: 75 MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYA GIRIQDLRTLGYSQQQQEKIKKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALET VQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGK GGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPQ AQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGS GSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFISGGPTPYPNYVSAGHVEGQSALFMRDNGISEGLVFHNNPKGTCGFCVNMIETLLPENAAMTVVPPEGSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML (The underlined parts correspond to TALE proteins containing the N-terminal domain and the C-terminal domain (including the half domain).)

[0161] SEQ ID NO:76 MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYA GIRIQDLRTLGYSQQQQEKIKKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASN GGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVL CQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASN GGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGS GSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFISGGPTPYPNYVSAGHVEGQSALFMRDNGISEGLVFHNNPKGTCGFCVNMIETLLPENAAMTVVPPEGSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML (The underlined parts correspond to TALE proteins containing the N-terminal domain and the C-terminal domain (including the half domain).)

[0162] SEQ ID NO:77 MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYA GIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFISGGPTPYPNYVSAGHVEGQSALFMRDNGISEGLVFHNNPKGTCGFCVNMIETLLPENAAMTVVPPEGSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML (The underlined parts correspond to TALE proteins containing the N-terminal domain and the C-terminal domain (including the half domain).)

[0163] The combinations of the first and second fusion proteins used in the base editor shown in Figure 1 are as follows:

[0164] [Table 1]

[0165] HEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN (The underlined parts correspond to TALE proteins containing the N-terminal domain and the C-terminal domain (including the half domain)).

[0166] SEQ ID NO: 95 MASVLTPLLLRGLTGSARRLVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDK GIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLLVGSAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN (The underlined parts correspond to TALE proteins containing the N-terminal domain and the C-terminal domain (including half domains).)

[0167] SEQ ID NO:96 MASVLTPLLLRGLTGSARRLVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDK GIRIQDLRTLGYSQQQQEKPKVRSTVAQHHEALVGHGFTFAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEAL LTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASN IGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGQALETVQRLLPVLCQAHGLTPAQVVAIASNNGKQALETVQRLLPVLCQAHGLTP DQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNNGGKQALETVQRLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVL CQDHGLTPAQVVAIASNGGGKQALETVQRLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLPVLCQAHGLTPDQVVAIASNGGGKQALE TVQRLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIAS NNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLPVLCQAHGLT PAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNGGGKQALESIVQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLLV GSAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN (The underlined parts correspond to TALE proteins containing the N-terminal domain and the C-terminal domain (including half domains).)

[0168] SEQ ID NO:97 MASVLTPLLLRGLTGSARRLVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDK GIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLLVGSAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN (The underlined parts correspond to TALE proteins containing the N-terminal domain and the C-terminal domain (including half domains).)

[0169] SEQ ID NO: 98 MASVLTPLLLRGLTGSARRLVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDK GIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKYHGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLLV GSAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN (The underlined parts correspond to TALE proteins containing the N-terminal domain and the C-terminal domain (including half domains).)

[0170] SEQ ID NO: 99 MASVLTPLLLRGLTGSARRLVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDK GIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKYHGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLLVGSAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN (The underlined parts correspond to TALE proteins containing the N-terminal domain and the C-terminal domain (including half domains).)

[0171] The combinations of the first and second fusion proteins used in the base editor shown in Figure 2 are as follows:

[0172] [Table 2]

[0173] Example 7: Construction of TALED and ZFD for correction of the m.G11778A mutation

[0174] TALEDs and ZFDs (ZF deaminase) were constructed to allow for an editing window of 1-20 bp, encompassing the G11778A point mutation in mitochondrial DNA ND4 to be corrected. The TALED construction method was similar to that described in Example 4. To construct the ZFD, the sequence encoding the zinc finger protein that binds to the target site was codon-optimized for human expression, and the non-stranded DNA sequence was synthesized from the gBlock DNA fragment by IDT (Integrated DNA Technologies). Using the synthesized gBlock DNA fragment and the expression vector backbone (containing MTS, HA tag, NES, DddAtox1397N or 1397C, and TadA8e) as templates, the DNA fragments required for Gibson assembly were amplified using PrimeSTAR® GXL DNA polymerase (TAKARA) and purified using the PCR SV mini kit (GeneAll). The purified DNA fragments were assembled using the HiFi DNA Assembly Kit (NEB). The procedure from transformation into competent DH5α (enzynomics) E. coli cells to confirmation of the nucleotide sequence by Sanger sequencing (Macrogen) was the same as that described in Example 4.

[0175] The first and second fusion proteins were linked in the following order: [MTS]-[tag]-[TALE protein]-[linker]-[DddAtox fragment], [MTS]-[tag]-[TALE protein]-[linker]-[DddAtox fragment]-[linker]-[TadA8e], [MTS]-[tag]-[NES]-[linker]-[DddAtox fragment]-[linker]-[ZF protein], or [MTS]-[tag]-[NES]-[linker]-[DddAtox fragment]-[linker]-[TadA8e]-[linker]-[ZF protein]. 3X HA (SEQ ID NO: 136) or 3X FLAG (SEQ ID NO: 137) was used as the tag for the fusion protein containing the TALE protein, and 1X HA (SEQ ID NO: 138) was used for the fusion protein containing the ZF protein. The proteins used are located between the CMV promoter and the T7 promoter and termination sequence.

[0176] Example 8: Transformation of fusion proteins into cells derived from urine of a patient carrying the G11778A mutation in the mitochondrial ND4 gene

[0177] Primary cells were isolated from the urine of a LHON patient with the G11778A point mutation in the mitochondrial ND4 gene, and were then aliquoted at passage 2 and cryopreserved.

[0178] The method for culturing urinary derived cells (UDC) derived from LHON patients with the G11778A point mutation was the same as that described in Example 5. 1 μg each of the plasmids encoding the first fusion protein and the second fusion protein, totaling 2 μg, was cultured in UDC cells (1.0 × 10 4UDC cells were transfected using the NEON 10 μL TRANSFECTION KIT (Invitrogen) and the NEON Transfection System (Invitrogen) by electroporation (1350 V, 30 ms, 1 pulse). The transfected UDC cells were placed in 0.1% gelatin-coated 8-well Clear TC-Treated Multiple Well Plates (Corning) pre-filled with culture medium. The transfected cells were maintained in a 5% CO2 atmosphere at 37°C with the culture medium replaced. After 6 days, the cells were harvested, the culture medium was removed, and cell lysis was performed using the same procedure as described in Example 5.

[0179] Example 9: Sequencing to confirm correction of m.G11778A mutation and measurement of base correction efficiency

[0180] The reaction product obtained in Example 8 was used as a template without purification, and the sequence was analyzed by targeted deep sequencing to analyze the base correction ratio of the target site. The method from library construction to deep sequencing was the same as that described in Example 6.

[0181] To screen the efficiency of the engineered base editors for m.A11778G correction, we transfected TALED-TALED combinations, TALED-ZFD hybrid combinations, and ZFD-ZFD combinations. The DNA binding sites and m.A11778G correction efficiencies of the 42 base editor combinations with the highest base correction efficiencies are shown in Figure 3.

[0182] The amino acid sequences of the first and second fusion proteins used in the combination of 42 base editors that were confirmed to have high base correction efficiency are as follows:

[0183] The first fusion protein

[0184] SEQ ID NO: 100 MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYA GIRIQDLRTLGYSQQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKYHGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAH QDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVL CQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGS GSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEG (The underlined parts correspond to TALE proteins containing the N-terminal domain and the C-terminal domain (including half domains).)

[0185] SEQ ID NO: 101 MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYA GIRIQDLRTLGYSQQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKYHGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLC QAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVL CQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGS GSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEG (The underlined parts correspond to TALE proteins containing the N-terminal domain and the C-terminal domain (including half domains).)

[0186] SEQ ID NO: 102 MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYA GIRIQDLRTLGYSQQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKYHGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLC QAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVL CQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGS GSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEG (The underlined parts correspond to TALE proteins containing the N-terminal domain and the C-terminal domain (including half domains).)

[0187] SEQ ID NO: 103 MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAGIRIQDLRTLGYSQQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKYHGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLC QDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPARADAVKKGLGGGS GSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEG (The underlined parts correspond to TALE proteins containing the N-terminal domain and the C-terminal domain (including half domains).)

[0188] SEQ ID NO: 104 MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYA GIRIQDLRTLGYSQQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKYHGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLC QAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNSIVAQLSRPDPALAALTNDHLVALACLGGRPARADAVKKGLGGGS GSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEG (The underlined parts correspond to TALE proteins containing the N-terminal domain and the C-terminal domain (including half domains).)

[0189] SEQ ID NO: 105 MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYA GIRIQDLRTLGYSQQQQEKIKKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNIGGKQALETVQRLLP CQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQHGLTPQVVAIASHDGGKQALETVQRLLPVL CQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGKQALETVQRLLPVL CQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGS GSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEG (The underlined parts correspond to TALE proteins containing the N-terminal domain and the C-terminal domain (including half domains).)

[0190] SEQ ID NO: 106 MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYA GIRIQDLRTLGYSQQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKYHGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLC QAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPQVVAIASHDGKQALETVQRLLPVLC CQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEG (The underlined parts correspond to TALE proteins containing the N-terminal domain and the C-terminal domain (including half domains).)

[0191] SEQ ID NO: 107 MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYA GIRIQDLRTLGYSQQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKYHGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLC QDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHD GSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEG (The underlined parts correspond to TALE proteins containing the N-terminal domain and the C-terminal domain (including half domains).)

[0192] SEQ ID NO: 108 MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYA GIRIQDLRTLGYSQQQQEKIKKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKYHGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPNLLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLC QAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLC QAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLLPVL CQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGS GSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEG (The underlined parts correspond to TALE proteins containing the N-terminal domain and the C-terminal domain (including half domains).)

[0193] SEQ ID NO: 109 MASVLTPLLLRGLTGSARRLVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDK GIRIQDLRTLGYSQQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKYHGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHD CQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNSIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLLVGSAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN (The underlined parts correspond to TALE proteins containing the N-terminal domain and the C-terminal domain (including half domains).)

[0194] SEQ ID NO: 110 MASVLTPLLLRGLTGSARRLVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDK GIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLLV GSAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN (The underlined parts correspond to TALE proteins containing the N-terminal domain and the C-terminal domain (including half domains).)

[0195] SEQ ID NO: 111 MLGFVGRVAAAPASGALRRLTPSASLPPAQLLLRAAPTAVHPVRDYAAQYPYDVPDYAVDEMTKKFGTLTIHDTEKAAEFGIHGVPAAMGGSYALGPYQISAPQLPAYNGQ TVGTFYYVNDAGGLESKVFSSGGPPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGSGTPHEVGVYTLSGTPHEVGVYTL YKCPECGKSFSSKKALTEHQRTHTGEKPYKCPECGKSFSTHLDLIRHQRTHTGEKPYKCPECGKSFSHTGHLLEHQRTHTGEKPFECKDCGKAFIQKSNLIRHQRTH (The underlined parts correspond to zinc finger proteins.)

[0196] SEQ ID NO: 112 MLGFVGRVAAAPASGALRRLTPSASLPPAQLLLRAAPTAVHPVRDYAAQYPYDVPDYAVDEMTKKFGTLTIHDTEKAAEFGIHGVPAAMGGSYALGPYQISAPQLPAYNGQ TVGTFYYVNDAGGLESKVFSSGGPPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGSGTPHEVGVYTLSGTPHEVGVYTL YSCGICGKSFSDSSAKRRHCILHTGEKPYKCPECGKSFSSPADLTRHQRTHLRQKDGERPYKCPECGKSFSTHLDLIRHQRTHTGEKPYKCPECGKSFSHTGHLLEHQRTHTGEKPFECKDCGKAFIQKSNLIRHQRTH (The underlined parts correspond to zinc finger proteins.)

[0197] SEQ ID NO: 113 MLGFVGRVAAAPASGALRRLTPSASLPPAQLLLRAAPTAVHPVRDYAAQYPYDVPDYAVDEMTKKFGTLTIHDTEKAAEFGIHGVPAAMGGSYALGPYQISAPQLPAYNGQ TVGTFYYVNDAGGLESKVFSSGGPPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGSGTPHEVGVYTLSGTPHEVGVYTL YKCPECGKSFSTHLDLIRHQRTHTGEKPYKCPECGKSFSHTGHLLEHQRTHTGEKPFECKDCGKAFIQKSNLIRHQRTHLRQKDGGGSERPYKCPECGKSFSTHLDLIRHQRTHTGEKPYKCDECGKNFTQSSNLIVHKRIHTGEKPYKCPECGKSFSTHLDLIRHQRTH (The underlined parts correspond to zinc finger proteins.)

[0198] SEQ ID NO: 114 MLGFVGRVAAAPASGALRRLTPSASLPPAQLLLRAAPTAVHPVRDYAAQYPYDVPDYAVDEMTKKFGTLTIHDTEKAAEFGIHGVPAAMGGSAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDERE VPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSINSGTPHEVGVYTLSGTPHEVGVYTL YKCPECGKSFSTHLDLIRHQRTHTGEKPYKCPECGKSFSHTGHLLEHQRTHTGEKPFECKDCGKAFIQKSNLIRHQRTHLRQKDGGGSERPYKCPECGKSFSTHLDLIRHQRTHTGEKPYKCDECGKNFTQSSNLIVHKRIHTGEKPYKCPECGKSFSTHLDLIRHQRTH (The underlined parts correspond to zinc finger proteins.)

[0199] The second fusion protein

[0200] SEQ ID NO: 115 MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYA GIRIQDLRTLGYSQQQQEKPKVRSTVAQHHEALVGHGFTFAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKYHGARALEALLTVAGELRG PPLQLDTGQLLKIAKRGGVTAAVHAWRNALTGAPLNLTPAQVVAIASNIGGKQALETVQRLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLC QAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLC QDHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLPVLC QDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVL CQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVL CQDHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVL CQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESVAQLSRPDPALALTNDHLVALACLGGRPALDAVKKGLGGS GSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEG (The underlined parts correspond to TALE proteins containing the N-terminal domain and the C-terminal domain (including half domains).)

[0201] SEQ ID NO: 116 MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYA GIRIQDLRTLGYSQQQQEKPKVRSTVAQHHEALVGHGFTFAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKYHGARALEALLTVAGELRG PPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPAQVVAIASNIGGKQALETVQRLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLC QDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLC QAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLPVLC QAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNNGGKQALETVQRLLPVL CQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTQVVAIASNNGGKQALETVQRLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVL CQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVL CQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESVAQLSRPDPALALTNDHLVALACLGGRPALDAVKKGLGGSGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEG (The underlined parts correspond to TALE proteins containing the N-terminal domain and the C-terminal domain (including half domains).)

[0202] SEQ ID NO: 117 MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYA GIRIQDLRTLGYSQQQQEKPKVRSTVAQHHEALVGHGFTFAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKYHGARALEALLTVAGELRG PPLQLDTGQLLKIAKRGGVTAAVHAWRNALTGAPLNLTPAQVVAIASNNGGKQALETVQRLPVLCQDHGLTPAQVVAIAASNGGGKQALETVQRLLPVLC QAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLC QAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTQVVAIASNNGGKQALETVQRLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLC QAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNNGGKQALETVQRLLPVL CQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVL CQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVL CQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNGGGKQALESIVQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGS GSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEG (The underlined parts correspond to TALE proteins containing the N-terminal domain and the C-terminal domain (including half domains).)

[0203] SEQ ID NO: 118 MASVLTPLLLRGLTGSARRLVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDK GIRIQDLRTLGYSQQQQEKPKVRSTVAQHHEALVGHGFTFAHIVALSQHPAALGTVAVKYQDMiaALPEATHEAIVGVGKQWSGARALEALLTVAGELR GPPLQLDTGQLLKIAKRGGVTAAVHAWRNALTGAPLNLTPEQVVAIASNNGGKQALETVQRLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVL CQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLLPVL CQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLPVL CQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPV LCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLPVLCQDHGLTPAQVVAIASNNGGKQALETVQRLLPV LCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLPV LCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLLV GSAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN (The underlined parts correspond to TALE proteins containing the N-terminal domain and the C-terminal domain (including half domains).)

[0204] SEQ ID NO: 119 MASVLTPLLLRGLTGSARRLVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDK GIRIQDLRTLGYSQQQQEKPKVRSTVAQHHEALVGHGFTFAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKYHGARALEALLTVAGELRG PPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPAQVVAIASNIGGKQALETVQRLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLC QDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVL CQAHGLTPEQVVAIASHDGGKQALETVQRLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLPVL CQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTQVVAIASNIGGKQALETVQRLPVLCQDHGLTPAQVVAIASNNGGKQALETVQRLLPV LCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLPV LCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLPV LCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVQLSRPDPALALTNDHLVALACLGGRPALDAVKKGLLVGSAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN (The underlined parts correspond to TALE proteins containing the N-terminal domain and the C-terminal domain (including half domains).)

[0205] SEQ ID NO: 120 MASVLTPLLLRGLTGSARRLVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDK GIRIQDLRTLGYSQQQQEKPKVRSTVAQHHEALVGHGFTFAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKYHGARALEALLTVAGELRG PPLQLDTGQLLKIAKRGGVTAAVHAWRNALTGAPLNLTPAQVVAIASNGGGKQALETVQRLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLC QAHGLTPEQVVAIASHDGGKQALETVQRLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVL CQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTQVVAIASNIGGKQALETVQRLPVLCQDHGLTPAQVVAIASNNGGKQALETVQRLLPVL CQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPV LCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLPV LCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLPVLCQAHGLTQVVAIASNIGGKQALETVQRLLPV LCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLLV GSAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN (The underlined parts correspond to TALE proteins containing the N-terminal domain and the C-terminal domain (including half domains).)

[0206] SEQ ID NO: 121 MASVLTPLLLRGLTGSARRLVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDK GIRIQDLRTLGYSQQQQEKPKVRSTVAQHHEALVGHGFTFAHIVALSQHPAALGTVAVKYQDMiaALPEATHEAIVGVGKQWSGARALEALLTVAGELR GPPLQLDTGQLLKIAKRGGVTAAVHAWRNALTGAPLNLTPAQVVAIASNGGGKQALETVQRLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVL CQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVL CQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLPVL CQDHGLTPAQVVAIASNGGGKQALETVQRLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLPV LCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLPV LCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPV LCQDHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLLVGSAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN (The underlined parts correspond to TALE proteins containing the N-terminal domain and the C-terminal domain (including half domains).)

[0207] SEQ ID NO: 122 MASVLTPLLLRGLTGSARRLVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDK GIRIQDLRTLGYSQQQQEKPKVRSTVAQHHEALVGHGFTFAHIVALSQHPAALGTVAVKYQDMiaALPEATHEAIVGVGKQWSGARALEALLTVAGELR GPPLQLDTGQLLKIAKRGGVTAAVHAWRNALTGAPLNLTPEQVVAIASNGGGKQALETVQRLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVL CQAHGLTPDQVVAIASNNGGQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVL CQDHGLTPDQVVAIASNNGGKQALETVQRLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVL CQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLPV LCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLPVLCQDHGLTPAQVVAIASNNGGKQALETVQRLPV LCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLPV LCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLLV GSAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN (The underlined parts correspond to TALE proteins containing the N-terminal domain and the C-terminal domain (including half domains).)

[0208] SEQ ID NO: 123 MASVLTPLLLRGLTGSARRLVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDK GIRIQDLRTLGYSQQQQEKPKVRSTVAQHHEALVGHGFTFAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKYHGARALEALLTVAGELRG PPLQLDTGQLLKIAKRGGVTAAVHAWRNALTGAPLNLTPAQVVAIASNIGGKQALETVQRLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLC QAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLPVL CQDHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVL CQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPV LCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPV LCQDHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLPV LCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVQLSRPDPALALTNDHLVALACLGGRPALDAVKKGLLVGSAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN (The underlined parts correspond to TALE proteins containing the N-terminal domain and the C-terminal domain (including half domains).)

[0209] SEQ ID NO: 124 MASVLTPLLLRGLTGSARRLVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDK GIRIQDLRTLGYSQQQQEKIKKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKYHGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLC QAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLLPVL CQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQHGLTPAQVVAIASNIGGKQALETVQRLLPVLC LCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLLV GSAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN (The underlined parts correspond to TALE proteins containing the N-terminal domain and the C-terminal domain (including the half domain).)

[0210] SEQ ID NO: 125 MASVLTPLLLRGLTGSARRLVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDK GIRIQDLRTLGYSQQQQEKPKVRSTVAQHHEALVGHGFTFAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKYHGARALEALLTVAGELRG PPLQLDTGQLLKIAKRGGVTAAVHAWRNALTGAPLNLTPAQVVAIASNNGGKQALETVQRLPVLCQDHGLTPAQVVAIAASNGGGKQALETVQRLLPVLC QAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLPVL CQAHGLTPAQVVAIASNGGGKQALETVQRLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVL CQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLPVLCQDHGLTPAQVVAIASNNGGKQALETVQRLLPV LCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLPV LCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPV LCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNGGGKQALESIVQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLLVGSAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN (The underlined parts correspond to TALE proteins containing the N-terminal domain and the C-terminal domain (including the half domain).)

[0211] The combinations of the first and second fusion proteins used in the base editor shown in Figure 3 are as follows:

[0212] [Table 3]

[0213] [Table 4]

Claims

1. A base-correcting composition having activity for correcting mitochondrial DNA mutations in patients with Leber's hereditary optic neuropathy (LHON), comprising: The base correcting composition comprises one or more fusion proteins, The one or more fusion proteins each independently comprise a DNA-binding protein that specifically binds to mitochondrial DNA of an LHON patient, Adenine deaminase and cytosine deaminase exist in full-length or two split forms. Base corrective composition.

2. In claim 1, The compound has an activity of correcting the adenine (A) base at position 3460 of mitochondrial DNA to guanine (G), and / or correcting the adenine (A) base at position 11778 of mitochondrial DNA to guanine (G), and / or correcting the cytosine (C) base at position 14484 of mitochondrial DNA to thymine (T) in LHON patients. Base corrective composition.

3. In claim 1 or claim 2, The cytosine deaminase is APOBEC (apolipoprotein B editing complex), AID (activation-induced deaminase), TadA (tRNA-specific adenosine deaminase), DddAtox, or a mutant thereof; Base corrective composition.

4. In claim 1 or claim 2, The cytosine deaminase is DddAtox, and is included in the form of a first segment and a second segment, and one or more amino acids located on the surface where the first segment and the second segment are bonded to each other are substituted with other amino acids. Mitochondrial DNA base corrector fusion protein.

5. In claim 4, The first segment of DddAtox comprises the amino acid sequence of SEQ ID NO: 5 or 139 or a variant thereof; The variant is one in which one or more amino acids selected from the group consisting of positions 87, 88, 91, 92, 95, 100, 101, 102, and 103 of the amino acid sequence of SEQ ID NO: 5 are substituted with other amino acids; The second fragment of DddAtox comprises the amino acid sequence of SEQ ID NO: 6 or 140 or a variant thereof; The variant has one or more amino acids selected from the group consisting of positions 13, 14, 15, and 16 of the amino acid sequence of SEQ ID NO: 6 substituted with other amino acids. Base corrective composition.

6. In claim 1 or claim 2, The adenine deaminase is APOBEC, AID, TadA, or a mutant thereof; Base corrective composition.

7. In claim 6, The adenine deaminase comprises the amino acid sequence of SEQ ID NO: 1 or a conservative amino acid substitution thereof. Base corrective composition.

8. In any one of claims 1 to 7, The DNA binding protein is selected from the group consisting of a zinc finger protein, a TALE protein, and a CRISPR-associated nuclease; Base corrective composition.

9. In any one of claims 1 to 8, The one or more fusion proteins each independently contain a uracil glycosylase inhibitor (UGI), Base corrective composition.

10. In any one of claims 1 to 9, The one or more fusion proteins each independently contain a nuclear export signal (NES), Base corrective composition.

11. In any one of claims 1 to 10, The one or more fusion proteins each independently contain a mitochondrial targeting sequence; Base corrective composition.

12. In any one of claims 1 to 8, Any one of the DNA binding proteins is in the mitochondrial ND1 DNA sequence: TACGGGCTACTACAACCCTTCGCTGACACCATAAAACTCTTCACCAAAGAGCCCCTAAA (5' to 3') or a part thereof; Base corrective composition.

13. 10. The method of claim 1, wherein the DNA binding protein is a mitochondrial ND4 DNA sequence: Binding to the nucleotide sequence CAAACTCAAACTACGAACGCACTCACAGTCACATCATAATCCTCTCTCAAGGACTTCAAAC (5' to 3') or a part thereof; Base corrective composition.

14. 10. The method of claim 1, wherein the DNA binding protein is a mitochondrial ND6 DNA sequence: TCGCTGTAGTATATCCAAAGACAACCACCATTCCCCCTAAATAAATTAAAAAAACT (5' to 3') or a part thereof; Base corrective composition.

15. In any one of claims 1 to 8, Any one of the DNA binding proteins comprises an amino acid sequence selected from the group consisting of SEQ ID NO: 7 to SEQ ID NO: 65, or a conservative amino acid substitution thereof; Base corrective composition.

16. In claim 1, The invention comprises two fusion proteins, each of which comprises a TALE protein that specifically binds to mitochondrial ND1 DNA and a DddAtox fragment, and one of the two fusion proteins additionally comprises TadA8e and has the activity of correcting the adenine (A) base at position 3460 of mitochondrial DNA to guanine (G) in LHON patients. Base corrective composition.

17. In claim 1, The invention comprises two fusion proteins, each of which comprises a TALE or zinc finger protein that specifically binds to mitochondrial ND4 DNA and a DddAtox fragment, and one of the two fusion proteins additionally comprises TadA8e and has the activity of correcting the adenine (A) base at position 11778 of mitochondrial DNA to guanine (G) in LHON patients. Base corrective composition.

18. In claim 1, The invention comprises two fusion proteins, each of which comprises a TALE protein that specifically binds to mitochondrial ND6 DNA and a DddAtox fragment, and which have the activity of correcting the cytosine (C) base at position 14484 of mitochondrial DNA to thymine (T) in LHON patients. Base corrective composition.

19. In claim 16, One fusion protein comprises a TALE protein comprising an amino acid sequence selected from the group consisting of SEQ ID NO: 10 to SEQ ID NO: 17 or a conservative amino acid substitute thereof, and the other fusion protein comprises a TALE protein comprising an amino acid sequence selected from the group consisting of SEQ ID NO: 18 to SEQ ID NO: 31 or a conservative amino acid substitute thereof. Base corrective composition.

20. 18. In claim 17, One fusion protein comprises a TALE protein comprising an amino acid sequence selected from the group consisting of SEQ ID NO:32 to SEQ ID NO:42 or a conservative amino acid substitute thereof, or a zinc finger protein comprising an amino acid sequence selected from the group consisting of SEQ ID NO:7 to SEQ ID NO:9 or a conservative amino acid substitute thereof, and another fusion protein comprises a TALE protein comprising an amino acid sequence selected from the group consisting of SEQ ID NO:43 to SEQ ID NO:53 or a conservative amino acid substitute thereof, or a zinc finger protein comprising an amino acid sequence selected from the group consisting of SEQ ID NO:7 to SEQ ID NO:9 or a conservative amino acid substitute thereof. Base corrective composition.

21. In claim 18, One fusion protein comprises a TALE protein comprising an amino acid sequence selected from the group consisting of SEQ ID NO:54 to SEQ ID NO:57 or a conservative amino acid substitute thereof, and the other fusion protein comprises a TALE protein comprising an amino acid sequence selected from the group consisting of SEQ ID NO:58 to SEQ ID NO:65 or a conservative amino acid substitute thereof. Base corrective composition.

22. In claim 16 or claim 19, The present invention relates to a fusion protein of an amino acid sequence selected from the group consisting of SEQ ID NO: 78 to SEQ ID NO: 85 and a fusion protein of an amino acid sequence selected from the group consisting of SEQ ID NO: 86 to SEQ ID NO:

99. Base corrective composition.

23. In claim 17 or claim 20, The present invention relates to a fusion protein of an amino acid sequence selected from the group consisting of SEQ ID NO: 100 to SEQ ID NO: 114 and a fusion protein of an amino acid sequence selected from the group consisting of SEQ ID NO: 115 to SEQ ID NO:

125. Base corrective composition.

24. In claim 18 or claim 21, The present invention relates to a fusion protein of an amino acid sequence selected from the group consisting of SEQ ID NO: 66 to SEQ ID NO: 69 and a fusion protein of an amino acid sequence selected from the group consisting of SEQ ID NO: 70 to SEQ ID NO:

77. Base corrective composition.

25. A polynucleotide encoding any one of one or more fusion proteins contained in the base composition according to any one of claims 1 to 24, or a combination of two or more of said polynucleotides.

26. 26. A combination of a polynucleotide encoding an amino acid sequence selected from the group consisting of SEQ ID NO:78 to SEQ ID NO:85 and a polynucleotide encoding an amino acid sequence selected from the group consisting of SEQ ID NO:86 to SEQ ID NO:99, as described in claim 25.

27. 26. A combination of a polynucleotide encoding an amino acid sequence selected from the group consisting of SEQ ID NO: 100 to SEQ ID NO: 114 and a polynucleotide encoding an amino acid sequence selected from the group consisting of SEQ ID NO: 115 to SEQ ID NO: 125, as described in claim 25.

28. 26. A combination of a polynucleotide encoding an amino acid sequence selected from the group consisting of SEQ ID NO: 66 to SEQ ID NO: 69 and a polynucleotide encoding an amino acid sequence selected from the group consisting of SEQ ID NO: 70 to SEQ ID NO: 77, as described in claim 25.

29. A method for correcting a mitochondrial DNA base mutation in an LHON patient, comprising contacting the mitochondrial DNA of an LHON patient with a base correction composition according to any one of claims 1 to 24, The correction is to correct the adenine (A) base at position 3460 of the mitochondrial DNA of the LHON patient to guanine (G), and / or to correct the adenine (A) base at position 11778 of the mitochondrial DNA to guanine (G), and / or to correct the cytosine (C) base at position 14484 of the mitochondrial DNA to thymine (T). A method for correcting mitochondrial DNA mutations in LHON patients.

30. 26. A method for preventing or treating LHON comprising administering to a patient in need thereof a base correcting composition according to any one of claims 1 to 24 or a composition comprising a polynucleotide or a combination of polynucleotides according to any one of claims 25 to 28. A method for preventing or treating LHON.

31. A base correction composition according to any one of claims 1 to 24 or a polynucleotide or a combination of polynucleotides according to any one of claims 25 to 28, A pharmaceutical composition for preventing or treating LHON.

32. 29. A polynucleotide or a combination of polynucleotides according to any one of claims 25 to 28. Gene vectors.

33. 33. The method of claim 32, wherein the vector is an adeno-associated virus vector. Gene vectors.

34. 33. The method according to claim 32, wherein the nanoparticle is a lipid nanoparticle or a polymeric nanoparticle. Gene vectors.

35. 35. A gene vector according to any one of claims 32 to 34. A gene therapy agent for preventing or treating LHON.