Novel base editor comprising zinc finger protein and nickase

A zinc finger protein-based DNA base editor addresses the delivery challenges of conventional tools by providing efficient and precise base corrections in organelle DNA, enhancing genetic disease treatment and crop improvement.

WO2025259055A1PCT designated stage Publication Date: 2025-12-18EDGENE INC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/008158
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-13
Filing Date
2025-06-13
Publication Date
2025-12-18

AI Technical Summary

Technical Problem

Conventional genome editing tools are unsuitable for correcting DNA sequences in organelles like mitochondria and chloroplasts due to difficulties in delivering guide RNAs and co-expressing components, limiting their effectiveness in genetic disease treatment and crop trait improvement.

Method used

A DNA base editor comprising a zinc finger protein, a nickase, and a deaminase is developed, allowing precise base correction in organelle DNA without generating DNA double-strand breaks, with a small molecular weight for efficient delivery and reduced bystander editing.

Benefits of technology

The DNA base editor achieves comparable editing efficiency to existing systems while minimizing vector size limitations and bystander corrections, enabling targeted base conversions in mitochondrial DNA.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025008158_18122025_PF_FP_ABST
    Figure KR2025008158_18122025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a DNA base editor and a base editing method using a nickase and a zinc finger protein. The present invention is useful for correcting bases in nuclear DNA or organellar DNA and, in particular, for correcting bases in DNA of organelles such as chloroplasts or mitochondria. The DNA base editor according to the present invention has a relatively low molecular weight, and thus can be delivered into an organism while being relatively loosely constrained in terms of vector size. In addition, correction can only occur at a single base within a spacer, thus having the advantage of lowering the possibility of bystander correction.
Need to check novelty before this filing date? Find Prior Art

Description

A novel base editor comprising a zinc finger protein and a nickase

[0001] The present invention relates to the field of gene editing technology. Specifically, the present invention relates to a DNA base editor comprising a zinc finger protein, a nickase, and a deaminase. More particularly, the present invention relates to a DNA base editor comprising one or more DNA binding proteins, a nickase, and a deaminase, wherein at least one of the DNA binding proteins is a zinc finger protein. The present invention also relates to a polynucleotide encoding the DNA base editor, a base editing composition and carrier comprising the DNA base editor or the polynucleotide, and a DNA base editing method using the same. In some embodiments, the present invention relates to DNA base editing in a cellular organelle.

[0002] Fusion proteins that link DNA binding proteins and deaminase enzymes enable the induction of DNA mutations, such as single nucleotide conversions in a targeted manner to replace nucleotides or correct bases in the genome without generating DNA double-strand breaks (DSBs), to correct point mutations that cause genetic disorders, or to introduce desired single nucleotide mutations in prokaryotes and eukaryotic cells such as humans.

[0003] Programmable genome editing tools, such as zinc-finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), clustered regularly interspaced short palindromic repeats (CRISPR) systems, and nucleotide editors composed of nucleotide deaminase proteins and CRISPR-inefficient CRISPR-associated protein 9 (Cas9) variants, have the potential to be used to treat genetic diseases involving mutations in mitochondrial DNA or to improve crop traits in plants by changing the base sequence. However, these conventional genome editing tools are not suitable for correcting the DNA sequence of organelles, including mitochondria and chloroplasts, mainly because of the inability to deliver the guide RNAs required to operate the most widely used CRISPR systems to these organelles, or the difficulty in co-expressing the two compounds in the organelles.

[0004] Prior to the present invention, base correction of organelle DNA was performed using DddA, a bacterial toxin derived from Burkholderia cenocepacia. tox It was possible by using DddA tox DddA is the enzymatic component of a bacterial toxin derived from Burkholderia cenocepacia, which can deaminate cytosine in double-stranded DNA. tox To avoid toxicity in host cells, DddA is toxic to cells. tox It is used by dividing into two inactive splits, and each half can be used as a cytosine base editor (DdCBE, DddA-derived cytosine base editor) that functions as a pair by linking to a DNA binding protein designed to bind to DNA. Meanwhile, DddA is attached to the DNA binding protein. toxBy linking cytosine deaminase with adenine deaminase, which can cause adenine to guanine (A-to-G) correction, adenine bases can be corrected (TALED, see International Patent Application Publication WO 2022 / 060185A).

[0005] Zinc finger proteins (ZFPs) are small DNA-binding protein modules that typically contain a Cys2-His2 structure and specifically recognize DNA sequences of 2-4 base pairs. These zinc finger modules can be repeatedly linked to program binding specificity for a desired base sequence, making them applicable to various gene regulation and editing technologies both in vivo and in vitro. ZFPs have a relatively small molecular weight and can bind to target sequences solely with their protein structure, without relying on RNA guide sequences. This makes them ideal for use in delivery systems with vector size limitations or in environments where RNA is unstable.

[0006] The inventors of the present invention are DddA tox The inventors of the present invention sought to develop a DNA base correction system capable of specifically inducing base correction in cellular DNA without using a base correction system. Furthermore, the inventors of the present invention sought to develop a base correction system that can be delivered into a living organism with relatively small molecular weights and thus relatively less subject to vector size limitations. Furthermore, the inventors of the present invention sought to develop a base correction system that has a low possibility of bystander correction and thus can be useful for the treatment of genetic diseases.

[0007] One aspect of the present invention provides a DNA base editor comprising (i) one or more DNA binding proteins, (ii) a nickase, and (iii) a deaminase, wherein at least one of the one or more DNA binding proteins is a zinc finger protein.

[0008] One aspect of the present invention provides a polynucleotide encoding the DNA base editor.

[0009] One aspect of the present invention provides a base correction composition and a carrier comprising the DNA base editor or polynucleotide.

[0010] One aspect of the present invention provides a DNA base correction method comprising introducing the DNA base editor, polynucleotide, base correction composition or carrier into a cell containing target DNA for base correction.

[0011] Specific embodiments of the present invention may include additional technical features not specifically described in the above summary. These additional features are either specifically disclosed elsewhere in this specification or, if not, are within the scope of common knowledge and comprehensible to those skilled in the art.

[0012] The DNA base editor according to the present invention is useful for specifically correcting DNA bases, and is particularly suitable for precisely correcting bases in organelle DNA, such as mitochondrial DNA. Specifically, it can correct adenine (A) in DNA to guanine (G), or cytosine (C) to thymine (T).

[0013] The DNA base editor of the present invention is DddA tox Even without using TALED, it can exhibit a level of mitochondrial DNA editing effect comparable to that of TALED. In addition, the DNA base editor of the present invention has a relatively small molecular weight, so it can be delivered into a living organism with relatively less restrictions on vector size. Furthermore, even if the length of the spacer region capable of base editing is adjusted to be short, editing can only occur at a single base within the spacer, which has the advantage of reducing the possibility of bystander editing.

[0014] Figure 1 illustrates the DNA sequences recognized by the DNA binding proteins of the base correction data used in the experiment and the spacer region where base correction takes place. The proteins listed in [ ] indicate that they constitute a single fusion protein. For example, [Left 407-TadA8e] represents a fusion protein (linker omitted) in which Left 407 (TALE protein) and TadA8e are linked, and [Right 425-MutH*] represents a fusion protein (linker omitted) in which Right 425 (TALE protein) and MutH* are linked. The outlined rectangles indicate the DNA to which each DNA binding protein used binds. For example, for the first displayed fusion protein pair, Left 407 (TALE protein) included in the [Left 407-1397C-TadA8e] fusion protein binds to DNA containing 5'-TCTAGCCTAGCCGTTT-3' and its complementary sequence, and Right 425 (TALE protein) included in the [Right 425-1397N] fusion protein binds to DNA containing 5'-TGAGTTTGATGCTCACCCT-3' and its complementary sequence. The above descriptions apply equally to other fusion proteins.

[0015] Figure 2 shows the results of base editing of the ND1 gene region of human mitochondria using base editing editors according to the present invention. For each base editing editor used, the correction efficiency at A5, A6, A7, A10, and A12, where A-to-G base editing is possible, is plotted.

[0016] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. Generally, the terms used herein are those well known and commonly used in the art.

[0017] The embodiments described in this specification and the configurations depicted in the drawings are only one embodiment of how the present invention is realized and do not fully represent the technical idea of ​​the present invention. Therefore, it should be understood that there may be various equivalents, modifications, and applicable examples that can replace them at the time of this application. In addition, the various aspects and embodiments described in this specification can be applied to other aspects and embodiments, and all combinations of the various elements described in the present invention fall within the scope of the present invention, and the scope of the present invention cannot be considered limited by the specific description described below.

[0018] In this specification, the use of the singular includes the plural unless specifically stated otherwise. As used herein, it should be noted that the singular form includes plural referents unless the context clearly dictates otherwise. In this application, the use of "or" means "and / or" unless otherwise stated.

[0019] The term “comprising” as used herein, unless otherwise specified, is understood to be an open-ended expression that essentially includes the described components, ingredients, steps, etc., but does not exclude the presence of other components, ingredients, steps, etc. Accordingly, the term “comprising” is interpreted to include the more limited meaning of “consisting of” or “consisting essentially of.”

[0020] The terms "correction," "editing," and "editing" as used herein are used interchangeably and refer to a method of altering a nucleic acid sequence by selectively modifying a specific genomic target. Such specific genomic target includes, but is not limited to, a gene, a promoter, an open reading frame, or any nucleic acid sequence.

[0021] As used herein, the terms "base editor," "base editing system," and "base correction system" are used interchangeably and mean a substance having an activity of changing a nucleic acid sequence by selective mutation of a genomic target, and include a combination of one or more different base editors. The terms "base editor," "base editing system," or "base correction system," as used herein, may be in the form of a polypeptide (which may be a fusion protein), a polynucleotide, or a combination thereof, depending on the context, and may be a composition comprising one or more polypeptides (which may be fusion proteins), polynucleotides, or a combination thereof. Accordingly, the terms "base correction system" or "base correction composition," as used herein, may include one base editor or a combination of two or more different base editors, wherein the different base editors may be used simultaneously or separately.

[0022] The term “base editor” as used herein refers to an artificial gene editing protein that can selectively correct a base in a specific base sequence to another base without double-strand cleavage of the target DNA sequence.

[0023] As used herein, the term "conservative amino acid substitution" refers to the replacement of some amino acids with amino acids of different properties while maintaining structural or functional similarity within a protein. Specifically, it refers to substitutions between amino acids with similar physicochemical properties (e.g., charge, size, hydrophobicity, polarity, etc.), and includes substitutions that substantially maintain the structural stability or biological function of the protein.

[0024] For example, substitutions within the following amino acid groups may be conservative substitutions:

[0025] Hydrophobic amino acid group: Ala, Val, Leu, Ile, Met

[0026] Polar uncharged amino acid group: Ser, Thr, Gln, Asn

[0027] Acidic amino acid group: Asp, Glu

[0028] Basic amino acid group: Lys, Arg, His

[0029] Aromatic amino acid group: Phe, Tyr, Trp

[0030] Such determination of substitution can be performed based on the standard amino acid classification that takes into account the charge, polarity, hydrophobicity, structural similarity, etc. of the amino acids, and a person skilled in the art can objectively determine whether the substitution is conservative by utilizing sequence alignment tools such as BLAST and Clustal Omega and conservation matrices (BLOSUM, PAM, etc.). When a specific amino acid sequence is described in this specification, it is interpreted that a variant in which one or more amino acids in the sequence are substituted with another amino acid corresponding to the conservative substitution is also included in the technical scope of the present invention.

[0031] The term “sequence” in this specification may be interpreted, depending on the context, as a nucleic acid (or polynucleotide) molecule or a protein (or polypeptide) molecule having a given sequence.

[0032] The term “sequence identity” or “sequence homology” as used herein means the number of residues present at the same position when two amino acid sequences or nucleic acid sequences are aligned, expressed as a percentage of the total length of the sequences. When it is said herein that a specific sequence “has at least X% sequence identity,” X can be, for example, 85%, 90%, 95%, 98%, or 99%. Sequence identity is typically calculated using BLAST (Basic Local Alignment Search Tool), ClustalW, EMBOSS, or other known sequence alignment algorithms, and is based on default parameters. For example, when aligning amino acid sequences using BLASTP, the identity value calculated using the BLOSUM62 matrix and gap penalty as default values ​​can be used as a basis. In addition, in this specification, “sequence identity” or “sequence homology” may include a value calculated according to an optimized global alignment or local alignment that takes into account insertions, deletions, substitutions, etc. during sequence alignment, and is also used as a standard for explaining the scope of functional equivalents that can maintain the technical effect of the invention.

[0033] The term "homolog" as used herein refers to a protein or nucleic acid that has homology or similarity to a specific protein or gene sequence and exhibits functionally similar biological activity. Homologs may perform the same or similar function, but may be of different species or may have some variation in the amino acid or base sequence. Homologs may be naturally occurring or artificially modified, and are considered to provide the same technical effect within the scope of the present invention.

[0034] The term "ortholog" as used herein refers to a gene or protein derived from two or more species that share a common ancestor, and thus corresponds to a corresponding gene or protein found in different species but with the same evolutionary origin. Generally, orthologs have a high degree of sequence homology between different species and are known to perform similar biological functions. As used herein, an ortholog of a specific protein may include a variant that performs essentially the same function as the original protein, even if a portion of the sequence contains amino acid substitutions, insertions, or deletions.

[0035] The term "functional variant" as used herein refers to a protein that substantially maintains the basic biological function of a polypeptide (e.g., a protein) having a specific amino acid sequence, or has an activity essentially similar to that of the protein, despite having one or more amino acids in the entire sequence conservatively or non-conservatively substituted, or modified, such as insertion, deletion, or substitution. For example, some differences in the sequence may be considered functional variants of the protein described herein, as long as the protein performs the effective function intended in the present invention, such as substrate recognition, catalytic activity, binding ability, or specificity for a target molecule. Such functional variants may be generated by spontaneous mutation, evolutionary modification, induced mutation, or genetic engineering methods.

[0036] As used herein, the terms "target" or "target site" refer to a pre-identified nucleic acid sequence of any composition and / or length. Such target sites include, but are not limited to, genes, promoters, or any nucleic acid sequence.

[0037] The term "fusion protein" as used herein refers to a protein in which two or more different protein (polypeptide) sequences or functional domains are combined into a single continuous polypeptide chain. Such a fusion protein may retain the original biological function of each component or be endowed with new functional properties, and may include a linker sequence between the components. When designating components of a fusion protein herein, unless otherwise specified, the left-to-right direction refers to the N-terminus to the C-terminus, respectively. Additionally, the linker sequence used may not be explicitly indicated.

[0038] The term "monomeric base editor" as used herein means a base editor that exists in the form of a single fusion protein.

[0039] The term "dimeric base editor" as used herein refers to a base editor that exists in the form of two fusion proteins. Among the two fusion proteins, the fusion protein that binds to a DNA sequence located 5' upstream of the spacer region may be referred to as a "first fusion protein" or "left fusion protein," and the fusion protein that binds to a DNA sequence located 3' downstream of the spacer region may be referred to as a "second fusion protein" or "right fusion protein." Similarly, the DNA binding protein included in the first fusion protein may be expressed by the modifier "first" or "left," and the DNA binding protein included in the second fusion protein may be expressed by the modifier "second" or "right."

[0040] In this specification, binding of a fusion protein to a given nucleotide sequence means that the DNA binding protein included in the fusion protein recognizes and binds to the nucleotide sequence.

[0041] The term “CRISPR-associated nuclease” as used herein, also called Cas, generally refers to a protein that is a component of the CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats) system, binds to a guide RNA, specifically binds to a target nucleotide sequence, and has the activity of cleaving or regulating the sequence. In this specification, “CRISPR-associated nuclease” and “Cas” are used interchangeably.

[0042] 1. DNA base editor

[0043] One aspect of the present invention relates to a DNA base editor comprising (i) one or more DNA binding proteins, (ii) a nickase, and (iii) a deaminase.

[0044] The DNA base editor according to the present invention is characterized in that (i) at least one of the one or more DNA binding proteins is a zinc finger protein.

[0045] A. DNA binding protein

[0046] A DNA base editor according to the present invention comprises one or more DNA binding proteins.

[0047] The "DNA binding protein" used in the DNA base editor according to the present invention means a protein that can recognize a specific base sequence and selectively bind to the sequence, and the binding specificity has a programmable characteristic according to the target sequence. In that sense, the DNA binding protein included in the DNA base editor according to the present invention is any programmable DNA binding protein suitable for use in DNA base editing. Those skilled in the art are well aware of the types of programmable DNA binding proteins suitable for use in DNA base editing. Examples of such DNA-binding proteins include zinc finger proteins and TALE (transcription activator-like effector, TALE) proteins, which can be designed to recognize specific base sequences through modular repeat sequences, dead Cas proteins (e.g., dCas9, dCas12a) that are targeted by guide RNA, and other artificially designed DNA recognition domains. These DNA-binding proteins can be fused to enzymatic proteins (e.g., nickases, cytosine deaminase, or adenine deaminase) to induce a desired biochemical reaction at a specific location in the genome.

[0048] The DNA binding protein used in the DNA base editor according to the present invention is not limited to a specific amino acid sequence or structure, and generally includes a programmable protein or functional variant thereof that has the ability to recognize a specific base sequence and selectively bind to that sequence. The DNA binding protein may be naturally occurring, a variant thereof, or an artificially designed protein. The DNA binding protein that can be used in the DNA base editor according to the present invention is understood to encompass all proteins capable of selectively binding to target DNA to achieve the purpose of the present invention, regardless of the specific sequence or origin.

[0049] In some embodiments, the DNA binding protein may be selected from the group consisting of a zinc finger protein, a TALE protein, a CRISPR-associated nuclease, or a combination thereof. Reference may be made to prior disclosures regarding zinc finger proteins, TALE proteins, and CRISPR-associated nucleases, including WO 2022 / 060185 and WO2022 / 017745, which are incorporated by reference herein in their entirety.

[0050] In the present invention, at least one of the one or more DNA binding proteins included in the DNA base editor is a zinc finger protein.

[0051] In some embodiments involving the use of two or more DNA binding proteins, at least one of the two or more DNA binding proteins is a zinc finger protein.

[0052] The above "zinc finger protein (ZFP)" generally refers to a protein or protein domain that forms a stabilized structure through the binding of zinc ions (Zn²) and has the function of binding to a specific DNA sequence. Such zinc finger proteins include one or more "zinc finger (ZF)" structures. Zinc finger proteins have sequence specificity that allows them to bind to a DNA sequence consisting of a specific 3-4 base pair, and by designing them by continuously combining multiple zinc finger domains, they can have high specificity and binding affinity for a long target DNA sequence.

[0053] The zinc finger protein used in the present invention includes a naturally occurring protein or an artificial recombinant, mutant, functional variant, or variant with improved specificity derived therefrom, and is not limited in origin or sequence composition as long as it can specifically bind to a desired target DNA sequence. Those skilled in the art can design and produce a desired zinc finger protein using previously disclosed ZFP libraries, genetic engineering methods, and techniques for analyzing DNA-binding specificity.

[0054] Zinc finger proteins have relatively small molecular weights and can bind to target sequences solely with their protein structure, without relying on RNA guide sequences. This makes them suitable for use in delivery systems with vector size limitations or in environments where RNA is unstable. Furthermore, the use of zinc finger proteins as DNA-binding proteins in DNA base editors allowed for a reduction in the spacer region, which serves as the window for DNA editing. Furthermore, it was confirmed that editing occurred at only a single base within the spacer region, thereby reducing the likelihood of potential bystander editing.

[0055] In some embodiments, one or more DNA binding proteins included in the DNA base editor may all be zinc finger proteins.

[0056] In some embodiments, one or more DNA binding proteins included in the DNA base editor may comprise a TALE protein in addition to a zinc finger protein.

[0057] In some embodiments involving the use of two or more DNA binding proteins, at least one of the two or more DNA binding proteins is a zinc finger protein and at least one of the other is a TALE protein.

[0058] The above "TALE protein" is generally based on a transcription activator-like factor derived from the plant pathogenic bacteria Xanthomonas genus, and refers to a DNA binding protein with sequence specificity that can bind to a specific DNA sequence. TALE proteins consist of a series of repeat modules, each module consisting of about 34 amino acids, of which two amino acid residues at positions 12 and 13 (so-called RVD, repeat-variable diresidue) determine binding specificity for a single DNA base. By designing a combination of these modules, a TALE protein with sequence specificity tailored to a desired target DNA sequence can be generated. As used herein, the TALE-repeat modules may be referred to as a "TALE array," and the expression "TALE protein" refers to a configuration in which an N-terminal domain and a C-terminal domain (which may include a half domain) are included on both sides of the TALE array, respectively.

[0059] TALE proteins that can be used in some embodiments of the present invention include naturally occurring TALE sequences or recombinants, mutants, functional variants, or forms with artificially controlled sequence specificity derived therefrom, and are not limited in sequence composition or origin as long as they can bind to a desired DNA sequence. Those skilled in the art can utilize published TALE libraries and TALE design algorithms to implement TALE proteins that specifically bind to various DNA target sequences.

[0060] TALE proteins have the advantage of not requiring guide RNA, precisely recognizing target DNA sequences by directly assembling repeat modules at the protein level, and being free from constraints on PAM sequences. Therefore, TALE proteins are particularly advantageous in complex genomic environments or situations requiring flexible targeting.

[0061] In some embodiments, one or more DNA binding proteins included in the DNA base editor may comprise a CRISPR-associated nuclease in addition to a zinc finger protein.

[0062] In some embodiments involving the use of two or more DNA binding proteins, at least one of the two or more DNA binding proteins is a zinc finger protein and at least one of the other is a CRISPR-associated nuclease.

[0063] The above "CRISPR-associated nuclease" is also called "Cas protein" and generally refers to a protein having nuclease activity capable of cleaving DNA or RNA derived from the CRISPR (clustered regularly interspaced short palindromic repeats)-Cas system, which is an acquired immune system of bacteria or archaea. These Cas proteins generally form a complex with a guide RNA, recognize a target nucleic acid sequence complementary to the base sequence of the guide RNA, and then induce cleavage (nicking or double-strand break) at the corresponding site. Representative examples include Cas9 (e.g., Streptococcus pyogenesCas9), Cas12a (Cpf1), Cas12b, Cas13, and Cas14.

[0064] CRISPR-associated nucleases that may be used in some embodiments of the present invention may include naturally occurring proteins, functional variants, conservative amino acid substitutions, truncated forms, or variants in which the enzymatic activity is altered or eliminated, and also include inactive forms (dead Cas or dCas), nickase forms (nCas), or forms that include fusions with various functional domains (e.g., deaminase, transcription factor, etc.).

[0065] B. Nickadze

[0066] The DNA base editor according to the present invention comprises a nickase.

[0067] The term "nickase" above refers to an endonuclease or modified nuclease that can cleave only one strand of double-stranded DNA. Unlike typical DNA cleavage enzymes, nickases induce single-strand breaks (nicks) rather than double-strand breaks.

[0068] The nickase used in the present invention is not limited to a specific protein, amino acid sequence, structure, origin, or mechanism of action, and generally refers to an enzyme having an activity that can selectively induce a nick only on one strand of a double-stranded DNA. It includes all proteins of natural origin, recombinant, or artificially modified, and may also include functional fragments, functional variants, conservative amino acid substitutions, truncated forms, or fusion proteins designed to have specificity for a specific target sequence. Therefore, in the present specification, “nickase” is used as a concept encompassing all proteins capable of selectively inducing a nick on one strand of a target double-stranded DNA, or a combination having such activity.

[0069] The nickase may be selected from the group consisting of MutH, BspD6I, FokI (including homodimers or heterodimers), BsaI, BsmBI, BsmAI, BsrDI, CviPII, BspQI, AlwI, and I-TevI, catalytically active fragments thereof, and mutants thereof, or conservative amino acid substitutions thereof. The nickase may be used in conjunction with a DNA binding protein and / or another protein as a single fusion protein, or may be used to be expressed as a separate protein. When more than one fusion protein is used in the base correction system according to the present invention, the nickase may be included in only one of the more than one fusion protein, or the same or different nickases may be included in the plurality of fusion proteins. Furthermore, when more than one fusion protein is used in the base correction system according to the present invention, the nickase may be included in a fusion protein comprising a DNA binding protein that binds to DNA located 5' upstream or 3' of the target DNA site for base correction.

[0070] In some embodiments, the nickase is MutH.

[0071] The above “MutH” is a DNA repair-related enzyme commonly found in some Gram-negative bacteria, including Escherichia coli, and is an endonuclease involved in the mismatch repair (MMR) pathway. MutH recognizes an unmethylated GATC sequence within a DNA double strand and induces a single-strand break (nick) at that site, thereby playing a role in the initiation step of mismatch repair. MutH that can be used in some embodiments is not necessarily limited to the naturally occurring wild-type MutH, but also includes functional variants, orthologs, or artificial proteins that have been improved for specific purposes while maintaining the function thereof.

[0072] In some embodiments, the nickase has an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, or 100% identical to MutH or an ortholog thereof having the following amino acid sequence, or is a conservative amino acid substitution thereof.

[0073] MutH:

[0074] SQPRPLLSPPETEEQLLAQAQQLSGYTLGELAALAGLVTPENLKRDKGWIGVLLEIWLGASAGSKPEQDFAALGVELKTIPVDSLGRPLETTFVCVAPLTGNSGVTWETSHVRH KLKRVLWIPVEGERSIPLAKRRVGSPLLWSPNEEEDRQLREDWEELMMIVLGQIERITARHGEYLQIRPKAANAKALTEAIGARGERILTLPRGFYLKKNFTSALLARHFLIQ (SEQ ID NO: 1)

[0075] In some embodiments, the nickase has an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, or 100% identical to MutH* having the following amino acid sequence, or is a conservative amino acid substitution thereof.

[0076] MutH*:

[0077] SQPRPLLSPPETEEQLLAQAQQLSGYTLGELAALAGLVTPENLKRDKGWIGVLLEIWLGASAGSKPEQDFAALGVELKTIPVDSLGRPLATTAVCVAPLTGNSGVTWETSHVRH KLKRVLWIPVEGERSIPLAKRRVGSPLLWSPNEEEDRQLREDWEELMMIVLGQIERITARHGEYLQIRPKAANAKALTEAIGARGERILTLPRGFYLKKNFTSALLARHFLIQ (SEQ ID NO: 2)

[0078] In some embodiments, the nickase is BspD6I.

[0079] The above “BspD6I” generally refers to a nuclease that is a restriction enzyme derived from a Bacillus species, which recognizes a specific DNA base sequence (e.g., G^TCTC or a similar sequence), and has the activity of cleaving double-stranded DNA or inducing a single-strand break (nick) within or adjacent to the recognition sequence. The BspD6I that can be used in some embodiments is not necessarily limited to the naturally occurring wild-type BspD6I, but also includes functional variants, orthologs, or artificial proteins that have been improved for a specific purpose while maintaining the function thereof.

[0080] In some embodiments, the nickase has an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, or 100% identical to BspD6I(C) or an ortholog thereof having the following amino acid sequence, or is a conservative amino acid substitution thereof.

[0081] BspD6I(C):

[0082] RQLEEVIDLLEVYHEKKNVIEEKIKARFIANKNTVFEWLTWNGFIILGNALEYKNNFVIDEELQPVTHAAGNQPDMEIIYEDFIVLGEVTTSKGATQFKMESEPVTRHYLN KKKELEKQGVEKELYCLFIAPEINKNTFEEFMKYNIVQNTRIIPLSLKQFNMLLMVQKKLIEKGRRLSSYDIKNLMVSLYRTTIECERKYTQIKAGLEETLNNWVVDKEVRF (SEQ ID NO: 3)

[0083] C. Deaminase

[0084] The DNA base editor according to the present invention comprises a deaminase.

[0085] The "deaminase" included in the DNA base editor according to the present invention refers to a biological catalyst protein that changes the chemical structure of a base by removing the amino group (-NH2) present in a specific nucleotide base. Through the deamination reaction, cytosine (C) is converted to uracil (U), and adenine (A) is converted to inosine (I). Since inosine is recognized as guanine (G) during DNA replication or transcription, the conversion of A·T → G·C becomes possible.

[0086] The deaminase included in the DNA base editor according to the present invention is not limited in its origin, structure, amino acid sequence, or substrate specificity, and includes, for example, (i) a cytosine deaminase capable of deaminating cytosine to convert it into uracil, and (ii) an adenine deaminase capable of deaminating adenine to convert it into inosine.

[0087] C1. Cytosine deaminase

[0088] In some embodiments, the deaminase included in the DNA base editor according to the present invention is a cytosine deaminase.

[0089] The above "cytosine deaminase" generally refers to an enzyme that catalyzes the deamination reaction that converts the cytosine base in DNA to uracil. This enzyme converts cytosine to uracil by removing the amino group (-NH2) of cytosine, which can result in a C:G → T:A conversion in the corresponding base pair.

[0090] The cytosine deaminase that can be used in some embodiments of the present invention is not limited to a specific amino acid sequence, structural characteristics, biological species or name, and includes all wild-type proteins, artificial or evolutionary modifications, functional variants, orthologs, conservative amino acid substitutions, etc., as long as the enzyme has a biological activity that can act on a cytosine base on DNA to induce a deamination reaction. Those skilled in the art can easily access numerous literatures and public databases (e.g., GenBank, UniProt, REBASE, etc.) that are already known regarding the sequence, structure, and function of such enzymes, and can implement a cytosine deaminase for implementing selective DNA editing according to a specific purpose.

[0091] For example, proteins of the APOBEC1 or AID (activation-induced cytidine deaminase) family are well-known representative cytosine deaminases, and they have the activity of inducing deamination reactions by acting on single-stranded DNA around specific base sequences.

[0092] In some embodiments, the cytosine deaminase is APOBEC1.

[0093] The above “APOBEC1” is an abbreviation for Apolipoprotein B mRNA Editing Catalytic Polypeptide 1, and refers to an enzyme that has the activity of deaminating cytosine to uracil. APOBEC1 is originally known as an enzyme involved in editing apolipoprotein B mRNA in mammalian hepatocytes, etc., but this enzyme can exhibit the activity of deaminating cytosine bases within a specific sequence of DNA or RNA. In the present invention, APOBEC1 is not necessarily limited to the naturally occurring wild type, and also includes functional variants, orthologs, or artificially improved proteins exhibiting similar activity.

[0094] In some embodiments, the cytosine deaminase has an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, or 100% identical to APOBEC1 or an ortholog thereof having the following amino acid sequence, or is a conservative amino acid substitution thereof.

[0095] APOBEC1:

[0096] MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSIWRHTSQNTNKHVEVNFIEKFTTERYFCPNTRCSITWFLSWSPCGECSRAITEFLSRYPHVTLFI YIARLYHHADPRNRQGLRDLISSGVTIQIMTEQESGYCWRNFVNYSPSNEAHWPRYPHLWVRLYVLELYCIILGLPPCLNILRRKQPQLTFFTIALQSCHYQRLPPHILWATGLK (SEQ ID NO: 4)

[0097] In some embodiments, the cytosine deaminase is AID.

[0098] The above “AID” ​​stands for Activation-Induced Cytidine Deaminase, which refers to a cytosine deaminase that plays an essential role in the generation of antibody diversity in the immune system. The AID that can be used in some embodiments of the present invention is not necessarily limited to the naturally occurring wild type, but also includes functional variants thereof, orthologs, or artificially improved proteins that exhibit similar activity.

[0099] In some embodiments, the cytosine deaminase has an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, or 100% identical to AID or an ortholog thereof having the following amino acid sequence, or a conservative amino acid substitution thereof.

[0100] AID:

[0101] MDSLLMNRRKFLYQFKNVRWAKGRRETYLCYVVKRRDSATSFSLDFGYLRNKNGCHVELLFLRYISDWDLDPGRCYRVTWFTSWSPCYDCARHVADFLRGNPNLSLRIFTARLYFCEDRKAEPEGLRRLHRAGVQIAIMTFKDYFYCWNTFVENHERTFKAWEGLHENSVRLSRQLRRILLPLYEVDDLRDAFRTLGL (SEQ ID NO. 5)

[0102] In some embodiments, the cytosine deaminase is a variant of TadA8e that has been mutated to have cytosine deaminase activity. Such variants may include, for example, one or more of amino acid residues 6, 27, 28, 46, 48, 49, 61, 74, 76, 77, 82, 96, 107, 108, 112, 114, 115, 119, 122, 127, 142, 143, 151, 154, and 158 of the amino acid sequence of TadA8e (SEQ ID NO: 7) being mutated to a different amino acid. With regard to the composition of such cytosine deaminase, reference may be made to WO 2022 / 060185, WO 2023 / 086953, etc., which are incorporated by reference in their entirety into this application, and which were already known prior to the present application.

[0103] Preferably, when a cytosine deaminase is used in the base correction system according to the present invention, the cytosine deaminase is not DddAtox, which is a bacterial cytosine deaminase.

[0104] Preferably, when a cytosine deaminase is used in the base correction system according to the present invention, the cytosine deaminase has deaminase activity for single-stranded DNA.

[0105] C2. Adenine deaminase

[0106] In some embodiments, the deaminase included in the DNA base editor according to the present invention is an adenine deaminase.

[0107] The above "adenine deaminase" generally refers to an enzyme that catalyzes the deamination reaction that converts the adenine base in DNA to inosine. Inosine acts similarly to guanine during DNA replication or transcription, resulting in an A:T → G:C conversion in the corresponding base pair.

[0108] The adenine deaminase that can be used in some embodiments of the present invention is not limited to a specific amino acid sequence, structural characteristics, biological species or name, and includes all wild-type proteins, artificial or evolutionary modifications, functional variants, orthologs, conservative amino acid substitutions, etc., as long as the enzyme has a biological activity that can act on an adenine base on DNA to induce a deamination reaction. Those skilled in the art can implement an adenine deaminase suitable for a specific purpose by referring to various literature and public databases (e.g., GenBank, UniProt, etc.) regarding the sequence, structure, and function of such enzymes.

[0109] For example, TadA (tRNA-specific adenosine deaminase) from Escherichia coli is a well-known representative adenine deaminase that originally has the activity of deaminating adenosine in tRNA, and has been improved to acquire the activity that can act on DNA through specific mutations or protein engineering.

[0110] In some embodiments, the adenine deaminase has an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, or 100% identical to TadA or an ortholog thereof having the following amino acid sequence, a functional variant thereof, or a conservative amino acid substitution thereof.

[0111] TadA:

[0112] MSEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRVVFGARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKAQKKAQSSTD (SEQ ID NO: 6)

[0113] In some embodiments, the adenine deaminase has an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, or 100% identical to TadA8e having the amino acid sequence below, or a functional variant thereof, or a conservative amino acid substitution thereof. Such variants may include, for example, one or more amino acids selected from the group consisting of positions 28, 30, 46, 48, 49, 82, 84, 106, 108, 110, and 111 of the amino acid sequence of TadA8e (SEQ ID NO: 7) in which one or more amino acids are mutated to another amino acid or a conservative amino acid substitution thereof. For example, the amino acid variant may include one or more amino acid substitutions selected from the group consisting of V28Q, V28R, A48W, F84M, V106A, K110S, K110T, K110V, R111F, R111Q, R111S, R111T, and R111Y. V28Q means that the 28th valine (V) is mutated to glutamine (Q), and a person skilled in the art who is familiar with amino acid symbols can easily understand the meaning of the mutant notations. With regard to the composition of an adenine deaminase that can be used in the present invention, contents already known prior to the present application, including WO 2022 / 060185, WO 2023 / 086953, etc., which are incorporated by reference in their entirety by this application, may be cited.

[0114] TadA8e:

[0115] MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN (SEQ ID NO: 7)

[0116] D. Additional polypeptide elements

[0117] The DNA base editor according to the present invention may further comprise additional polypeptide or protein components in addition to (i) a DNA binding protein, (ii) a nickase, and (iii) a deaminase.

[0118] In some embodiments, the DNA base editor of the present invention may comprise a uracil-DNA glycosylase inhibitor (UGI).

[0119] The above "UGI" generally refers to a protein or polypeptide that inhibits the activity of uracil-DNA glycosylase (UDG), an enzyme that removes uracil from DNA. This UGI is used to induce stable base substitutions by preventing uracil from being removed by UDG after cytosine is deaminated to uracil during the DNA base proofreading process.

[0120] The UGI that can be used in some embodiments of the present invention is not limited to a specific amino acid sequence, origin, or structural characteristic, and is understood to include any protein or functional variant thereof that has the function of inhibiting the activity of UDG. For example, functional variants or conservative amino acid substitutions based on a representative UGI derived from the Bacillus subtilis PBS2 phage (e.g., GenBank Accession No. AAA22209.1 or UniProtKB P14739) may also be included in the UGI of the present disclosure.

[0121] In some embodiments, the DNA base editor of the present invention may comprise UDG.

[0122] The above "UDG" refers to a DNA repair enzyme that recognizes and excises uracil bases within DNA. UDG typically plays an early role in the base excision repair (BER) pathway, recognizing and excising uracil, which has been aberrantly inserted into DNA due to cytosine deamination. This enzyme is conserved across diverse organisms, including bacteria, eukaryotes, archaea, and viruses.

[0123] The UDG used in some embodiments of the present invention is not limited to a specific amino acid sequence, structure, origin, or biological species, and is generally understood as a concept encompassing all proteins or functional variants thereof having the activity of recognizing and removing uracil bases in DNA. For example, human UDG (human UNG, UniProtKB: P13051), E. coli UDG (UniProtKB: P13035), UDG derived from archaea, or UDG derived from phage, although they have different sequences and structures, all perform the common biological function of removing uracil, and are therefore included in the "UDG" referred to herein.

[0124] In some embodiments, the DNA base editor of the present invention may comprise an NLS.

[0125] The above "NLS (nuclear localization signal)" refers to an amino acid sequence motif required for protein translocation from the cytoplasm to the nucleus. NLSs are generally composed of short sequences rich in basic amino acids such as lysine or arginine, and mediate protein entry into the nucleus through interaction with nuclear transport receptors (e.g., importins).

[0126] The NLSs that can be used in some embodiments of the present invention are not limited to a specific sequence, structure, or origin, and various forms of NLSs can be used as long as they maintain their functional characteristics. Information regarding the function and sequence of NLSs is widely known through prior literature and publicly available materials, and those skilled in the art can select or modify an NLS sequence suitable for the intended purpose based on this information.

[0127] In some embodiments, the DNA base editor of the present invention may comprise a NES.

[0128] The above "NES (nuclear export signal)" refers to a peptide sequence or a functional variant thereof capable of inducing protein movement into the cytoplasm. NES generally has the function of binding to a nuclear export receptor (exportin) to transport a protein from the nucleus to the cytoplasm, and typically has a structural characteristic in which 4 to 5 hydrophobic amino acids are arranged at specific intervals. The NES used in some embodiments of the present invention is not limited to a specific sequence, structure, or origin, and may include various forms of nuclear export sequences as long as the functional characteristics are maintained.

[0129] The NES used in some embodiments of the present invention may be derived from various proteins existing in nature, and artificially designed sequences may also be utilized. For example, the NES derived from the HIV-1 Rev protein (e.g., LQLPPLERLTL), the NES derived from the PKI protein (e.g., LALKLAGLDI), or the NES derived from the NS2 protein of the mouse minivirus (MVM) (e.g., VDEMTKKFGTLTIHDTEK (SEQ ID NO: 8)) are widely known as representative sequences, and various variants having a similar nuclear export function are also well known to those skilled in the art.

[0130] In some embodiments, the DNA base editor of the present invention may comprise MTS.

[0131] The above "MTS (mitochondrial targeting sequence)" refers to an amino acid sequence that induces the transport of a protein to the mitochondria after translation. MTS is generally located at the N-terminus and forms a unique α-helical structure with a repeating arrangement of basic and hydrophobic amino acids, which interacts with the mitochondrial inner membrane transport complex to achieve transport. The MTS that can be used in some embodiments of the present invention is not limited to a specific sequence, length, structure, or origin, and is understood to include any functional sequence that can effectively direct a protein to the mitochondria.

[0132] Information on MTS has been identified from various biological proteins, and relevant sequences are widely available through public databases such as UniProt, NCBI, and MitoCarta. Using this publicly available sequence information, skilled artisans can select or combine MTS sequences appropriate for the desired protein to design the protein.

[0133] The MTS used in some embodiments of the present invention may be derived from various mitochondrial proteins existing in nature, and sequences artificially designed to have specific mitochondrial transport functions may also be utilized. For example, MTS derived from human superoxide dismutase 2 (SOD2) protein (e.g., MALSRAVCGTSRQLAPVLGYLGSRQKHSLPD (SEQ ID NO: 9)), MTS derived from cytochrome c oxidase subunit 8A (COX8A) protein (e.g., MASVLTPLLLRGLTGSARRLPVPRAKIHSL (SEQ ID NO: 10)), or MTS derived from human mitochondrial ATP synthase F1β subunit (e.g., MLGFVGRVAAAPASGALRRLTPSASLPPAQLLLRAAPTAVHPVRDYAAQ (SEQ ID NO: 11)) are widely known representative sequences, and similarly, various variants having the function of protein transport into mitochondria are well known to those skilled in the art.

[0134] In some embodiments, the DNA base editor of the present invention may comprise a CTS.

[0135] The above “chloroplast targeting sequence (CTS)” refers to an amino acid sequence that induces transport of a protein translated in the cytoplasm to the chloroplast. It is generally located at the N-terminus of the protein, and forms a structure in which positively charged amino acids and hydrophobic amino acids are repeatedly arranged to interact with a transport complex that passes through the outer and inner membranes of the chloroplast, thereby mediating transport to the chloroplast stroma. The CTS in the present specification is not limited to a specific sequence, length, structure, or biological origin, and is understood to include any functional sequence that can effectively direct a protein to the chloroplast.

[0136] CTS sequences have been identified from various plant-derived proteins (e.g., RuBisCO subunits, ferredoxin, plastocyanin, etc.) and are widely available in public databases such as UniProt, TAIR, and NCBI. Those skilled in the art can design a CTS suitable for a target protein by referencing or combining these published sequences. Therefore, the CTS that can be used in some embodiments of the present invention is not limited to a specific CTS sequence, but may include various homologs, functional analogs, or variants that can functionally induce chloroplast transport.

[0137] In some embodiments, the DNA base editor of the present invention may include additional sequences for biotechnology techniques, such as tags.

[0138] E. fusion protein

[0139] The DNA base editor according to the present invention is a fusion protein. That is, (i) one or more DNA binding proteins (at least one of which is a zinc finger protein), (ii) a nickase, and (iii) a deaminase used in the DNA base editor according to the present invention are present in the form of a fusion protein.

[0140] The sequence of DNA binding proteins, nickases, and deaminase enzymes included in a fusion protein can vary.

[0141] In some embodiments, the DNA binding protein, nickase, and deaminase are configured in an array selected from the group consisting of the following array sequences:

[0142] (N-terminal) [DNA binding protein] - [nickase] - [deaminase] (C-terminal)

[0143] (N-terminal) [DNA binding protein] - [deaminase] - [nickase] (C-terminal)

[0144] (N-terminal) [nickase] - [DNA binding protein] - [deaminase] (C-terminal)

[0145] (N-terminal) [nickase] - [deaminase] - [DNA binding protein] (C-terminal)

[0146] (N-terminal) [deaminase] - [DNA binding protein] - [nickase] (C-terminal)

[0147] (N-terminal) [deaminase] - [nickase] - [DNA binding protein] (C-terminal)

[0148] The protein components mentioned in the above arrays may be directly connected to each other or connected via one or more linkers.

[0149] Various additional polypeptide components described in the "Additional Polypeptide Elements" section above may be added to the arrangement described above. These may also be directly linked to other protein components or linked via one or more linkers.

[0150] The term "linker" as used herein refers to an amino acid linker, which is an amino acid sequence that covalently connects two or more functional protein domains, peptides, or other biological molecular elements. Such linkers serve to provide sufficient flexibility, length, or spatial separation so that each connected component can maintain its own structural or functional activity, and may sometimes be designed to include a specific secondary structure (e.g., an α-helix) or recognition sequence. For example, a repeating sequence based on glycine (G) and serine (S) (e.g., GGGGS)n is known as a representative example that confers high flexibility and water solubility.

[0151] The linker that can be used in the DNA editing editor according to the present invention is not limited to a specific amino acid sequence, length, or structure. It can be a naturally occurring sequence or an artificially designed sequence, and any amino acid sequence having a variety of lengths, sequence combinations, or structural characteristics is encompassed, as long as the function of each component connected via the linker is substantially maintained. Those skilled in the art can select or design an appropriate linker based on the characteristics of the target domain and the intended application.

[0152] In some embodiments, the DNA editing editor according to the present invention may comprise one or more linkers selected from the following linkers:

[0153] 2a.a. Linker: GS

[0154] 5a.a. Linker: TGEKQ (SEQ ID NO: 12)

[0155] 10a.a. Linker: SGAQGSTLDF (SEQ ID NO: 13)

[0156] 13a.a. Linker: AAEFGIRIPGEKP (SEQ ID NO: 14)

[0157] 14a.a. Linker: AAEFGIHGVPAAMG (SEQ ID NO: 15)

[0158] 16a.a. Linker: SGSETPGTSESATPES (SEQ ID NO: 16)

[0159] 24a.a. Linker: SGTPHEVGVYTLSGTPHEVGVYTL (SEQ ID NO: 17)

[0160] 32a.a. Linker: SGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID NO: 18)

[0161] In some embodiments, a DNA base editor according to the present invention may comprise, as part of a fusion protein, "additional polypeptide elements" as described above, in addition to one or more DNA binding proteins (at least one of which is a zinc finger protein), a nickase, and a deaminase.

[0162] When UGI is included as an additional polypeptide element, the UGI may be located at any position N-terminal to the DNA binding protein or at any position C-terminal to the DNA binding protein. Preferably, the UGI is located C-terminal to the DNA binding protein.

[0163] When UDG is included as an additional polypeptide element, UDG may be located at any position N-terminal to the DNA binding protein or at any position C-terminal to the DNA binding protein. Preferably, UDG is located C-terminal to the DNA binding protein.

[0164] When the additional polypeptide element comprises one or more signal sequences selected from NLS, NES, MTS and CTS, it is preferred that such signal sequence(s) be located N-terminally relative to the DNA binding protein.

[0165] In some embodiments, the DNA base editor according to the present invention may take the form of a single fusion protein. This may be referred to as a monomeric base editor.

[0166] In some embodiments, the DNA base editor according to the present invention may take the form of two fusion proteins (a first fusion protein and a second fusion protein). This may be referred to as a dimeric base editor.

[0167] In some embodiments wherein the DNA base editor according to the present invention is a dimeric base editor, the two fusion proteins (the first fusion protein and the second fusion protein) each independently comprise one DNA binding protein, and one of the two fusion proteins comprises a nickase and the other comprises a deaminase.

[0168] In some embodiments wherein the DNA base editor according to the present invention is a dimeric base editor, the two fusion proteins (the first fusion protein and the second fusion protein) comprise the same type of DNA binding protein, one of the two fusion proteins comprises a nickase and the other comprises a deaminase. In such a case, the two fusion proteins (the first fusion protein and the second fusion protein) may have the following structure, and may further comprise an "additional polypeptide element" as described herein.

[0169] First fusion proteinSecond fusion protein[zinc finger protein] - [nickase][zinc finger protein] - [cytosine deaminase][zinc finger protein] - [nickase][zinc finger protein] - [adenine deaminase][zinc finger protein] - [cytosine deaminase][zinc finger protein] - [nickase][zinc finger protein] - [adenine deaminase][zinc finger protein] - [nickase]

[0170] In some embodiments wherein the DNA base editor according to the present invention is a dimeric base editor, the two fusion proteins (the first fusion protein and the second fusion protein) comprise different types of DNA binding proteins, one of the two fusion proteins comprising a nickase and the other comprising a deaminase. In such a case, the two fusion proteins (the first fusion protein and the second fusion protein) may have the following structure, and may further comprise an "additional polypeptide element" as described herein.

[0171] First fusion proteinSecond fusion protein[Zinc finger protein] - [Nickase][TALE protein] - [Cytosine deaminase][Zinc finger protein] - [Nickase][Cas protein] - [Cytosine deaminase][TALE protein] - [Nickase][Zinc finger protein] - [Cytosine deaminase][Cas protein] - [Nickase][Zinc finger protein] - [Cytosine deaminase][Zinc finger protein] - [Cytosine deaminase][TALE protein] - [Nickase][Zinc finger protein] - [Cytosine deaminase][Cas protein] - [Nickase][TALE protein] - [Cytosine deaminase][Zinc finger protein] - [Nickase][Cas protein] - [Cytosine deaminase][Zinc finger protein] - [Nickase][Zinc finger protein] - [nickase][TALE protein] - [adenine deaminase][zinc finger protein] - [nickase][Cas protein] - [adenine deaminase][TALE protein] - [nickase][zinc finger protein] - [adenine deaminase][Cas protein] - [nickase][zinc finger protein] - [adenine deaminase][zinc finger protein] - [adenine deaminase][TALE protein] - [nickase][zinc finger protein] - [adenine deaminase][Cas protein] - [nickase][TALE protein] - [adenine deaminase][zinc finger protein] - [nickase][Cas protein] - [adenine deaminase][zinc finger protein] - [nickase]

[0172] In some embodiments where the DNA base editor according to the present invention is a dimeric base editor, DNA editing occurs in the region between the DNA sequences to which the two DNA binding proteins included in the two fusion proteins each bind, said region being referred to as a "spacer."

[0173] In some embodiments where the DNA base editor according to the present invention is a dimeric base editor, the fusion protein comprising a nickase binds to DNA located 5' upstream of the spacer, and the fusion protein comprising a deaminase binds to DNA located 3' downstream of the spacer. In other words, the fusion protein comprising a nickase binds to DNA located 5' upstream of the target site for base correction, and the fusion protein comprising a deaminase binds to DNA located 3' downstream of the target site for base correction.

[0174] In some embodiments where the DNA base editor according to the present invention is a dimeric base editor, the fusion protein comprising a deaminase binds to DNA located 5' upstream of the spacer, and the fusion protein comprising a nickase binds to DNA located 3' downstream of the spacer. In other words, the fusion protein comprising a deaminase binds to DNA located 5' upstream of the target site for base correction, and the fusion protein comprising a nickase binds to DNA located 3' downstream of the target site for base correction.

[0175] DNA editing occurs in the spacer region through the action of enzyme proteins, such as deaminase. Therefore, a certain amount of space (base length) must be secured within this spacer region for the desired base editing to occur. However, if this region is too large, unwanted base editing can occur within it, a phenomenon known as bystander editing.

[0176] Some embodiments of the DNA base editor according to the present invention, wherein the DNA base editor is a dimeric base editor, are characterized by a relatively short spacer region where proofreading is induced. In some embodiments, the spacer region can be adjusted to have a length of 10 or fewer, 9 or fewer, 8 or fewer, 7 or fewer, 6 or fewer, 5 or fewer, 4 or fewer, or 3 or fewer base pairs. This narrowing of the spacer region where proofreading is induced significantly reduces the likelihood of bystander proofreading, and has the advantage of allowing proofreading reactions to occur at specific bases. This enables more precise and specific base proofreading, which is particularly advantageous in the development of therapeutics through DNA base proofreading. For example, when a zinc finger protein was used as at least one of the two DNA-binding proteins included in the dimeric base editor, base proofreading occurred only at a single base.

[0177] 2. Polynucleotide

[0178] One aspect of the present invention relates to a polynucleotide encoding the DNA base editor of the present invention.

[0179] The polynucleotide according to the present invention is a polynucleotide encoding the DNA base editor of the present invention as described above, and the polynucleotide may be DNA or RNA. The DNA or RNA includes a sequence contained in mRNA, cDNA, synthetic DNA, plasmid DNA, linear DNA, or a viral vector.

[0180] The polynucleotide according to the present invention may be a single polynucleotide encoding a monomeric base editor as described above, or may be separate polynucleotides each encoding two fusion proteins constituting a dimeric base editor.

[0181] A person skilled in the art can easily obtain the amino acid sequence of each constituent protein by referring to the contents disclosed in the present application specification and the published protein sequences registered in databases such as NCBI GenBank and UniProt, and produce a polynucleotide according to the present invention. Based on the obtained amino acid sequences, a nucleotide sequence can be designed considering the codon usage frequency suitable for the host organism, and further, it is obvious to a person skilled in the art or within the scope of routine experimental techniques to produce a polynucleotide containing the sequence using commercial services for custom production of synthetic genes (e.g., IDT, GenScript, etc.) and vector cloning, PCR amplification, and DNA assembly technologies (Gibson assembly, Golden Gate, etc.).

[0182] The polynucleotide(s) may include expression control sequences (e.g., promoter) in addition to the coding region so as to sufficiently express the function of the encoded protein(s).

[0183] In some embodiments of the present invention, the polynucleotide may be codon optimized to suit the type of expression system (e.g., bacteria, plants, mammalian cells, etc.). Codon optimization is a general technique for increasing gene expression efficiency, and can increase protein expression by adjusting the nucleotide sequence according to the tRNA utilization frequency of the target organism.

[0184] 3. Base correction composition

[0185] One aspect of the present invention relates to a base correction composition comprising the DNA base editor of the present invention or a polynucleotide encoding the same.

[0186] The base correction composition according to the present invention may comprise the DNA base editor of the present invention or a polynucleotide encoding the same as described above and a biocompatible carrier.

[0187] The above "biocompatible carrier" refers to a material that can be included in the composition according to the present invention, and which does not significantly inhibit physiological functions when in contact with a biological system—e.g., cells, tissues, organs, or entire organisms of humans, animals, or plants—and does not induce toxic or immune reactions, etc. Such a carrier can be selected in various ways depending on pharmaceutical, biological, or agricultural use, and can play a role in improving the stability, permeability, and delivery efficiency of a base correction enzyme or a protein, nucleic acid, or auxiliary molecule related thereto.

[0188] The biocompatible carrier that can be used in the base correction composition according to the present invention is not limited to a specific chemical structure, physical form or origin, and may include, for example, a buffer solution, a surfactant, a liposome, a nanoparticle, a hydrogel, a polymeric material (e.g., PEG, PVA, PLA, PLGA), a natural or synthetic polysaccharide (e.g., dextran, hyaluronic acid, chitosan), a plant- or microbial-derived polymer, a liposome, a micelle, a biopolymer, or a combination thereof. In addition, the biocompatible carrier may be appropriately selected or combined depending on a specific delivery route or application target (e.g., human tissue, animal tissue, plant tissue, etc.), and may also include a biologically acceptable solvent, preservative, stabilizer, buffer, surfactant, reducing agent, etc.

[0189] A person skilled in the art can easily select and combine a carrier suitable for a given application purpose, delivery route, or target organism based on information already widely known through various literature and public databases regarding the types, properties, and application methods of the biocompatible carriers.

[0190] The base correction composition according to the present invention may be provided in various physical forms. The physical form may be selected based on the intended application, route of administration, stability, storage conditions, or manufacturing process, and falls within the general pharmaceutical design criteria for enhancing the efficacy and ease of use of the composition.

[0191] In some embodiments, the base correction composition according to the present invention may be in the form of a liquid, suspension, gel, powder, lyophilisate, tablet, capsule, or injectable composition. It may also be formulated as a liposome, nanoparticle, lipid nanoparticle (LNP), or other delivery particle.

[0192] The base correction composition aspect of the present invention is not limited to the above physical form, and all formulation modifications that can be appropriately selected and manufactured by a person skilled in the art according to the purpose are included in the scope of the present invention.

[0193] The base correction composition according to the present invention can be applied in various ways to ensure effective delivery to the target cell, tissue, or organism. The application method may vary depending on the target species (e.g., animal, plant, microorganism), cell type, delivery route, or formulation characteristics, and is selected based on the stability, efficacy, and biological compatibility of the composition.

[0194] In some embodiments, the base correction composition of the present invention can be administered in vivo to animals such as mammals by intravenous or intramuscular administration, subcutaneous or intradermal injection, transdermal delivery, oral or nasal administration, eye drops, etc. For example, the base editor can be delivered into the body using lipid nanoparticles (LNPs), liposomes, viral vectors (AAV, LV, etc.), plasmid DNA, or mRNA. In addition, a method of performing base correction outside the cell in an ex vivo manner and then injecting the corrected cell into the body can also be used. This can be utilized, for example, by genetically modifying and then transplanting a specific cell group such as stem cells or immune cells.

[0195] Meanwhile, the base correction composition of the present invention can be applied to plants or plant cells. Methods for applying to plants may include agroinfiltration, Agrobacterium-mediated delivery, gene gun delivery, electroporation, or direct intratissue injection. The composition to be applied may be prepared in the form of protein, DNA, mRNA, or ribonucleoprotein (RNP), and may be used with a delivery vehicle capable of penetrating plant cell walls (e.g., cell-penetrating peptide, non-targeting nanoparticle, etc.) depending on the purpose.

[0196] The base correction composition of the present invention can be flexibly applied to various biological subjects, and any in vivo, ex vivo, or in vitro delivery method that can be selected by a person skilled in the art to achieve the desired effect is included in the scope of application of the present invention.

[0197] The base correction composition of the present invention can be used to directly manipulate cells in an extracellular environment (in vitro), and can also be used to directly induce base correction in an intracellular environment (in vivo). Depending on the need, cells can be isolated and manipulated ex vivo and then reinjected into the patient, or the carrier can be directly injected into the body in vivo. Various delivery routes can be designed, including intravenous, intramuscular, epidermal, ocular, pulmonary, cerebral, and intra-body administration.

[0198] The base correction composition according to the present invention can be usefully applied to base correction technology that can precisely manipulate genetic sequences by selectively converting specific DNA bases in vivo or in vitro. In particular, the present composition can be used to precisely correct a desired target base sequence through A-to-G correction, which converts adenine (A) to guanine (G), or C-to-T correction, which converts cytosine (C) to thymine (T).

[0199] Because this correction reaction occurs without DNA double-strand breaks, it has the advantage of being less mutagenic than existing gene editing technologies and enhancing genome stability. Therefore, the composition of the present invention can be widely utilized in various fields, including human disease treatment, plant breeding, animal model production, industrial microorganism improvement, and biotechnology research.

[0200] The composition according to the present invention, unlike the existing DdCBE, contains DddA tox It has the advantage of being able to implement C-to-T correction through the activity of cytosine deaminase without using DddA. Accordingly, tox It can avoid toxicity problems, difficulties in optimizing expression, and intracellular stability problems.

[0201] Additionally, when the DNA binding protein used in the present composition is based on a zinc finger or TALE protein, the window (spacer) region within which base correction occurs is greatly narrowed, thereby enhancing the precision and specificity of correction. This serves to suppress unintentional correction of bystander bases and favorably influences the selective conversion of only the target base.

[0202] Furthermore, the base editor of the present invention has a smaller overall molecular size compared to a CRISPR / Cas system-based editor, which is advantageous in satisfying the loading limitations of a delivery vector (e.g., AAV), and has high usability in terms of intracellular delivery efficiency.

[0203] Based on these technical advantages, the composition of the present invention can be widely used in various fields, including the treatment of human genetic diseases, plant breeding, animal model production, microbial improvement, and gene function research. In particular, it can play a crucial role in clinical applications requiring accurate and safe base editing.

[0204] 4. Transmitter

[0205] One aspect of the present invention relates to a delivery vehicle comprising the DNA base editor of the present invention or a polynucleotide encoding the same.

[0206] The delivery vehicle according to the present invention refers to a means for effectively delivering the DNA base editor of the present invention or one or more polynucleotides (e.g., DNA, mRNA, etc.) encoding the same to a target site in a cell or a living body. Such a delivery vehicle may include various components to enhance the cell penetration efficiency, intracellular stability, organelle targeting ability, or in vivo distribution characteristics of the editor component, and preferably has acceptable properties such as biocompatibility, biodegradability, and non-immunogenicity. The delivery vehicle of the present invention aims to increase the efficiency and specificity of gene correction, while minimizing cytotoxicity and reducing the possibility of affecting non-target tissues.

[0207] The term "vector" or "delivery vehicle" as used herein refers to a biological or non-biological means capable of effectively delivering the DNA base editor of the present invention or one or more polynucleotides encoding it into cells. These delivery vehicles may be categorized into various types based on their structure, origin, mechanism of action, etc., and may be appropriately selected depending on the intended application target (e.g., human cells, animal cells, plant cells, bacteria, etc.) and administration method.

[0208] In some embodiments, the vector may be a viral vector, including but not limited to adeno-associated virus (AAV), lentivirus, adenovirus, retrovirus, bacteriophage-based vector, and the like.

[0209] In other embodiments, the carrier may be a non-viral carrier, including, for example, lipid nanoparticles (LNPs), polymeric nanoparticles, cationic liposomes, lipofectins, peptide-based carriers, electroporation, or nanoneedle-based systems.

[0210] Additionally, in some embodiments for application to plants, the carrier may comprise a physical delivery means implemented by Agrobacterium tumefaciens strains, plant virus-based vectors, protoplast delivery systems, or gene gun technology.

[0211] A person skilled in the art can select and combine appropriate carriers based on known techniques, depending on the characteristics of a specific base editor system, the delivery route, and the type of cell or organism to which it is applied.

[0212] The delivery system according to the present invention is applicable to various cells and organisms and can be selectively adjusted according to the purpose. Specifically, the delivery target includes eukaryotic cells or prokaryotic cells, and may include human and animal somatic cells or germ cells, embryonic stem cells, induced pluripotent stem cells (iPSCs), cancer cell lines, stromal cells, immune cells (e.g., T cells, B cells, macrophages), neural cells, hepatocytes, muscle cells, etc. In some embodiments, the delivery target may be a plant tissue, cell, or embryo.

[0213] 5. Base correction method

[0214] One aspect of the present invention relates to a DNA base correction method, comprising introducing the DNA base editor, base correction composition or carrier of the present invention as described above into a cell containing target DNA for base correction.

[0215] The above cells include, but are not limited to, human cells, animal cells, plant cells, and microbial cells, and may include various cell types such as somatic cells, germ cells, embryonic stem cells, induced pluripotent stem cells (iPSCs), plant tissue cells, or cultured cells.

[0216] The above base correction method can be performed in vitro, ex vivo, or in vivo, and can be designed according to various application purposes such as therapeutic purposes, research purposes, and trait improvement purposes. In particular, the DNA base editor of the present invention can efficiently correct adenine bases in target DNA to guanine (A:T → G:C) or cytosine bases to thymine (C:G → T:A), and therefore has wide applicability such as treatment of genetic diseases, introduction of agriculturally useful traits, and exploration of functional genes.

[0217] The components used in the base correction method of the present invention may include one or more of the DNA base editor described above, a polynucleotide encoding the same (e.g., mRNA or DNA), and a carrier (e.g., adeno-associated virus vector, lipid nanoparticle, etc.) containing the components. The components may be introduced into cells singly or in combination, and when present in the form of a fusion protein, stable expression and base correction activity can be provided through optimized binding between the component proteins. In addition, the components may additionally include an organelle targeting sequence (MTS, CTS, etc.), NES, or NLS, as needed.

[0218] The base correction method of the present invention is characterized by selectively correcting a specific base on DNA to another base, and includes, for example, A:T→G:C correction that corrects adenine (A) to guanine (G), or C:G→T:A correction that corrects cytosine (C) to thymine (T). The target of correction may be diverse, such as a mutation causing a genetic disease, an abnormal expression control region, or an artificial mutation for conferring a specific trait. The base correction method of the present invention enables high-precision base correction, particularly within a narrow spacer window, and thus can selectively edit only the target base while minimizing off-target effects.

[0219] In the base correction method of the present invention, the DNA base editor or a composition comprising the same can be introduced into cells through various physical or chemical methods. For example, lipofection, electroporation, microinjection, viral vector delivery, nanoparticle delivery, Agrobacterium-mediated delivery, etc. can be used, and an appropriate delivery method can be selected depending on the type of cell being introduced, the target gene location, the target organism species, etc. In addition, the introduction conditions (e.g., pH, temperature, incubation time, introduction amount, etc.) can be easily optimized by a person skilled in the art, taking into account the desired base correction efficiency and cell viability.

[0220] The present invention may be as follows based on the above-described contents, but is not limited thereto.

[0221] 1. A DNA base editor comprising (i) one or more DNA binding proteins, (ii) a nickase, and (iii) a deaminase, wherein at least one of the one or more DNA binding proteins is a zinc finger protein.

[0222] 2. In the above paragraph 1, (i) one or more DNA binding proteins, (ii) nickase, and (iii) deaminase are present in the form of two fusion proteins,

[0223] One of the two fusion proteins above (i) comprises a zinc finger protein as a DNA binding protein,

[0224] A DNA base editor, wherein one of the two fusion proteins comprises (ii) a nickase and the other comprises (iii) a deaminase.

[0225] 3. A DNA base editor according to claim 1 or 2, wherein both fusion proteins comprise a zinc finger protein as a DNA binding protein.

[0226] 4. A DNA base editor according to any one of claims 1 to 3, wherein one of the two fusion proteins comprises a zinc finger protein as a DNA binding protein, and the other fusion protein comprises a TALE protein as a DNA binding protein.

[0227] 5. A DNA base editor according to any one of claims 1 to 4, wherein the fusion protein comprising a nickase binds to DNA located 5' upstream of a target site for base correction, and the fusion protein comprising a deaminase binds to DNA located 3' downstream of the target site for base correction.

[0228] 6. A DNA base editor according to any one of claims 1 to 5, wherein the fusion protein comprising a deaminase binds to DNA located 5' upstream of a target site for base correction, and the fusion protein comprising a nickase binds to DNA located 3' downstream of the target site for base correction.

[0229] 7. A DNA base editor according to any one of claims 1 to 6, wherein the spacer region located between the DNA binding sites recognized by the two fusion proteins has a length of 10 base pairs or less.

[0230] 8. A DNA base editor in the first paragraph, wherein (i) a zinc finger protein as a DNA binding protein, (ii) a nickase, and (iii) a deaminase are present in the form of a single fusion protein.

[0231] 9. A DNA base editor according to any one of the above clauses 1 to 8, wherein the nickase is MutH or a functional variant thereof.

[0232] 10. A DNA base editor according to any one of claims 1 to 8, wherein the nickase is BspD6I or a catalytically active fragment thereof, or a functional variant thereof.

[0233] 11. A DNA base editor according to any one of claims 1 to 10, wherein the deaminase is an adenine deaminase.

[0234] 12. A DNA base editor according to any one of claims 1 to 11, wherein the adenine deaminase is Tad8Ae or a functional variant thereof.

[0235] 13. A DNA base editor according to any one of claims 1 to 12, wherein the deaminase is a cytosine deaminase.

[0236] 14. A DNA base editor according to claim 13, wherein the cytosine deaminase is APOBEC1 or a functional variant thereof.

[0237] 15. A DNA base editor according to any one of claims 1 to 14, further comprising a cell organelle targeting sequence and / or a nuclear export signal (NES).

[0238] 16. A DNA base editor according to the above 15th paragraph, wherein the DNA is mitochondrial DNA and the organelle targeting sequence is a mitochondrial targeting sequence (MTS).

[0239] 17. A polynucleotide(s) encoding a base editor according to any one of claims 1 to 16.

[0240] 18. A DNA base correction composition comprising a base editor according to any one of claims 1 to 16 or a polynucleotide(s) encoding the same.

[0241] 19. A carrier comprising a base editor according to any one of the above clauses 1 to 16 or a polynucleotide(s) encoding the same.

[0242] 20. The delivery vehicle of claim 19, which is an adeno-associated virus vector.

[0243] 21. A carrier according to the above 19th paragraph, which is a lipid nanoparticle or a polymer nanoparticle.

[0244] 22. A DNA base correction method comprising introducing a DNA base correction composition according to any one of claims 1 to 16, claim 18, or a carrier according to any one of claims 19 to 21 into a cell containing target DNA for base correction.

[0245] Hereinafter, the present invention will be described in detail with reference to the following examples. However, the following examples are provided only to illustrate the present invention and the present invention is not limited thereto.

[0246]

[0247] Example 1: Development of a base editor targeting the mitochondrial ND1 gene.

[0248] A nickase-based base editor capable of correcting the ND1 site of mitochondrial DNA was constructed using zinc finger proteins and TALE proteins as DNA binding proteins.

[0249] To construct a base editor fusion protein (ZFDN fusion protein) containing a zinc finger protein, a zinc finger protein capable of binding to the target site was designed, and the base sequence encoding the zinc finger proteins was codon-optimized for expression in humans. The double-stranded DNA sequence was synthesized by IDT (Integrated DNA Technologies) to synthesize a gBlock DNA fragment. The synthesized gBlock DNA fragment and the expression plasmid (containing CMV and T7 promoters, mitochondrial targeting sequences, Flag or HA tags, nickases (MutH or MutH* or BspD6I(C)) and / or TadA8e) were used as templates to construct a DNA sequence encoding the fusion protein through the Gibson assembly system. Primers for each template were custom-designed, and the DNA fragments required for Gibson assembly were amplified using PrimeSTAR® GXL DNA Polymerase (TAKARA). The PCR SV mini kit (GeneAll) was used to purify the amplified DNA fragments, and the purified DNA fragments were assembled using the HiFi DNA assembly kit (NEB). The reassembled DNA was transformed into competent DH5α (enzynomics) Escherichia coli cells by heat shock (42°C), and a single colony selectively grown on LB solid medium containing antibiotics was cultured in LB liquid medium (containing antibiotics) by shaking overnight at 37°C. Afterwards, the plasmid DNA was purified by mini-prep using the Plasmid SV mini kit (GeneAll) according to the manufacturer's protocol. The base sequence of the purified plasmid DNA was confirmed by Sanger sequencing (Macrogen) to confirm whether it was properly cloned.

[0250] To construct a TALE protein-containing nucleotide editor fusion protein (TAELDN), 424 TAL effector array plasmids and expression plasmids were constructed using the Golden-Gate assembly system. The expression plasmids contained DNA encoding CMV and T7 promoters, MTS, Flag or HA tags, nickase (MutH or MutH* or BspD6I(C)), and / or TadA8e, and were constructed using the Gibson assembly system. The process up to transformation of competent DH5α (enzynomics) Escherichia coli cells and subsequent sequence confirmation through Sanger sequencing (Macrogen) was the same as the process for constructing the plasmid DNA encoding the ZFDN fusion protein above.

[0251] ZFDN fusion proteins were linked in the following order: [MTS]-[Tag]-[NES]-[Linker]-[ZF]-[Linker]-[Nickase or TadA8e] or [MTS]-[Tag]-[NES]-[Linker]-[Nickase or TadA8e]-[Linker]-[ZF]. TALEDN fusion proteins were linked in the following order: [MTS]-[Tag]-[TALE protein]-[Linker]-[Nickase or TadA8e]. The proteins used were located between the CMV promoter and the T7 promoter and terminator.

[0252] TALED, used as a positive control, was [MTS]-[Tag]-[TALE protein]-[Linker]-[DddA tox The first fusion protein linked in the order of [1397C]-[linker]-[TadA8e] and [MTS]-[tag]-[TALE protein]-[linker]-[DddA tox 1397N] was used in the form of a second fusion protein linked in sequence.

[0253] The amino acid sequence of TadA8e used is as follows.

[0254] TadA8e:

[0255] SEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN

[0256] The above sequence is the amino acid sequence of sequence number 7, but without the initiating methionine.

[0257] The following amino acid sequences were used for Nickazero:

[0258] MutH (SEQ ID NO: 1):

[0259] SQPRPLLSPPETEEQLLAQAQQLSGYTLGELAALAGLVTPENLKRDKGWIGVLLEIWLGASAGSKPEQDFAALGVELKTIPVDSLGRPLETTFVCVAPLTGNSGVTWETSHVRH KLKRVLWIPVEGERSIPLAKRRVGSPLLWSPNEEEDRQLREDWEELMMIVLGQIERITARHGEYLQIRPKAANAKALTEAIGARGERILTLPRGFYLKKNFTSALLARHFLIQ

[0260] MutH* (SEQ ID NO: 2):

[0261] SQPRPLLSPPETEEQLLAQAQQLSGYTLGELAALAGLVTPENLKRDKGWIGVLLEIWLGASAGSKPEQDFAALGVELKTIPVDSLGRPLATTAVCVAPLTGNSGVTWETSHVRH KLKRVLWIPVEGERSIPLAKRRVGSPLLWSPNEEEDRQLREDWEELMMIVLGQIERITARHGEYLQIRPKAANAKALTEAIGARGERILTLPRGFYLKKNFTSALLARHFLIQ

[0262] BspD6I(C)(SEQ ID NO: 3):

[0263] RQLEEVIDLLEVYHEKKNVIEEKIKARFIANKNTVFEWLTWNGFIILGNALEYKNNFVIDEELQPVTHAAGNQPDMEIIYEDFIVLGEVTTSKGATQFKMESEPVTRHYLN KKKELEKQGVEKELYCLFIAPEINKNTFEEFMKYNIVQNTRIIPLSLKQFNMLLMVQKKLIEKGRRLSSYDIKNLMVSLYRTTIECERKYTQIKAGLEETLNNWVVDKEVRF

[0264] The following amino acid sequences were used for MTS.

[0265] COX8A-derived MTS (SEQ ID NO: 10):

[0266] MASVLTPLLLRGLTGSARRRLPVPRAKIHSL

[0267] SOD2-derived MTS (SEQ ID NO: 9):

[0268] MALSRAVCGTSRQLAPVLGYLGSRQKHSLPD

[0269] MTS (SEQ ID NO: 11) derived from human mitochondrial ATP synthase F1β subunit:

[0270] MLGFVGRVAAAPASGALLRRLTPSASLPPAQLLLRAAPTAVHPVRDYAAQ

[0271] The following amino acid sequences were used as tags.

[0272] 3X Flag (SEQ ID NO: 19):

[0273] DYKDHDGDYKDHDIDYKDDDDK

[0274] 3X HA (SEQ ID NO: 20):

[0275] YPYDVPDYAGYPYDVPDYAGYPYDVPDYA

[0276] 1X FLAG (SEQ ID NO: 21):

[0277] DYKDDDDK

[0278] 1X HA (SEQ ID NO: 22):

[0279] YPYDVPDYA

[0280] The following amino acid sequence was used as NES.

[0281] NES (SEQ ID NO: 8):

[0282] VDEMTKKFGTLTIHDTEK

[0283] The following amino acid sequences were used as linkers.

[0284] 2aa linker:

[0285] GS

[0286] 13aa linker (SEQ ID NO: 14):

[0287] AAEFGIRIPGEKP

[0288] 14aa linker (SEQ ID NO: 15):

[0289] AAEFGIHGVPAAMG

[0290] 16aa linker (SEQ ID NO: 16):

[0291] SGSETPGTSESATPES

[0292] 24aa linker (SEQ ID NO: 17):

[0293] SGTPHEVGVYTLSGTPHEVGVYTL

[0294] DddA used tox The amino acid sequences of the fragments are as follows.

[0295] DddA tox 1397N (SEQ ID NO: 23):

[0296] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEG

[0297] DddA tox 1397C (SEQ ID NO: 24):

[0298] AIPVKRGATGETKVFTGNNSNSPKSPTKGGC

[0299] The amino acid sequences of the zinc finger (ZF) proteins used are as follows.

[0300] ZF_L1 (SEQ ID NO: 25):

[0301] FQCRICMRNFSDSGNLRVHIRTHTGEKPYKCPDCGKSFSQSSSLIRHQRTHTGEKPYECDHCGKSFSQSSHLNVHKRTHTGEKPYRCKYCDRSFSISSNLQRHVRNIHLR

[0302] ZF_R1 (SEQ ID NO: 26):

[0303] YKCPECGKSFSTKNSLTEHQRTHTGEKPYKCPECGKSFSSKKALTEHQRTHTGEKPYECNYCGKTFSVSSTLIRHQRIHTGEKPYRCKYCDRSFSISSNLQRHVRNIHLRSGGS

[0304] The amino acid sequences of the TALE arrays used are as follows.

[0305] Left TALE 407 (SEQ ID NO: 27):

[0306] NLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLPVLCQDHGLTPDQVVAIASNNGGK QALETVQRLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDH GLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKALETVQRLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGK QALETVQRLPVLCQAHGLTPEQVVAIASNNGGQALETVQRLPVLCQAHGLTPAQVVAIASNGGGQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAH

[0307] Right TALE 425(서열번호 28):

[0308] NLTPDQVVAIASNNGGQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVV AIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLPVLCQAHGLTPAQVVAIASNNGGQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIAASNGGGK QALETVQRLPVLCQAHGLTQVVAIASNNGGQALETVQRLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVQDHGLTPDQVVAIASNGGKQALETVQRLLPVLCQAHLTQGLTQVVAIASHDGGKQALETVQR LLPVLCQAHGLTQVVAIASNIGGQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAH

[0309] Right TALE 427(서열번호 29):

[0310] NLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDDGGKQALETVQRLLPVLCQDH GLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAH

[0311]

[0312] example 2: HEK293T cell transfection

[0313] HEK293T cells were cultured in DMEM (Welgene1) supplemented with 10% FBS and 1% penicillin antibiotics in 12-well Clear TC-Treated Multiple Well Plates (Corning) at 37°C in 5% CO2. UDC cells were seeded at a density of 0.75 x 105 cells per well in 48-well Clear TC-Treated Multiple Well Plates (Corning). After 20-24 hours, 500 ng each of the plasmids encoding the first and second fusion proteins (1 μg in total) were mixed with Lipofectamine 2000 (Invitrogen) and Opti-MEM (Gibco) and added to the cells pre-seeded in the 48-well plates for transfection. The transfected cells were cultured at 37°C in 5% CO2 while replacing the culture medium. After 3 days, the cells were harvested, the culture medium was removed, and 100 μL of cell lysis buffer (50 mM Tris-HCl pH 7.4 (Welgene), 1 mM EDTA pH 8.0 (Welgene), 0.05% sodium dodecyl sulfate (Welgene), 5 μL Proteinase K (Qiagen)) was added to each well, and incubated in a PCR machine at 50°C for 1 hour and 80°C for 20 minutes.

[0314]

[0315] Example 3: Sequencing and base correction efficiency analysis to confirm base correction for the ND1 target site.

[0316] The reactants obtained in Example 2 were used as templates without purification, and the sequences were analyzed using the targeted deep sequencing technique to analyze the base correction ratio of the target region. To construct a deep sequencing library, nested first PCR and second PCR were performed using PrimeSTAR® GXL DNA Polymerase (TAKARA) as templates, and a third PCR was performed using index-containing primers to add the final index sequence. The third PCR reaction product with the added index sequence was purified using a PCR SV mini kit (GeneAll), and paired-end sequencing was performed using a MiniSeq Mid Output Kit (Illumina) using the MiniSeq system (Illumina).

[0317] The proofreading efficiency of the mitochondrial ND1 DNA of the developed base editors was screened, and the proofreading efficiency of the ND1 target gene of 17 base editor combinations with relatively high target proofreading efficiency and low off-target proofreading efficiency is shown in Fig. 2. The DNA binding sites and spacer regions of these base editors are shown in Fig. 1. DddA, which has been previously reported to have A-to-G proofreading activity of the ND1 gene tox TALED was used as a positive control, and a negative control that was not treated with the base editor was also tested.

[0318] As a result of the experiment, all 17 base editors produced showed significant proofreading efficiency at the base site where A-to-G proofreading was performed through TALED (A base or T base (when the complementary strand of DNA is A base) as shown in Fig. 2).

[0319] Thus, when using the fusion protein according to the present invention, DddA, which has been recognized as an essential component of conventional TALED, toxWe confirmed that A-to-G correction effect can be effectively achieved in mitochondrial DNA even without using .

[0320] In particular, when a fusion protein comprising a zinc finger protein was used, excellent proofreading efficiency was achieved even when the spacer region located between the DNA binding sites recognized by the two fusion proteins had a length of 7-8 base pairs. In particular, it was confirmed that when the fusion protein according to the present invention was used, proofreading occurred only at a single base, thereby reducing the effect of undesirable bystander proofreading.

[0321] The amino acid sequences of the individual fusion proteins used in the experiments are as follows. The underlined amino acid sequence represents the linker sequence.

[0322] Left 407-1397C-TadA8e (SEQ ID NO: 30):

[0323] MASVLTPLLLRGLTGSARRLPVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDKGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLLVGSAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN

[0324] Right 425-1397N(서열번호 31):

[0325] MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEG

[0326] Left 407-MutH (SEQ ID NO: 32):

[0327] MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSGSSQPRPLLSPPETEEQLLAQAQQLSGYTLGELAALAGLVTPENLKRDKGWIGVLLEIWLGASAGSKPEQDFAALGVELKTIPVDSLGRPLETTFVCVAPLTGNSGVTWETSHVRHKLKRVLWIPVEGERSIPLAKRRVGSPLLWSPNEEEDRQLREDWEELMDMIVLGQIERITARHGEYLQIRPKAANAKALTEAIGARGERILTLPRGFYLKKNFTSALLARHFLIQ

[0328] Left 407-BspD6I(C)(SEQ ID NO: 33):

[0329] MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSGSRQLEEVIDLLEVYHEKKNVIEEKIKARFIANKNTVFEWLTWNGFIILGNALEYKNNFVIDEELQPVTHAAGNQPDMEIIYEDFIVLGEVTTSKGATQFKMESEPVTRHYLNKKKELEKQGVEKELYCLFIAPEINKNTFEEFMKYNIVQNTRIIPLSLKQFNMLLMVQKKLIEKGRRLSSYDIKNLMVSLYRTTIECERKYTQIKAGLEETLNNWVVDKEVRF

[0330] Left 407-TadA8e(서열번호 34):

[0331] MASVLTPLLLRGLTGSARRLPVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDKGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN

[0332] Right 425-MutH*(SEQ ID NO: 35):

[0333]

[0334] Right 425-BspD6I(C)(SEQ ID NO: 36):

[0335]

[0336] Right 425-TadA8e (SEQ ID NO: 37):

[0337]

[0338] Right 427-MutH*(SEQ ID NO: 38):

[0339] MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSGSSQPRPLLSPPETEEQLLAQAQQLSGYTLGELAALAGLVTPENLKRDKGWIGVLLEIWLGASAGSKPEQDFAALGVELKTIPVDSLGRPLATTAVCVAPLTGNSGVTWETSHVRHKLKRVLWIPVEGERSIPLAKRRVGSPLLWSPNEEEDRQLREDWEELMDMIVLGQIERITARHGEYLQIRPKAANAKALTEAIGARGERILTLPRGFYLKKNFTSALLARHFLIQ

[0340] Right 427-BspD6I(C)(SEQ ID NO: 39):

[0341] MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSGSRQLEEVIDLLEVYHEKKNVIEEKIKARFIANKNTVFEWLTWNGFIILGNALEYKNNFVIDEELQPVTHAAGNQPDMEIIYEDFIVLGEVTTSKGATQFKMESEPVTRHYLNKKKELEKQGVEKELYCLFIAPEINKNTFEEFMKYNIVQNTRIIPLSLKQFNMLLMVQKKLIEKGRRLSSYDIKNLMVSLYRTTIECERKYTQIKAGLEETLNNWVVDKEVRF

[0342] Right 427-TadA8e(서열번호 40):

[0343] MASVLTPLLLRGLTGSARRLPVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDKGIRIQDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSSGSETPGTSESATPESSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN

[0344] Left ZF_L1-MutH(서열번호 41):

[0345] MLGFVGRVAAAPASGALRRLTPSASLPPAQLLLRAAPTAVHPVRDYAAQDYKDDDDKVDEMTKKFGTLTIHDTEKAAEFGIRIPGEKPFQCRICMRNFSDSGNLRVHIRTHTGEKPYKCPDCGKSFSQSSSLIRHQRTHTGEKPYECDHCGKSFSQSSHLNVHKRTHTGEKPYRCKYCDRSFSISSNLQRHVRNIHLRSGTPHEVGVYTLSGTPHEVGVYTLSQPRPLLSPPETEEQLLAQAQQLSGYTLGELAALAGLVTPENLKRDKGWIGVLLEIWLGASAGSKPEQDFAALGVELKTIPVDSLGRPLETTFVCVAPLTGNSGVTWETSHVRHKLKRVLWIPVEGERSIPLAKRRVGSPLLWSPNEEEDRQLREDWEELMDMIVLGQIERITARHGEYLQIRPKAANAKALTEAIGARGERILTLPRGFYLKKNFTSALLARHFLIQ

[0346] Left ZF_L1-MutH*(서열번호 42):

[0347] MLGFVGRVAAAPASGALRRLTPSASLPPAQLLLRAAPTAVHPVRDYAAQDYKDDDDKVDEMTKKFGTLTIHDTEKAAEFGIRIPGEKPFQCRICMRNFSDSGNLRVHIRTHTGEKPYKCPDCGKSFSQSSSLIRHQRTHTGEKPYECDHCGKSFSQSSHLNVHKRTHTGEKPYRCKYCDRSFSISSNLQRHVRNIHLRSGTPHEVGVYTLSGTPHEVGVYTLSQPRPLLSPPETEEQLLAQAQQLSGYTLGELAALAGLVTPENLKRDKGWIGVLLEIWLGASAGSKPEQDFAALGVELKTIPVDSLGRPLATTAVCVAPLTGNSGVTWETSHVRHKLKRVLWIPVEGERSIPLAKRRVGSPLLWSPNEEEDRQLREDWEELMDMIVLGQIERITARHGEYLQIRPKAANAKALTEAIGARGERILTLPRGFYLKKNFTSALLARHFLIQ

[0348] Left ZF_L1-TadA8e(서열번호 43):

[0349] MLGFVGRVAAAPASGALRRLTPSASLPPAQLLLRAAPTAVHPVRDYAAQDYKDDDDKVDEMTKKFGTLTIHDTEKAAEFGIRIPGEKPFQCRICMRNFSDSGNLRVHIRTHTGEKPYKCPDCGKSFSQSSSLIRHQRTHTGEKPYECDHCGKSFSQSSHLNVHKRTHTGEKPYRCKYCDRSFSISSNLQRHVRNIHLRSGTPHEVGVYTLSGTPHEVGVYTLSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN

[0350] Right ZF_R1-MutH(서열번호 44):

[0351] MLGFVGRVAAAPASGALRLTPSASLPPAQLLLRAAPTAVHPVRDYAAAQYPYDVPDYAVDEMTKKFGTLTIHDTEKAAEFGIHGVPAAMGSQPRPLLSPPETEEQLLAQAQQLSGYTLGELAALAGLVTPENLKRDKGWIGVLLEIWLGASGSKPEQDFAALGVELKTIPVDSLGRPLETTFVCVAPLTGNSGVTWETSHVRHKLKRVLWIPVEGERSIPLAKRRVG SPLLWSPNEEEDRQLREDWEELMDMIVLGQIERITARHGEYLQIRPKAANAKALTEAIGARGERILTLPRGFYLKKNFTSALLARHFLIQSGTPHEVGVYTLSGTPHEVVYTLYKCPECGKSFSTKNSLTEHQRTHTGEKPYKCPECGKSFSSKKALTEHQRTHTGEKPYECNYCGKTFSVSSLIRHQRIHTGEKPYRCKYCDRSFSSSLQRHVRNIHLRSGGS

[0352] Right ZF_R1-MutH*(서열번호 45):

[0353] MLGFVGRVAAAPASGALRLTPSASLPPAQLLLLRAAPTAVHPVRDYAAAQYPYDVPDYAVDEMTKKFGTLTIHDTEKAAEFGIHGVPAAMGSQPRPLLSPPETEEQLLAQAQQLSGYTLGELAALAGLVTPENLKRDKGWIGVLLEIWLGASGSKPEQDFAALGVELKTIPVDSLGRPLATTAVCVAPLTGNSGVTWETSHVRHKLKRVLWIPVEGERSIPLAKRRVG SPLLWSPNEEEDRQLREDWEELMDMIVLGQIERITARHGEYLQIRPKAANAKALTEAIGARGERILTLPRGFYLKKNFTSALLARHFLIQSGTPHEVGVYTLSGTPHEVVYTLYKCPECGKSFSTKNSLTEHQRTHTGEKPYKCPECGKSFSSKKALTEHQRTHTGEKPYECNYCGKTFSVSSLIRHQRIHTGEKPYRCKYCDRSFSSSLQRHVRNIHLRSGGS

[0354] Right ZF_R1-TadA8e(서열번호 46):

[0355] MLGFVGRVAAAPASGALRRLTPSASLPPAQLLLRAAPTAVHPVRDYAAQYPYDVPDYAVDEMTKKFGTLTIHDTEKAAEFGIHGVPAAMGSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSINSGTPHEVGVYTLSGTPHEVGVYTLYKCPECGKSFSTKNSLTEHQRTHTGEKPYKCPECGKSFSSKKALTEHQRTHTGEKPYECNYCGKTFSVSSTLIRHQRIHTGEKPYRCKYCDRSFSISSNLQRHVRNIHLRSGGS

Claims

1. A DNA base editor comprising (i) one or more DNA binding proteins, (ii) a nickase, and (iii) a deaminase, wherein at least one of the one or more DNA binding proteins is a zinc finger protein.

2. In paragraph 1, (i) one or more DNA binding proteins, (ii) nickase, and (iii) deaminase are present in the form of two fusion proteins, One of the two fusion proteins above (i) comprises a zinc finger protein as a DNA binding protein, A DNA base editor, wherein one of the two fusion proteins comprises (ii) a nickase and the other comprises (iii) a deaminase.

3. A DNA base editor in the second paragraph, wherein both fusion proteins comprise a zinc finger protein as a DNA binding protein.

4. A DNA base editor according to claim 2, wherein one of the two fusion proteins comprises a zinc finger protein as a DNA binding protein, and the other fusion protein comprises a TALE protein as a DNA binding protein.

5. A DNA base editor in the second paragraph, wherein the fusion protein containing a nickase binds to DNA located 5' upstream of the target site for base correction, and the fusion protein containing a deaminase binds to DNA located 3' downstream of the target site for base correction.

6. A DNA base editor in the second paragraph, wherein a fusion protein containing a deaminase binds to DNA located 5' upstream of a target site for base correction, and a fusion protein containing a nickase binds to DNA located 3' downstream of the target site for base correction.

7. A DNA base editor according to claim 2, wherein the spacer region located between the DNA binding sites recognized by the two fusion proteins has a length of 10 base pairs or less.

8. A DNA base editor in paragraph 1, wherein (i) a zinc finger protein as a DNA binding protein, (ii) a nickase, and (iii) a deaminase are present in the form of a single fusion protein.

9. A DNA base editor according to claim 1, wherein the nickase is MutH or a functional variant thereof.

10. A DNA base editor according to claim 1, wherein the nickase is BspD6I or a catalytically active fragment thereof, or a functional variant thereof.

11. A DNA base editor according to claim 1, wherein the deaminase is an adenine deaminase.

12. A DNA base editor according to claim 11, wherein the adenine deaminase is Tad8Ae or a functional variant thereof.

13. A DNA base editor according to claim 1, wherein the deaminase is cytosine deaminase.

14. A DNA base editor according to claim 13, wherein the cytosine deaminase is APOBEC1 or a functional variant thereof.

15. A DNA base editor according to claim 1, further comprising a cell organelle targeting sequence and / or a nuclear export signal (NES).

16. A DNA base editor according to claim 15, wherein the DNA is mitochondrial DNA and the organelle targeting sequence is a mitochondrial targeting sequence (MTS).

17. Polynucleotide(s) encoding a base editor according to paragraph 1.

18. A DNA base correction composition comprising a base editor according to paragraph 1 or a polynucleotide(s) encoding the same.

19. A carrier comprising a base editor according to paragraph 1 or a polynucleotide(s) encoding the same.

20. A vector according to claim 19, which is an adeno-associated virus vector.

21. In claim 19, a carrier that is a lipid nanoparticle or a polymer nanoparticle.

22. A DNA base correction method comprising introducing the DNA base editor described in paragraph 1, the DNA base correction composition described in paragraph 18, or the carrier described in paragraph 19 into a cell containing target DNA for base correction.

Citation Information

Patent Citations

  • Substrate transfer module and method for manufacturing substrate transfer module

    KR1020250112686A

  • Method for deoxidation of off-grade titanium sponge using magnesium

    KR102300837B1

  • Exhaust structure of mold

    KR102756855B1

  • Base editors, compositions, and methods for modifying the mitochondrial genome

    WO2021155065A1

  • Novel zinc finger fusion proteins for nucleobase editing

    WO2023122722A1