Fusion protein comprising nucleic acid-binding protein and ssda, and use thereof

A fusion protein combining a nucleic acid-binding protein and SsdA enables efficient genome editing in organelles, addressing delivery limitations of CRISPR systems and facilitating disease models and trait improvement.

WO2025211767A1PCT designated stage Publication Date: 2025-10-09THE ASAN FOUND +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/004383
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-02
Filing Date
2025-04-02
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

Existing genome editing technologies, such as CRISPR-based systems, are not suitable for efficiently editing DNA in cellular organelles like mitochondria and chloroplasts due to limitations in delivering or expressing Cas9 protein and guide RNA.

Method used

A fusion protein comprising a nucleic acid-binding protein (NABP) and a bacterial cytosine deaminase, SsdA, which is smaller and easier to design as a vector, enabling efficient correction of cytosine bases in organelle DNA.

Benefits of technology

The fusion protein allows for precise genome editing in organelles, facilitating the development of mitochondrial disease models and improving plant traits by correcting point mutations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025004383_09102025_PF_FP_ABST
    Figure KR2025004383_09102025_PF_FP_ABST
Patent Text Reader

Abstract

A fusion protein, according to one aspect, enables effective base editing within organelles, and provides an effect of enabling genetic base editing of both cell nucleus and organelles by using a single fusion protein.
Need to check novelty before this filing date? Find Prior Art

Description

Fusion protein comprising nucleic acid binding protein and SSDA and use thereof

[0001] The present invention relates to a fusion protein comprising a nucleic acid binding protein and SsdA and to uses thereof.

[0002] Genome editing is a technology that targets and modifies the genetic information of living organisms. It is being applied to various species of organisms, including humans, animals, plants, and microorganisms, dramatically expanding its scope of application.

[0003] Gene scissors, a key element of this technology, are molecular tools designed to precisely cut or correct a desired base sequence. Recently, fusion proteins that enable targeted base substitution (base editing) rather than cutting have been attracting attention.

[0004] In particular, base editors, which fuse a DNA-binding protein with a cytosine deaminase, can induce targeted cytosine (C) to thymine (T) or adenine (A) to guanine (G) substitutions at specific nucleotides within the genome, enabling the precise correction of point mutations that cause genetic defects. This technology can be utilized to generate single nucleotide polymorphisms (SNPs) or to build genetic disease models in various systems, including cultured cells, animals, and plants.

[0005] Meanwhile, mitochondrial DNA contains key genetic information involved in cellular respiration, and mutations in these genes can lead to serious dysfunction in tissues with high energy demands (e.g., muscles, nerves, etc.). Indeed, patients with mitochondrial disease harbor a mixture of normal and mutant mtDNA, and the ratio of these genes has a critical impact on disease progression. Therefore, mitochondrial base editing technology has diverse potential applications, including studying the pathogenesis of mitochondrial diseases, developing animal models, and designing therapeutics.

[0006] Similarly, plant organelles like chloroplasts and mitochondria contain genes crucial for vital functions like photosynthesis and respiration. Therefore, targeted modification of these genes is crucial for introducing useful traits into crops, such as improving traits, inducing antibiotic resistance, and conferring sterility. For example, targeted mutations in the mitochondrial atp6 gene can induce male sterility, and specific point mutations in the chloroplast 16SrRNA gene can confer antibiotic resistance.

[0007] Against this backdrop, there is a need for the development of a compact, high-efficiency base editor capable of precisely editing the genome of cellular organelles.

[0008] Representative fusion-type base editors developed to date include the base editor (BE), which fuses the rat-derived APOBEC1 cytidine deaminase to the catalytically inactive Cas9 (dCas9) derived from Streptococcus pyogenes or nCas9 with the D10A mutation, and the Target-AID system, which fuses the sea lamprey-derived AID analog PmCDA1 or human-derived AID to dCas9 or nCas9, is also known. In addition, the CRISPR-X system, which induces deamination by recruiting hyperactivated AID mutants using dCas9 and MS2 RNA hairpin, also exists.

[0009] However, these systems are not suitable for editing DNA in organelles such as mitochondria and chloroplasts because of technical limitations in simultaneously delivering or expressing the Cas9 protein and guide RNA (gRNA) into the organelles.

[0010] Accordingly, the present inventors developed a base-editing fusion protein based on bacterial cytosine deaminase (single-stranded DNA deaminase toxin A, SsdA), which is smaller in size and easier to design as a vector than existing APOBEC or AID series. The fusion protein of the present invention can efficiently correct cytosine bases in organelle DNA and can provide a genome editing platform that can be utilized in various life sciences and biotechnology fields, such as developing mitochondrial disease models, establishing treatment strategies, and improving plant traits.

[0011] One aspect is to provide a fusion protein comprising a nucleic acid-binding protein (NABP) and a single-stranded DNA deaminase toxin A (SsdA).

[0012] Another aspect is to provide a polynucleotide encoding the fusion protein.

[0013] Another aspect is to provide a vector comprising the polynucleotide.

[0014] Another aspect provides a composition for base correction comprising at least one from the group consisting of the fusion protein; a polynucleotide encoding the fusion protein; and a vector comprising a polynucleotide encoding the fusion protein.

[0015] Another aspect provides a method for correcting a nucleic acid, comprising the step of contacting a target nucleic acid molecule with the base correction composition.

[0016] Another aspect provides a base editing system comprising a fusion protein comprising a nucleic acid-binding protein (NABP); and a bacterial toxin or a polynucleotide encoding the fusion protein; and a guide polynucleotide.

[0017] One aspect is to provide a fusion protein comprising a nucleic acid-binding protein (NABP) and a bacterial toxin.

[0018] As used herein, the term “bacterial toxin” refers to a deaminase that removes an amine from an amino group in an amino acid or nucleotide. The bacterial toxin may be bound to the terminal of a nucleic acid-binding protein (NABP). For example, the bacterial toxin may be bound to the C-terminus, N-terminus, or both the C-terminus and N-terminus of a TALE (Transcription Activator-like Effector) protein, a ZFP (Zinc Finger Protein), or a Cas protein, but is not limited thereto.

[0019] In one specific example, the bacterial toxin may be single-stranded DNA deaminase toxin A (SsdA).

[0020] The term "SsdA (single-stranded DNA deaminase toxin A)" in this specification refers to a bacterial cytidine deaminase, an enzyme that deamidates cytidine bases into uridine. The above SsdA belongs to the DYW-like cytidine deaminase family and may be, but is not limited to, a strain of Pseudomonas sp., for example, Pseudomonas syringae, Pseudomonas congelans, Pseudomonas savastanoi, Pseudomonas viridiflava, Pseudomonas coronafaciens, Pseudomonas fluorescens, Pseudomonas sp. MPC6, Pseudomonas sp. GL-R-26, or Pseudomonas sp. GL-RE-26.

[0021] In one specific example, the SsdA may be wild-type.

[0022] The polypeptide of the wild-type SsdA may have a length of 400 to 500 amino acids, for example, a length of 400 to 500 amino acids, a length of 410 to 500 amino acids, a length of 420 to 500 amino acids, a length of 430 to 500 amino acids, a length of 440 to 500 amino acids, a length of 450 to 500 amino acids, a length of 460 to 500 amino acids, a length of 470 to 500 amino acids, a length of 480 to 500 amino acids, or a length of 490 to 500 amino acids. The SsdA may comprise the amino acid sequence of SEQ ID NO: 1.

[0023] Meanwhile, the polypeptide of the wild-type SsdA may have a length of 100 to 200 amino acids, for example, a length of 120 to 180 amino acids, a length of 120 to 170 amino acids, a length of 130 to 160 amino acids, or a length of 140 to 160 amino acids. The SsdA may comprise the amino acid sequence of SEQ ID NO: 2.

[0024] In one specific example, the wild-type SsdA may have a PAAR domain at the N-terminus, and the SsdA may have a deaminase domain having deamination activity at the C-terminus. Specifically, the SsdA may include the sequence of SEQ ID NO: 1, and the deaminase domain may include the sequence of SEQ ID NO: 2.

[0025] In one specific example, the SsdA may be an SsdA mutant.

[0026] The term "variant" as used herein refers to a protein having a modified sequence in which one or more amino acids in the amino acid sequence of a specific protein are substituted, inserted, or deleted. The variant may retain biological activity similar to that of the original protein, or may have an enhanced or altered function. For example, the variant may be designed to have characteristics such as enhanced deamination efficiency, increased target DNA binding affinity, enhanced selectivity for specific bases, enhanced intracellular stability, or reduced toxicity, while preserving the structural or functional characteristics of the wild-type protein.

[0027] In one specific example, the SsdA variant may have an amino acid substituted at a position capable of improving binding affinity with target DNA with another amino acid.

[0028] In one specific example, the position capable of enhancing binding affinity with the target DNA may be V289, H333, Y335, P282, K392 or a position having a corresponding function among the amino acid sequences of SEQ ID NO: 1.

[0029] In one specific example, the SsdA variant comprises an amino acid sequence having at least 80% identity with the amino acid sequence of SEQ ID NO: 1, and may comprise a mutation at at least one position selected from the group consisting of V289, H333, Y335, P282, K392, and positions having corresponding functions within the amino acid sequence, but is not limited thereto.

[0030] For example, the SsdA variant may comprise an amino acid sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with the amino acid sequence of SEQ ID NO: 1, and may comprise a mutation at at least one position selected from the group consisting of V289, H333, Y335, P282, K392, and positions having corresponding functions thereto within the amino acid sequence.

[0031] In one specific example, the position capable of enhancing binding affinity with the target DNA may be V32, H76, Y78, P25, K135 or a position having a corresponding function among the amino acid sequences of SEQ ID NO: 2.

[0032] In one specific example, the SsdA variant comprises an amino acid sequence having at least 80% identity with the amino acid sequence of SEQ ID NO: 2, and may comprise a mutation at at least one position selected from the group consisting of V32, H76, Y78, P25, K135, and positions having corresponding functions within the amino acid sequence, but is not limited thereto.

[0033] For example, the SsdA variant may comprise an amino acid sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the amino acid sequence of SEQ ID NO: 2, and may comprise a mutation at at least one position selected from the group consisting of V32, H 76, Y78, P25, K135, and positions having corresponding functions thereto within the amino acid sequence.

[0034] In one specific example, the other amino acid may be any one selected from the group consisting of arginine (R), histidine (H), lysine (K), aspartic acid (D), glutamic acid (E), serine (S), threonine (T), asparagine (N), glutamine (Q), cysteine ​​(C), selenocysteine ​​(U), glycine (G), proline (P), alanine (A), valine (V), isoleucine (I), leucine (L), methionine (M), phenylalanine (F), tyrosine (Y), tryptophan (W), and all variants of the above amino acids, excluding the amino acid that the wild-type protein has at the variant position.

[0035] In one specific example, the other amino acid may be a basic amino acid, and the basic amino acid may be, but is not limited to, arginine (R), histidine (H), or lysine (K).

[0036] In one specific example, the SsdA variant comprises an amino acid sequence having at least 80% identity with the amino acid sequence of SEQ ID NO: 1, and may comprise a mutation at at least one position selected from the group consisting of Y335, P282, K392, and positions having corresponding functions within the amino acid sequence, but is not limited thereto.

[0037] For example, the SsdA variant may comprise an amino acid sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with the amino acid sequence of SEQ ID NO: 1, and may comprise a mutation at at least one position selected from the group consisting of Y335, P282, K392, and positions having corresponding functions thereto within the amino acid sequence.

[0038] In one specific example, the SsdA variant comprises an amino acid sequence having at least 80% identity with the amino acid sequence of SEQ ID NO: 4, and may include mutations at positions Y335, P282, and K392 within the amino acid sequence, but is not limited thereto.

[0039] For example, the SsdA variant may comprise an amino acid sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the amino acid sequence of SEQ ID NO: 4, and may comprise mutations at positions Y335, P282, and K392 within the amino acid sequence.

[0040] In one specific example, the SsdA variant may include a mutation at at least one position selected from the group consisting of Y335, P282, K392 and positions having corresponding functions within the amino acid sequence of SEQ ID NO: 1, but is not limited thereto.

[0041] In one specific example, the SsdA variant comprises an amino acid sequence having at least 80% identity with the amino acid sequence of SEQ ID NO: 2, and may comprise a mutation at at least one position selected from the group consisting of P25, Y78, K135, and positions having corresponding functions within the amino acid sequence, but is not limited thereto.

[0042] For example, the SsdA variant may comprise an amino acid sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the amino acid sequence of SEQ ID NO: 2, and may comprise a mutation at at least one position selected from the group consisting of P25, Y78, K135, and positions having corresponding functions thereto within the amino acid sequence.

[0043] In one specific example, the SsdA variant comprises an amino acid sequence having at least 80% identity with the amino acid sequence of SEQ ID NO: 5, and may include mutations at positions P25, Y78, and K135 within the amino acid sequence, but is not limited thereto.

[0044] For example, the SsdA variant may comprise an amino acid sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the amino acid sequence of SEQ ID NO: 5, and may comprise mutations at positions P25, Y78, and K135 within the amino acid sequence.

[0045] In one specific example, the SsdA variant may include a mutation at at least one position selected from the group consisting of P25, Y78, K135 and positions having corresponding functions within the amino acid sequence of SEQ ID NO: 5, but is not limited thereto.

[0046] In one specific example, the SsdA variant may be a fusion protein in which an amino acid in the toxin domain of the SsdA protein is substituted with another amino acid. The toxin domain is a portion having deaminase activity in the SsdA protein and may have a length of 100 to 200 amino acids, for example, a length of 120 to 180 amino acids, a length of 120 to 170 amino acids, a length of 130 to 160 amino acids, or a length of 140 to 160 amino acids. Specifically, the toxin domain may include the amino acid sequence of SEQ ID NO: 2. The amino acid sequence of SEQ ID NO: 2 of the SsdA may include a catalytic active site.

[0047] In one specific example, the SsdA variant may be one in which an amino acid in the catalytic active site of the wild-type SsdA protein is substituted with another amino acid. For example, the SsdA variant may be one in which an amino acid in the catalytic active site of the wild-type SsdA protein having the base sequence of SEQ ID NO: 1 is substituted with another amino acid.

[0048] In one specific example, the SsdA may be an inactivated SsdA. The inactivated SsdA may have an amino acid mutation in the catalytic active site of wild-type SsdA. The inactivated SsdA may have lower cytotoxicity than the wild-type SsdA. Specifically, the catalytic active site may be, but is not limited to, T291, K365, or a position having a corresponding function in the amino acid sequence of SEQ ID NO: 1.

[0049] According to one specific example, the inactivated SsdA may have T291, K365 or a corresponding amino acid mutation thereof in the amino acid sequence of SEQ ID NO: 1. Specifically, the substituted other amino acid may be any one selected from the group consisting of arginine (R), histidine (H), lysine (K), aspartic acid (D), glutamic acid (E), serine (S), threonine (T), asparagine (N), glutamine (Q), cysteine ​​(C), selenocysteine ​​(U), glycine (G), proline (P), alanine (A), valine (V), isoleucine (I), leucine (L), methionine (M), phenylalanine (F), tyrosine (Y), tryptophan (W) and all variants of the above amino acids, excluding the amino acid that the wild-type protein has at the mutation position.

[0050] In one specific example, the SsdA variant may comprise an amino acid sequence having at least 80% identity with the amino acid sequence of SEQ ID NO: 1, and may comprise a mutation at at least one position selected from the group consisting of T291, K365, and positions having a corresponding function within the amino acid sequence.

[0051] In one specific example, the SsdA variant may have another amino acid substituted at any one or more of T291, K365 or a position having a corresponding function among the amino acid sequences of SEQ ID NO: 1. Specifically, the other amino acid may be any one selected from the group consisting of arginine (R), histidine (H), lysine (K), aspartic acid (D), glutamic acid (E), serine (S), threonine (T), asparagine (N), glutamine (Q), cysteine ​​(C), selenocysteine ​​(U), glycine (G), proline (P), alanine (A), valine (V), isoleucine (I), leucine (L), methionine (M), phenylalanine (F), tyrosine (Y), tryptophan (W) and all variants of the above amino acids, and may be an amino acid other than an amino acid that the wild-type protein has at a variant position.

[0052] In one specific example, the catalytic active site may be, but is not limited to, T34, K108 or a position having a corresponding function in the amino acid sequence of SEQ ID NO: 2.

[0053] In one specific example, the SsdA variant may comprise an amino acid sequence having at least 80% identity with the amino acid sequence of SEQ ID NO: 2, and may comprise a mutation at at least one position selected from the group consisting of T34, K108, and positions having a corresponding function within the amino acid sequence.

[0054] In one specific example, the SsdA variant may be a variant having another amino acid substituted at any one or more of positions V289, H333, Y335, P282, K392 or a position having a corresponding function among the amino acid sequence of SEQ ID NO: 1, and further having an amino acid substituted at T291, K365 or a position having a corresponding function among the amino acid sequence of SEQ ID NO: 1 with another amino acid.

[0055] In one specific example, the SsdA variant may be one in which an amino acid within the deaminase domain of the wild-type SsdA protein is substituted with another amino acid. For example, the SsdA variant may be one in which an amino acid within the wild-type deaminase domain comprising the amino acid sequence of SEQ ID NO: 2 is substituted with another amino acid.

[0056] In one specific example, the SsdA mutant in which an amino acid in the deaminase domain of the wild-type SsdA protein is substituted with another amino acid may be inactivated.

[0057] In one specific example, the SsdA variant may have another amino acid substituted at any one or more of V32, H76, Y78, P25, K135, T34, K108 or a position having a corresponding function in the amino acid sequence of SEQ ID NO: 2. Specifically, the other amino acid may be any one selected from the group consisting of arginine (R), histidine (H), lysine (K), aspartic acid (D), glutamic acid (E), serine (S), threonine (T), asparagine (N), glutamine (Q), cysteine ​​(C), selenocysteine ​​(U), glycine (G), proline (P), alanine (A), valine (V), isoleucine (I), leucine (L), methionine (M), phenylalanine (F), tyrosine (Y), tryptophan (W), and all variants of the above amino acids, excluding the amino acid that the wild-type protein has at the mutant position.

[0058] In one specific example, the nucleic acid binding protein may be any one selected from the group consisting of, but not limited to, TALE (Transcription Activator-like Effector) protein, ZFP (Zinc Finger Protein), Cas9 (CRISPR-associated protein 9), Cpf1, dCpf1, dCas12, dCas13, Cas14, Cas12b, Cas12f and variants thereof.

[0059] For example, the fusion protein may comprise a TALE (Transcription Activator-like Effector) protein and a bacterial toxin, wherein the bacterial toxin may be a wild-type SsdA (single-stranded DNA deaminase toxin A) or an SsdA mutant.

[0060] For example, the fusion protein may comprise ZFP (Zinc Finger Protein) and a bacterial toxin, wherein the bacterial toxin may be wild-type SsdA (single-stranded DNA deaminase toxin A) or an SsdA mutant.

[0061] For example, the fusion protein may comprise a Cas protein (CRISPR-associated protein) and a bacterial toxin, wherein the bacterial toxin may be a wild-type SsdA (single-stranded DNA deaminase toxin A) or an SsdA mutant.

[0062] The term “nucleic acid-binding protein (NABP)” as used herein refers to a group of proteins that directly bind to DNA or RNA and play a key role in the replication, transcription, translation, repair, processing, degradation, and regulation of genetic information, and can recognize a specific base sequence or recognize and bind to various secondary structures (e.g., stem-loop, G-cytostine rich structure, etc.) including structural features of a double helix (dsDNA, dsRNA) or a single strand (ssDNA, ssRNA). For example, the nucleic acid-binding protein may bind to a nucleic acid in a sequence-specific manner.

[0063] The term “TALE (Transcription Activator-like Effector)” used herein refers to a bacterial DNA-binding protein derived from the plant pathogenic bacteria Xanthomonas that has the ability to recognize and bind to a specific base sequence on DNA. The TALE protein can recognize DNA through a repeat sequence located in the center, bind to the target DNA sequence, and perform transcriptional regulation or genome editing functions, and can also be utilized as a base editing tool by fusing with an external deaminase or an activation domain. Specifically, the TALE protein is composed of a repeated form of 33 to 34 repeated amino acid sequences, and about 9 or more domains (RVD, Repeated variant diresidues) are repeated. Each domain can recognize one nucleotide, and can bind to a specific DNA sequence depending on the 12th to 13th amino acid sequence (HD->Cytosine, NI->Adenine, NG->Thymine, NN->Guanine). TALE proteins recognize a single DNA strand within a target region. The distance between target regions can be 12-14 nucleotides. The TALE domain refers to a protein domain that binds to a nucleotide in a sequence-specific manner by a combination of one or more TALE-Repeat. It includes, but is not limited to, at least one TALE-Repeat, specifically 1 to 30 TALE-Repeat. A TALE-Repeat is a portion that recognizes a specific nucleotide sequence within the TALE domain.

[0064] In one specific example, the bacterial toxin may be bound to a terminal of the TALE protein. For example, the bacterial toxin may be bound to the C-terminus, the N-terminus, or both the C-terminus and the N-terminus of the TALE protein.

[0065] In one specific example, a single module TALE may be coupled to the N-terminus of the wild-type SsdA (single-stranded DNA deaminase toxin A) or SsdA mutant. A single TALE module and the wild-type SsdA (single-stranded DNA deaminase toxin A) or SsdA mutant may be included in the NC direction. A dual module TALE may be included, wherein a first TALE is coupled to the N-terminus of the wild-type SsdA (single-stranded DNA deaminase toxin A) or SsdA mutant, and a second TALE may be included separately. The first TALE module and the wild-type SsdA (single-stranded DNA deaminase toxin A) or SsdA mutant in the NC direction are included, and may have the structures of N'-TALE-SsdA (single-stranded DNA deaminase toxin A)-C' and N'-TALE-C'.

[0066] The term “ZFP (Zinc Finger Protein)” in this specification refers to zinc (Zn 2+ ) refers to a group of proteins that interact with DNA, RNA, or proteins through a unique finger structure formed by combining with ions. In particular, ZFP has the characteristic of being able to recognize and bind to a specific DNA sequence, and is utilized as a gene targeting tool in a base editing system. Specifically, by fusing the DNA binding domain of ZFP with wild-type SsdA (single-stranded DNA deaminase toxin A) or an SsdA mutant, a ZFP-based base editing system (ZFP-Base Editor, ZFP-BE) that induces a base substitution of cytosine (C) in a specific base sequence can be constructed.

[0067] As used herein, the term "Cas protein (CRISPR-associated protein)" may be a CRISPR-binding endonuclease. The Cas protein may cleave all or part of a specific target polynucleotide sequence.

[0068] The above Cas protein forms an active endonuclease, or nickase, when it forms a complex with two RNAs called CRISPR RNA (crRNA) and trans-activating crRNA (tracrRNA). Non-limiting examples of the above Cas proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, a homolog thereof or a homolog thereof. Includes modified versions.

[0069] The above Cas protein may be Cas9, Cpf1, or a Cas9 variant protein, and the Cas9 (or Cpf1) protein refers to an essential protein element in the CRISPR / Cas9 system, and information on the Cas9 (or Cpf1) gene and protein can be obtained from GenBank of the National Center for Biotechnology Information (NCBI), but is not limited thereto. CRISPR-associated genes encoding Cas (or Cpf1) proteins are known to have more than 40 different Cas (or Cpf1) protein families, and eight CRISPR subtypes (Ecoli, Ypest, Nmeni, Dvulg, Tneap, Hmari, Apern, and Mtube) can be defined according to specific combinations of cas genes and repeat structures. Therefore, each of the above CRISPR subtypes can form a repeat unit to form a polyribonucleotide-protein complex.

[0070] In one specific example, the Cas protein may be at least one selected from the group consisting of a Cas9 protein derived from Streptococcus pyogenes, a Cas9 protein derived from Campylobacter jejuni, a Cas9 protein derived from Streptococcus thermophiles, a Cas9 protein derived from Streptococcus aureus, a Cas9 protein derived from Neisseria meningitidis, and a Cpf1 protein, but is not limited thereto.

[0071] In one embodiment of the present invention, the nucleic acid binding protein and the bacterial toxin can be fused via a linker. The linker can be positioned at the C-terminus, N-terminus, or both the C-terminus and N-terminus of the nucleic acid binding protein, and the bacterial toxin can be bound to the nucleic acid binding protein via the linker. Suitable linker motifs and linker configurations include those described in the literature [Chen et al., Fusion protein linkers: property, design, and functionality. Adv Drug Deliv Rev. 2013; 65(10):1357-69], the entire contents of which are incorporated herein by reference.

[0072]

[0073] In one specific example, the fusion protein may further comprise a DNA glycosylase inhibitor.

[0074] In one specific example, the DNA glycosylase inhibitor may be linked to one or more of the N-terminus or C-terminus of the fusion protein.

[0075] In one specific example, the DNA glycosylase inhibitor may be, but is not limited to, a thymine glycosylase inhibitor, a uracil glycosylase inhibitor, an oxoguanine glycosylase inhibitor, or an alkylguanine DNA glycosylase inhibitor.

[0076] In one specific example, the uracil DNA glycosylase inhibitor (UGI) may be, but is not limited to, a uracil DNA glycosylase inhibitor derived from Bacillus subtilis bacteriophage, PBS1, a uracil DNA glycosylase inhibitor derived from Bacillus subtilis bacteriophage, or PBS2.

[0077] In one specific example, the fusion protein may further comprise an organelle signal peptide domain. Specifically, the organelle signal peptide domain may be, but is not limited to, a nuclear localization signal (NLS), a nuclear export signal (NES), a mitochondrial transfer signal (MTS), or a chloroplast transit peptide (CTP). For example, the fusion protein may optionally further comprise a nuclear localization signal (NLS) for base correction of nuclear DNA. In another example, the fusion protein may optionally further comprise a mitochondrial transfer signal (MTS) or a nuclear export signal (NES) for base correction of mitochondrial DNA. As another example, the fusion protein may be for base correction of chloroplast DNA and may optionally additionally include a chloroplast transit peptide (CTP).

[0078] The term "organelle" as used herein refers to a small-unit organelle with a specific function existing within a cell, which contributes to the survival, growth, metabolism, and various biological functions of the cell. For example, it may include the nucleus of all eukaryotic cells, mitochondria, and chloroplasts of plant cells. The mitochondria are organelles existing in eukaryotic cells, which function to supply intracellular energy in the form of ATP, and have their own DNA. The chloroplasts are organelles existing in higher plants and marine algae, which perform photosynthesis, and have their own DNA.

[0079] In one specific example, the target nucleic acid molecule may be located in an organelle. For example, the organelle may be a nucleus, a mitochondrion, or a chloroplast.

[0080]

[0081] Another aspect provides a polynucleotide encoding the fusion protein.

[0082]

[0083] Another aspect provides a vector comprising the polynucleotide.

[0084] The term "vector" as used herein may refer to a nucleic acid molecule capable of transporting another nucleic acid to which it is linked. A vector may include a nucleic acid molecule that is single-stranded, double-stranded, or partially double-stranded; a nucleic acid molecule comprising one or more free ends, a nucleic acid molecule without free ends (e.g., circular); a nucleic acid molecule comprising DNA, RNA, or both; and various other polynucleotides known in the art. One type of vector is a "plasmid," which may refer to a circular double-stranded DNA loop into which additional DNA segments may be inserted, for example, by standard molecular cloning techniques. Another type of vector is a viral vector, in which viral-derived DNA or RNA sequences may be present for packaging into a virus (e.g., a retrovirus, a replication-defective retrovirus, an adenovirus, a replication-defective adenovirus, and an adeno-associated virus). A recombinant expression vector may comprise a nucleic acid of the invention in a form suitable for expression of the nucleic acid in a host cell, which may mean that the recombinant expression vector comprises one or more regulatory elements, which one or more regulatory elements may be selected based on the host cell to be used for expression, and which may be operably linked to the nucleic acid sequence to be expressed. Within a recombinant expression vector, “operably linked” may mean that the nucleotide sequence of interest is linked to the regulatory element(s) in a manner that permits expression of the nucleotide sequence (e.g., in an in vitro transcription / translation system or in a host cell if the vector has been introduced into the host cell).

[0085] The "regulatory elements" may include promoters, enhancers, internal ribosomal entry sites (IRES), and other expression control elements (e.g., transcription termination signals such as polyadenylation signals and poly-U sequences). Regulatory elements are described, for example, in Goeddel, GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY 185, Academic Press, San Diego, Calif. (1990). Regulatory elements may include those that direct constitutive expression of a nucleotide sequence in many types of host cells and those that direct expression of a nucleotide sequence only in certain host cells (e.g., tissue-specific regulatory sequences). In some embodiments, the vector can comprise one or more pol III promoters (e.g., one, two, three, four, five or more pol III promoters), one or more pol II promoters (e.g., one, two, three, four, five or more pol II promoters), one or more pol I promoters (e.g., one, two, three, four, five or more pol I promoters), or a combination thereof. Examples of pol III promoters can include, but are not limited to, U6 and H1 promoters. Examples of pol II promoters can include, but are not limited to, the retrovirus Rous sarcoma virus (RSV) LTR promoter (optionally with an RSV enhancer), the cytomegalovirus (CMV) promoter (optionally with a CMV enhancer), the SV40 promoter, the dihydrofolate reductase promoter, the β-actin promoter, the phosphoglycerol kinase (PGK) promoter, and the EF1α promoter.For example, vectors may include lentiviruses and adeno-associated viruses (AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, or AAV9), and the type of vector may also be selected for targeting specific types of cells.

[0086] Additionally, multiple nucleic acid molecules within the vector system may be located on the same or different vectors.

[0087] The above fusion protein, the polynucleotide encoding it, and SsdA (single-stranded DNA deaminase toxin A) are as described above.

[0088]

[0089] Another aspect provides a composition for base correction comprising at least one from the group consisting of the fusion protein; a polynucleotide encoding the fusion protein; and a vector comprising a polynucleotide encoding the fusion protein.

[0090] In one specific example, the base correction composition may substitute at least one nucleotide in the nucleotide sequence of a target nucleic acid molecule.

[0091] In one specific example, the target nucleic acid molecule may be located in an organelle, for example, but not limited to, a nucleus, a mitochondria, or a chloroplast.

[0092] The fusion protein, the polynucleotide encoding the fusion protein, the nucleic acid-binding protein (NABP), the single-stranded DNA deaminase toxin A (SsdA), and the organelle are as described above.

[0093]

[0094] Another aspect provides a method of correcting a target nucleic acid, comprising the step of contacting a target nucleic acid molecule with the base correction composition.

[0095] In one specific example, the method for correcting the target nucleic acid may be to substitute at least one nucleotide in the nucleotide sequence of the target nucleic acid molecule, and the target nucleic acid molecule may be located in an organelle, for example, the organelle may be a nucleus, a mitochondrion, or a chloroplast, but is not limited thereto.

[0096] In one specific example, the correction of the target nucleic acid may be from cytosine to uracil.

[0097] The nucleic acid may be RNA or DNA.

[0098] The contacting step may refer to treating a plant cell with the base correction composition. For example, it may refer to treating a plant cell with a composition comprising at least one selected from the group consisting of the fusion protein; a polynucleotide encoding the fusion protein; and a vector comprising a polynucleotide encoding the fusion protein.

[0099] In one specific example, the contacting step may mean treating a plant cell with a composition comprising any one of the fusion protein; a nucleotide encoding the fusion protein; and a vector comprising a polynucleotide encoding the fusion protein; and a nuclear localization signal (NLS) peptide or a nucleotide encoding the same.

[0100] In one specific example, the contacting step may mean treating a plant cell with a composition comprising any one of the fusion protein; a nucleotide encoding the fusion protein; and a vector comprising a polynucleotide encoding the fusion protein; and a chloroplast transit peptide or a nucleotide encoding the same.

[0101] In one specific example, the contacting step may include treating a plant cell with a composition comprising any one of the fusion protein; a nucleotide encoding the fusion protein; and a vector comprising a polynucleotide encoding the fusion protein; and a mitochondrial targeting signal or a nucleotide encoding the same.

[0102] Additionally, the contacting step may mean treating an animal cell with the base correction composition. For example, it may mean treating an animal cell with a composition comprising at least one selected from the group consisting of the fusion protein; a nucleotide encoding the fusion protein; and a vector comprising a polynucleotide encoding the fusion protein.

[0103] In one specific example, the contacting step may mean treating an animal cell with a composition comprising any one of the fusion protein; a nucleotide encoding the fusion protein; and a vector comprising a polynucleotide encoding the fusion protein; and a nuclear localization signal (NLS) peptide or a nucleotide encoding the same.

[0104] In one specific example, the contacting step may mean treating an animal cell with a composition comprising any one of the fusion protein; a nucleotide encoding the fusion protein; and a vector comprising a polynucleotide encoding the fusion protein; and a mitochondrial targeting signal or a nucleotide encoding the same.

[0105] In one specific example, the contacting step may mean treating a plant cell or an animal cell with a composition comprising any one of the fusion protein; a nucleotide encoding the fusion protein; and a vector comprising a polynucleotide encoding the fusion protein; and a nuclear export signal (NES) or a nucleotide encoding the same.

[0106]

[0107] Another aspect provides a base editing system comprising a fusion protein comprising a nucleic acid-binding protein (NABP); and a bacterial toxin or a polynucleotide encoding the fusion protein; and a guide polynucleotide.

[0108] The above fusion protein, the polynucleotide encoding the fusion protein, the nucleic acid-binding protein (NABP) and the single-stranded DNA deaminase toxin A (SsdA) are as described above.

[0109] According to the fusion protein according to the aspect, effective base correction can be performed within the cell organelle, and this has the effect of enabling base correction of genes in the cell nucleus and cell organelles with a single fusion protein.

[0110] Figure 1 is a schematic diagram showing the structure of a fusion protein combining a wild-type SsdA protein and a UGI protein with a TALE protein binding to mitochondrial nucleic acid.

[0111] Figure 2 is a schematic diagram illustrating the base correction process in the mitochondrial ND1 gene using a TALE-SRE expression plasmid of one aspect.

[0112] Figure 3 is a graph showing the results of evaluating the base correction efficiency for mitochondrial genes by delivering a TALE-SsdA(Y335R)-UGI fusion protein of one aspect into cells in the form of a monomer or dimer.

[0113] Figure 4 shows the results of analyzing the base correction efficiency in mitochondrial genes induced by one aspect of TALE-SRE: Figure 4a is a schematic diagram showing the process in which a TALE-SRE fusion protein binds to the target base sequence of mitochondrial genes ND1, CYB, COX3, and ATP6 to induce cytosine base correction, and the result visualizing the base correction frequency in the form of a heatmap, and Figure 4b is a graph quantitatively comparing the base correction efficiency between the conditions in which a TALE-SRE fusion protein was introduced in the form of a monomer (left or right) or dimer (left+right) targeting the ND1, ATP6-1, ATP6-2, COX3-1, COX3-2, COX3-3, CYB-1, and CYB-2 gene regions and the untreated condition.

[0114] Figure 5 is a graph showing the results of introducing point mutations into the mitochondrial ND1 gene using a ZFP-SRE expression plasmid.

[0115] The present invention will be described in more detail through the following examples. However, these examples are provided for illustrative purposes only and the scope of the present invention is not limited to these examples.

[0116]

[0117] Example 1. Method for producing SsdA protein

[0118] The method for producing the SsdA protein was performed according to the method described in Republic of Korea Patent Application No. 10-2023-0140009. The entire contents of the above document are incorporated herein by reference.

[0119] Briefly, gBlock double-stranded DNA fragments encoding His6-SsdA and SsdAI (Integrated DNA Technologies) and the bacterial expression vector pET-28b(+) DNA (Novagen) were treated with XbaI and XhoI restriction enzymes (New England Biolabs), respectively, at 37°C for 3 h. The linearized pET-28b vector was then purified using agarose gel extraction (Geneall) and ligated to the gBlock double-stranded DNA fragments using Quick ligase (New England Biolabs).

[0120] Then, using pCMV plasmid DNA, the coding sequence of UGI was obtained through PCR amplification, and the SsdA sequence was amplified from gBlock using Gibson Assembly Master Mix (New England Biolabs) and subcloned into the pCMV plasmid. The sequences of the SsdA wild-type sequence (PAAR domain-containing protein), SsdA deaminase domain, and SsdAI are shown in Table 1.

[0121] SsdA protein sequence SEQ ID NO: PAAR-domain containing proteinMSAAARVNDPIEHTGSLTGLLAGLAIGAIGAALVVGTGGLAAVAIVGASAATGAGVGQLIGSLSCCNHQTGQIVSGSSNVYINGEPAARAHADQAKCDEHTSRPQVIAQGSSNVYINGHPAARVGDRTACDAKIVVGSSNVFIGGGTETTDPINPEVPELLERSILLVGLASAVVLASPVIVIAGLVGGIAGGTVGSMGGAQLFGE GTDGQKLMAFGGALLGGGLGAKGGKWFDTRYDIKVQGVGSNLGNLKITPKGAAKVSNIAESEAALGRASQARADLPQSKELKVKTVSSNDKKTLSGWGNKKPEGYER ISAEQVKAKSEEIGHEVKSHPYDRDYKGQYFSSHAEKQMSIASPNHPLGVSKPMCTDCQGYFSQLAKYSKVEQTVADPKAIRIFKTDGSVETIMRSE1deaminase DomainKVSNIAESEA ALGRASQARA DLPQSKELKV KTVSSNDKKT LSGWGNKKPE GYERISAEQV KAKSEEIGHE VKSHPYDRDY KGQYFSSHAE KQMSIASPNH PLGVSKPMCT DCQGYFSQLA KYSKVEQTVA DPKAIRIFKT DGSVETIMRSE2SsdAIMNNKSKVLIEKLLLEVAKSPEGELILPLRKLLWNTITEDETAAKKKAILTALDVMCVRQGVNFWIKKFGDNEPLNYILNIALETAEGK FDESKALGLRDEFYVSIVEDQEYEVEEYPAMFVGHAAANTIARAVDDFQFEPYDHRVDRDLDPEGFESSYLVASAFAGGLSEDGDPKLRRAFWEWYLSIAVPQVV3

[0122] The cloned SsdA protein was purified using Escherichia coli BL21. The protein purification results were confirmed by SDS-PAGE. In addition, deamination of cytosine in single-stranded DNA of SsdA was confirmed.

[0123]

[0124] Example 2. Preparation of SsdA mutant protein

[0125] The SsdA mutants and the method for producing them were performed according to the method described in Republic of Korea Patent Application No. 10-2025-0033494. The entire contents of the above document are incorporated herein by reference. Examples of SsdA mutants are briefly described below.

[0126]

[0127] 2.1 Preparation of SsdA mutants with amino acid substitutions at positions that can enhance binding affinity with target DNA and confirmation of their base correction effects.

[0128] Except for the amino acid sequence of the SsdA mutant, a mutant was prepared by substituting amino acids at positions (V289, H333, Y335, P282, K392) that can improve the binding affinity of the SsdA protein (PAAR domain-containing protein, sequence number 1) with target DNA, in the same manner as in Example 1, and the base correction effect of the mutant was confirmed.

[0129]

[0130] 2.2 Preparation of SsdA mutants with amino acids substituted in the catalytic active site and confirmation of their base correction effect

[0131] A mutant in which amino acids in the catalytically active site of the PAAR domain-containing protein (SEQ ID NO: 1) were substituted was prepared in the same manner as in Example 1 except for the amino acid sequence of the SsdA mutant (a mutant in which amino acids were substituted at positions T291 and K365 of the PAAR domain-containing protein (SEQ ID NO: 1)), and the base correction effect of the mutant was confirmed.

[0132] The above catalytic active site corresponds to a portion 260 to 410 of the PAAR-domain containing protein (SEQ ID NO: 1), and is used interchangeably as the SsdA toxin domain or deaminase domain in the present specification, and the catalytic active site includes the amino acid sequence of SEQ ID NO: 2.

[0133]

[0134] 2.3 Preparation of SsdA mutants with mutations in Y335R and verification of their base correction efficiency

[0135] An SsdA mutant (SsdA-Y335R) was prepared by introducing an R mutation at the Y335 position of the PAAR domain-containing protein (SEQ ID NO: 1) using the method of Example 1, except for the amino acid sequence of the SsdA mutant, and the base correction effect was confirmed.

[0136]

[0137] 2.4 Preparation of SsdA mutants with additional mutations introduced into the Y335R mutant and verification of their base correction efficiency

[0138] An SsdA mutant (SsdA-SRE) was prepared by introducing S, R, and E mutations at positions P282, Y335, and K392 of the PAAR domain-containing protein (SEQ ID NO: 1) using the method of Example 1, except for the amino acid sequence of the SsdA mutant, and the base correction effect of the mutant was confirmed.

[0139] The amino acid sequence of the above SsdA-SRE is shown in Table 2.

[0140] SsdA protein sequence Sequence number SsdA-SRE (PARR containing protein_SRE) MSAAARVNDPIEHTGSLTGLLAGLAIGAIGAALVVGTGGLAAVAIVGASAATGAGVGQLIGSLSCCNHQTGQIVSGSSNVYINGEPAARAHADQAKCDEHTSRPQVIAQGSSNVYINGHPAARVGDRTACDAKIVVGSSNVFIGGGTETTDPINPEVPELLERSILLVGLASAVVLASPVIVIAGLVGGIAGGTVGSMGGAQLFGEGT DGQKLMAFGGALLGGGLGAKGGKWFDTRYDIKVQGVGSNLGNLKITPKGAAKVSNIAESEAALGRASQARADLSQSKELKVKTVSSNDKKTLSGWGNKKPEGYERISAEQ VKAKSEEIGHEVKSHPRDRDYKGQYFSSHAEKQMSIASPNHPLGVSKPMCTDCQGYFSQLAKYSKVEQTVADPEAIRIFKTDGSVETIMRSE4SsdA-SRE(deaminase domain_SRE)KVSNIAESEAALGRASQARADLSQSKELKVKTVSSNDKKTLSGWGNKKPEGYERISAEQVKAKSEEIGHEVKSHPRDRDYKGQYFSSHAEKQMSIASPNHPLGVSKPMCTDCQGYFSQLAKYSKVEQTVADPEAIRIFKTDGSVETIMRSE5

[0141] Sequence number 4 is an SsdA mutant (SsdA-SRE) that has introduced S, R, and E mutations at positions P282, Y335, and K392, respectively, of sequence number 1, and sequence number 5 is an SsdA mutant (SsdA-SRE) that has introduced S, R, and E mutations at positions P25, Y78, and K135, respectively, of sequence number 2.

[0142]

[0143] Example 3. Construction of a base editing system using TALE proteins and SsdA mutants.

[0144] 3.1 Construction of a base editing system using TALE proteins and wild-type SsdA

[0145] To construct a cytosine base correction system using TALE protein and SsdA (wild-type) of Example 1, a Golden Gate cloning system consisting of a total of four expression plasmids (including protein construct expression for transferring the produced protein to mitochondria, Mitochondrial transfer signal, MTS) and 424 TALE sub-array plasmids was developed.

[0146] In order to efficiently introduce point mutations into mitochondrial DNA, the above system cloned the MTS sequence at the N-terminus of the expression plasmid, and sequentially linked the SsdA (wild-type), one UGI, and NES (nuclear export signal) sequences at the C-terminus.

[0147] To construct a TALE array that specifically binds to a target sequence, a TALE subarray plasmid corresponding to the target sequence and an expression plasmid were assembled by reacting them with BsaI restriction enzyme and T4 DNA ligase. The assembled TALE-SsdA (wild-type) expression plasmid was transformed into bacteria, and the exact composition of the plasmid was confirmed through colony PCR and base sequence analysis.

[0148] Figure 1 is a schematic diagram showing the structure of a fusion protein combining a wild-type SsdA protein and a UGI protein with a TALE protein binding to mitochondrial nucleic acid.

[0149]

[0150] 3.2 Construction of a base editing system using TALE proteins and SsdA mutants

[0151] To construct a cytosine base correction system using TALE array, a Golden Gate cloning system consisting of four expression plasmids (including a protein construct expression plasmid for delivering the produced protein to the mitochondria, a Mitochondrial transfer signal, MTS) and 424 TALE subarray plasmids was developed. In this system, MTS was cloned at the N-terminus of the expression plasmids to efficiently introduce point mutations into mitochondrial DNA, and the SsdA-SRE of Example 2.4, one UGI, and an NES signal were added to the C-terminus. To synthesize a TALE array that recognizes the target base sequence, the TALE subarray plasmid that recognizes the target base sequence and the expression plasmid were reacted with BsaI restriction enzyme and T4 DNA ligase to complete the TALE-SRE expression plasmid. Meanwhile, TALE-Y335R was completed using the same method as the TALE-SRE expression plasmid production method described above, except that SsdA-Y335R of Example 2.3 was introduced to the C-terminus of the expression plasmid.

[0152] After the assembled vector was transformed into bacteria, it was verified through colony PCR and sequencing.

[0153] Figure 2 is a schematic diagram illustrating the base correction process in the mitochondrial ND1 gene using a TALE-SRE expression plasmid of one aspect.

[0154]

[0155] Example 4. Construction of a base correction system using zinc finger proteins and SsdA mutants.

[0156] 4.1 Construction of a Base Correction System Using Zinc Finger Protein and Wild-Type SsdA

[0157] To construct a base editing system (ZFP-SsdA) using Zinc Finger Protein (ZFP), a DNA binding protein, and SsdA (wild-type), a ZFP-SsdA expression plasmid containing a mitochondrial transfer signal (MTS) and SsdA (wild-type) and expressing a zinc finger protein obtained from a publicly available zinc finger resource was constructed.

[0158]

[0159] 4.2 Construction of a Base Correction System Using Zinc Finger Protein and SsdA Mutants

[0160] To construct a base editing system (ZFP-SRE) using a zinc finger protein known as a DNA binding protein and a SsdA protein variant (SsdA-SRE), a ZFP-SRE expression plasmid was constructed that contains a mitochondrial transfer signal (MTS) and a SsdA protein variant (SsdA-SRE) and expresses the zinc finger protein obtained from a publicly available zinc finger resource.

[0161]

[0162] Experimental Example 1. Confirmation of cytosine deamination in mitochondria of wild-type SsdA.

[0163] To determine whether cytosine base correction is induced in mitochondrial genes when wild-type SsdA and UGI proteins are fused to TALE proteins and delivered into cells, a TALE-SsdA (wild-type)-UGI fusion protein produced by the method of Example 3.1 was introduced into HEK293T / 17 cells, a human cell line. The fusion proteins were designed to have Left (L) and Right (R) TALE structures, respectively, and were delivered into cells singly (monomer) or in pairs (dimer).

[0164] The occurrence of the base correction reaction was confirmed by NGS (Next Generation Sequencing) to confirm the efficiency of the reaction in which SsdA (wild-type) substitutes the cytosine base in mitochondrial DNA with uracil.

[0165] The above TALE-SsdA(wild-type)-UGI fusion protein was confirmed to induce base correction in which the target cytosine base of a mitochondrial gene was substituted with thymine under both monomer and dimer conditions.

[0166]

[0167] Experimental Example 2. Confirmation of the mitochondrial base-correction effect of SsdA mutants.

[0168] 2.1 Confirmation of the mitochondrial base correction effect using the TALE-Y335R expression plasmid

[0169] To determine whether cytosine base correction is induced in mitochondrial genes when TALE protein is fused with SsdA mutant (Y335R) and UGI protein and delivered into cells, a TALE-SsdA(Y335R)-UGI fusion protein produced by the method of Example 3.2 was introduced into HEK293T / 17 cells, a human cell line. The fusion proteins were designed to have Left (L) and Right (R) TALE structures, respectively, and were delivered into cells singly (monomer) or in pairs (dimer).

[0170] The occurrence of the base correction reaction was confirmed by NGS (Next Generation Sequencing) to determine the efficiency of the reaction in which SsdA (wild-type) substitutes the cytosine base in mitochondrial DNA with uracil, and the results are shown in Fig. 3.

[0171] Figure 3 is a graph showing the results of evaluating the base correction efficiency for mitochondrial genes by delivering a TALE-SsdA(Y335R)-UGI fusion protein of one aspect into cells in the form of a monomer or dimer.

[0172] As shown in Fig. 3, the TALE-SsdA(Y335R)-UGI fusion protein was confirmed to induce base correction in which the target cytosine base of a mitochondrial gene was substituted with thymine under both monomer and dimer conditions. In particular, base correction was observed to be induced within mitochondria even under L-monomer conditions, and a selectively high correction frequency was observed under monomer conditions at some specific base positions. This suggests that a single TALE-SsdA(Y335R)-UGI protein can bind to the target base sequence and effectively correct the cytosine base.

[0173] Meanwhile, base correction was also induced under L / R dimer conditions, in which case base correction occurred at a wider range of cytosine positions.

[0174] These results indicate that the fusion protein can bind to a target sequence within a mitochondrial gene and effectively correct cytosine bases under various conditions.

[0175]

[0176] 2.2. Confirmation of the effect of mitochondrial base correction using TALE-SRE expression plasmid

[0177] The TALE-SRE expression plasmid of Example 3.2 was used to introduce point mutations into mitochondrial DNA. Specifically, the TALE-SRE expression plasmid was delivered to HEK293T cells, and 4 days later, genomic DNA was extracted from the cells, and the target sequence was amplified using PCR. The amplified DNA was then analyzed using next-generation sequencing (NGS), and the target sequences were the ND1 gene, ATP6-1 gene, ATP6-2 gene, COX3-1 gene, COX3-2 gene, COX3-3 gene, CYB-1 gene, and CYB-2 gene. The TALE-SRE expression plasmids for introducing point mutations into each target sequence are as follows:

[0178] Point mutations were introduced into the CYB gene using TALE-SRE-CYB-Left (SEQ ID NO: 6) and TALE-SRE-CYB-Right (SEQ ID NO: 7). The components of TALE-SRE-CYB-Left and TALE-SRE-CYB-Right are shown in Table 3 below.

[0179] TALE-SRE-CYB-LeftComponentsAmino acid sequencesMTSMASVLTPLLLRGLTGSARRLPVPRAKIHSLFLAG tagDYKDHDGDYKDHDIDYKDDDDKTALEN-terminal domainDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNCYB-leftbinding domainLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLHalf-domainLPVLCQAHGLTPEQVVAIASNGGGKQALETALEC-terminaldomainDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNSsdA-SREKVSNIAESEAALGRASQARADLSQSKELKVKTVSSNDKKTLSGWGNKKPEGYERISAEQVKAKSEEIGHEVKSHPRDRDYKGQYFSSHAEKQMSIASPNHPLGVSKPMCTDCQGYFSQLAKYSKVEQTVADPEAIRIFKTDGSVETIMRSE1xUGISGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKMLNESVDEMTKKFGTLTIHDTEK*TALE-SRE-CYB-RightComponentsAmino acid sequencesMTSMASVLTPLLRGLTGSARRLPVPRAKIHSLFLAG tagDYKDHDGDYKDHDIDYKDDDDKTALEN-terminal domainDLRTLGYSQQQQEKIKKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPNLCYB-rightbinding-domainLPVLCQAHGLTPEQVVAIASNGGGKQALETALEC-terminaldomainDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNSsdA-SREKVSNIAESEAALGRASQARADLSQSKELKVKTVSSNDKKTLSGWGNKKPEGYERISAE QVKAKSEEIGHEVKSHPRDRDYKGQYFSSHAEKQMSIASPNHPLGVSKPMCTDCQGYFSQLAKYSKVEQTVADPEAIRIFKTDGSVETIMRSE1xUGISGGST NLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKMLNESVDEMTKKFGTLTIHDTEK*

[0180] Point mutations were introduced into the ND1 gene using TALE-SRE-ND1-Left (SEQ ID NO: 8) and TALE-SRE-ND1-right (SEQ ID NO: 9). The components of TALE-SRE-ND1-Left and TALE-SRE-ND1-right are shown in Table 4.

[0181] TALE-SRE-ND1-LeftComponentsAmino acid sequencesMTSMASVLTPLLLRGLTGSARRLPVPRAKIHSLFLAG tagDYKDHDGDYKDHDIDYKDDDDKTALEN-terminal domainDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNND1-leftbinding domainLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLHalf-domainLPVLCQAHGLTPEQVVAIASNGGGKQALETALEC-terminaldomainDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNSsdA-SREKVSNIAESEAALGRASQARADLSQSKELKVKTVSSNDKKTLSGWGNKKPEGYERISAEQVKAKSEEIGHEVKSHPRDRDYKGQYFSSHAEKQMSIASPNHPLGVSKPMCTDCQGYFSQLAKYSKVEQTVADPEAIRIFKTDGSVETIMRSE1xUGISGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKMLNESVDEMTKKFGTLTIHDTEK*TALE-SRE-ND1-rightComponentsAmino acid sequencesMTSMASVLTPLLLRGLTGSARRLPVPRAKIHSLFLAG tagDYKDHDGDYKDHDIDYKDDDDKTALEN-terminal domainDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNND1-rightbinding-domainLPVLCQAHGLTPEQVVAIASNGGGKQALETALEC-terminaldomainDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNSsdA-SREKVSNIAESEAALGRASQARADLSQSKELKVKTVSSNDKKTLSGWGNKKPEGYERISAE QVKAKSEEIGHEVKSHPRDRDYKGQYFSSHAEKQMSIASPNHPLGVSKPMCTDCQGYFSQLAKYSKVEQTVADPEAIRIFKTDGSVETIMRSE1xUGISGGST NLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKMLNESVDEMTKKFGTLTIHDTEK*

[0182] Point mutations were introduced into the ATP6 gene using TALE-SRE-ATP6-left (SEQ ID NO: 10) and TALE-SRE-ATP6-right (SEQ ID NO: 11). The components of TALE-SRE-ATP6-left and TALE-SRE-ATP6-right are shown in Table 5 below.

[0183] TALE-SRE-ATP6-leftComponentsAmino acid sequencesMTSMASVLTPLLLRGLTGSARRLPVPRAKIHSLFLAG tagDYKDHDGDYKDHDIDYKDDDDKTALEN-terminal domainDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNATP6-leftbinding domainLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGHalf-domainLPVLCQAHGLTPEQVVAIASNGGGKQALETALEC-terminaldomainDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNSsdA-SREKVSNIAESEAALGRASQARADLSQSKELKVKTVSSNDKKTLSGWGNKKPEGYERISAEQVKAKSEEIGHEVKSHPRDRDYKGQYFSSHAEKQMSIASPNHPLGVSKPMCTDCQGYFSQLAKYSKVEQTVADPEAIRIFKTDGSVETIMRSE1xUGISGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKMLNESVDEMTKKFGTLTIHDTEK*TALE-SRE-ATP6-rightComponentsAmino acid sequencesMTSMASVLTPLLLRGLTGSARRLPVPRAKIHSLFLAG tagDYKDHDGDYKDHDIDYKDDDDKTALEN-terminal domainDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNATP6-rightbinding-domainLPVLCQAHGLTPEQVVAIASNGGGKQALETALEC-terminaldomainDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNSsdA-SREKVSNIAESEAALGRASQARADLSQSKELKVKTVSSNDKKTLSGWGNKKPEGYERISAE QVKAKSEEIGHEVKSHPRDRDYKGQYFSSHAEKQMSIASPNHPLGVSKPMCTDCQGYFSQLAKYSKVEQTVADPEAIRIFKTDGSVETIMRSE1xUGISGGST NLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKMLNESVDEMTKKFGTLTIHDTEK*

[0184] Point mutations were introduced into the COX3 gene using TALE-SRE-COX3-left (SEQ ID NO: 12) and TALE-SRE-COX3-right (SEQ ID NO: 13). The components of TALE-SRE-COX3-left and TALE-SRE-COX3-right are shown in Table 6 below.

[0185] TALE-SRE-COX3-leftComponentsAmino acid sequencesMTSMASVLTPLLLRGLTGSARRLPVPRAKIHSLFLAG tagDYKDHDGDYKDHDIDYKDDDDKTALEN-terminal domainDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNCOX3-leftbinding domainLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGHalf-domainLPVLCQAHGLTPEQVVAIASNGGGKQALETALEC-terminaldomainDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNSsdA-SREKVSNIAESEAALGRASQARADLSQSKELKVKTVSSNDKKTLSGWGNKKPEGYERISAEQVKAKSEEIGHEVKSHPRDRDYKGQYFSSHAEKQMSIASPNHPLGVSKPMCTDCQGYFSQLAKYSKVEQTVADPEAIRIFKTDGSVETIMRSE1xUGISGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKMLNESVDEMTKKFGTLTIHDTEK*TALE-SRE-COX3-rightComponentsAmino acid sequencesMTSMASVLTPLLLRGLTGSARRLPVPRAKIHSLFLAG tagDYKDHDGDYKDHDIDYKDDDDKTALEN-terminal domainDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNCOX3-rightbinding-domainLPVLCQAHGLTPEQVVAIASNGGGKQALETALEC-terminaldomainDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNSsdA-SREKVSNIAESEAALGRASQARADLSQSKELKVKTVSSNDKKTLSGWGNKKPEGYERISAE QVKAKSEEIGHEVKSHPRDRDYKGQYFSSHAEKQMSIASPNHPLGVSKPMCTDCQGYFSQLAKYSKVEQTVADPEAIRIFKTDGSVETIMRSE1xUGISGGST NLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKMLNESVDEMTKKFGTLTIHDTEK*

[0186] The results of base correction of mitochondrial genes using the above TALE-SRE expression plasmid are shown in Fig. 4 (data are expressed as mean + - standard error of the mean (sem) obtained from n=3 biologically independent samples). Fig. 4 shows the results of analyzing the base correction efficiency in mitochondrial genes induced by one aspect of TALE-SRE: Fig. 4a is a schematic diagram showing the process in which TALE-SRE fusion proteins bind to target base sequences of mitochondrial genes ND1, CYB, COX3 and ATP6 to induce cytosine base correction, and the result visualizing the base correction frequency in the form of a heatmap, and Fig. 4b is a diagram showing the process in which TALE-SRE fusion proteins bind to target base sequences of mitochondrial genes ND1, CYB, COX3 and ATP6 to induce cytosine base correction, and the result visualizing the base correction frequency in the form of a heatmap, targeting ND1, ATP6-1, ATP6-2, COX3-1, COX3-2, COX3-3, CYB-1 and CYB-2 gene regions, and TALE-SRE fusion proteins as monomers (left or right) or This is a graph quantitatively comparing the base correction efficiency between the conditions introduced in the dimer (left+right) form and the untreated conditions.

[0187] As shown in Fig. 4a, significant base-editing activity was observed in multiple mitochondrial target genes, including ND1, CYB, COX3, and ATP6, even under conditions where only a single (monomer) TALE-SRE fusion protein was introduced (Left or Right). This result demonstrates that efficient cytosine base-editing within mitochondria is possible even with a simpler protein composition compared to the existing dimer-based TALE system.

[0188] In addition, as shown in Fig. 4b, it was confirmed that the base correction system using TALE protein and SsdA mutant (SsdA-SRE) exhibited a point mutation introduction efficiency of up to 11.2% for mitochondrial genes ND1, ATP6, CYB, and COX3.

[0189] These results suggest that the TALE-SRE system enables base editing within mitochondria using only a single (monomer) TALE array. This differs from the existing mitochondrial base editing technology, DdCBE, which requires a pair of (dimer) TALE arrays using segmented DddAtox, demonstrating that the TALE-SRE system can reliably perform mitochondrial base editing with fewer components. Furthermore, the TALE-SRE system has the advantage of facilitating the gene delivery of the base editing system into mitochondria because it enables effective base editing with fewer components than existing technologies.

[0190]

[0191] 2.2. Confirmation of the effect of mitochondrial base correction using the ZFP-SRE expression plasmid.

[0192] A point mutation was introduced into the ND1 gene in mitochondria using the ZFP-SRE expression plasmid constructed using the method of Example 4.2. Specifically, the ZFP-SRE expression plasmid was delivered to HEK293T cells, and 4 days later, genomic DNA was extracted from the cells and the target sequence was amplified using PCR. The amplified DNA was analyzed using next-generation sequencing (NGS). The ZFP-SRE expression plasmid for introducing a point mutation into the ND1 gene in mitochondria is as follows:

[0193] Point mutations were introduced into the ND1 gene using mitoZFD-SRE-ND1-left (SEQ ID NO: 14) and mitoZFD-SRE-ND1-right (SEQ ID NO: 15). The components of mitoZFD-SRE-ND1-left and mitoZFD-SRE-ND1-right are shown in Table 7.

[0194] mitoZFD-SRE-ND1-leftComponentsAmino acid sequencesMTSMLGFVGRVAAAPASGALRRLTPSASLPPAQLLLRAAPTAVHPVRDYAAQHA tagYPYDVPDYANESVDEMTKKFGTLTIHDTEKSsdA-SREKVSNIAESEAALGRASQARADLSQSKELKVKTVSSNDKKTLSGWGNKKPEGYERISAEQVKAKSEEIGHEVKSHPRDRDYKGQYFSSHAEKQMSIASPNHPLGVSKPMCTDCQGYFSQLAKYSKVEQTVADPEAIRIFKTDGSVETIMRSEND1-leftZinc-fingerbinding domainFQCRICMRNFSDSGNLRVHIRTHTGEKPYKCPDCGKSFSQSSSLIRHQRTHTGEKPYECDHCGKSFSQSSHLNVHKRTHTGEKPYRCKYCDRSFSISSNLQRHVRNIH1xUGITNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML*mitoZFD-SRE-ND1-rightComponentsAmino acid sequencesMTSMLGFVGRVAAAPASGALRRLTPSASLPPAQLLLRAAPTAVHPVRDYAAQHA tagYPYDVPDYANESVDEMTKKFGTLTIHDTEKSsdA-SREKVSNIAESEAALGRASQARADLSQSKELKVKTVSSNDKKTLSGWGNKKPEGYERISAEQVKAKSEEIGHEVKSHPRDRDYKGQYFSSHAEKQMSIASPNHPLGVSKPMCTDCQGYFSQLAKYSKVEQTVADPEAIRIFKTDGSVETIMRSEND1-rightZinc-fingerbindingdomainYKCPECGKSFSTKNSLTEHQRTHTGEKPYKCPECGKSFSSKKALTEHQRTHTGEKPYECNYCGKTFSVSSTLIRHQRIHTGEKPYRCKYCDRSFS ISSNLQRHVRNIH1xUGITNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML*

[0195] The results of introducing point mutations into the ND1 gene in mitochondria using the ZFP-SRE expression plasmid are shown in Fig. 5 (data are expressed as the mean + - standard error of the mean (sem) obtained from n=3 biologically independent samples). Fig. 5 is a graph showing the results of introducing point mutations into the mitochondrial ND1 gene using the ZFP-SRE expression plasmid.

[0196] As shown in Fig. 5, the base correction system (ZFP-SRE) using Zinc Finger Protein and SsdA mutant (SsdA-SRE) was confirmed to exhibit a point mutation introduction efficiency of up to 12% for the mitochondrial ND1 gene.

[0197] These results imply that the base editing system (ZFP-SRE) using Zinc Finger Protein and SsdA mutant (SsdA-SRE) can perform base editing in mitochondria with only a single (monomer) TALE array. This is different from the existing base editing technology using ZFP and DddAtox that required a pair (dimer) of TALE arrays, and demonstrates that the ZFP-SRE system can perform mitochondrial base editing stably with fewer components. In addition, the ZFP-SRE system has the advantage of easier gene delivery of the base editing system into mitochondria because it enables effective base editing with fewer components than existing technologies.

[0198]

[0199] Experimental Example 3. Confirmation of the chloroplast base-correction effect of SsdA.

[0200] 3.1. Confirmation of cytosine deamination in chloroplasts by TALE-SsdA fusion protein

[0201] An expression vector encoding a TALE-SsdA fusion protein, which fuses a TALE protein and an SsdA protein, or a TALE-SRE fusion protein, which fuses a TALE protein and an SsdA_SRE mutant, was constructed and introduced into plant cells. The vector contained a chloroplast targeting peptide (CTP) sequence at the N-terminus. Specifically, a TALE array was designed to recognize a specific target base sequence of the chloroplast gene psbA or rbcL, and a cytosine deamination reaction was induced using the TALE-SsdA fusion protein that binds to the sequence.

[0202] NGS analysis confirmed a deamination reaction in which cytosine at that position is converted to uracil.

[0203]

[0204] 3.2. Confirmation of cytosine deamination in chloroplasts by ZFP-SsdA fusion protein

[0205] An expression vector encoding a ZFP-SsdA fusion protein, which fuses Zinc Finger Protein (ZFP) and SsdA protein, or a ZFP-SRE fusion protein, which fuses ZFP and SsdA_SRE mutant, was constructed and introduced into plant cells. The vector contained a chloroplast targeting signal sequence (CTP, chloroplast targeting peptide) at the N-terminus to induce protein delivery to chloroplasts, and was designed to bind to specific target base sequences of the chloroplast genes psbA and rbcL to induce cytosine deamination.

[0206] NGS analysis of chloroplast DNA confirmed that a base correction reaction occurred in which the target cytosine was substituted with uracil in the groups treated with ZFP-SsdA and ZFP-SRE fusion proteins.

Claims

1. A fusion protein containing nucleic acid-binding protein (NABP) and single-stranded DNA deaminase toxin A (SsdA).

2. A fusion protein according to claim 1, wherein the SsdA has an amino acid sequence of sequence number 1 or 2.

3. A fusion protein according to claim 1, wherein the SsdA is an SsdA mutant.

4. In claim 3, the SsdA variant is a fusion protein in which an amino acid at a position capable of improving binding affinity with target DNA is replaced with another amino acid.

5. A fusion protein according to claim 4, wherein the position capable of enhancing binding affinity with the target DNA is V289, H333, Y335, P282, K392 in the amino acid sequence of SEQ ID NO: 1 or a position having a corresponding function thereto or a position having a corresponding function thereto.

6. A fusion protein according to claim 4, wherein the position capable of enhancing binding affinity with the target DNA is V32, H76, Y78, P25, K135 in the amino acid sequence of SEQ ID NO: 2 or a position having a corresponding function or a position having a corresponding function.

7. In claim 4, the other amino acid is any one selected from the group consisting of arginine (R), histidine (H), lysine (K), aspartic acid (D), glutamic acid (E), serine (S), threonine (T), asparagine (N), glutamine (Q), cysteine ​​(C), selenocysteine ​​(U), glycine (G), proline (P), alanine (A), valine (V), isoleucine (I), leucine (L), methionine (M), phenylalanine (F), tyrosine (Y), tryptophan (W), and all variants of the above amino acids, and is an amino acid excluding an amino acid that the wild-type protein has at a mutant position.

8. A fusion protein according to claim 4, wherein the other amino acid is a basic amino acid.

9. A fusion protein according to claim 8, wherein the basic amino acid is arginine (R), histidine (H), or lysine (K).

10. A fusion protein according to claim 3, wherein the SsdA variant comprises an amino acid sequence having at least 80% identity with the amino acid sequence of SEQ ID NO: 1, and comprises a mutation at at least one position selected from the group consisting of Y335, P282, K392, and positions having a corresponding function within the amino acid sequence.

11. A fusion protein according to claim 1, wherein the SsdA variant comprises an amino acid sequence having at least 80% identity with the amino acid sequence of SEQ ID NO: 2, and comprises a mutation at at least one position selected from the group consisting of P25, Y78, K135, and positions having a corresponding function within the amino acid sequence.

12. In claim 3, the SsdA mutant is a fusion protein in which an amino acid in the catalytic active site of the wild-type SsdA protein is replaced with another amino acid.

13. A fusion protein according to claim 12, wherein the SsdA mutant is inactivated.

14. A fusion protein according to claim 12, wherein the catalytically active site is T291, K365 or a position having a corresponding function among the amino acid sequences of SEQ ID NO.

1.

15. A fusion protein according to claim 12, wherein the other amino acid is any one selected from the group consisting of arginine (R), histidine (H), lysine (K), aspartic acid (D), glutamic acid (E), serine (S), threonine (T), asparagine (N), glutamine (Q), cysteine ​​(C), selenocysteine ​​(U), glycine (G), proline (P), alanine (A), valine (V), isoleucine (I), leucine (L), methionine (M), phenylalanine (F), tyrosine (Y), tryptophan (W), and all variants of the above amino acids, excluding the amino acid that the wild-type protein has at the variant position.

16. In claim 5, the SsdA variant is a fusion protein in which the amino acid at a position having a function corresponding thereto, such as T291, K365, or a position having a corresponding function, in the amino acid sequence of SEQ ID NO. 1 is additionally substituted with another amino acid.

17. In claim 3, the SsdA mutant is a fusion protein in which an amino acid in the deaminase domain of the wild-type SsdA protein is replaced with another amino acid.

18. A fusion protein according to claim 17, wherein the SsdA variant has a different amino acid substituted at any one or more of positions having a corresponding function, such as V32, H76, Y78, P25, K135, T34, K108, or a position having a corresponding function in the amino acid sequence of SEQ ID NO:

2.

19. A fusion protein according to claim 1, wherein the nucleic acid binding protein is any one selected from the group consisting of TALE (Transcription Activator-like Effector) protein, ZFP (Zinc Finger Protein), Cas9 (CRISPR-associated protein 9), Cpf1, dCpf1, dCas12, dCas13, Cas14, Cas12b, Cas12f, and variants thereof.

20. A fusion protein according to claim 1, wherein the fusion protein further comprises a DNA glycosylase inhibitor.

21. A fusion protein according to claim 20, wherein the DNA glycosylase inhibitor is a thymine glycosylase inhibitor, a uracil glycosylase inhibitor, an oxoguanine glycosylase inhibitor, or an alkylguanine DNA glycosylase inhibitor.

22. A fusion protein according to claim 21, wherein the DNA glycosylase inhibitor is linked to at least one N-terminus or C-terminus of the fusion protein.

23. A fusion protein according to claim 1, further comprising an organelle signal peptide domain.

24. A fusion protein according to claim 23, wherein the organelle signal peptide domain is a nuclear localization signal (NLS), a nuclear export signal (NES), a mitochondrial transfer signal (MTS), or a chloroplast transit peptide (CTP).

25. A fusion protein according to claim 1, wherein the nucleic acid binding protein and the SsdA (single-stranded DNA deaminase toxin A) variant are connected by a linker.

26. A polynucleotide encoding a fusion protein of any one of claims 1 to 25.

27. A vector comprising the polynucleotide of claim 26.

28. A composition for base correction comprising at least one selected from the group consisting of the fusion protein of any one of claims 1 to 25; the polynucleotide of claim 26; and the vector of claim 27.

29. A base correction composition according to claim 28, wherein the base correction composition substitutes at least one nucleotide among the nucleotide sequence of a target nucleic acid molecule.

30. A composition for base correction according to claim 29, wherein the target nucleic acid molecule is located in an organelle.

31. A composition for base correction according to claim 30, wherein the cell organelle is a nucleus, a mitochondria, or a chloroplast.

32. A method for correcting a target nucleic acid, comprising the step of contacting a target nucleic acid molecule with a base correction composition according to any one of claims 28 to 31.

33. A method for correcting a target nucleic acid according to claim 32, wherein the nucleic acid molecule is located in an organelle.

34. A base editing system comprising a fusion protein comprising a nucleic acid-binding protein (NABP); and a bacterial toxin, or a polynucleotide encoding the fusion protein; and a guide polynucleotide.

Citation Information

Patent Citations

  • Method of identifying base editing by cytosine deaminase in DNA

    KR102026421B1

  • Bacterial DNA cytosine deaminases for mapping DNA methylation sites

    WO2022212584A1

  • New tale protein scaffolds with improved on-target / off-target activity ratios

    WO2023094435A1

  • KR20240055677A