Bacterial toxin variant and use thereof

A base-editing system using a smaller bacterial toxin, SsdA mutant, addresses delivery and efficiency issues of conventional systems by enhancing nucleotide editing in cellular organelles.

WO2025192997A1PCT designated stage Publication Date: 2025-09-18THE ASAN FOUND +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/003339
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-14
Filing Date
2025-03-14
Publication Date
2025-09-18

AI Technical Summary

Technical Problem

Existing base-editing systems using APOBEC or AID proteins are challenging to deliver as vectors and inefficient for nucleotide editing, particularly in cellular organelles like mitochondria and chloroplasts due to their size and delivery difficulties.

Method used

Development of a base-editing system utilizing a smaller bacterial toxin, SsdA mutant, fused with a nucleic acid-binding protein (NABP) to enhance delivery and efficiency for nucleotide editing in cellular organelles.

Benefits of technology

The system enables efficient nucleotide editing in cellular organelles with improved vector delivery and editing efficiency, overcoming the limitations of conventional APOBEC or AID proteins.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025003339_18092025_PF_FP_ABST
    Figure KR2025003339_18092025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a bacterial toxin variant and use thereof. According to an aspect, a variant, a fusion protein, a polypeptide thereof, and a base editing system including same facilitate binding with Cas proteins because the polypeptide size is small, and compared to an existing base editing system using Ssda proteins, base editing efficacy is significantly improved, enabling more effective base editing, and base editing is also possible in cell organelles.
Need to check novelty before this filing date? Find Prior Art

Description

Bacterial toxin variants and uses thereof

[0001] It relates to bacterial toxin variants and uses thereof.

[0002] Genome editing is a technology that allows for the free modification of a living organism's genetic information. Genome editing technology can dramatically expand its scope of application by altering the genetic information of humans, animals, plants, and microorganisms.

[0003] Gene scissors, molecular tools designed to precisely cut desired genetic information, play a key role in genome editing technology. Like next-generation sequencing (NGSE) technology, which has advanced the field of genetic sequencing, CRISPR is becoming a key technology that expands the speed and scope of genetic information utilization and creates new industries.

[0004] ZFN (zinc-finger nuclease), TALEN (transcription activator-like effector nuclease), CRISPR (clustered regularly interspaced short palindromic repeat) system, CRISPR-associated protein variants that do not have nucleic acid degradation efficiency, and base editing technology using nucleotide deaminase proteins are helping to develop treatments for various diseases that have not been attempted so far.

[0005] In particular, fusion proteins linking Cas proteins to deaminase enable single nucleotide conversion in a targeted manner to achieve targeted nucleotide substitutions or base editing in the genome, correct point mutations causing genetic disorders, or introduce desired single nucleotide mutations into prokaryotes, humans, and other eukaryotic cells without generating DNA double-strand breaks (DSBs).

[0006] However, APOBEC or AID proteins, deaminases commonly used for cytosine base correction using Cas proteins, are relatively large, making them difficult to deliver as vectors for use in cell therapeutics. Furthermore, it is technically challenging to deliver complexes consisting of APOBEC or AID proteins, Cas9, and guide RNA (gRNA) to cellular organelles, or to simultaneously express these components within organelles. Therefore, using base-editing systems containing APOBEC or AID proteins to correct DNA sequences in cellular organelles, including mitochondria and chloroplasts, is technically challenging.

[0007] Therefore, there is a pressing need for a base-editing system that is easily delivered via vectors, exhibits superior nucleotide editing efficiency, and can also function effectively in cellular organelles. To achieve this, we developed a base-editing system utilizing a bacterial toxin smaller than conventional APOBEC or AID proteins. This system exhibits ease of vector delivery, superior nucleotide editing efficiency, and enables base editing in cellular organelles.

[0008] One aspect is to provide a single-stranded DNA deaminase toxin A (SsdA) mutant.

[0009] Another aspect is to provide a fusion protein comprising a nucleic acid-binding protein (NABP) and a single-stranded DNA deaminase toxin A (SsdA) variant.

[0010] Another aspect is to provide a polynucleotide encoding the fusion protein.

[0011] Another aspect is to provide a vector comprising the polynucleotide.

[0012] Another aspect provides a composition for base correction comprising a fusion protein comprising a nucleic acid-binding protein (NABP) and a single-stranded DNA deaminase toxin A (SsdA) variant; or a polynucleotide encoding the fusion protein.

[0013] Another aspect is to provide a base editing system comprising the fusion protein or a polynucleotide encoding the fusion protein; and a guide polynucleotide.

[0014] Another aspect provides a method of editing a nucleic acid comprising the step of contacting a nucleic acid molecule with the base editing system.

[0015] One aspect is to provide a single-stranded DNA deaminase toxin A (SsdA) mutant.

[0016] The term "SsdA (single-stranded DNA deaminase toxin A)" in the present specification may mean a wild-type of SsdA. The SsdA may be derived from a Pseudomonas sp. strain. The Pseudomonas sp. strains include Pseudomonas syringae, Pseudomonas congelans, Pseudomonas savastanoi, Pseudomonas viridiflava, Pseudomonas coronafaciens, Pseudomonas fluorescens, Pseudomonas sp. MPC6, and Pseudomonas sp. GL-R-26, or Pseudomonas sp. GL-RE-26, but is not limited thereto.

[0017] The N-terminus of the SsdA may have a PAAR domain, and the C-terminus of the SsdA may have a deaminase domain. Specifically, the SsdA may include a sequence of SEQ ID NO: 1, and the deaminase domain may include a sequence of SEQ ID NO: 6. The SsdA may induce deamination of single-stranded DNA, and the SsdA may be cytidine deaminase. The cytidine deaminase is an enzyme that converts cytidine into uridine.

[0018] In one specific example, the SsdA mutant may be a wild-type SsdA protein in which an amino acid at a position capable of improving binding affinity with target DNA is substituted with another amino acid.

[0019] In one specific example, the wild-type SsdA protein may have an amino acid sequence of SEQ ID NO: 1.

[0020] In one specific example, the position capable of enhancing binding affinity with the target DNA may be V289, H333, Y335, P282, K392 or a position having a corresponding function.

[0021] In one specific example, the SsdA variant may have another amino acid substituted at any one or more of V289, H333, Y335, P282, K392 or a position having a corresponding function.

[0022] In one specific example, the other amino acid may be any one selected from the group consisting of arginine (R), histidine (H), lysine (K), aspartic acid (D), glutamic acid (E), serine (S), threonine (T), asparagine (N), glutamine (Q), cysteine ​​(C), selenocysteine ​​(U), glycine (G), proline (P), alanine (A), valine (V), isoleucine (I), leucine (L), methionine (M), phenylalanine (F), tyrosine (Y), tryptophan (W), and all variants of the above amino acids, excluding the amino acid that the wild-type protein has at the variant position.

[0023] In one specific example, the other amino acid may be a basic amino acid, and the basic amino acid may be, but is not limited to, arginine (R), histidine (H), or lysine (K).

[0024] In one specific example, the SsdA variant comprises an amino acid sequence having at least 90% identity with the amino acid sequence of SEQ ID NO: 1, and may comprise a mutation at at least one position selected from the group consisting of Y335, P282, K392, and positions having corresponding functions within the amino acid sequence, but is not limited thereto.

[0025] For example, the SsdA variant may comprise an amino acid sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with the amino acid sequence of SEQ ID NO: 1, and may comprise a mutation at at least one position selected from the group consisting of Y335, P282, K392, and positions having corresponding functions thereto within the amino acid sequence.

[0026] In one specific example, the SsdA variant comprises an amino acid sequence having at least 90% identity with the amino acid sequence of SEQ ID NO: 7, and may include mutations at positions Y335, P282, and K392 within the amino acid sequence, but is not limited thereto.

[0027] For example, the SsdA variant may comprise an amino acid sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the amino acid sequence of SEQ ID NO: 7, and may comprise mutations at positions Y335, P282, and K392 within the amino acid sequence.

[0028] In one specific example, the SsdA variant may include a mutation at at least one position selected from the group consisting of Y335, P282, K392 and positions having corresponding functions within the amino acid sequence of SEQ ID NO: 1, but is not limited thereto.

[0029] In one specific example, the SsdA variant comprises an amino acid sequence having at least 80% identity with the amino acid sequence of SEQ ID NO: 6, and may comprise a mutation at at least one position selected from the group consisting of P25, Y78, K135, and positions having corresponding functions within the amino acid sequence, but is not limited thereto.

[0030] For example, the SsdA variant may comprise an amino acid sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the amino acid sequence of SEQ ID NO: 6, and may comprise a mutation at at least one position selected from the group consisting of P25, Y78, K135, and positions having corresponding functions within the amino acid sequence.

[0031] In one specific example, the SsdA variant comprises an amino acid sequence having at least 80% identity with the amino acid sequence of SEQ ID NO: 8, and may include mutations at positions P25, Y78, and K135 within the amino acid sequence, but is not limited thereto.

[0032] For example, the SsdA variant may comprise an amino acid sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the amino acid sequence of SEQ ID NO: 8, and may comprise mutations at positions P25, Y78, and K135 within the amino acid sequence.

[0033] In one specific example, the SsdA variant may include a mutation at at least one position selected from the group consisting of P25, Y78, K135 and positions having corresponding functions within the amino acid sequence of SEQ ID NO: 6, but is not limited thereto.

[0034] In one specific example, the SsdA variant may be one in which an amino acid in the catalytic active site of the wild-type SsdA protein is substituted with another amino acid. For example, the SsdA variant may be one in which an amino acid in the catalytic active site of the wild-type SsdA protein having the base sequence of SEQ ID NO: 1 is substituted with another amino acid.

[0035] In one specific example, the catalytic active site may be, but is not limited to, T291, K365 or a position having a corresponding function.

[0036] In one specific example, the SsdA variant may have another amino acid substituted at any one or more of T291, K365 or a position having a corresponding function. Specifically, the other amino acid may be any one selected from the group consisting of arginine (R), histidine (H), lysine (K), aspartic acid (D), glutamic acid (E), serine (S), threonine (T), asparagine (N), glutamine (Q), cysteine ​​(C), selenocysteine ​​(U), glycine (G), proline (P), alanine (A), valine (V), isoleucine (I), leucine (L), methionine (M), phenylalanine (F), tyrosine (Y), tryptophan (W) and all variants of the above amino acids, excluding the amino acid that the wild-type protein has at the variant position.

[0037] In one specific example, the SsdA variant may be a variant having another amino acid substituted at any one or more of V289, H333, Y335, P282, K392 or a position having a corresponding function, and further having an amino acid substituted at T291, K365 or a position having a corresponding function with another amino acid.

[0038] In one specific example, the SsdA variant may be one in which an amino acid within the deaminase domain of the wild-type SsdA protein is substituted with another amino acid. For example, the SsdA variant may be one in which an amino acid within the wild-type deaminase domain comprising the amino acid sequence of SEQ ID NO: 6 is substituted with another amino acid.

[0039] In one specific example, the SsdA mutant may be inactivated. For example, an SsdA mutant in which an amino acid in the deaminase domain of the wild-type SsdA protein is substituted with another amino acid may be inactivated.

[0040] In one specific example, the deaminase domain may have another amino acid substituted at any one or more of V32, H76, Y78, P25, K135, T34, K108 or positions having a corresponding function. Specifically, the other amino acid may be any one selected from the group consisting of arginine (R), histidine (H), lysine (K), aspartic acid (D), glutamic acid (E), serine (S), threonine (T), asparagine (N), glutamine (Q), cysteine ​​(C), selenocysteine ​​(U), glycine (G), proline (P), alanine (A), valine (V), isoleucine (I), leucine (L), methionine (M), phenylalanine (F), tyrosine (Y), tryptophan (W) and all variants of the above amino acids, excluding the amino acid that the wild-type protein has at the variant position.

[0041]

[0042] Another aspect provides a fusion protein comprising a nucleic acid-binding protein (NABP) and a single-stranded DNA deaminase toxin A (SsdA) variant.

[0043] The term nucleic acid-binding protein (NABP) used herein refers to a group of proteins that directly bind to DNA or RNA and play a key role in the replication, transcription, translation, repair, processing, degradation, and regulation of genetic information, and can recognize a specific base sequence or recognize and bind to various secondary structures (e.g., stem-loop, G-cytostine rich structure, etc.) including structural features of a double helix (dsDNA, dsRNA) or a single strand (ssDNA, ssRNA). For example, the nucleic acid-binding protein may bind to a nucleic acid in a sequence-specific manner.

[0044] Specifically, the nucleic acid binding protein may be any one selected from the group consisting of, but not limited to, TALE (Transcription Activator-like Effector) protein, ZFP (Zinc Finger Protein), Cas9 (CRISPR-associated protein 9), Cpf1, dCpf1, dCas12, dCas13, Cas14, Cas12b, Cas12f and variants thereof.

[0045] The term TALE (Transcription Activator-like Effector) protein used herein refers to a bacterial protein that recognizes and binds to a specific DNA sequence through repeated amino acid sequence modules (repeat domains). Specifically, the TALE protein is composed of a repeated form of 33 to 34 repeated amino acid sequences, and about 9 or more domains (RVD, Repeated Variant Diresidues) are repeated. Each domain can recognize one nucleotide, and can bind to a specific DNA sequence according to the 12th to 13th amino acid sequence (HD->Cytosine, NI->Adenine, NG->Thymine, NN->Guanine). The TALE protein recognizes a single DNA strand within the target site. The distance between target sites can be 12-14 nucleotides. The TALE domain refers to a protein domain that binds to a nucleotide in a sequence-specific manner by a combination of one or more TALE-Repeat. At least one TALE-Repeat, specifically, but not limited to, 1 to 30 TALE-Repeat(s). A TALE-Repeat is a site that recognizes a specific nucleotide sequence within a TALE domain.

[0046] In one specific example, a single module TALE may be coupled to the N-terminus of the SsdA (single-stranded DNA deaminase toxin A) variant. A single TALE module and the SsdA (single-stranded DNA deaminase toxin A) variant may be included in the NC direction. A dual module TALE may be included, wherein a first TALE is coupled to the N-terminus of the SsdA (single-stranded DNA deaminase toxin A) variant, and a second TALE may be included separately. The first TALE module and the SsdA (single-stranded DNA deaminase toxin A) variant may be included in the NC direction, and may have a structure of N'-TALE-SsdA (single-stranded DNA deaminase toxin A) variant-C' and N'-TALE-C'.

[0047] The term ZFP (Zinc Finger Protein) used herein refers to a group of proteins that interact with DNA, RNA, or proteins through a unique finger structure formed by binding to zinc (Zn²) ions. In particular, ZFPs have the property of recognizing and binding to specific DNA sequences, and are therefore utilized as gene targeting tools in base editing systems. Specifically, by fusing the DNA binding domain of ZFP with a single-stranded DNA deaminase toxin A (SsdA) mutant, a ZFP-based base editing system (ZFP-Base Editor, ZFP-BE) that induces a base substitution of cytosine (C) within a specific base sequence can be constructed.

[0048] In one specific example, the fusion protein comprises a TALE (Transcription Activator-like Effector) protein and a bacterial toxin, wherein the bacterial toxin may be a single-stranded DNA deaminase toxin A (SsdA) mutant.

[0049] In one specific example, the fusion protein comprises ZFP (Zinc Finger Protein) and a bacterial toxin, wherein the bacterial toxin may be a single-stranded DNA deaminase toxin A (SsdA) mutant.

[0050] In one specific example, the fusion protein comprises a Cas protein (CRISPR-associated protein) and a bacterial toxin, and the bacterial toxin may be a single-stranded DNA deaminase toxin A (SsdA) mutant.

[0051] As used herein, the term "Cas protein (CRISPR-associated protein)" may be a CRISPR-binding endonuclease. The Cas protein may cleave all or part of a specific target polynucleotide sequence.

[0052] The above Cas protein forms an active endonuclease, or nickase, when it forms a complex with two RNAs called CRISPR RNA (crRNA) and trans-activating crRNA (tracrRNA). Non-limiting examples of the above Cas proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, a homolog thereof or a homolog thereof. Includes modified versions.

[0053] The above Cas protein may be Cas9, Cpf1, or a Cas9 variant protein, and the Cas9 (or Cpf1) protein refers to an essential protein element in the CRISPR / Cas9 system, and information on the Cas9 (or Cpf1) gene and protein can be obtained from GenBank of the National Center for Biotechnology Information (NCBI), but is not limited thereto. CRISPR-associated genes encoding Cas (or Cpf1) proteins are known to have more than 40 different Cas (or Cpf1) protein families, and eight CRISPR subtypes (Ecoli, Ypest, Nmeni, Dvulg, Tneap, Hmari, Apern, and Mtube) can be defined according to specific combinations of cas genes and repeat structures. Therefore, each of the above CRISPR subtypes can form a repeat unit to form a polyribonucleotide-protein complex.

[0054] In one specific example, the Cas protein may be at least one selected from the group consisting of a Cas9 protein derived from Streptococcus pyogenes, a Cas9 protein derived from Campylobacter jejuni, a Cas9 protein derived from Streptococcus thermophiles, a Cas9 protein derived from Streptococcus aureus, a Cas9 protein derived from Neisseria meningitidis, and a Cpf1 protein, but is not limited thereto.

[0055] As used herein, the term "bacterial toxin" refers to a deaminase that removes an amine from an amino group in an amino acid or nucleotide. The bacterial toxin may be bound to the terminus of a nucleic acid-binding protein (NABP). For example, the bacterial toxin may be bound to the C-terminus, N-terminus, or both the C-terminus and N-terminus of a Cas protein. For example, the bacterial toxin may be bound to the terminus of a Cas protein. For example, the bacterial toxin may be bound to the C-terminus, N-terminus, or both the C-terminus and N-terminus of a TALE (Transcription Activator-like Effector) protein, a ZFP (Zinc Finger Protein), or a Cas protein.

[0056]

[0057] In one embodiment of the present invention, the nucleic acid-binding protein (NABP) and the bacterial toxin can be fused via a linker. The linker can be positioned at the C-terminus, N-terminus, or both the C-terminus and N-terminus of the nucleic acid-binding protein (NABP), and the bacterial toxin can be bound to the Cas protein via the linker. Suitable linker motifs and linker configurations include those described in the literature [Chen et al., Fusion protein linkers: property, design and functionality. Adv Drug Deliv Rev. 2013; 65(10):1357-69], the entire contents of which are incorporated herein by reference.

[0058] In one specific example, the bacterial toxin may be a single-stranded DNA deaminase toxin A (SsdA) mutant.

[0059] The term "SsdA (single-stranded DNA deaminase toxin A)" in the present specification may mean a wild-type of SsdA. The SsdA may be derived from a Pseudomonas sp. strain. The Pseudomonas sp. strains include Pseudomonas syringae, Pseudomonas congelans, Pseudomonas savastanoi, Pseudomonas viridiflava, Pseudomonas coronafaciens, Pseudomonas fluorescens, Pseudomonas sp. MPC6, and Pseudomonas sp. GL-R-26, or Pseudomonas sp. GL-RE-26, but is not limited thereto.

[0060] The N-terminus of the SsdA may have a PAAR domain, and the C-terminus of the SsdA may have a deaminase domain. Specifically, the SsdA may include a sequence of SEQ ID NO: 1, and the deaminase domain may include a sequence of SEQ ID NO: 6. The SsdA may induce deamination of single-stranded DNA, and the SsdA may be cytidine deaminase. The cytidine deaminase is an enzyme that converts cytidine into uridine.

[0061] The term "SsdA (single-stranded DNA deaminase toxin A) mutant" in this specification means a mutant in which a specific amino acid sequence is substituted with another amino acid in wild-type SsdA.

[0062] In one specific example, the wild-type SsdA protein may have a base sequence of SEQ ID NO: 1.

[0063] In one specific example, the SsdA variant may be a wild-type SsdA protein in which an amino acid at a position capable of enhancing binding affinity to target DNA is replaced with another amino acid. Specifically, the position capable of enhancing binding affinity to target DNA may be V289, H333, Y335, P282, K392, or a position having a corresponding function in the amino acid sequence of SEQ ID NO: 1.

[0064] In one specific example, the SsdA variant may have another amino acid substituted at any one or more of V289, H333, Y335, P282, K392 or a position having a corresponding function.

[0065] In one specific example, the other amino acid substituted at a position capable of enhancing binding affinity with the target DNA may be any one selected from the group consisting of arginine (R), histidine (H), lysine (K), aspartic acid (D), glutamic acid (E), serine (S), threonine (T), asparagine (N), glutamine (Q), cysteine ​​(C), selenocysteine ​​(U), glycine (G), proline (P), alanine (A), valine (V), isoleucine (I), leucine (L), methionine (M), phenylalanine (F), tyrosine (Y), tryptophan (W), and all variants of the above amino acids, excluding the amino acid that the wild-type protein has at the variant position. Specifically, the other amino acid substituted may be a basic amino acid, and more specifically, the other amino acid may be, but is not limited to, arginine (R), histidine (H), or lysine (K).

[0066] In one specific example, the SsdA mutant may have an amino acid in the catalytic active site of the wild-type SsdA protein substituted with another amino acid. Specifically, the wild-type SsdA protein may have the base sequence of SEQ ID NO: 1.

[0067] In one specific example, the catalytic active site may be T291, K365 or a position having a corresponding function.

[0068] In one specific example, the SsdA variant may have another amino acid substituted at any one or more of T291, K365 or positions having a corresponding function. Specifically, the other amino acid may be any one selected from the group consisting of arginine (R), histidine (H), lysine (K), aspartic acid (D), glutamic acid (E), serine (S), threonine (T), asparagine (N), glutamine (Q), cysteine ​​(C), selenocysteine ​​(U), glycine (G), proline (P), alanine (A), valine (V), isoleucine (I), leucine (L), methionine (M), phenylalanine (F), tyrosine (Y), tryptophan (W) and all variants of the above amino acids, excluding the amino acid that the wild-type protein has at the variant position.

[0069] In one specific example, the SsdA variant may be an SsdA variant having another amino acid substituted at any one or more of V289, H333, Y335, P282, K392 or a position having a corresponding function, in which the amino acid at T291, K365 or a position having a corresponding function is further substituted with another amino acid.

[0070] In one specific example, the SsdA mutant may have an amino acid in the deaminase domain of the wild-type SsdA protein substituted with another amino acid. Specifically, the wild-type deaminase domain may include the amino acid sequence of SEQ ID NO: 6. More specifically, the amino acid sequence of SEQ ID NO: 6 of the SsdA may include a catalytic active site.

[0071] In one specific example, the SsdA variant may be inactivated.

[0072] In one specific example, the SsdA variant may have another amino acid substituted at any one or more of V32, H76, Y78, P25, K135, T34, K108 or a position having a corresponding function within the deaminase domain. Specifically, the substituted other amino acid may be any one selected from the group consisting of arginine (R), histidine (H), lysine (K), aspartic acid (D), glutamic acid (E), serine (S), threonine (T), asparagine (N), glutamine (Q), cysteine ​​(C), selenocysteine ​​(U), glycine (G), proline (P), alanine (A), valine (V), isoleucine (I), leucine (L), methionine (M), phenylalanine (F), tyrosine (Y), tryptophan (W), and all variants of the above amino acids, excluding the amino acid that the wild-type protein has at the mutant position.

[0073] In one specific example, the fusion protein may further comprise a DNA glycosylase inhibitor. Specifically, the DNA glycosylase inhibitor may be, but is not limited to, a thymine glycosylase inhibitor, a uracil glycosylase inhibitor, an oxoguanine glycosylase inhibitor, or an alkylguanine DNA glycosylase inhibitor.

[0074] The above uracil DNA glycosylase inhibitor may be, but is not limited to, a uracil DNA glycosylase inhibitor derived from Bacillus subtilis bacteriophage, PBS1, a uracil DNA glycosylase inhibitor derived from Bacillus subtilis bacteriophage, or PBS2.

[0075] In one specific example, the DNA glycosylase inhibitor may be linked to at least one N-terminus or C-terminus of the fusion protein.

[0076] In one specific example, the fusion protein may further comprise an organelle signal peptide domain. Specifically, the organelle signal peptide domain may be, but is not limited to, a nuclear localization signal (NLS), a nuclear export signal (NES), a mitochondrial transfer signal (MTS), or a chloroplast transit peptide (CTP). For example, the fusion protein may optionally further comprise a nuclear localization signal (NLS) for base correction of nuclear DNA. In another example, the fusion protein may optionally further comprise a mitochondrial transfer signal (MTS) or a nuclear export signal (NES) for base correction of mitochondrial DNA. As another example, the fusion protein may be for base correction of chloroplast DNA and may optionally additionally include a chloroplast transit peptide (CTP).

[0077]

[0078] Another aspect provides a polynucleotide encoding the fusion protein.

[0079]

[0080] Another aspect provides a vector comprising the polynucleotide.

[0081] The term "vector" as used herein may refer to a nucleic acid molecule capable of transporting another nucleic acid to which it is linked. A vector may include a nucleic acid molecule that is single-stranded, double-stranded, or partially double-stranded; a nucleic acid molecule comprising one or more free ends, a nucleic acid molecule without free ends (e.g., circular); a nucleic acid molecule comprising DNA, RNA, or both; and various other polynucleotides known in the art. One type of vector is a "plasmid," which may refer to a circular double-stranded DNA loop into which additional DNA segments may be inserted, for example, by standard molecular cloning techniques. Another type of vector is a viral vector, in which viral-derived DNA or RNA sequences may be present for packaging into a virus (e.g., a retrovirus, a replication-defective retrovirus, an adenovirus, a replication-defective adenovirus, and an adeno-associated virus). A recombinant expression vector may comprise a nucleic acid of the invention in a form suitable for expression of the nucleic acid in a host cell, which may mean that the recombinant expression vector comprises one or more regulatory elements, which one or more regulatory elements may be selected based on the host cell to be used for expression, and which may be operably linked to the nucleic acid sequence to be expressed. Within a recombinant expression vector, “operably linked” may mean that the nucleotide sequence of interest is linked to the regulatory element(s) in a manner that permits expression of the nucleotide sequence (e.g., in an in vitro transcription / translation system or in a host cell if the vector has been introduced into the host cell).

[0082] The "regulatory elements" may include promoters, enhancers, internal ribosomal entry sites (IRES), and other expression control elements (e.g., transcription termination signals such as polyadenylation signals and poly-U sequences). Regulatory elements are described, for example, in Goeddel, GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY 185, Academic Press, San Diego, Calif. (1990). Regulatory elements may include those that direct constitutive expression of a nucleotide sequence in many types of host cells and those that direct expression of a nucleotide sequence only in certain host cells (e.g., tissue-specific regulatory sequences). In some embodiments, the vector can comprise one or more pol III promoters (e.g., one, two, three, four, five or more pol III promoters), one or more pol II promoters (e.g., one, two, three, four, five or more pol II promoters), one or more pol I promoters (e.g., one, two, three, four, five or more pol I promoters), or a combination thereof. Examples of pol III promoters can include, but are not limited to, U6 and H1 promoters. Examples of pol II promoters can include, but are not limited to, the retrovirus Rous sarcoma virus (RSV) LTR promoter (optionally with an RSV enhancer), the cytomegalovirus (CMV) promoter (optionally with a CMV enhancer), the SV40 promoter, the dihydrofolate reductase promoter, the β-actin promoter, the phosphoglycerol kinase (PGK) promoter, and the EF1α promoter.For example, vectors may include lentiviruses and adeno-associated viruses (AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, or AAV9), and the type of vector may also be selected for targeting specific types of cells.

[0083] Additionally, multiple nucleic acid molecules within the vector system may be located on the same or different vectors.

[0084]

[0085] Another aspect provides a composition for base correction comprising a fusion protein comprising a nucleic acid-binding protein (NABP) and a single-stranded DNA deaminase toxin A (SsdA) variant; or a polynucleotide encoding the fusion protein.

[0086] The above fusion protein, the polynucleotide encoding the same, the nucleic acid-binding protein (NABP) and the single-stranded DNA deaminase toxin A (SsdA) variants are as described above.

[0087]

[0088] Another aspect provides a base correction system comprising a fusion protein comprising a nucleic acid-binding protein (NABP) and a bacterial toxin or a polynucleotide encoding the fusion protein; and a guide polynucleotide, wherein the bacterial toxin is a single-stranded DNA deaminase toxin A (SsdA) variant. The fusion protein, the polynucleotide encoding the fusion protein, the nucleic acid-binding protein (NABP) and the single-stranded DNA deaminase toxin A (SsdA) variant are as described above.

[0089] As used herein, the term "guide polynucleotide" refers to a short synthetic RNA molecule used in genome editing based on the CRISPR system. A "guide polynucleotide" comprises a spacer sequence that binds to target DNA, a scaffold sequence for Cas binding, and an endonuclease. The guide nucleotide comprises a guide RNA, gRNA, or guide RNA. Furthermore, the guide polynucleotide comprises a crRNA. The crRNA is a sequence within the crRNA that binds and / or interacts with a tracrRNA and / or an effector protein. The crRNA may be a wild-type crRNA or an engineered crRNA. The crRNA may comprise a direct repeat sequence and a spacer, and the direct repeat sequence may be located at the 5' end of the spacer. Furthermore, the crRNA may be located at the 3' end of the tracrRNA. The guide polynucleotide comprises a tracrRNA. The above tracrRNA scaffold sequence is the entire or a portion of the tracrRNA sequence that binds and / or interacts with the crRNA and / or effector protein.

[0090] The tracrRNA may be a wild-type tracrRNA or an engineered tracrRNA. The engineered crRNA or tracrRNA may be a sequence in which a part of the nucleotide sequence of the wild-type crRNA or tracrRNA is artificially modified (substituted, deleted, or inserted), or modified to be shorter than the wild-type crRNA or tracrRNA sequence. The guide polynucleotide may further include a linker. The linker is a sequence that serves to connect the tracrRNA and crRNA. The linker may be a sequence of 1 to 30 nucleotides. In one embodiment, the linker may be a sequence of 1 to 5, 5 to 10, 10 to 15, 15 to 20, 20 to 25, or 25 to 30 nucleotides. For example, the linker may be a 5'-GAAA-3' sequence, but is not limited thereto. The above guide polynucleotide may be an engineered guide RNA in which one or more nucleotide sequences are deleted, substituted, or added to a wild-type guide polynucleotide.

[0091] In one specific example, the guide polynucleotide may be at least one selected from the group consisting of SEQ ID NOs: 44 to 103.

[0092] The above guide polynucleotide may comprise a targeting sequence and / or an activating sequence.

[0093] As used herein, the term "targeting sequence" may refer to a polynucleotide comprising DNA, or a mixture of DNA and RNA, complementary to a sequence within a target nucleic acid. In certain embodiments, the targeting sequence may also comprise other nucleic acids, nucleic acid analogs, or combinations thereof. In certain embodiments, the targeting sequence may consist solely of DNA, as such a construct is less likely to be degraded within the host cell. In some embodiments, such a construct may increase target sequence recognition specificity and / or reduce the occurrence of off-target binding / hybridization. The targeting sequence may comprise a guide sequence or a spacer sequence. The length of the domain of the above targeting sequence may be at least 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29 or 30 nucleotides in length.

[0094] As used herein, the term "activation sequence" may refer to a portion of a polynucleotide comprising RNA or DNA, or a mixture of DNA and RNA, that can interact, associate, or bind to a Cas protein. In one embodiment, the activation domain may also comprise another nucleic acid, or a nucleic acid analog, or a combination thereof. In one embodiment, the activation sequence may be adjacent to or linked to a target sequence. In one embodiment, the activation domain may be downstream of the targeting domain. In one embodiment, the activation domain may be upstream of the targeting domain. The activation sequence may comprise a direct repeat sequence, a CRISPR RNA (crRNA), and / or a trans-activating RNA (tracrRNA).

[0095] In one specific embodiment, the guide polynucleotide, guide RNA, mature crRNA, and immature crRNA may comprise or consist of a direct repeat sequence and a guide sequence or spacer sequence. For example, the guide RNA or mature crRNA may comprise or consist of a direct repeat sequence linked to a guide sequence or spacer sequence. Specifically, the direct repeat sequence may be located upstream (i.e., 5') from the guide sequence or spacer sequence.

[0096] In one specific example, the guide polynucleotide may comprise a crRNA and a tracrRNA.

[0097] In one specific example, the guide polynucleotide may be a dual guide RNA or a single-chain guide RNA (sgRNA).

[0098] The above base correction system can substitute at least one nucleotide in the nucleotide sequence of a target nucleic acid molecule. Specifically, the nucleotide substitution may be from cytosine to uracil.

[0099] The nucleic acid may be RNA or DNA.

[0100] According to one specific example, the base correction system has an efficiency of forming a substitution of a sequence of nucleotides of 1% to 100%, for example, 1% to 100%, 1% to 99%, 1% to 98%, 1% to 97%, 1% to 96%, 1% to 95%, 1% to 94%, 1% to 93%, 1% to 92%, 1% to 91%, 1% to 89%, 1% to 88%, 1% to 87%, 1% to 86%, 1% to 85%, 1% to 84%, 1% to 83%, 1% to 82%, 1% to 81%, 1% to 80%, 1% to 79%, 1% to 78%, 1% to 77%, 1% to 76%, 1% to 75%, 1% to 74%, 1% to 73%, 1% to 72%, 1% to 71%, 1% to 70%, 1% to 68%, 1% to 66%, 1% to 64%, 1% to 62%, 1% to 60%, 1% to 58%, 1% to 56%, 1% to 54%, 1% to 52%, 1% to 50%, 1% to 48%, 1% to 46%, 1% to 44%, 1% to 42%, 1% to 40%, 1% to 38%, 1% to 36%, 1% to 34%, 1% to 32%, 1% to It can be 30%, 1% to 28%, 1% to 26%, 1% to 24%, 1% to 22%, 1% to 20%, 1% to 18%, 1% to 16%, 1% to 14%, 1% to 12%, 1% to 10%, 1% to 8%, 1% to 6%, 1% to 4%, 1% to 2%.

[0101] The efficiency of forming a substitution in the sequence of the above nucleotides, i.e., the base correction efficiency, may vary depending on the target position.

[0102] Another aspect provides a method of editing a nucleic acid comprising the step of contacting a nucleic acid molecule with the base editing system.

[0103] The nucleic acid and base correction system are as described above.

[0104] In one specific example, the editing may be a substitution of at least one nucleotide in the nucleotide sequence of the nucleic acid molecule.

[0105] According to one specific example, the method of editing a nucleic acid has an efficiency of forming a substitution of a nucleotide sequence of from 1% to 100%, for example, from 1% to 100%, from 1% to 99%, from 1% to 98%, from 1% to 97%, from 1% to 96%, from 1% to 95%, from 1% to 94%, from 1% to 93%, from 1% to 92%, from 1% to 91%, from 1% to 89%, from 1% to 88%, from 1% to 87%, from 1% to 86%, from 1% to 85%, from 1% to 84%, from 1% to 83%, from 1% to 82%, from 1% to 81%, from 1% to 80%, from 1% to 79%, from 1% to 78%, from 1% to 77%, 1% to 76%, 1% to 75%, 1% to 74%, 1% to 73%, 1% to 72%, 1% to 71%, 1% to 70%, 1% to 68%, 1% to 66%, 1% to 64%, 1% to 62%, 1% to 60%, 1% to 58%, 1% to 56%, 1% to 54%, 1% to 52%, 1% to 50%, 1% to 48%, 1% to 46%, 1% to 44%, 1% to 42%, 1% to 40%, 1% to 38%, 1% to 36%, 1% to 34%, 1% to 32%, 1% to It can be 30%, 1% to 28%, 1% to 26%, 1% to 24%, 1% to 22%, 1% to 20%, 1% to 18%, 1% to 16%, 1% to 14%, 1% to 12%, 1% to 10%, 1% to 8%, 1% to 6%, 1% to 4%, 1% to 2%.

[0106] The efficiency of forming a substitution in the sequence of the above nucleotides, i.e., the base correction efficiency, may vary depending on the target position.

[0107] The term "product purity" as used herein refers to a measure of the performance of a base editing technology to determine whether the desired base correction has been accurately performed. Specifically, it refers to the rate at which unintended base corrections occur in a base editing technology. For example, when a cytosine base editing technology is used to induce base correction from cytosine (C) to thymine (T) within a target position, this refers to the rate at which unintended changes from cytosine (C) to guanine (G) and cytosine (C) to adenine (A) occur. It can be expressed as the ratio of thymine, adenine, and guanine at each target cytosine base.

[0108] The term "off-target effect" as used herein refers to the occurrence of mutations at undesired locations within a cell when using base editing technology. This off-target effect occurs not only in DNA-based genomes but also in RNA. Specifically, off-target effects that can occur in the genome can be broadly divided into guide RNA-dependent off-target effects based on target sequence similarity and guide RNA-independent off-target effects that occur regardless of the target sequence.

[0109] The above guide RNA-dependent off-target effect refers to the off-target effect caused by a mismatch between the guide RNA and the target sequence. In order to measure the above guide RNA-dependent off-target effect, one or two or three or four mutations are inserted into the guide RNA at the same target position and the base correction efficiency at the target position is confirmed, or the off-target effect that appears in a sequence similar to the target base sequence throughout the genome is confirmed and a candidate position to analyze the off-target effect is selected using a prediction model through computer analysis, or the method of finding a position where a mutation actually appears in genomic DNA through an analysis experiment such as Digenome-seq can be used for measurement, but is not limited thereto.

[0110] The above guide RNA-independent off-target effect refers to the deaminase causing mutations in DNA single-stranded molecules naturally exposed within cells. Specifically, by artificially exposing DNA single-stranded molecules within cells where deaminase can function, mutations can be observed at these sites.

[0111] The term "R-loop assay" in this specification refers to a method in which, when a dSaCas9 protein that inhibits the activity of SaCas9 and a guide RNA that binds to it are delivered to a cell, dSaCas9 and the guide RNA bind to the dSaCas9 at the target site in the cell, thereby forming an R-loop, unwinding the DNA double helix and exposing a single helix, and at this time, a base-editing protein is delivered to the cell together to randomly cause mutations in the exposed DNA single helix at the target site of dSaCas9.

[0112] The term "whole genome sequencing (WGS)" used herein refers to a method of analyzing the entire genome sequence by analyzing and connecting DNA fragments. Generally, when WGS is performed, the DNA fragments are continuously connected as a result, but when the SsdA mutant is processed, it can be confirmed that the sequences of the sequenced DNA fragments appear in a row at the same location.

[0113] The above fusion protein, the polynucleotide encoding the same, and the SsdA (single-stranded DNA deaminase toxin A) mutant are as described above.

[0114] According to one aspect of the present invention, a variant, a fusion protein, a polypeptide thereof, and a base correction system comprising the same enable effective base correction. Furthermore, the present invention facilitates vector production due to the small size of the polypeptide, exhibits higher performance than existing BE4max technology at specific target sites, and enables base correction even in cellular organelles.

[0115] Figure 1 is a graph showing the base correction efficiency of a CRISPR-Cas system comprising a guide polynucleotide and a fusion protein of a SsdA mutant with amino acid substitutions at positions V289, H333, and Y335 of the Cas protein and the PAAR domain-containing protein (SEQ ID NO: 1) at the target site of the intracellular target site (HEK2, HEK3, and RNF2).

[0116] Figure 2 is a graph showing the base correction efficiency of a CRISPR-Cas system comprising a Cas protein, an SsdA mutant with amino acid substitutions at positions T291 and K365 of SEQ ID NO: 1, and a guide polynucleotide at the target site (SEQ ID NO: 3 to 5) of an intracellular target site (HEK2, HEK3, and RNF2).

[0117] Figure 3 is a graph showing the base correction efficiency of a CRISPR-Cas system comprising a Cas protein, an SsdA mutant with amino acid substitutions at positions P282, Y335, and K392, and a guide polynucleotide at the target site (SEQ ID NOs: 3 to 5) of an intracellular target site (HEK2, HEK3, and RNF2).

[0118] Figure 4 is a graph analyzing the accuracy of the product of the cytosine base correction technology using SsdA (wild-type) and an SsdA-SRE mutant, which is a specific example of Example 7, at the target site (SEQ ID NO: 3 to 5) of the intracellular target site (HEK2, HEK3, and RNF2).

[0119] Figure 5 is a graph showing the cytotoxicity of each of the Cas9(D10A)-SsdA fusion protein, Cas9(D10A)-SsdA-UGI fusion protein, dCas9-SsdA fusion protein, and dCas9-SsdA-UGI fusion protein.

[0120] Figure 6a is a graph showing the base correction efficiency of SsdA-SRE-UGI-C2, SsdA-SRE-UGI-N2, and BE4max, and Figure 6b is a graph comparing the base correction efficiency of fusion proteins (SsCBE-UGI-C2, SsCBE-UGI-N2) that combine UGI with SsdA-SRE mutants and BE4max at about 30 intracellular target sites (SEQ ID NOs: 9 to 39).

[0121] Figure 7 is a graph showing the indel formation frequency of SsdA-SRE-UGI-C2, SsdA-SRE-UGI-N2, and BE4max.

[0122] Figure 8 is a graph showing the base editing range (editing window) of SsdA-SRE-UGI-C2, SsdA-SRE-UGI-N2, and BE4max for cytosine base correction at the intracellular target site shown in Table 4.

[0123] Figure 9 is a graph showing the cytosine base correction efficiency of SsdA mutants in cell lines other than HEK293T cells.

[0124] Figure 10 is a schematic diagram showing the binding structure of cjBE4max and cjCas9(D8A)-SsdA-SRE (cjCBE-SRE).

[0125] Figure 11 is a graph measuring the base correction efficiency at intracellular target sites (EPAS-E2-1, HIF-E9-4, HPD-1 HPD-2) using cjCas9(D8A)-SsdA-SRE fusion protein and cjBE4max.

[0126] Figure 12a is a graph showing the base correction efficiency according to guide RNA mutation at the intracellular target site HEK3 using SsdA mutant (SsdA-SRE) or BE4max, and Figure 12b is a graph showing the base correction efficiency according to guide RNA mutation at the intracellular target site RNF2 using SsdA mutant (SsdA-SRE) or BE4max.

[0127] Figure 13 is a schematic diagram showing the Digenome-seq experimental process for analyzing the off-target effect of SsdA mutants.

[0128] Figure 14a shows the results of a Digenome-seq experiment using an SsdA mutant fusion protein and a HEK2 target guide RNA, and Figure 14b is a graph showing the base mutation efficiency at each position in cells transfected with an SsdA mutant and a HEK2 target guide RNA to confirm whether the off-target effect actually occurs at each off-target effect position.

[0129] Figure 15 is a graph showing the results of an experiment analyzing the guide RNA-independent off-target effect.

[0130] Figure 16a shows the results of analyzing base mutations in transcripts in cells delivered with SsdA mutants or BE4max and RNF2 targeting guide RNA, and Figure 16b shows the results of analyzing base mutations in transcripts in cells delivered with SsdA mutants or BE4max and HEK2 targeting guide RNA.

[0131] Figure 17a is a schematic diagram of mitochondrial ND1 base editing using a TALE-SRE expression plasmid, and Figure 17b is a graph showing the base editing effect after delivering a TALE-SRE expression plasmid in HEK293T cells.

[0132] Figure 18 is a graph showing the results of introducing point mutations into the mitochondrial ND1 gene using a ZFP-SRE expression plasmid.

[0133] The present invention will be described in more detail through the following examples. However, these examples are provided for illustrative purposes only and the scope of the present invention is not limited to these examples.

[0134] Example 1. Plasmid cloning

[0135] gBlock double-stranded DNA fragments encoding His6-SsdA and SsdAI (Integrated DNA Technologies) and bacterial expression vector pET-28b(+) DNA (Novagen) were treated with XbaI and XhoI restriction enzymes (New England Biolabs) at 37°C for 3 h, respectively. The linearized pET-28b vector was then purified using agarose gel extraction (Geneall) and ligated to the gBlock double-stranded DNA fragments using Quick ligase (New England Biolabs).

[0136] Then, using pCMV plasmid DNA, the coding sequences of Cas9 and UGI were obtained through PCR amplification, and the SsdA sequence was amplified from gBlock using Gibson Assembly Master Mix (New England Biolabs) and subcloned into pCMV plasmid.

[0137] The sequences of the above PAAR domain-containing protein and SsdAI are shown in Table 1.

[0138] SsdA protein sequence SEQ ID NO: PAAR-domain containing proteinMSAAARVNDPIEHTGSLTGLLAGLAIGAIGAALVVGTGGLAAVAIVGASAATGAGVGQLIGSLSCCNHQTGQIVSGSSNVYINGEPAARAHADQAKCDEHTSRPQVIAQGSSNVYINGHPAARVGDRTACDAKIVVGSSNVFIGGGT ETTDPINPEVPELLERSILLVGLASAVVLASPVIVIAGLVGGIAGGTVGSMGGAQLFGEGTDGQKLMAFGGALLGGGLGAKGGKWFDTRYDIKVQGVGSNLGNLKITPKGAAKVSNIAESEAALGRASQARADLPQSKELKVKTVSSNDKKTLS GWGNKKPEGYERISAEQVKAKSEEIGHEVKSHPYDRDYKGQYFSSHAEKQMSIASPNHPLGVSKPMCTDCQGYFSQLAKYSKVEQTVADPKAIRIFKTDGSVETIMRSE1SsdAIMNNKSKVLIEKLLLEVAKSPEGELILPLRKLLWNTITED ETAAKKKAILTALDVMCVRQGVNFWIKKFGDNEPLNYILNIALETAEGKFDESKALGLRDEFYVSIVEDQEYEVEEYPAMFVGHAAANTIARAVDDFQFEPYDHRVDRDLDPEGFESSYLVASAFAGGLSEDGDPKLRRAFWEWYLSIAVPQVV2

[0139]

[0140] Example 2. Purification of SsdA

[0141] The SsdA protein cloned in Example 1 was purified using Escherichia coli BL21. More specifically, pET-28b-His6-SsdA-SsdAI was introduced into E. coli BL21 using 0.5 mM IPTG, and the His6-SsdA and SsdAI complex proteins were purified using Ni-NTA agarose beads (Qiagen). To separate His6-SsdA from SsdAI, the His6-SsdA and SsdAI complex proteins were denatured with a denaturing buffer (8 M urea, 50 mM Tris-HCl pH 7.5, 500 mM NaCl, and 1 mM DTT) and then incubated at 4°C for 16 hours. The suspension buffer containing the denatured protein complex was mixed with Ni-NTA agarose beads (Qiagen) and loaded onto a gravity-flow column to remove unbound SsdAI. Subsequently, denaturing buffers containing decreasing concentrations of urea (6 M, 4 M, 2 M, 1 M, and 0 M) were treated for refolding of SsdA. The refolded protein bound to the Ni-NTA agarose beads was eluted with an elution buffer containing 300 mM imidazole, and the eluted protein was dialyzed against 20 mM Tris-HCl pH 7.5, 200 mM NaCl, 1 mM DTT, and 40% (w / v) glycerol and concentrated using an Amicon Ultra-15 Centrifugal Filter Unit (Millipore). The concentration of the His6-SsdA protein was analyzed by SDS-PAGE.

[0142]

[0143] Example 3. Cell culture and transfection

[0144] HEK293T cells (ATCC CRL-11268) were maintained in Dulbecco's modified Eagle's medium (DMEM) containing 10% fetal bovine serum (FBS) and 1% penicillin / streptomycin (Welgene), and the HEK293T cells were seeded at a density of 6 x 10 per well. 4 Cells were seeded in TC-treated 48-well plates (Corning Life Sciences). Twenty-four hours after seeding, transfection was performed at approximately 60% cell confluency using 500 ng of plasmids (250 ng of Cas9-SsdA expression plasmid and 250 ng of gRNA expression plasmid) and 1.5 μL of Lipofectamine 2000 (Thermo Fisher Scientific). The transfected cells were incubated at 37°C for 3 days, and genomic DNA was prepared by directly lysing the cells using lysis buffer (10 mM Tris-HCl, pH 7.5, 0.05% SDS, 100 mg / mL proteinase K; QIAGEN). The cell lysate was incubated at 56°C for 30 minutes and then further incubated at 99°C for 15 minutes to inactivate Proteinase K.

[0145]

[0146] Example 4. Target deep sequencing and data analysis

[0147] The target region was amplified by PCR (a total of three times) and sequenced using the Illumina MiniSeq or iSeq 100 sequencing system.

[0148] More specifically, 3 mL of cell lysate or 1 mL of isolated genomic DNA was applied to the first PCR, and then 1 mL of the first PCR product was used for the second PCR. Illumina TruSeq HT dual-index adapter sequences were attached as index PCR primer pairs using 1 mL of the second PCR product. The size of the PCR amplicons was confirmed on a 2% agarose gel, and the amplicons were used for Illumina MiniSeq or iSeq 100 sequencing system. Targeted deep sequencing analysis was performed using MAUND (https: / github.com / ibscge / maund), and all results were confirmed with Cas-Analyzer (http: / www.rgenome.net / cas-analyzer / ).

[0149]

[0150] Example 5. Preparation of SsdA mutants with amino acid substitutions at positions that can enhance binding affinity with target DNA and verification of their base correction efficiency.

[0151] A mutant was prepared by substituting amino acids at positions (V289, H333, Y335) that can enhance binding affinity of the SsdA protein (PAAR domain-containing protein, SEQ ID NO: 1) with target DNA by the method of Example 1. In addition, in order to confirm whether the SsdA of the CRISPR-Cas system effectively causes cytosine deamination in eukaryotic cells, a plasmid expressing a protein combined with UGI and Cas9 was constructed and transfected into HEK293 cells by the method of Example 3. Then, the base sequence of the target position (HEK2, HEK3, RNF2 site) was amplified by PCR from the cells transfected by the method of Example 4, and the introduction of mutations and the formation of indels were analyzed using next-generation sequencing (NGS), and the results are shown in Fig. 1.

[0152] Table 2 shows the base sequences of the target sites (HEK2, HEK3, RNF2 sites).

[0153] Target DNA target site sequence number HEK2GAACACAAAGCATAGACTGC3HEK3GGCCCAGACTGAGCACGTGA4RNF2GTCATCTTAGTCATTACCTG5

[0154] Figure 1 is a graph showing the base editing efficiency of a CRISPR-Cas system comprising a guide polynucleotide and a fusion protein of a Cas protein and an SsdA mutant with amino acid substitutions at positions V289, H333, and Y335 of SEQ ID NO: 1 (PAAR domain-containing protein) at the target site of the intracellular target site (HEK2, HEK3, and RNF2). Specifically, the three alphabets of the items on the X-axis represent the amino acids at positions V289, H333, and Y335 of SEQ ID NO: 1, in that order. For example, VHR means a mutant in which no mutation occurs at positions V289 and H333 of SEQ ID NO: 1, but a mutation from Y to R occurs at position Y335. The SsdA wild-type is indicated as VHY. Meanwhile, the shading indicated on the X-axis is a heat-map showing the efficiency of unwanted indels.

[0155] As shown in Fig. 1, it was confirmed that the substitution efficiency of the SsdA fusion protein in which the amino acids at positions V289 and Y335 of sequence number 1 were substituted was significantly improved in all of HEK2, HEK3, and RNF2 compared to the wild-type SsdA fusion protein.

[0156] This means that base correction efficiency can be improved by introducing mutations at positions V289 and Y335 of wild-type SsdA (PAAR domain-containing protein, sequence number 1).

[0157]

[0158] Example 6 Preparation of SsdA mutant with amino acid substitution in catalytic active site and verification of base correction efficiency thereof

[0159] A mutant having a substituted amino acid in the catalytically active site of a PAAR domain-containing protein (SEQ ID NO: 1) was prepared using the method of Example 5. The catalytically active site corresponds to a portion 260 to 410 of the PAAR domain containing protein (SEQ ID NO: 1), and is used interchangeably as the SsdA toxin domain or deaminase domain in the present specification, and the catalytically active site includes the amino acid sequence of SEQ ID NO: 6.

[0160] The base correction efficiency of a mutant in which amino acids were substituted at positions T291 and K365 of the PAAR domain-containing protein (SEQ ID NO: 1) using the method of Example 5 was measured, and is shown in Fig. 2.

[0161] Figure 2 is a graph showing the base correction efficiency of a CRISPR-Cas system comprising a Cas protein, an SsdA mutant with amino acid substitutions at positions T291 and K365 of SEQ ID NO: 1, and a guide polynucleotide at the target site (SEQ ID NO: 3 to 5) of the intracellular target site (HEK2, HEK3, and RNF2). Specifically, the two alphabets of the items on the X-axis represent the amino acids at positions T291 and K365 of SEQ ID NO: 1, in that order. The SsdA wild-type is indicated as TK. Meanwhile, the shading indicated on the X-axis is a heat-map showing the efficiency of unwanted indels.

[0162] As shown in Fig. 2, in the case of fusion proteins having T291S (SK, SH, SA, SN), K365H (TH, SH, CH, AN), K365A (TA, SA, CA, AA), and K365N (TN, SN, CN, AN) mutations, it was confirmed that the base correction efficiency increased compared to wild-type SsdA in all of HEK2, HEK3, and RNF2.

[0163] This means that base correction efficiency can be improved by introducing mutations at positions T291 and K365 of wild-type SsdA (PAAR domain-containing protein, SEQ ID NO: 1).

[0164]

[0165] Example 7. Preparation of an SsdA mutant with additional mutations introduced into the Y335R mutant and verification of its base correction efficiency.

[0166] An SsdA mutant was prepared by introducing additional mutations at positions P282 and K392 of the SsdA (Y335R) mutant prepared by the method of Example 5, and the base correction efficiency of the SsdA mutant with the additional mutations introduced by the method of Example 4 was measured, and is shown in Figure 3.

[0167] Figure 3 is a graph showing the base correction efficiency of the CRISPR-Cas system comprising Cas protein, SsdA mutants with amino acid substitutions at positions P282, Y335, and K392, and a guide polynucleotide at the target site (SEQ ID NOs: 3 to 5) of the intracellular target site (HEK2, HEK3, and RNF2). Specifically, the three alphabetic items on the X-axis represent the amino acids at positions P282, Y335, and K392 of SEQ ID NO: 1, in that order. The SsdA wild-type is indicated as PYK. Meanwhile, the shading indicated on the X-axis is a heat-map showing the efficiency of unwanted indels.

[0168] As shown in Fig. 3, even a single mutation (PYE, PRK, SYK) of P282S, Y335R, and K392E in HEK2, HEK3, and RNF2 improved the base correction efficiency compared to the wild-type SsdA, and the combination of P282, Y335, and K392 mutations (PRE, SYE, SRK, SRE) also improved the base correction efficiency compared to the wild-type SsdA. In particular, the SsdA mutant (SsdA-SRE) that introduced S, R, and E mutations at positions P282, Y335, and K392 of SEQ ID NO: 1, respectively, showed the highest base correction efficiency in HEK2 and RNF2, and was also confirmed to show a high level of base correction efficiency in HEK3. The amino acid sequence of SsdA-SRE is shown in Table 3. The bolded parts in the sequences in Table 3 correspond to the parts where S, R, and E mutations were introduced, respectively. Sequence number 7 is an SsdA mutant (SsdA-SRE) that introduced S, R, and E mutations at positions P282, Y335, and K392 of sequence number 1, respectively, and sequence number 8 is an SsdA mutant (SsdA-SRE) that introduced S, R, and E mutations at positions P25, Y78, and K135 of sequence number 6, respectively.

[0169] SsdA protein sequence Sequence number SsdA-SRE (PARR containing protein) MSAAARVNDPIEHTGSLTGLLAGLAIGAIGAALVVGTGGLAAVAIVGASAATGAGVGQLIGSLSCCNHQTGQIVSGSSNVYINGEPAARAHADQAKCDEHTSRPQVIAQGSSNVYINGHPAARVGDRTACDAKIVVGSSNVFIGGGTETTDPINPEVPELLERSILLVGLASAVVLASPVIVIAGLVGGIAGGTVGSMGGAQLFGEGTDG QKLMAFGGALLGGGLGAKGGKWFDTRYDIKVQGVGSNLGNLKITPKGAAKVSNIAESEAALGRASQARADLSQSKELKVKTVSSNDKKTLSGWGNKKPEGYERISAEQV KAKSEEIGHEVKSHPRDRDYKGQYFSSHAEKQMSIASPNHPLGVSKPMCTDCQGYFSQLAKYSKVEQTVADPEAIRIFKTDGSVETIMRSE7SsdA-SRE(deaminase domain)KVSNIAESEAALGRASQARADLSQSKELKVKTVSSNDKKTLSGWGNKKPEGYERISAEQVKAKSEEIGHEVKSHPRDRDYKGQYFSSHAEKQMSIASPNHPLGVSKPMCTDCQGYFSQLAKYSKVEQTVADPEAIRIFKTDGSVETIMRSE8

[0170]

[0171] Example 8. Confirmation of product purity of SsdA-SRE mutants

[0172] The product purity of the SsdA-SRE variant, which is an example of the present invention, was confirmed by the method of Example 4 at the target site (SEQ ID NO: 3 to 5) of representative intracellular target sites (HEK2, HEK3, and RNF2).

[0173] Figure 4 is a graph analyzing the accuracy of the product of the cytosine base correction technology using SsdA (wild-type) and an SsdA-SRE mutant, which is a specific example of Example 7, at the target site (SEQ ID NO: 3 to 5) of the intracellular target site (HEK2, HEK3, and RNF2).

[0174] As shown in Fig. 4, it was confirmed that the SsdA mutant (SsdA-SRE) had improved product purity compared to the cytosine base correction technology using wild-type SsdA.

[0175] This means that not only can the base correction efficiency of the wild-type SsdA increase by introducing mutations at positions P282, Y335, and K392 of sequence number 1 or P25, Y78, and K135 of sequence number 6, but also the product purity can be improved.

[0176]

[0177] Example 9. Evaluation of cytotoxicity according to the presence or absence of UGI

[0178] To determine whether SsdA of the CRISPR-Cas system is toxic in eukaryotic cells, a plasmid expressing a protein bound to Cas9 (D10A) was constructed. Then, HEK293 cells were seeded at 6 x 10 in a 48-well plate. 4 After dispensing at a cell / well concentration, the plasmid was transfected into HEK293 cells. 48 hours after transfection, live cells were trypsinized and the number of HEK293 cells was counted using a hemocytometer, and the results are shown in Figure 5.

[0179] Figure 5 is a graph showing the cytotoxicity of each of the Cas9(D10A)-SsdA fusion protein, Cas9(D10A)-SsdA-UGI fusion protein, dCas9-SsdA fusion protein, and dCas9-SsdA-UGI fusion protein.

[0180] As shown in Fig. 5, it was confirmed that the fusion protein containing Cas protein, SsdA (wild-type), and UGI (Uracil Glycosylase Inhibitor) had almost no cytotoxicity compared to the fusion protein containing Cas protein and SsdA (wild-type).

[0181]

[0182] Example 10. Confirmation of base correction efficiency and indel formation frequency according to UGI inclusion location.

[0183] 10.1 Confirmation of base correction efficiency according to target position of SsdA mutant (SsdA-SRE) and comparison with base correction efficiency of BE4max

[0184] Based on the results of Example 8, an experiment was conducted to compare and analyze the base correction efficiency at the intracellular target site by binding UGI (Uracil Glycosylase Inhibitor) to SsdA mutants (SsdA-SRE) in which S, R, and E mutations were introduced at positions P282, Y335, and K392 of SEQ ID NO: 1 or P25, Y78, and K135 of SEQ ID NO: 6, which are specific examples. Specifically, the base correction efficiency of two cases (SsdA-SRE-UGI-C2, SsdA-SRE-UGI-N2) in which UGI was positioned at the C-terminus or N-terminus of the SsdA-SRE mutant and BE4max, which is widely used for conventional cytosine base correction, for the intracellular target site was confirmed by the method of Example 4, and this was statistically analyzed. The results of analyzing the cytosine base correction efficiency of each fusion protein at the intracellular target site are presented in a box and whisker plot, which is shown in Figure 6. The intracellular target sites are shown in Table 4.

[0185] 표적 DNA타겟 사이트(target site)서열 번호HEK2GAACACAAAGCATAGACTGCGGG9HEK2-1CCAGCCCGCTGGCCCTGTAAAGG10HEK2-2GCTGGCCCTGTAAAGGAAACTGG11HEK2-3GTTTCCTTTACAGGGCCAGCGGG12HEK2-4GCACTTGTTTGCAGCTATTCAGG13HEK3GGCCCAGACTGAGCACGTGATGG14HEK3-1CTGCTTCTCCAGCCCTGGCCTGG15HEK3-2CCCTGGCCTGGGTCAATCCTTGG16HEK3-3GACTGAGCACGTGATGGCAGAGG17HEK3-6CTTCCTCCAGAGGGCGTCGCAGG18HEK3-7CAGGACAGCTTTTCCTAGACAGG19HEK3-8CAGCTCCTGCACCGGGATACTGG20HEK4GGCACTGCGGCTGGAGGTGGGGG21HEK4-1GGGGCACCGCGGCGCCCCGGTGG22HEK4-2GCGGCGCCCCGGTGGCACTGCGG23HEK4-3CGCCCCGGTGGCACTGCGGCTGG24HEK4-4TCCCTTCCTTCCACCCAGCCCGG25HEK4-5CCCTGCCTGTCATCCTGCTTTGG26HEK4-6GCAGTGCCACCGGGGCGCCGCGG27HEK4-7CTCCAGCCGCAGTGCCACCGGGG28HEK4-8ACCTCCAGCCGCAGTGCCACCGG29RNF2GTCATCTTAGTCATTACCTGAGG30RNF2-3TACACGTCTCATATGCCCCTTGG31RNF2-4TCAACCATTAAGCAAAACATGGG32EMX1GTCACCTCCAATGACTAGGGTGG33FANCFGGAATCCCTTCTGCAGCACCTGG34TYRO3GGCCACACTAGCGTTGCTGCTGG35CCR5TGACATCAATTATTATACATCGG36HEK2-1CCAGCCCGCTGGCCCTGTAAAGG37AAVS1GCTGACTCAGAGACCCTGAGTGG38CUL3GTAAACCTGGAATAACACGATGG39

[0186] Figure 6 is a graph comparing the base correction efficiency of the SsdA-SRE mutant according to the binding position of UGI with the base correction efficiency of BE4max:

[0187] Figure 6a is a graph showing the base correction efficiency of SsdA-SRE-UGI-C2, SsdA-SRE-UGI-N2, and BE4max, and Figure 6b is a graph comparing the base correction efficiency of fusion proteins (SsCBE-UGI-C2, SsCBE-UGI-N2) that combine UGI with SsdA-SRE mutants and BE4max at approximately 30 intracellular target sites (SEQ ID NOs: 9 to 39). Meanwhile, the shading indicated on the X-axis of Figure 6b is a heat-map showing the efficiency of unwanted indels.

[0188] As shown in Figures 6a and 6b, it was confirmed that the base correction efficiency varied depending on the location of UGI.

[0189] In addition, as shown in Fig. 6b, it was confirmed that the SsdA-SRE mutant fusion protein (SsCBE-UGI-C2, SsCBE-UGI-N2) of the present invention exhibits higher base correction efficiency than the existing cytosine base correction technology, BE4max, depending on the target position. This means that although it varies depending on the target position (context-dependency), it exhibits higher performance than the existing BE4max technology at certain positions.

[0190]

[0191] 10.2 Confirmation of indel formation frequency according to target location of SsdA variant (SsdA-SRE)

[0192] Based on the results of Example 8, the frequency of indel formation at the intracellular target site of Table 4 was confirmed by combining UGI (Uracil Glycosylase Inhibitor) with an SsdA mutant (SsdA-SRE) in which S, R, and E mutations were introduced at positions P282, Y335, and K392 of SEQ ID NO: 1 or P25, Y78, and K135 of SEQ ID NO: 6, which is a specific example of the present invention.

[0193] Figure 7 is a graph showing the indel formation frequency of SsdA-SRE-UGI-C2, SsdA-SRE-UGI-N2, and BE4max.

[0194] As shown in Fig. 7, it was confirmed that there was no significant difference in the indel formation frequency of SsdA-SRE-UGI-C2, SsdA-SRE-UGI-N2, and BE4max.

[0195]

[0196] Example 11. Construction of a base editing system using TALE protein and SsdA mutant (SsdA-SRE)

[0197] To construct a cytosine base correction system using TALE arrays, we developed a Golden Gate cloning system consisting of four expression plasmids (including a protein construct expression plasmid, a Mitochondrial transfer signal (MTS), for delivering the resulting proteins to mitochondria) and 424 TALE subarray plasmids. In this system, MTS was cloned at the N-terminus of the expression plasmids to efficiently introduce point mutations into mitochondrial DNA, and SsdA-SRE, a UGI, and a NES signal were added to the C-terminus. To synthesize a TALE array that recognizes the target sequence, the TALE subarray plasmid that recognizes the target sequence and the expression plasmid were reacted with BsaI restriction enzyme and T4 DNA ligase to complete the TALE-SRE expression plasmid. The assembled vector was transformed into bacteria and verified through colony PCR and sequencing.

[0198]

[0199] Example 12. Construction of a Base Correction System Using Zinc Finger Protein and SsdA Mutant (SsdA-SRE)

[0200] To construct a base editing system (ZFP-SRE) using a zinc finger protein known as a DNA binding protein and a SsdA protein variant (SsdA-SRE), a ZFP-SRE expression plasmid was constructed that contains a mitochondrial transfer signal (MTS) and a SsdA protein variant (SsdA-SRE) and expresses the zinc finger protein obtained from a publicly available zinc finger resource.

[0201]

[0202] Experimental Example 1. Confirmation of the base editing range (editing window) according to the target position of the SsdA mutant (SsdA-SRE).

[0203] Figure 8 is a graph showing the base editing range (editing window) of SsdA-SRE-UGI-C2, SsdA-SRE-UGI-N2, and BE4max for cytosine base correction at the intracellular target site shown in Table 4.

[0204] As shown in Fig. 8, the base editing ranges (editing windows) of SsdA-SRE-UGI-C2, SsdA-SRE-UGI-N2, and BE4max were compared at the target positions shown in Table 4, and it was confirmed that high base editing efficiency was shown at positions between 4 and 8 bp from the 5' end of the gRNA target base sequence. In addition, it was confirmed that SsdA-SRE-UGI-C2, SsdA-SRE-UGI-N2, and BE4max all showed similar base editing ranges (editing windows).

[0205]

[0206] Experimental Example 2. Confirmation of base editing efficiency of SsdA mutant (SsdA-SRE) in K562, SKOV3, and HeLa cell lines.

[0207] The base correction efficiency of SsdA-SRE-UGI-C2 and BE4max, which are specific examples of the present invention, in cell lines other than HEK293T cells (K562, SKOV3, HeLa) was compared. Specifically, the base correction efficiency was measured at the target sites (SEQ ID NOs: 3 to 5, SEQ ID NO: 21) of the intracellular target sites (HEK2, HEK3, RNF2, and HEK4) of K562, SKOV3, and HeLa cell lines by the method of Example 4, and the results are shown in Fig. 9.

[0208] Figure 9 is a graph showing the cytosine base correction efficiency of SsdA mutant (SsdA-SRE) in cell lines other than HEK293T cells.

[0209] As shown in Fig. 9, it was confirmed that other human cell lines besides HEK293T cells, such as K562, SKOV3, and HeLa, also exhibited base correction efficiency at a level similar to BE4max.

[0210]

[0211] Experimental Example 3. Confirmation of the base-editing efficiency of the SsdA mutant (SsdA-SRE) and cjCas9 (D8A) fusion protein.

[0212] Using the method of Example 4, the base correction efficiency at the target site of the intracellular target site (EPAS-E2-1, HIF-E9-4, HPD-1 HPD-2) was measured using the cjCas9(D8A)-SsdA-SRE fusion protein and cjBE4max, which are specific examples of the present invention. Specifically, cjBE4max comprising APOBEC1, cjCas9(D8A), and UGI; and cjCas9(D8A)-SsdA-SRE (cjCBE-SRE) comprising SsdA, cjCas9(D8A), and UGI; were produced. A schematic diagram showing the binding structure of cjBE4max and cjCas9(D8A)-SsdA-SRE (cjCBE-SRE) is shown in Fig. 10. The base correction efficiency was measured at the intracellular target sites EPAS-E2-1 (SEQ ID NO: 40), HIF-E9-4 (SEQ ID NO: 41), HPD-1 (SEQ ID NO: 42), and HPD-2 (SEQ ID NO: 43) using the above cjCas9(D8A)-SsdA-SRE fusion protein and cjBE4max, and is shown in Fig. 11.

[0213] Figure 10 is a schematic diagram showing the binding structure of cjBE4max and cjCas9(D8A)-SsdA-SRE (cjCBE-SRE).

[0214] Figure 11 is a graph measuring the base correction efficiency at intracellular target sites (EPAS-E2-1, HIF-E9-4, HPD-1 HPD-2) using cjCas9(D8A)-SsdA-SRE fusion protein and cjBE4max.

[0215] As shown in Fig. 11, it was confirmed that the cjCas9(D8A)-SsdA-SRE fusion protein exhibited higher base correction efficiency than the existing cytosine base correction technology using cjBE4max.

[0216]

[0217] Experimental Example 4. Confirmation of off-target effects due to gRNA mismatch in the SsdA mutant (SsdA-SRE).

[0218] 4.1 Confirmation of off-target effects due to gRNA mismatch in SsdA mutant (SsdA-SRE)

[0219] In order to confirm the off-target effect of the SsdA mutant (SsdA-SRE), which is a specific example of the present invention, due to gRNA mismatch, 30 types of guide RNAs were produced in which 1, 2, 3, or 4 mutations were introduced in the 20 nucleotide base sequence of the guide RNA targeting the intracellular target sites HEK3 and RNF2 (SEQ ID NO: 3 and SEQ ID NO: 5), and an experiment was performed to compare and analyze the base correction efficiency of the SsdA mutant (SsdA-SRE) or BE4max for each guide RNA.

[0220] Guide RNA sequences with introduced mutations are shown in Tables 5 and 6. Table 5 shows the HEK3 target guide RNA sequence with introduced mutations, and Table 6 shows the RNF2 target guide RNA sequence with introduced mutations. In the guide RNA sequences shown in Tables 5 and 6, lowercase letters indicate sequences with inserted mismatch bases.

[0221] Guide RNA Sequence Sequence Number HEK3Target Guide RNA 1GatttAGACTGAGCACGTGATGG 4HEK3Target Guide RNA 2GGCCtgagCTGAGCACGTGATGG 45HEK3Target Guide RNA 3GGCCCAGAtcagGCACGTGATGG 46HEK3Target Guide RNA 4GGCCCAGACTGAatgtGTGATGG 47HEK3Target Guide RNA 5GGCCCAGACTGAGCACacagTGG 48HEK3Target Guide RNA 6GGtttAGACTGAGCACGTGATGG 49HEK3Target Guide RNA 7GGCCCgagCTGAGCACGTGATGG 50HEK3Target Guide RNA 8GGCCCAGAtcaAGCACGTGATGG 51HEK3Target Guide RNA 9GGCCCAGACTGgatACGTGATGG 52HEK3Target Guide RNA 10GGCCCAGACTGAGCgtaTGATGG53HEK3Target Guide RNA 11GGCCCAGACTGAGCACGcagTGG54HEK3Target Guide RNA 12GGttCAGACTGAGCACGTGATGG55HEK3Target Guide RNA 13GGCCtgGACTGAGCACGTGATGG56HEK3Target Guide RNA 14GGCCCAagCTGAGCACGTGATGG57HEK3Target Guide RNA 15GGCCCAGAtcGAGCACGTGATGG58HEK3Target Guide RNA 16GGCCCAGACTagGCACGTGATGG59HEK3Target Guide RNA 17GGCCCAGACTGAatACGTGATGG60HEK3Target Guide RNA 18GGCCCAGACTGAGCgtGTGATGG61HEK3Target Guide RNA 19GGCCCAGACTGAGCACacGATGG62HEK3Target Guide RNA 20GGCCCAGACTGAGCACGTagTGG63HEK3Target guide RNA 21GaCCCAGACTGAGCACGTGATGG64HEK3Target guide RNA 22GGCtCAGACTGAGCACGTGATGG65HEK3Target guide RNA 23GGCCCgGACTGAGCACGTGATGG66HEK3Target guide RNA24GGCCCAGgCTGAGCACGTGATGG67HEK3Target guide RNA 25GGCCCAGACcGAGCACGTGATGG68HEK3Target guide RNA 26GGCCCAGACTGgGCACGTGATGG69HEK3Target guide RNA 27GGCCCAGACTGAGtACGTGATGG70HEK3Target guide RNA 28GGCCCAGACTGAGCAtGTGATGG71HEK3Target guide RNA 29GGCCCAGACTGAGCACGcGATGG72HEK3Target guide RNA 30GGCCCAGACTGAGCACGTGgTGG73

[0222] Guide RNA Sequence Sequence Number RNF2 Target Guide RNA 1 GctgcCTTAGTCATTACCTGAGG74 RNF2 Target Guide RNA 2 GTCActccAGTCATTACCTGAGG75 RNF2 Target Guide RNA 3 GTCATCTTgactATTACCTGAGG76 RNF2 Target Guide RNA 4 GTCATCTTAGTCgccgCCTGAGG77 RNF2 Target Guide RNA 5 GTCATCTTAGTCATTAttcaAGG78 RNF2 Target Guide RNA 6 GTtgcCTTAGTCATTACCTGAGG79 RNF2 Target Guide RNA 7 GTCATtccAGTCATTACCTGAGG80 RNF2 Target Guide RNA 8 GTCATCTTgacCATTACCTGAGG81 RNF2 Target Guide RNA 9 GTCATCTTAGTtgcTACCTGAGG82 RNF2 Target Guide RNA 10GTCATCTTAGTCATcgtCTGAGG83RNF2 Target guide RNA 11GTCATCTTAGTCATTACtcaAGG84RNF2 Target guide RNA 12GTtgTCTTAGTCATTACCTGAGG85RNF2 Target guide RNA 13GTCActTTAGTCATTACCTGAGG86RNF2 Target guide RNA 14GTCATCccAGTCATTACCTGAGG87RNF2 Target guide RNA 15GTCATCTTgaTCATTACCTGAGG88RNF2 Target guide RNA 16GTCATCTTAGctATTACCTGAGG89RNF2 Target guide RNA 17GTCATCTTAGTCgcTACCTGAGG90RNF2 Target guide RNA 18GTCATCTTAGTCATcgCCTGAGG91RNF2 Target guide RNA 19GTCATCTTAGTCATTAttTGAGG92RNF2 Target guide RNA 20GTCATCTTAGTCATTACCcaAGG93RNF2 Target guide RNA 21GcCATCTTAGTCATTACCTGAGG94RNF2 Target guide RNA 22GTCgTCTTAGTCATTACCTGAGG95RNF2 Target guide RNA 23GTCATtTTAGTCATTACCTGAGG96RNF2Target guide RNA 24GTCATCTcAGTCATTACCTGAGG97RNF2 Target guide RNA 25GTCATCTTAaTCATTACCTGAGG98RNF2 Target guide RNA 26GTCATCTTAGTtATTACCTGAGG99RNF2 Target guide RNA 27GTCATCTTAGTCAcTACCTGAGG100RNF2 Target guide RNA 28GTCATCTTAGTCATTgCCTGAGG101RNF2 Target guide RNA 29GTCATCTTAGTCATTACtTGAGG102RNF2 Target guide RNA 30GTCATCTTAGTCATTACCTaAGG103

[0223] The gRNA and SsdA mutant (SsdA-SRE) or BE4max produced as shown in Tables 5 and 6 above were transfected into HEK293T cells, and 3 days later, genomic DNA was extracted and the mutation efficiency of the intracellular target sites HEK3 and RNF2 (SEQ ID NO: 3 and SEQ ID NO: 5) was analyzed using next-generation sequencing, and the results are shown in Fig. 12.

[0224] Figure 12 is a graph showing the base correction efficiency according to guide RNA mutation at the intracellular target site using SsdA-SRE or BE4max, which are specific examples of SsdA mutants:

[0225] Figure 12a is a graph showing the base correction efficiency according to guide RNA mutation at the intracellular target site HEK3 using SsdA mutant (SsdA-SRE) or BE4max, and Figure 12b is a graph showing the base correction efficiency according to guide RNA mutation at the intracellular target site RNF2 using SsdA mutant (SsdA-SRE) or BE4max.

[0226] As shown in Figure 12, the cytosine base correction technology using the SsdA mutant (SsdA-SRE) fusion protein was found to be sensitive to guide RNA mutations, and the base correction efficiency was significantly reduced when two or more mismatched mutations were present. Furthermore, compared to BE4max, the sensitivity according to guide RNA mutations was confirmed to be similar or slightly higher.

[0227]

[0228] 4.2 Identification of locations where SsdA variants (SsdA-SRE) can cause mutations across the entire human genome using Digenome-seq

[0229] In order to identify the location where the SsdA variant can cause mutations in the entire human genome, SsdA-SRE, which is an example of the present invention, was used. Specifically, the Digenome-seq method was used to identify the location where the SsdA variant (SsdA-SRE) can actually cause mutations in the entire human genome. To find the target location where the SsdA variant causes mutations in the genome in vitro, the SsdA variant (SsdA-SRE) and guide RNA were treated with human DNA in vitro, and the location where the mutation occurred was analyzed using the whole genome sequencing (WGS) technique. Whole genome sequencing (WGS) was performed by treating the SsdA variant (SsdA-SRE), and the location where the sequenced DNA fragments appear in a row at the same location was analyzed to identify the location on the genome where the off-target effect of the SsdA variant (SsdA-SRE) appears. The Digenome-seq experimental process for analyzing the off-target effect of SsdA mutants is shown in Figure 13.

[0230] Figure 13 is a schematic diagram showing the Digenome-seq experimental process for analyzing the off-target effect of SsdA mutants.

[0231] Figure 14 shows the results of the analysis of the off-target effect of the SsdA mutant:

[0232] Figure 14a shows the results of a Digenome-seq experiment using an SsdA mutant fusion protein and a HEK2 target guide RNA, and Figure 14b is a graph showing the base mutation efficiency at each position in cells transfected with an SsdA mutant and a HEK2 target guide RNA to confirm whether the off-target effect actually occurs at each off-target effect position.

[0233] As shown in Figure 14a, Digenome-seq experiments identified target sites where SsdA variants induce mutations in the in vitro genome. The base sequence information for sites predicted to exhibit off-target effects, identified through Digenome-seq, is shown in Table 7. In the base sequences shown in Table 7, lowercase letters indicate mismatches, and dashes indicate bulges.

[0234] Base sequence information at the location where the off-target effect appears Sequence number 1 GAACACAAAGCATAGACTGCGGG 1042 GAACACA-tGCATAGACTGCTAG 1053 GAACACAAtGCATAGAtTGCCGG 1064 aActcCAAAGCATAtACTGCTGG 1075 acACACAAAGCAT-GACTGCAGG 1086 aAACACAgAGCAcAGACTGCTGA 1097 GAACAC-AAGCAcAGACTGaAGG 1108 GtAaACAAAGCATAGACTGaGGG 1119 G AACACAtA-CATAGACaGCTGG11210GAAttCAAAGCATAGAtTGCAGG11311tcACACAAAcCATAGACTGaGGG11412tAACAaAtAGCATAGACTGtGTG1151 3GAACACAgtaCATAGACTGgCAG11614aAACAtAAAGaATAGACTGCAAG11715GAAtACtAAGCATAGACTcCAGG11816GAACtCAAAGCATAGAaTaaTGG119

[0235] In addition, as shown in Fig. 14b, the SsdA mutant did not exhibit off-target effects for each candidate off-target effect site on the genome obtained from the Digenome-seq results in transfected cells, and this was confirmed to be at a level similar to that of the previously known BE4max.

[0236]

[0237] 4.3 Confirmation of the guide RNA-independent off-target effect of the SsdA mutant (SsdA-SRE)

[0238] To confirm the guide RNA-independent off-target effect of SsdA-SRE, a specific example of an SsdA mutant, we performed an R-loop assay using dSaCas9 and a guide RNA targeting its target sequence, sites 1 to 6. Specifically, the dSaCas9 protein, which inhibits SaCas9 activity, and the guide RNA binding to it were delivered to cells. dSaCas9 and the guide RNA bind to the target site in the cell, forming an R-loop, unwinding the DNA double helix and exposing a single helix. At this time, a base-editing protein was also delivered to the cells to randomly induce mutations in the exposed DNA single helix at the target site of dSaCas9. The results of the R-loop assay are shown in Fig. 15.

[0239] Figure 15 is a graph showing the results of an experiment analyzing the guide RNA-independent off-target effect.

[0240] As shown in Fig. 15, when the SsdA mutant fusion protein is compared with BE4max, it can be seen that the SsdA mutant has a lower guide RNA-independent off-target effect than the conventional BE4max.

[0241]

[0242] 4.4 Confirmation of RNA off-targeting effect of SsdA mutant (SsdA-SRE)

[0243] To confirm the RNA off-targeting effect of SsdA-SRE, a specific example of an SsdA mutant, guide RNA for the target site was delivered into cells, and base mutations were analyzed. Specifically, the SsdA mutant (SsdA-SRE) and guide RNAs for HEK2 and RNF2 (SEQ ID NOs: 3 and 5) were delivered into cells. RNA from cells introduced with the guide RNA was extracted, and the base sequence of the entire transcriptome was obtained using next-generation sequencing, and base mutations appearing in this base sequence were analyzed.

[0244] Figure 16 is a graph analyzing the RNA off-targeting effect of SsdA mutants:

[0245] Figure 16a shows the results of analyzing base mutations in transcripts in cells delivered with SsdA mutants or BE4max and RNF2 targeting guide RNA, and Figure 16b shows the results of analyzing base mutations in transcripts in cells delivered with SsdA mutants or BE4max and HEK2 targeting guide RNA.

[0246] As shown in Fig. 16, it can be seen that the SsdA mutant exhibits an RNA off-target effect at a lower level than the previously known BE4max.

[0247]

[0248] Experimental Example 5. Confirmation of the mitochondrial base-editing effect of the SsdA mutant (SsdA-SRE).

[0249] 5.1. Confirmation of the effect of mitochondrial base correction using TALE-SRE expression plasmid

[0250] The TALE-SRE expression plasmid of Example 11 was used to introduce point mutations into mitochondrial DNA. Specifically, the TALE-SRE expression plasmid was delivered to HEK293T cells, and 4 days later, genomic DNA was extracted from the cells, and the target sequence was amplified using PCR. The amplified DNA was then analyzed using next-generation sequencing (NGS), and the target sequences were the ND1 gene, ATP6-1 gene, ATP6-2 gene, COX3-1 gene, COX3-2 gene, COX3-3 gene, CYB-1 gene, and CYB-2 gene. The TALE-SRE expression plasmids for introducing point mutations into each target sequence are as follows:

[0251] Point mutations were introduced into the CYB gene using TALE-SRE-CYB-Left (SEQ ID NO: 125) and TALE-SRE-CYB-Right (SEQ ID NO: 126). The components of TALE-SRE-CYB-Left and TALE-SRE-CYB-Right are shown in Table 8 below.

[0252] TALE-SRE-CYB-LeftComponentsAmino acid sequencesMTSMASVLTPLLLRGLTGSARRLPVPRAKIHSLFLAG tagDYKDHDGDYKDHDIDYKDDDDKTALEN-terminal domainDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNCYB-leftbinding domainLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLHalf-domainLPVLCQAHGLTPEQVVAIASNGGGKQALETALEC-terminaldomainDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNSsdA-SREKVSNIAESEAALGRASQARADLSQSKELKVKTVSSNDKKTLSGWGNKKPEGYERISAEQVKAKSEEIGHEVKSHPRDRDYKGQYFSSHAEKQMSIASPNHPLGVSKPMCTDCQGYFSQLAKYSKVEQTVADPEAIRIFKTDGSVETIMRSE1xUGISGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKMLNESVDEMTKKFGTLTIHDTEK*TALE-SRE-CYB-RightComponentsAmino acid sequencesMTSMASVLTPLLRGLTGSARRLPVPRAKIHSLFLAG tagDYKDHDGDYKDHDIDYKDDDDKTALEN-terminal domainDLRTLGYSQQQQEKIKKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPNLCYB-rightbinding-domainLPVLCQAHGLTPEQVVAIASNGGGKQALETALEC-terminaldomainDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNSsdA-SREKVSNIAESEAALGRASQARADLSQSKELKVKTVSSNDKKTLSGWGNKKPEGYERISAE QVKAKSEEIGHEVKSHPRDRDYKGQYFSSHAEKQMSIASPNHPLGVSKPMCTDCQGYFSQLAKYSKVEQTVADPEAIRIFKTDGSVETIMRSE1xUGISGGST NLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKMLNESVDEMTKKFGTLTIHDTEK*

[0253] Point mutations were introduced into the ND1 gene using TALE-SRE-ND1-Left (SEQ ID NO: 127) and TALE-SRE-ND1-right (SEQ ID NO: 128). The components of TALE-SRE-ND1-Left and TALE-SRE-ND1-right are shown in Table 9 below.

[0254] TALE-SRE-ND1-LeftComponentsAmino acid sequencesMTSMASVLTPLLLRGLTGSARRLPVPRAKIHSLFLAG tagDYKDHDGDYKDHDIDYKDDDDKTALEN-terminal domainDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNND1-leftbinding domainLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLHalf-domainLPVLCQAHGLTPEQVVAIASNGGGKQALETALEC-terminaldomainDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNSsdA-SREKVSNIAESEAALGRASQARADLSQSKELKVKTVSSNDKKTLSGWGNKKPEGYERISAEQVKAKSEEIGHEVKSHPRDRDYKGQYFSSHAEKQMSIASPNHPLGVSKPMCTDCQGYFSQLAKYSKVEQTVADPEAIRIFKTDGSVETIMRSE1xUGISGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKMLNESVDEMTKKFGTLTIHDTEK*TALE-SRE-ND1-rightComponentsAmino acid sequencesMTSMASVLTPLLLRGLTGSARRLPVPRAKIHSLFLAG tagDYKDHDGDYKDHDIDYKDDDDKTALEN-terminal domainDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNND1-rightbinding-domainLPVLCQAHGLTPEQVVAIASNGGGKQALETALEC-terminaldomainDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNSsdA-SREKVSNIAESEAALGRASQARADLSQSKELKVKTVSSNDKKTLSGWGNKKPEGYERISAE QVKAKSEEIGHEVKSHPRDRDYKGQYFSSHAEKQMSIASPNHPLGVSKPMCTDCQGYFSQLAKYSKVEQTVADPEAIRIFKTDGSVETIMRSE1xUGISGGST NLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKMLNESVDEMTKKFGTLTIHDTEK*

[0255] Point mutations were introduced into the ATP6 gene using TALE-SRE-ATP6-left (SEQ ID NO: 129) and TALE-SRE-ATP6-right (SEQ ID NO: 130). The components of TALE-SRE-ATP6-left and TALE-SRE-ATP6-right are shown in Table 10 below.

[0256] TALE-SRE-ATP6-leftComponentsAmino acid sequencesMTSMASVLTPLLLRGLTGSARRLPVPRAKIHSLFLAG tagDYKDHDGDYKDHDIDYKDDDDKTALEN-terminal domainDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNATP6-leftbinding domainLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGHalf-domainLPVLCQAHGLTPEQVVAIASNGGGKQALETALEC-terminaldomainDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNSsdA-SREKVSNIAESEAALGRASQARADLSQSKELKVKTVSSNDKKTLSGWGNKKPEGYERISAEQVKAKSEEIGHEVKSHPRDRDYKGQYFSSHAEKQMSIASPNHPLGVSKPMCTDCQGYFSQLAKYSKVEQTVADPEAIRIFKTDGSVETIMRSE1xUGISGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKMLNESVDEMTKKFGTLTIHDTEK*TALE-SRE-ATP6-rightComponentsAmino acid sequencesMTSMASVLTPLLLRGLTGSARRLPVPRAKIHSLFLAG tagDYKDHDGDYKDHDIDYKDDDDKTALEN-terminal domainDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNATP6-rightbinding-domainLPVLCQAHGLTPEQVVAIASNGGGKQALETALEC-terminaldomainDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNSsdA-SREKVSNIAESEAALGRASQARADLSQSKELKVKTVSSNDKKTLSGWGNKKPEGYERISAE QVKAKSEEIGHEVKSHPRDRDYKGQYFSSHAEKQMSIASPNHPLGVSKPMCTDCQGYFSQLAKYSKVEQTVADPEAIRIFKTDGSVETIMRSE1xUGISGGST NLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKMLNESVDEMTKKFGTLTIHDTEK*

[0257] Point mutations were introduced into the COX3 gene using TALE-SRE-COX3-left (SEQ ID NO: 131) and TALE-SRE-COX3-right (SEQ ID NO: 132). The components of TALE-SRE-COX3-left and TALE-SRE-COX3-right are shown in Table 11 below.

[0258] TALE-SRE-COX3-leftComponentsAmino acid sequencesMTSMASVLTPLLLRGLTGSARRLPVPRAKIHSLFLAG tagDYKDHDGDYKDHDIDYKDDDDKTALEN-terminal domainDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNCOX3-leftbinding domainLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGHalf-domainLPVLCQAHGLTPEQVVAIASNGGGKQALETALEC-terminaldomainDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNSsdA-SREKVSNIAESEAALGRASQARADLSQSKELKVKTVSSNDKKTLSGWGNKKPEGYERISAEQVKAKSEEIGHEVKSHPRDRDYKGQYFSSHAEKQMSIASPNHPLGVSKPMCTDCQGYFSQLAKYSKVEQTVADPEAIRIFKTDGSVETIMRSE1xUGISGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKMLNESVDEMTKKFGTLTIHDTEK*TALE-SRE-COX3-rightComponentsAmino acid sequencesMTSMASVLTPLLLRGLTGSARRLPVPRAKIHSLFLAG tagDYKDHDGDYKDHDIDYKDDDDKTALEN-terminal domainDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNCOX3-rightbinding-domainLPVLCQAHGLTPEQVVAIASNGGGKQALETALEC-terminaldomainDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNSsdA-SREKVSNIAESEAALGRASQARADLSQSKELKVKTVSSNDKKTLSGWGNKKPEGYERISAE QVKAKSEEIGHEVKSHPRDRDYKGQYFSSHAEKQMSIASPNHPLGVSKPMCTDCQGYFSQLAKYSKVEQTVADPEAIRIFKTDGSVETIMRSE1xUGISGGST NLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKMLNESVDEMTKKFGTLTIHDTEK*

[0259] The results of base correction of mitochondrial genes using the above T ALE-SRE expression plasmid are shown in Figure 17 (data are expressed as mean + - standard error of the mean (sem) obtained from n=3 biologically independent samples).

[0260] Figure 17 shows the results of introducing point mutations into mitochondrial genes using the TALE-SRE system:

[0261] Figure 17a is a schematic diagram of mitochondrial ND1 base editing using a TALE-SRE expression plasmid, and Figure 17b is a graph showing the base editing effect after delivering a TALE-SRE expression plasmid in HEK293T cells.

[0262] As shown in Fig. 17, the base correction system using TALE protein and SsdA mutant (SsdA-SRE) was confirmed to exhibit a point mutation introduction efficiency of up to 11.2% for mitochondrial genes ND1, ATP6, CYB, and COX3.

[0263] These results suggest that the TALE-SRE system enables base editing within mitochondria using only a single (monomer) TALE array. This differs from the existing mitochondrial base editing technology, DdCBE, which requires a pair of (dimer) TALE arrays using segmented DddAtox, demonstrating that the TALE-SRE system can reliably perform mitochondrial base editing with fewer components. Furthermore, the TALE-SRE system has the advantage of facilitating the gene delivery of the base editing system into mitochondria because it enables effective base editing with fewer components than existing technologies.

[0264] 5.2. Confirmation of the effect of mitochondrial base correction using the ZFP-SRE expression plasmid.

[0265] A point mutation was introduced into the ND1 gene in mitochondria using the ZFP-SRE expression plasmid constructed according to the method of Example 12. Specifically, the ZFP-SRE expression plasmid was delivered to HEK293T cells, and 4 days later, genomic DNA was extracted from the cells and the target sequence was amplified using PCR. The amplified DNA was analyzed using next-generation sequencing (NGS). The ZFP-SRE expression plasmid for introducing a point mutation into the ND1 gene in mitochondria is as follows:

[0266] Point mutations were introduced into the ND1 gene using mitoZFD-SRE-ND1-left (SEQ ID NO: 133) and mitoZFD-SRE-ND1-right (SEQ ID NO: 134). The components of mitoZFD-SRE-ND1-left and mitoZFD-SRE-ND1-right are shown in Table 12.

[0267] mitoZFD-SRE-ND1-leftComponentsAmino acid sequencesMTSMLGFVGRVAAAPASGALRRLTPSASLPPAQLLLRAAPTAVHPVRDYAAQHA tagYPYDVPDYANESVDEMTKKFGTLTIHDTEKSsdA-SREKVSNIAESEAALGRASQARADLSQSKELKVKTVSSNDKKTLSGWGNKKPEGYERISAEQVKAKSEEIGHEVKSHPRDRDYKGQYFSSHAEKQMSIASPNHPLGVSKPMCTDCQGYFSQLAKYSKVEQTVADPEAIRIFKTDGSVETIMRSEND1-leftZinc-fingerbinding domainFQCRICMRNFSDSGNLRVHIRTHTGEKPYKCPDCGKSFSQSSSLIRHQRTHTGEKPYECDHCGKSFSQSSHLNVHKRTHTGEKPYRCKYCDRSFSISSNLQRHVRNIH1xUGITNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML*mitoZFD-SRE-ND1-rightComponentsAmino acid sequencesMTSMLGFVGRVAAAPASGALRRLTPSASLPPAQLLLRAAPTAVHPVRDYAAQHA tagYPYDVPDYANESVDEMTKKFGTLTIHDTEKSsdA-SREKVSNIAESEAALGRASQARADLSQSKELKVKTVSSNDKKTLSGWGNKKPEGYERISAEQVKAKSEEIGHEVKSHPRDRDYKGQYFSSHAEKQMSIASPNHPLGVSKPMCTDCQGYFSQLAKYSKVEQTVADPEAIRIFKTDGSVETIMRSEND1-rightZinc-fingerbindingdomainYKCPECGKSFSTKNSLTEHQRTHTGEKPYKCPECGKSFSSKKALTEHQRTHTGEKPYECNYCGKTFSVSSTLIRHQRIHTGEKPYRCKYCDRSFS ISSNLQRHVRNIH1xUGITNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML*

[0268] The results of introducing a point mutation into the ND1 gene in mitochondria using the above ZFP-SRE expression plasmid are shown in Figure 18 (data are expressed as mean + - standard error of the mean (sem) obtained from n=3 biologically independent samples).

[0269] Figure 18 is a graph showing the results of introducing point mutations into the mitochondrial ND1 gene using a ZFP-SRE expression plasmid.

[0270] As shown in Fig. 18, the base correction system (ZFP-SRE) using Zinc Finger Protein and SsdA mutant (SsdA-SRE) was confirmed to exhibit a point mutation introduction efficiency of up to 12% for the mitochondrial ND1 gene.

[0271] These results imply that the base-editing system (ZFP-SRE) using zinc finger proteins and SsdA mutants (SsdA-SRE) can perform base-editing in mitochondria with only a single (monomer) TALE array. This is comparable to the existing ZFP and DddA toxUnlike previous base-editing techniques that required a single pair (dimer) of TALE arrays, the ZFP-SRE system demonstrates that it can reliably perform mitochondrial base editing with fewer components. Furthermore, the ZFP-SRE system offers the advantage of facilitating the gene transfer of the base-editing system into mitochondria, as it enables effective base editing with fewer components than existing technologies.

Claims

1. SsdA (single-stranded DNA deaminase toxin A) mutant.

2. In claim 1, the SsdA mutant is a mutant in which an amino acid at a position capable of improving binding affinity with target DNA is substituted with another amino acid in the wild-type SsdA protein.

3. A mutant according to claim 2, wherein the wild-type SsdA protein has an amino acid sequence of sequence number 1.

4. A variant according to claim 2, wherein the position capable of improving binding affinity with the target DNA is V289, H333, Y335, P282, K392 or a position having a corresponding function.

5. A variant according to claim 1, wherein the SsdA variant has another amino acid substituted at one or more of V289, H333, Y335, P282, K392 or a position having a corresponding function.

6. In claim 2, the other amino acid is a mutant selected from the group consisting of arginine (R), histidine (H), lysine (K), aspartic acid (D), glutamic acid (E), serine (S), threonine (T), asparagine (N), glutamine (Q), cysteine ​​(C), selenocysteine ​​(U), glycine (G), proline (P), alanine (A), valine (V), isoleucine (I), leucine (L), methionine (M), phenylalanine (F), tyrosine (Y), tryptophan (W), and all variants of the above amino acids, excluding the amino acid that the wild-type protein has at the variant position.

7. A variant according to claim 2, wherein the other amino acid is a basic amino acid.

8. A variant according to claim 7, wherein the basic amino acid is arginine (R), histidine (H), or lysine (K).

9. In claim 1, the SsdA variant comprises an amino acid sequence having at least 80% identity with the amino acid sequence of SEQ ID NO: 1, and comprises a mutation at at least one position selected from the group consisting of Y335, P282, K392, and positions having a corresponding function within the amino acid sequence.

10. In claim 1, the SsdA variant comprises an amino acid sequence having at least 80% identity with the amino acid sequence of SEQ ID NO: 6, and comprises a mutation at at least one position selected from the group consisting of P25, Y78, K135, and positions having a corresponding function within the amino acid sequence.

11. In claim 1, the SsdA mutant is a mutant in which an amino acid in the catalytic active site of the wild-type SsdA protein is substituted with another amino acid.

12. A mutant according to claim 11, wherein the wild-type SsdA protein has a base sequence of sequence number 1.

13. A variant according to claim 11, wherein the catalytic active site is T291, K365 or a position having a corresponding function.

14. A variant according to claim 1, wherein the SsdA variant has another amino acid substituted at one or more of T291, K365 or positions having a corresponding function.

15. In claim 11, the other amino acid is a mutant selected from the group consisting of arginine (R), histidine (H), lysine (K), aspartic acid (D), glutamic acid (E), serine (S), threonine (T), asparagine (N), glutamine (Q), cysteine ​​(C), selenocysteine ​​(U), glycine (G), proline (P), alanine (A), valine (V), isoleucine (I), leucine (L), methionine (M), phenylalanine (F), tyrosine (Y), tryptophan (W), and all variants of the above amino acids, excluding the amino acid that the wild-type protein has at the variant position.

16. In claim 5, the SsdA variant is a variant in which the amino acid at a position having a function corresponding thereto, such as T291, K365, or the like, is further substituted with another amino acid.

17. In claim 1, the SsdA mutant is a mutant in which an amino acid in the deaminase domain of the wild-type SsdA protein is substituted with another amino acid.

18. A variant according to claim 11, wherein the wild-type deaminase domain comprises the amino acid sequence of SEQ ID NO:

6.

19. A mutant according to claim 17, wherein the SsdA mutant is inactivated.

20. A variant according to claim 17, wherein the deaminase domain has another amino acid substituted at any one or more of V32, H76, Y78, P25, K135, T34, K108 or a position having a corresponding function.

21. In claim 17, the other amino acid is a mutant selected from the group consisting of arginine (R), histidine (H), lysine (K), aspartic acid (D), glutamic acid (E), serine (S), threonine (T), asparagine (N), glutamine (Q), cysteine ​​(C), selenocysteine ​​(U), glycine (G), proline (P), alanine (A), valine (V), isoleucine (I), leucine (L), methionine (M), phenylalanine (F), tyrosine (Y), tryptophan (W), and all variants of the above amino acids, excluding the amino acid that the wild-type protein has at the variant position.

22. A fusion protein comprising a nucleic acid-binding protein (NABP) and a single-stranded DNA deaminase toxin A (SsdA) variant.

23. A fusion protein according to claim 22, wherein the SsdA variant is any one of the variants of claims 1 to 21.

24. A fusion protein according to claim 22, wherein the nucleic acid binding protein binds to a nucleic acid in a sequence-specific manner.

25. A fusion protein according to claim 22, wherein the nucleic acid binding protein is any one selected from the group consisting of TALE (Transcription Activator-like Effector) protein, ZFP (Zinc Finger Protein), Cas9 (CRISPR-associated protein 9), Cpf1, dCpf1, dCas12, dCas13, Cas14, Cas12b, Cas12f, and variants thereof.

26. A fusion protein according to claim 22, wherein the fusion protein further comprises a DNA glycosylase inhibitor.

27. A fusion protein according to claim 26, wherein the DNA glycosylase inhibitor is a thymine glycosylase inhibitor, a uracil glycosylase inhibitor, an oxoguanine glycosylase inhibitor, or an alkylguanine DNA glycosylase inhibitor.

28. A fusion protein according to claim 27, wherein the DNA glycosylase inhibitor is linked to at least one N-terminus or C-terminus of the fusion protein.

29. A fusion protein according to claim 22, further comprising an organelle signal peptide domain.

30. A fusion protein according to claim 29, wherein the organelle signal peptide domain is a nuclear localization signal (NLS), a nuclear export signal (NES), a mitochondrial transfer signal (MTS), or a chloroplast transit peptide (CTP).

31. A fusion protein according to claim 22, wherein the nucleic acid binding protein and the SsdA (single-stranded DNA deaminase toxin A) variant are connected by a linker.

32. A polynucleotide encoding a fusion protein of any one of claims 22 to 31.

33. A vector comprising the polynucleotide of claim 32.

34. A fusion protein comprising a nucleic acid-binding protein (NABP) and a single-stranded DNA deaminase toxin A (SsdA) variant; or A composition for base correction comprising a polynucleotide encoding the above fusion protein.

35. A base editing system comprising a fusion protein comprising a nucleic acid-binding protein (NABP) and a bacterial toxin, or a polynucleotide encoding the fusion protein; and a guide polynucleotide, wherein the bacterial toxin is a single-stranded DNA deaminase toxin A (SsdA) variant.

36. A base correction system according to claim 35, wherein the system substitutes at least one nucleotide in the nucleotide sequence of a target nucleic acid molecule.

37. A base correction system according to claim 36, wherein the substitution of the nucleotide is from cytosine to uracil.

38. A base correction system according to claim 35, wherein the guide polynucleotide is at least one selected from the group consisting of sequence numbers 44 to 103.

39. A method of editing a nucleic acid, comprising the step of contacting a nucleic acid molecule with a base editing system, The above editing is a substitution of at least one nucleotide in the nucleotide sequence of the nucleic acid molecule, A method for editing a nucleic acid, wherein the base editing system comprises a fusion protein comprising a nucleic acid-binding protein (NABP) and a bacterial toxin or a polynucleotide encoding the fusion protein; and a guide polynucleotide, wherein the bacterial toxin is a single-stranded DNA deaminase toxin A (SsdA) variant.

40. A method for editing a nucleic acid, wherein the substitution of the nucleotide in claim 37 is from cytosine to uracil.