Improved base editing method using UDG variant

The use of inactive uracil DNA glycosylase (dUDG) with a DNA binding protein and cytosine deaminase in base editing systems addresses the limitations of CRISPR in organelles, achieving efficient C-to-T base correction in mitochondria and chloroplasts.

WO2026010436A1PCT designated stage Publication Date: 2026-01-08GREENGENE INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/009633
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-10-29
Filing Date
2025-07-04
Publication Date
2026-01-08

AI Technical Summary

Technical Problem

Conventional genome editing tools, such as CRISPR systems, are ineffective for correcting DNA bases in organelles like mitochondria and chloroplasts due to the inability to deliver guide RNAs, limiting their use in treating genetic diseases and improving crop traits.

Method used

A base editing method using a DNA binding protein, cytosine deaminase, and inactive uracil DNA glycosylase (dUDG) for C-to-T base conversion, which maintains uracil stability through dUDG's binding without catalytic activity, enhancing editing efficiency.

Benefits of technology

The method achieves high specificity and efficiency in C-to-T base correction in organelle DNA, surpassing the performance of uracil DNA glycosylase inhibitors (UGI), particularly in chloroplasts and mitochondria.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025009633_08012026_PF_FP_ABST
    Figure KR2025009633_08012026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to base editing using dead uracil DNA glycosylase (dUDG). The present invention is useful for C-to-T base editing in nuclear DNA or organelle DNA, and in particular, is useful for C-to-T base editing in organelle DNA such as chloroplasts or mitochondria. The present invention also relates to a UDG variant and a novel DNA base editing use thereof.
Need to check novelty before this filing date? Find Prior Art

Description

An improved base editing method using UDG variants

[0001] The present invention relates to DNA base editing. Specifically, the present invention relates to a base editor comprising a programmable DNA binding protein and a deaminase, and a base editing method using the same. More specifically, the present invention relates to base editing using dead uracil DNA glycosylase (dUDG). The present invention is useful for C-to-T base editing in nuclear DNA or organelle DNA, and is particularly useful for C-to-T base editing in organelle DNA such as chloroplasts or mitochondria. The present invention also relates to a mutant UDG and novel uses thereof for DNA base editing.

[0002] Fusion proteins that link DNA binding proteins and deaminase enzymes enable the induction of DNA mutations, such as single nucleotide conversions in a targeted manner to replace nucleotides or correct bases in the genome without generating DNA double-strand breaks (DSBs), to correct point mutations that cause genetic disorders, or to introduce desired single nucleotide mutations in prokaryotes and eukaryotic cells such as humans.

[0003] Programmable genome editing tools, such as zinc-finger nuclease (ZFN), transcription activator-like effector nuclease (TALEN), clustered regularly interspaced short palindromic repeat (CRISPR) systems, and base editors composed of CRISPR-associated protein 9 (Cas9) variants and nucleotide deaminase proteins, have the potential to treat genetic diseases and improve crop traits through base sequence changes. However, these conventional genome editing tools are not suitable for correcting DNA bases in organelles such as mitochondria and chloroplasts, particularly because they cannot deliver the guide RNAs required to activate the most widely used CRISPR systems to these organelles. Mitochondria and chloroplasts encode several essential genes required for photosynthesis and cellular respiration. Methods or tools for correcting genes in these organelles would be useful for studying the function of these genes, treating mitochondrial genetic diseases, and improving crop productivity and traits.

[0004] C-to-T base editing is a gene editing technology that selectively converts cytosine (C) to thymine (T). Its core lies in the enzyme activity of cytosine deaminase. Cytosine deaminase deamidates cytosine within a DNA chain to uracil, which is then recognized as thymine during DNA replication or base excision repair (BER), ultimately leading to a C-to-T base conversion.

[0005] Currently, widely used cytosine deaminases are mainly mammalian APOBEC1 (apolipoprotein B mRNA editing enzyme catalytic subunit 1) or AID (activation-induced cytidine deaminase), which are enzymes that originally deaminate cytosine bases in RNA or single-stranded DNA to uracil. These deaminases are usually fused to DNA-binding proteins such as Cas9 nickase (nCas9), transcription activator-like effector (TALE), or zinc finger protein (ZFP), and after specifically binding to the desired genomic site, the cytosine base at that position is selectively converted to uracil.

[0006] Meanwhile, double-stranded DNA-specific cytosine deaminase, unlike APOBEC1 or AID, is an enzyme that can directly deaminate cytosine bases on double-stranded DNA, providing a novel platform that can expand the scope of existing base editing systems (see WO 2022 / 060185). Typically, double-stranded DNA-specific cytosine deaminase is split into two halves to reduce cytotoxicity, each fused to a DNA-binding protein. Activity is induced through dimerization on the target DNA. This enables C-to-T base editing with high specificity and efficiency, even in double-stranded DNA.

[0007] However, uracil, generated by cytosine deamination, can be recognized and cleaved by uracil DNA glycosylase (UDG), one of the cells' inherent DNA damage detection and repair mechanisms. UDG detects the abnormal uracil base within the DNA chain and cleaves the N-glycosidic bond between the base and the sugar-phosphate backbone, creating an abasic (AP) site, thereby inducing DNA repair pathways. This process can reduce the efficiency of base editing by removing uracil, an intermediate for base correction.

[0008] To prevent this, C-to-T base editing systems are typically combined with a uracil DNA glycosylase inhibitor (UGI). UGI potently inhibits the activity of UDG, ensuring that deamidated uracil is not removed from DNA and remains stable. Consequently, the combined use of UGI enhances the efficiency of C-to-T base editing, and UGI is now recognized as an essential component of the system.

[0009] The inventors of the present invention surprisingly confirmed that when an inactive UDG (dead UDG, dUDG), which can bind to DNA containing uracil but does not have a catalytic activity to remove uracil, was used together with the base correction system, C-to-T base correction was achieved with considerable efficiency.

[0010] UDG is a well-known enzyme that naturally detects damaged DNA and selectively removes only uracil bases within it. Meanwhile, WO 2024 / 186134 reported that adding UDG to an A-to-G base editing editor consisting of a DNA binding protein, cytosine deaminase, and adenine deaminase suppressed the editing of cytosine bases and selectively allowed A-to-G editing to occur.

[0011] In light of this conventional technical understanding, the discovery of the present invention that C-to-T base correction is possible and more efficiently achieved by using inactive UDG (dUDG) is a completely unexpected result.

[0012] Accordingly, one aspect of the present invention relates to a novel use of inactive UDG for C-to-T base correction.

[0013] Another aspect of the present invention relates to a DNA base editor for C-to-T base correction comprising an inactive UDG. Specifically, the base editor comprises a DNA binding protein, a cytosine deaminase, and dUDG.

[0014] Another aspect of the present invention relates to a polynucleotide(s) encoding a DNA base editor comprising an inactive UDG.

[0015] Another aspect of the present invention relates to a DNA base editing composition comprising a DNA base editor comprising an inactive UDG or a polynucleotide(s) encoding the same.

[0016] Another aspect of the present invention relates to a C-to-T base editing method using a DNA base editor comprising an inactive UDG. The method may include introducing the base editor into a cell containing target DNA or expressing the DNA base editor within a cell containing target DNA.

[0017] Another aspect of the present invention relates to a C-to-T base correction method comprising introducing a DNA base editor or base correction composition comprising an inactive UDG, or a carrier comprising the same, into a cell comprising a target DNA for base correction, or expressing the same in a cell comprising the target DNA.

[0018] Another aspect of the present invention relates to a cell or organism in which C-to-T base correction has been performed using the above-described base correction method. In some embodiments, the present invention relates to a plant cell or protoplast in which C-to-T base correction has been performed, a plant grown or cultured therefrom (including its progeny or clone) or part thereof, or a seed obtained from such a plant.

[0019] The inventors of the present invention have confirmed that even when dUDG is used instead of UGI in a C-to-T base correction system, C-to-T correction efficiency is equivalent to or greater than that achieved using UGI. This result is completely different from previous expectations, suggesting the potential for implementing a new C-to-T correction platform with high specificity and efficiency.

[0020] Figure 1 is a schematic diagram of the proofreading principle of a C-to-T base editor containing UGI. Cytosine deaminase (DddA) tox ) when cytosine (C) is deaminated to uracil (U), the uracil can generally be removed by UDG. However, since UGI included in the editor inhibits the activity of UDG, uracil (U) is maintained on DNA and is subsequently corrected to thymine (T) through DNA repair and replication processes. Figure 2 illustrates the C-to-T base correction principle of the base editor including dUDG according to the present invention. Without wishing to be bound by theory, it is believed that cytosine (C) is removed by DddA. tox After being converted to uracil (U) by dUDG, the uracil can be removed by normal UDG, but dUDG prevents normal UDG from accessing or excising uracil, so that uracil is not removed and is maintained, and as a result, correction to thymine (T) can be made possible by subsequent DNA repair and replication processes.

[0021] Figures 3 and 4 show the results of C-to-T base correction performed on the psaA gene in chloroplast DNA of Arabidopsis thaliana using a base editor containing dUDG according to the present invention. Figure 3 is a schematic diagram of the base editor used, in which CTS represents a chloroplast transit signal, NTD represents the N-terminal domain of TALE protein, and CTD represents the C-terminal domain of TALE protein. 1397N and 1397C represent the cytosine deaminase (DddA) used. tox ) is a fragment of TALE, and “Right TALE repeats” and “Left TALE repeats” represent each TALE array. The underlined base sequence indicates the target site to which the TALE protein binds. Figure 4 shows that when dUDG is used, the C-to-T correction efficiency at the base pair at the G2 position is superior to that when UGI is used. Figure 4 shows the results using the 1397N / 1397C fragment as a cytosine deaminase and DddA, respectively. tox This is the result of using GSVG, a full-length non-toxic mutant of .

[0022] Figures 5 to 9 illustrate the results of C-to-T base correction performed on the 16S rRNA gene (site 1) in Arabidopsis thaliana chloroplasts using a base editor including dUDG according to the present invention. Figure 5 schematically illustrates the DNA sequence to which the base editor used binds and the spacer region where base correction occurs. Figure 6 shows that the C-to-T base correction efficiency at the C8 position is higher when dUDG is used than when UGI is used. Figures 7 and 8 show the results of comparing the correction efficiency at the C11 position, and in both figures, the correction efficiency is higher when dUDG is used than when UGI is used. The TALE proteins used here are different in Figures 7 and 8. Figure 9 shows the results of comparing the correction efficiency at the C14 position, and the C-to-T base correction efficiency is higher when dUDG is used than when UGI is used.

[0023] Figures 10 and 11 show the results of C-to-T base correction performed on the 16S rRNA gene (site 2) in Arabidopsis thaliana chloroplasts using a base editor containing dUDG according to the present invention. Figure 10 is a diagram illustrating the DNA sequence to which the base editor used binds and the spacer region where base correction occurs. Figure 11 shows the results of GSGV (DddA) when 1397N / 1397C splitters were used as cytosine deaminase, respectively. tox (full-length non-toxic mutant of dUDG) was used. In both cases, the efficiency of C-to-T base correction at the C9 position was significantly higher when dUDG was used than when UGI was used.

[0024] Figures 12 and 13 show the results of C-to-T base correction performed on the atpB gene in Arabidopsis thaliana chloroplasts using a base editor containing dUDG according to the present invention. Figure 12 is a diagram illustrating the DNA sequence to which the base editor used binds and the spacer region where base correction occurs. Figure 13 shows the result of GSVG (DddA) by cytosine deaminase. tox The results for the case where dUDG was used (a full-length non-toxic mutant) were shown, showing that the C-to-T base correction efficiency at the G7 position was significantly higher when dUDG was used than when UGI was used.

[0025] Figures 14 to 16 illustrate the results of C-to-T base correction performed on the atp1 gene in Arabidopsis thaliana mitochondria using a base editor comprising dUDG according to the present invention. Figure 14 schematically illustrates the DNA sequence to which the base editor used binds and the spacer region where base correction occurs. Figures 15 and 16 compare the base correction results at positions C3 and C4 and positions C7 and C10, respectively. In all positions, the use of dUDG showed significant C-to-T base correction efficiency.

[0026] Figures 17 and 18 show the results of C-to-T base correction in nuclear DNA of Arabidopsis thaliana using a base editor including dUDG according to the present invention. In the case of experiments targeting the PDS3 gene (Figure 17) or the CESA3 gene (Figure 18), significant C-to-T base correction results were observed in both cases.

[0027] Figures 19 and 20 are schematic diagrams of the Golden Gate technique used to produce a polynucleotide encoding a base editor according to the present invention. Figure 19 is a schematic diagram of the process of assembling a TALE array included in a base editor according to the present invention using the Golden Gate assembly method. Figure 20 is a schematic diagram of the process of constructing a plasmid encoding a base editor according to the present invention using the Golden Gate assembly method.

[0028] Figure 21 is a comparative alignment of the amino acid sequences of human UNG1 (UDG1) and UNG2 (UDG2), and the amino acid residues indicated by arrows indicate exemplary amino acid mutation positions that induce dUDG (inactive UDG).

[0029] Figure 22 shows an alignment of the amino acid sequences of UDG enzymes from several different species (Arabidopsis thaliana, Homo sapiens, Saccharomyces cerevisiae, and E. coli).

[0030] Figure 23 shows the amino acid sequences of motifs A and B of several different UDG families.

[0031] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. Generally, the terms used herein are those well known and commonly used in the art.

[0032] The embodiments described in this specification and the configurations depicted in the drawings are only one embodiment of how the present invention is realized and do not fully represent the technical idea of ​​the present invention. Therefore, it should be understood that there may be various equivalents, modifications, and applicable examples that can replace them at the time of this application. In addition, the various aspects and embodiments described in this specification can be applied to other aspects and embodiments, and all combinations of the various elements described in the present invention fall within the scope of the present invention, and the scope of the present invention cannot be considered limited by the specific description described below.

[0033] In this specification, the use of the singular includes the plural unless specifically stated otherwise. As used herein, it should be noted that the singular form includes plural referents unless the context clearly dictates otherwise. In this application, the use of "or" means "and / or" unless otherwise stated.

[0034] The term “comprising” as used herein, unless otherwise specified, is understood to be an open-ended expression that essentially includes the described components, ingredients, steps, etc., but does not exclude the presence of other components, ingredients, steps, etc. Accordingly, the term “comprising” is interpreted to include the more limited meaning of “consisting of” or “consisting essentially of.”

[0035] The expression “for” as used in this specification and claims not only means that the composition or material is designed to be used for a particular use, but also includes the meaning that it has functions or characteristics that are suitable or useful for that use even if it is not actually used for that use.

[0036] The terms "correction," "editing," and "editing" as used herein are used interchangeably and refer to a method of altering a nucleic acid sequence by selectively modifying a specific genomic target. Such specific genomic target includes, but is not limited to, a gene, a promoter, an open reading frame, or any nucleic acid sequence.

[0037] As used herein, the terms "base editor," "base editing system," "base editor system," and "base correction system" are used interchangeably and refer to a material or composition suitable for modifying a nucleic acid sequence by inducing selective mutations in a genomic target. As used herein, the terms "base editor," "base editing system," "base editor system," or "base correction system" may be in the form of a polypeptide (which may be a fusion protein) or a polynucleotide, or a combination thereof, depending on the context, and may be a composition comprising one or more polypeptides (which may be fusion proteins) or polynucleotides, or a combination thereof.

[0038] As used herein, the term "conservative amino acid substitution" refers to the replacement of some amino acids with amino acids of different properties while maintaining structural or functional similarity within a protein. Specifically, it refers to substitutions between amino acids with similar physicochemical properties (e.g., charge, size, hydrophobicity, polarity, etc.), and includes substitutions that substantially maintain the structural stability or biological function of the protein.

[0039] For example, substitutions within the following amino acid groups may be conservative substitutions:

[0040] Hydrophobic amino acid group: Ala, Val, Leu, Ile, Met

[0041] Polar uncharged amino acid group: Ser, Thr, Gln, Asn

[0042] Acidic amino acid group: Asp, Glu

[0043] Basic amino acid group: Lys, Arg, His

[0044] Aromatic amino acid group: Phe, Tyr, Trp

[0045] Such determination of substitution can be performed based on the standard amino acid classification that takes into account the charge, polarity, hydrophobicity, structural similarity, etc. of the amino acid, and a person skilled in the art can objectively determine whether the substitution is conservative by utilizing sequence alignment tools such as BLAST and Clustal Omega and conservation matrices (BLOSUM, PAM, etc.). When a specific amino acid sequence is described in this specification, it is interpreted that a variant in which one or more amino acids in the sequence are substituted with another amino acid corresponding to the conservative substitution is also included in the technical scope of the present invention.

[0046] The term “sequence” in this specification may be interpreted as a nucleic acid (or polynucleotide) molecule or a protein (or polypeptide) molecule having a given sequence, depending on the context.

[0047] The term "other amino acid" as used herein means an amino acid selected from among alanine, isoleucine, leucine, methionine, phenylalanine, proline, tryptophan, valine, aspartic acid, cysteine, glutamine, glycine, serine, threonine, tyrosine, aspartic acid, glutamic acid, arginine, histidine, lysine, and all known variants of the above amino acids, excluding the amino acid that the wild-type protein originally has at the mutation position.

[0048] The term “sequence identity” or “sequence homology” as used herein means the number of residues present at the same position when two amino acid sequences or nucleic acid sequences are aligned, expressed as a percentage of the total length of the sequences. When it is said herein that a specific sequence “has at least X% sequence identity,” X can be, for example, 85%, 90%, 95%, 98%, or 99%. Sequence identity is typically calculated using BLAST (Basic Local Alignment Search Tool), ClustalW, EMBOSS, or other known sequence alignment algorithms, and is based on default parameters. For example, when aligning amino acid sequences using BLASTP, the identity value calculated using the BLOSUM62 matrix and gap penalty as default values ​​can be used as a basis. In addition, in this specification, “sequence identity” or “sequence homology” may include a value calculated according to an optimized global alignment or local alignment that takes into account insertions, deletions, substitutions, etc. during sequence alignment, and is also used as a standard for explaining the scope of functional equivalents that can maintain the technical effect of the invention.

[0049] The term "homolog" as used herein refers to a protein or nucleic acid that has homology or similarity to a specific protein or gene sequence and exhibits functionally similar biological activity. Homologs may perform the same or similar function, but may be of different species or may have some variation in the amino acid or base sequence. Homologs may be naturally occurring or artificially modified, and are considered to provide the same technical effect within the scope of the present invention.

[0050] The term "ortholog" as used herein refers to a gene or protein derived from two or more species that share a common ancestor, and thus corresponds to a corresponding gene or protein found in different species but with the same evolutionary origin. Generally, orthologs have a high degree of sequence homology between different species and are known to perform similar biological functions. As used herein, an ortholog of a specific protein may include a variant that performs essentially the same function as the original protein, even if a portion of the sequence contains amino acid substitutions, insertions, or deletions.

[0051] The term "functional variant" as used herein refers to a protein that substantially maintains the basic biological function of a polypeptide (e.g., a protein) having a specific amino acid sequence, or has an activity essentially similar to that of the protein, despite having one or more amino acids in the entire sequence conservatively or non-conservatively substituted, or modified, such as insertion, deletion, or substitution. For example, some differences in the sequence may be considered functional variants of the protein described herein, as long as the protein performs the effective function intended in the present invention, such as substrate recognition, catalytic activity, binding ability, or specificity for a target molecule. Such functional variants may be generated by spontaneous mutation, evolutionary modification, induced mutation, or genetic engineering methods.

[0052] As used herein, the terms "target" or "target site" refer to a pre-identified nucleic acid sequence of any composition and / or length. Such target sites include, but are not limited to, genes, promoters, or any nucleic acid sequence.

[0053] The term "fusion protein" as used herein refers to a protein in which two or more different protein (polypeptide) sequences or functional domains are combined into a single continuous polypeptide chain. Such a fusion protein may retain the original biological function of each component or be endowed with new functional properties, and may include a linker sequence between the components. When designating components of a fusion protein herein, unless otherwise specified, the left-to-right direction refers to the N-terminus to the C-terminus, respectively. Additionally, the linker sequence used may not be explicitly indicated.

[0054] The term “expression” as used herein refers to the process by which a polynucleotide (e.g., DNA or mRNA) encoding a DNA base editor or a component thereof is introduced into a cell, and the base editor or a component protein thereof is produced through the cell’s transcription and / or translation mechanisms. The expression may be transient, or stable when integrated into the genome of a target cell. The expression product may be a single protein, or may be a fusion protein in which multiple functional domains are fused. In addition, the expression may be performed in the cytoplasm or organelles (e.g., chloroplasts, mitochondria, nuclei, etc.), and for this purpose, the expression product may additionally include an organelle targeting sequence (MTS, CTS, etc.), a nuclear export signal (NLS), or an extranuclear export signal (NES). The expression level and location may vary depending on the sequence of the polynucleotide being introduced, the promoter, codon optimization, target cell type, introduction method, etc., and a person skilled in the art can appropriately adjust it according to the purpose.

[0055] The term "monomeric base editor" as used herein refers to a base editor comprising a DNA binding protein and a deaminase as a single fusion protein. The monomeric base editor does not exclude the presence of additional polypeptide or polynucleotide components other than the single fusion protein.

[0056] The term "dimeric base editor" as used herein refers to a base editor comprising a DNA binding protein and a deaminase as two fusion proteins. Among the two fusion proteins, the fusion protein that binds to a DNA sequence located 5' upstream of the spacer region may be referred to as a "first fusion protein" or a "left fusion protein," and the fusion protein that binds to a DNA sequence located 3' downstream of the spacer region may be referred to as a "second fusion protein" or a "right fusion protein." Similarly, the DNA binding protein included in the first fusion protein may be expressed by the modifier "first" or "left," and the DNA binding protein included in the second fusion protein may be expressed by the modifier "second" or "right." The dimeric base editor does not exclude the presence of additional polypeptide or polynucleotide components other than the two fusion proteins.

[0057] In some embodiments, the base editor may be a monomeric base editor, and in other embodiments, it may be a dimeric base editor comprising two fusion proteins.

[0058] In this specification, binding of a fusion protein to a given nucleotide sequence means that the DNA binding protein included in the fusion protein recognizes and binds to the nucleotide sequence.

[0059] The term “CRISPR-associated nuclease” as used herein, also called Cas, generally refers to a protein that is a component of the CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats) system, binds to a guide RNA, specifically binds to a target nucleotide sequence, and has the activity of cleaving or regulating the sequence. In this specification, “CRISPR-associated nuclease” and “Cas” are used interchangeably.

[0060] The term "organelle DNA" as used herein refers to the genetic material present in organelles other than the nucleus within a eukaryotic cell, including mitochondrial DNA or plastid DNA.

[0061] The term “plastid” as used herein refers to an intracellular organelle present in plant cells and some algal cells, including chloroplasts, chromoplasts, and leucoplasts.

[0062] 1. dUDG

[0063] One aspect of the present invention relates to an inactive UDG for use in C-to-T base correction.

[0064] The term "inactive UDG" as used herein has the same meaning as "dUDG" and refers to a mutant or inactive form of UDG in which the catalytic activity is eliminated, and may also be referred to as "dead," "catalytically inactive," "inactivated," or "catalytically dead" UDG. Such dUDG can recognize or bind to uracil residues on DNA in the same manner as wild-type UDG, but does not have the glycosylation enzyme activity to remove uracil bases.

[0065] The UDG used in the present invention is not limited to a specific amino acid sequence, structure, origin, or biological species, and is generally understood as a concept including all proteins or functional variants thereof having an activity of recognizing and removing uracil bases in DNA. For example, human UDG (human UNG1, NCBI Reference Sequence: NP_003353; human UNG2, NCBI Reference Sequence: NP_550433), Escherichia coli UDG (E. coli UDG, NCBI Reference Sequence: NP_417075), Arabidopsis thaliana UDG (Arabidopsis thaliana UDG, NCBI Reference Sequence: NP_188493), archaeal UDG, or phage-derived UDG have different sequences and structures, but all perform a common biological function of removing uracil, and are therefore included in the "UDG" referred to herein. UDG is also called UNG (uracil-N-glycosylase).

[0066] The structure and function of UDG are already well known (see Schormann, N., et al. "Uracil-DNA glycosylases―structural and functional perspectives on an essential family of DNA repair enzymes." Protein science 23.12 (2014): 1667-1685., Cordoba-Canero, D., et al. "Arabidopsis uracil DNA glycosylase (UNG) is required for base excision repair of uracil and increases plant sensitivity to 5-fluorouracil." Journal of Biological Chemistry 285.10 (2010): 7475-7483.).

[0067] Inactive UDG (dUDG), which retains the DNA binding ability of UDG but has no catalytic activity, has already been used for research on the structure and function of UDG, and related concepts are also widely known in the literature.

[0068] As used herein, dUDG is not limited to a specific amino acid sequence, three-dimensional structure, origin, or biological species, and is understood as a concept that includes all UDGs that can recognize and bind to uracil bases on DNA but do not have glycosylation activity to remove uracil, or truncated forms thereof (e.g., truncated forms in which the amino acid sequence corresponding to the N-terminal 1-106 amino acid sequence of Arabidopsis thaliana UDG is removed), or functional variants thereof.

[0069] Based on the known knowledge regarding the mechanism of action and active site of UDG, a person skilled in the art can implement UDG with eliminated catalytic activity by substituting, removing or introducing specific amino acid residues, and can easily practice the present invention without excessive experimental burden when applying such dUDG to the technical configuration of the present invention.

[0070] Although not bound by theory, dUDG may play a role in protecting uracil deaminated by cytosine deaminase from normal (wild-type) UDG, thereby preventing access and excision of normal UDG and increasing the retention time of uracil, ultimately increasing the efficiency of C-to-T correction (see Figures 1 and 2).

[0071] UDG is known to have two functional modules or subdomains, motif A and motif B, involved in catalytic activity.

[0072] Figure 21 shows a comparative alignment of the amino acid sequences of human UNG1 (UDG1) and UNG2 (UDG2), with the amino acid residues indicated by arrows indicating exemplary amino acid mutation positions leading to dUDG (inactive UDG). Figures 22 and 23 show the amino acid sequences of motifs A and B of UDG enzymes from several different families and species.

[0073] In some embodiments, the dUDG used in the present invention may be a variant comprising an amino acid mutation in motif A or motif B that is not present in the wild-type UDG. Such embodiments also include cases where both motif A and motif B comprise such amino acid mutations.

[0074] In some embodiments, the dUDG used in the present invention is a variant having an amino acid mutation in an amino acid sequence corresponding to the amino acid sequence KTVIIGQDPYHGPGQAMGLSF contained in UDG of Arabidopsis thaliana or an amino acid sequence corresponding to HPSGLSA contained in UDG of Arabidopsis thaliana. This includes variants having amino acid mutations in both the amino acid sequence corresponding to the amino acid sequence KTVIIGQDPYHGPGQAMGLSF contained in UDG of Arabidopsis thaliana and the amino acid sequence corresponding to HPSGLSA contained in UDG of Arabidopsis thaliana.

[0075] As used herein, the term "corresponding amino acid sequence" refers to a sequence within a functional motif that is conserved across different species, protein families, orthologous proteins, or artificially evolved variants, and that is recognized as structural and / or functionally equivalent. Such sequences are typically located within conserved regions, and functional homology can be confirmed through protein alignment analysis (e.g., Clustal Omega, MUSCLE, etc.). Therefore, amino acid substitutions present at such positions (e.g., conservative substitutions or substitutions with similar biochemical properties) are understood to be within the scope of the present invention.

[0076] In some embodiments, the dUDG according to the present invention is a variant comprising an amino acid corresponding to an aspartic acid (D), tyrosine (Y), phenylalanine (F) or glutamine (Q) residue in the amino acid sequence KTVIIGQDPYHGPGQAMGLSF included in Arabidopsis thaliana UDG, which is mutated to another amino acid, or an amino acid corresponding to a histidine (H) or leucine (L) residue in the amino acid sequence HPSGLSA included in Arabidopsis thaliana UDG, which is mutated to another amino acid. This includes a variant in which an amino acid corresponding to an aspartic acid (D), tyrosine (Y), phenylalanine (F) or glutamine (Q) residue in the amino acid sequence KTVIIGQDPYHGPGQAMGLSF contained in Arabidopsis thaliana UDG is mutated to another amino acid, and an amino acid corresponding to a histidine (H) or leucine (L) residue in the amino acid sequence HPSGLSA contained in Arabidopsis thaliana UDG is also mutated to another amino acid.

[0077] The term "corresponding amino acid" as used herein refers to a structurally and / or functionally equivalent amino acid present at a position where structural and / or functional equivalence is recognized, as a corresponding amino acid within a functional motif conserved between different species, protein families, orthologous proteins, or artificially evolved variants. Whether an amino acid is a corresponding amino acid can be confirmed through protein alignment analysis (e.g., Clustal Omega, MUSCLE, etc.).

[0078] In some embodiments, the dUDG according to the present invention is a variant comprising an amino acid corresponding to an aspartic acid (D) residue in the amino acid sequence KTVIIGQDPYHGPGQAMGLSF included in Arabidopsis thaliana UDG, which is mutated to another amino acid, or an amino acid corresponding to a histidine (H) residue in the amino acid sequence HPSGLSA included in Arabidopsis thaliana UDG, which is mutated to another amino acid. This includes variants comprising an amino acid corresponding to an aspartic acid (D) residue in the amino acid sequence KTVIIGQDPYHGPGQAMGLSF included in Arabidopsis thaliana UDG, which is mutated to another amino acid, and an amino acid corresponding to a histidine (H) residue in the amino acid sequence HPSGLSA included in Arabidopsis thaliana UDG, which is also mutated to another amino acid.

[0079] In some embodiments, the dUDG according to the present invention is a variant comprising an amino acid corresponding to an aspartic acid (D) residue in the amino acid sequence KTVIIGQDPYHGPGQAMGLSF included in Arabidopsis thaliana UDG, which is mutated to asparagine (N), alanine (A), or a conservative amino acid substitution thereof, or an amino acid corresponding to a histidine (H) residue in the amino acid sequence HPSGLSA included in Arabidopsis thaliana UDG, which is mutated to alanine (A) or a conservative substitution thereof. This includes a variant in which an amino acid corresponding to an aspartic acid (D) residue in the amino acid sequence KTVIIGQDPYHGPGQAMGLSF contained in Arabidopsis thaliana UDG is mutated to asparagine (N), alanine (A), or a conservative amino acid substitution thereof, and an amino acid corresponding to a histidine (H) residue in the amino acid sequence HPSGLSA contained in Arabidopsis thaliana UDG is also mutated to alanine (A) or a conservative substitution thereof.

[0080] In some embodiments, the dUDG according to the present invention is derived from UDG of Arabidopsis thaliana.

[0081] In some embodiments, the dUDG according to the present invention has the amino acid sequence of UDG of Arabidopsis thaliana, but has D173N or H295A, or both of these amino acid mutations.

[0082] In some embodiments, the dUDG according to the present invention comprises an amino acid sequence of SEQ ID NO: 1 or an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, or at least 90% sequence identity thereto.

[0083] In some embodiments, the dUDG according to the present invention has the amino acid sequence of human UNG1, but has D145N or H268L, or both of these amino acid mutations.

[0084] In some embodiments, the dUDG according to the present invention has the amino acid sequence of human UNG2, but has D154N or H277A, or both of these amino acid mutations.

[0085] In some embodiments, the dUDG according to the present invention has the amino acid sequence of E. coli UNG, but has D64N or H187A, or both of these amino acid mutations.

[0086] In some embodiments, the dUDG according to the present invention may be used in a fused form with a programmable DNA binding protein and a cytosine deaminase, or may be used independently without being fused with them. When used in a fused form, it is preferably located at the C-terminus of the programmable DNA binding protein and the cytosine deaminase.

[0087] When dUDG was used in conjunction with a programmable DNA binding protein and cytosine deaminase, C-to-T correction was confirmed. Furthermore, the resulting C-to-T correction efficiency was confirmed to be equivalent to that achieved using UGI. This provides the advantageous technical benefit of achieving effective C-to-T correction without the use of UGI.

[0088] Furthermore, the use of dUDG for C-to-T proofreading of chloroplast DNA demonstrated a significant improvement in proofreading efficiency. This demonstrates that higher C-to-T proofreading efficiency can be achieved with chloroplast DNA than with UGI, providing technical advantages.

[0089] 2. DNA base editor

[0090] One aspect of the present invention relates to a DNA base editor for correcting cytosine (C) bases to thymine (T) bases, comprising a DNA binding protein, cytosine deaminase, and dUDG. The cytosine deaminase exists in a single full-length form or in the form of two fragments, and when present in the form of two fragments, the cytosine deaminase activity is exerted through dimerization.

[0091] The term "dimerization," as used herein, refers to the process by which two protein fragments are physically placed in close proximity to each other and form a functional protein complex through noncovalent interactions. For example, the cleaved form of cytosine deaminase exerts its overall enzymatic activity by dimerizing each fragment.

[0092] These DNA base editors have the ability to selectively correct specific DNA bases, and are particularly useful for converting cytosine (C) bases to thymine (T) (C-to-T).

[0093] The DNA base editor of the present invention may comprise one or more fusion proteins, and in some embodiments, two fusion proteins. These fusion proteins comprise one or more DNA binding proteins and a cytosine deaminase, and may comprise dUDG. A linker sequence may be included between the components included in the fusion proteins.

[0094] A. DNA binding protein

[0095] A DNA base editor according to the present invention comprises one or more DNA binding proteins.

[0096] The "DNA binding protein" used in the DNA base editor according to the present invention means a protein that can recognize a specific base sequence and selectively bind to the sequence, and the binding specificity has a programmable characteristic according to the target sequence. In that sense, the DNA binding protein included in the DNA base editor according to the present invention is any programmable DNA binding protein suitable for use in DNA base editing. Those skilled in the art are well aware of the types of programmable DNA binding proteins suitable for use in DNA base editing. Examples of such DNA-binding proteins include zinc finger proteins and TALE (transcription activator-like effector, TALE) proteins, which can be designed to recognize specific base sequences through modular repeat sequences, dead Cas proteins (e.g., dCas9, dCas12a) that are targeted by guide RNA, and other artificially designed DNA recognition domains. These DNA-binding proteins can be fused to enzymatic proteins (e.g., nickases, cytosine deaminase, or adenine deaminase) to induce a desired biochemical reaction at a specific location in the genome.

[0097] The DNA binding protein used in the DNA base editor according to the present invention is not limited to a specific amino acid sequence or structure, and generally includes a programmable protein or functional variant thereof that has the ability to recognize a specific base sequence and selectively bind to that sequence. The DNA binding protein may be naturally occurring, a variant thereof, or an artificially designed protein. The DNA binding protein that can be used in the DNA base editor according to the present invention is understood to encompass all proteins capable of selectively binding to target DNA to achieve the purpose of the present invention, regardless of the specific sequence or origin.

[0098] The DNA base editor according to the present invention comprises one or more DNA binding proteins. The DNA binding proteins are proteins capable of selectively binding to a specific DNA sequence, enabling specific recognition of the target site for base editing.

[0099] In some embodiments, the DNA binding protein may be selected from the group consisting of a zinc finger protein, a TALE protein, a CRISPR-associated nuclease, or a combination thereof. Reference may be made to prior disclosures regarding zinc finger proteins, TALE proteins, and CRISPR-associated nucleases, including WO 2022 / 060185 and WO2022 / 017745, which are incorporated by reference herein in their entirety.

[0100] The above "zinc finger protein (ZFP)" generally refers to a protein or protein domain that forms a stabilized structure through the binding of zinc ions (Zn²) and has the function of binding to a specific DNA sequence. Such zinc finger proteins include one or more "zinc finger (ZF)" structures. Zinc finger proteins have sequence specificity that allows them to bind to a DNA sequence consisting of a specific 3-4 base pair, and by designing them by continuously combining multiple zinc finger domains, they can have high specificity and binding affinity for a long target DNA sequence.

[0101] Zinc finger proteins that can be used in some embodiments of the present invention include naturally occurring proteins or artificial recombinants, mutants, functional variants, or variants with improved specificity derived therefrom, and are not limited in origin or sequence composition, as long as they can specifically bind to a desired target DNA sequence. Those skilled in the art can design and produce a desired zinc finger protein using previously disclosed ZFP libraries, genetic engineering methods, and techniques for analyzing DNA-binding specificity.

[0102] Zinc finger proteins have a relatively small molecular weight and can bind to target sequences with only their pure protein structure without relying on an RNA guide sequence, so they have the advantage of being easy to apply even in delivery systems with vector size limitations or in environments where RNA is unstable.

[0103] The above "TALE protein" is generally based on a transcription activator-like factor derived from the plant pathogenic bacteria Xanthomonas genus, and refers to a DNA binding protein with sequence specificity that can bind to a specific DNA sequence. The TALE protein is composed of a series of repeat modules, each module consisting of about 34 amino acids, of which two amino acid residues at positions 12 and 13 (so-called RVD, repeat-variable diresidue) determine the binding specificity for a single DNA base. By designing a combination of these modules, a TALE protein with sequence specificity tailored to a desired target DNA sequence can be generated. As used herein, the TALE-repeat modules may be referred to as a "TALE array", a "TALE repeat sequence", etc., and the expression "TALE protein" means a configuration in which an N-terminal domain and a C-terminal domain (which may include a half domain) are included on both sides of the TALE array, respectively.

[0104] The term "N-terminal domain (NTD)" as used herein refers to a region located at the amino terminus of a TALE protein, which includes a sequence that contributes to the alignment of DNA binding sites or maintenance of protein stability. For example, in a TALE protein derived from Xanthomonas, the N-terminal domain may be composed of a sequence approximately between amino acids 1 and 150. However, such sequences are merely examples, and the "N-terminal domain" in the present specification also includes variants, homologs, or artificially designed sequences of the sequence as long as the sequence can perform the above function. A person skilled in the art can select or design a suitable N-terminal domain sequence based on the structure and function of a known TALE protein.

[0105] The term "C-terminal domain (CTD)" used herein refers to a region located at the carboxy terminus of a TALE protein, which includes a sequence that performs protein stability or other regulatory functions. For example, in Xanthomonas TALE, a sequence corresponding to amino acids 800 to 900 may be included. However, the "C-terminal domain" in the present specification is not limited to such sequence, and also includes other biological sequences or artificial sequences that can perform the same or similar functions. Such sequences can be easily selected or designed by a person skilled in the art based on publicly available TALE protein information.

[0106] TALE proteins that can be used in some embodiments of the present invention include naturally occurring TALE sequences or recombinants, mutants, functional variants, or forms with artificially controlled sequence specificity derived therefrom, and are not limited in sequence composition or origin, as long as they can bind to a desired DNA sequence. Those skilled in the art can utilize published TALE libraries and TALE design algorithms to create TALE proteins that specifically bind to various DNA target sequences. For example, Kim, Yongsub, et al. "A library of TAL effector nucleases spanning the human genome." Nature biotechnology 31.3 (2013): 251-258.

[0107] TALE proteins have the advantage of not requiring guide RNA, precisely recognizing target DNA sequences by directly assembling repeat modules at the protein level, and being free from constraints on PAM sequences. Therefore, TALE proteins are particularly advantageous in complex genomic environments or situations requiring flexible targeting.

[0108] The above "CRISPR-associated nuclease" is also called "Cas protein" and generally refers to a protein having nuclease activity capable of cleaving DNA or RNA derived from the CRISPR (clustered regularly interspaced short palindromic repeats)-Cas system, which is an acquired immune system of bacteria or archaea. These Cas proteins generally form a complex with a guide RNA, recognize a target nucleic acid sequence complementary to the base sequence of the guide RNA, and then induce cleavage (nicking or double-strand break) at the corresponding site. Representative examples include Cas9 (e.g., Streptococcus pyogenesCas9), Cas12a (Cpf1), Cas12b, Cas13, and Cas14.

[0109] CRISPR-associated nucleases that may be used in some embodiments of the present invention may include naturally occurring proteins, functional variants, conservative amino acid substitutions, truncated forms, or variants in which the enzymatic activity is altered or eliminated, and also include inactive forms (dead Cas or dCas), nickase forms (nCas), or forms that include fusions with various functional domains (e.g., deaminase, transcription factor, etc.).

[0110] When the base editor according to the present invention has two fusion proteins, the two fusion proteins each comprise a DNA binding protein, which may be the same or different from each other. That is, one of the two fusion proteins may be, for example, a zinc finger protein, a TALE protein, or a CRISPR-associated nuclease, and the other may independently be, for example, a zinc finger protein, a TALE protein, or a CRISPR-associated nuclease.

[0111] B. Cytosine deaminase

[0112] The DNA base editor according to the present invention comprises cytosine deaminase.

[0113] The "cytosine deaminase" included in the DNA base editor according to the present invention generally refers to an enzyme that catalyzes a deamination reaction that converts the cytosine base in DNA to uracil. This enzyme converts cytosine to uracil by removing the amino group (-NH2) of cytosine, thereby inducing a C:G → T:A conversion in the corresponding base pair.

[0114] The cytosine deaminase that can be used in some embodiments of the present invention is not limited to a specific amino acid sequence, structural characteristics, biological species of origin, or name, and includes all wild-type proteins, artificial or evolutionary modifications, functional variants, orthologs, conservative amino acid substitutions, etc., as long as the enzyme has a biological activity that can act on a cytosine base on DNA to induce a deamination reaction. Those skilled in the art can easily access numerous literatures and public databases (e.g., GenBank, UniProt, REBASE, etc.) that are already known regarding the sequence, structure, and function of such enzymes, and can implement a cytosine deaminase for implementing selective DNA editing according to a specific purpose.

[0115] In some embodiments, the cytosine deaminase can be an apolipoprotein B editing complex (APOBEC), an activation-induced deaminase (AID), a double strand DNA-specific cytosine deaminase, or a tRNA-specific adenosine deaminase (TadA) with C-to-T proofreading activity.

[0116] B-1. Double-stranded DNA-specific cytosine deaminase

[0117] In some embodiments, the cytosine deaminase is a double-stranded DNA-specific cytosine deaminase.

[0118] As used herein, the term "double-stranded DNA-specific cytosine deaminase" refers to an enzyme that deamines the cytosine base in double-stranded DNA, converting it to uracil. Unlike typical cytosine deaminases, this enzyme has the characteristic of directly recognizing and reacting with cytosine within the normal double-stranded DNA structure, rather than single-stranded DNA, as its substrate.

[0119] An example of such a double-stranded DNA-specific deaminase is DddA, a toxin protein from Burkholderia cenocepacia (DddA tox ) is known to have the activity of selectively recognizing cytosine in double-stranded DNA and converting it to uracil. Since the enzyme can generally cause cytotoxicity, it may be desirable to divide it into two split forms as needed and use it by fusing it with a DNA binding protein.

[0120] The double-stranded DNA-specific cytosine deaminase used in the present invention is not limited to the above examples, and is understood to include all proteins or functional variants thereof that have the activity of deaminating cytosine using double-stranded DNA as a substrate, regardless of its origin, amino acid sequence, structure, or name. Such double-stranded DNA-specific cytosine deaminase and variants thereof have already been described in various documents. For example, WO 2022 / 060185, WO 2022 / 221337, WO 2022 / 155265, WO 2023 / 081855, WO 2023 / 097226, WO 2024 / 112441, WO 2024 / 107263, etc., which are incorporated by reference in their entirety by this application, and contents already known prior to the present application may be cited.

[0121] In some embodiments, the cytosine deaminase is DddA, a cytosine deaminase from Burkholderia cenocepacia. tox Or its variant. DddA tox The amino acid sequence is as follows.

[0122] wild-type DddA tox :

[0123] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNNSNSPKSPTKGGC (SEQ ID NO: 2)

[0124] DddA toxAs variants of DddA, various variants characterized by the context of the cytosine (C) base to be base-edited, the editing efficiency, the improved specificity, etc. are also known, and these include DddA2, DddA3, DddA4, DddA5, DddA6, DddA7, DddA8, DddA9, DddA10, DddA11, etc. With regard to these variants, reference may be made to the contents already known prior to the present application, including, for example, WO 2022 / 221337, which is incorporated by reference in its entirety herein.

[0125] In some embodiments, the present invention provides DddA as a cytosine deaminase. tox Full-length or split forms (e.g., 1333N / 1333C or 1397N / 1397C) may be used, and the split forms may each be contained in separate fusion proteins and function cooperatively. As used herein, "cooperatively functioning" with respect to split forms of cytosine deaminase means that although they exist as separate proteins, they exert cytosine deaminase activity through dimerization.

[0126] DddA tox It can be used alone, but because of its high toxicity and strong enzymatic activity, it is divided into N-terminal and C-terminal fragments (split DddA) for safer and more efficient DNA base editing. tox It is preferable to use it in the form of DddA. In this case, DddA toxIt exists as two splits, and each is incorporated into or linked to an independent fusion protein to function. When the cytosine deaminase used in the present invention is used in the form of a first split and a second split, the first split and the second split do not have deamination activity, and the deaminization activity is exhibited only when the two splits are adjacent to each other. That is, in order for the cytosine deaminase to be used in the form of two splits, the full-length sequence of the cytosine deaminase must be formed when the sequences of the two splits are combined, and a person skilled in the art of base correction technology using cytosine deaminase is well aware of this point.

[0127] DddA tox When present in the form of two fragments, one of the two fragments may include a sequence from the N-terminus to the 33rd, 44th, 54th, 68th, 82nd, 98th or 108th amino acid of the amino acid sequence of SEQ ID NO: 2, and the other of the two fragments may include a sequence from the 34th, 45th, 55th, 69th, 83rd, 99th or 109th amino acid of the amino acid sequence of SEQ ID NO: 2 to the C-terminus.

[0128] In some embodiments, when cytosine deaminase is used in the form of two fragments, it may be in the form of 1333N and 1333C having the following amino acid sequences.

[0129] 1333N:

[0130] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGG (SEQ ID NO: 3)

[0131] 1333C:

[0132] PTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNNSNSPKSPTKGGC (SEQ ID NO: 4)

[0133] In some embodiments, when cytosine deaminase is used in the form of two fragments, it may be in the form of 1397N and 1397C having the following amino acid sequences.

[0134] 1397N:

[0135] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEG (SEQ ID NO: 5)

[0136] 1397C:

[0137] AIPVKRGATGETKVFTGNSNSPKSPTKGGC (SEQ ID NO: 6)

[0138] In some embodiments, when cytosine deaminase is used in the form of two fragments, it may be in the form of DddA11 1333N and DddA11 1333C having the following amino acid sequences.

[0139] DddA11 1333N:

[0140] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFISGG

[0141] DddA11 1333C:

[0142] PTPYPNYVSAGHVEGQSALFMRDNGISEGLVFHNNPKGTCGFCVNMIETLLPENAKMTVVPPEGAIPVKRGATGETKVFIGNNSNSPKSPTKGGC

[0143] In some embodiments, when cytosine deaminase is used in the form of two fragments, it may be in the form of DddA11 1397N and DddA11 1397C having the following amino acid sequences.

[0144] DddA11 1397N:

[0145] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFISGGPTPYPNYVSAGHVEGQSALFMRDNGISEGLVFHNNPKGTCGFCVNMIETLLPENAKMTVVPPEG

[0146] DddA11 1397C:

[0147] AIPVKRGATGETKVFIGNNSNSPKSPTKGGC

[0148] In some embodiments, when cytosine deaminase is used in the form of two fragments, it may be in the form of DddA11 1333N and DddA11 1333C having the following amino acid sequences.

[0149] In some embodiments, when cytosine deaminase is used in the form of two fragments, it may be in the form of FZY2 100N and FZY2 100C having the following amino acid sequences.

[0150] FZY2 100N:

[0151] MSLPEYDGTTTHGVLVLDDGTQIGFTSGNGDPRYTNYRNNGHVEQKSALYMRENNISNATVYHNNTNGTCGYCNTMIATFLPEGATLTVVPPENAVANNS

[0152] FZY2 100C:

[0153] RAIDYVKTYTGTSNDPKISPRYKGN

[0154] In some embodiments, DddA toxOr, when its variant exists in the form of splits of 1333N and 1333C, one or more amino acids selected from the group consisting of positions 3, 5, 10, 11, 13, 14, 15, 16, 17, 18, 19, 28, 30 and 31 of 1333N or one or more amino acids selected from the group consisting of positions 13, 16, 17, 20, 21, 28, 29, 30, 31, 32, 33, 56, 57, 58 and 60 of 1333C may be substituted with another amino acid. DddA tox Or, when its variant exists in the form of splits of 1397N and 1397C, one or more amino acids selected from the group consisting of positions 87, 88, 91, 92, 95, 100, 101, 102 and 103 of 1397N) or one or more amino acids selected from the group consisting of positions 13, 14, 15 and 16 of 1397C may be substituted with another amino acid. In some embodiments, the other amino acid is alanine.

[0155] In some embodiments, DddA tox Or, if its variant exists in the form of splits of 1333N and 1333C, the amino acid at position 56, 57 or 58 of 1333C may be substituted with another amino acid. DddA tox Alternatively, if its variant exists in the form of a split of 1397N and 1397C, the amino acid at position 100, 101 or 102 of 1397N may be substituted with another amino acid. In some embodiments, the other amino acid is alanine.

[0156] Using these mutants, each linked to a DNA binding protein, DddA toxAlternatively, if two fragment pairs derived from its variant fail to bind to DNA, they may not function properly, resulting in highly efficient and precise C-to-T editing without causing undesirable off-target C-to-T editing. For this purpose, reference may be made to prior disclosures, including, for example, WO 2022 / 060185, which is incorporated herein by reference in its entirety.

[0157] In some embodiments, the cytosine deaminase used in the present invention may be used in a full-length form, and the full-length cytosine deaminase used at this time (e.g., DddA) tox ) are amino acid sequences that have been modified to have no or only low toxicity.

[0158] DddA tox The C-terminus of DNA has a specific concentration of positively charged amino acids. Since DNA is negatively charged, it binds to the positively charged amino acids of proteins. By replacing these positively charged amino acids, DddA is formed. tox By weakening the binding force of DddA to DNA, intracellular toxicity can be reduced or eliminated. In other words, if a positively charged amino acid is substituted to make it non-toxic, cloning using E. coli is possible, resulting in full-length DddA. tox can be secured. Based on this, the non-toxic full-length cytosine deaminase is DddA of sequence number 1. tox It can be provided by replacing one or more, two or more, three or more, four or more, or five or more amino acids in the amino acid sequence with other amino acids (e.g., alanine), and in this regard, reference may be made to contents already known prior to the present application, including, for example, WO 2022 / 060185, which is incorporated by reference in its entirety by this application.

[0159] In some embodiments, when a full-length cytosine deaminase is used as the cytosine deaminase, such cytosine deaminase may have one or more amino acid substitutions selected from the group consisting of a substitution of S at position 37 with G, a substitution of G at position 59 with S, a substitution of A at position 109 with V, and a substitution of S at position 129 with G in the amino acid sequence of SEQ ID NO: 2.

[0160] A cytosine deaminase mutant having all of the following substitutions: S to G at position 37, G to S at position 59, A to V at position 109, and S to G at position 129 in the amino acid sequence of SEQ ID NO: 2 is commonly referred to as "GSVG."

[0161] GSVG:

[0162] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLEGKVFSSGGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGVIPVKRGATGETKVFTGNSNGPKSPTKGGC (SEQ ID NO: 7)

[0163] In addition, non-toxic battlefield DddA tox may comprise an amino acid sequence selected from the group consisting of the following amino acid sequences.

[0164] A1341D KRKKA variant:

[0165] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYDNAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNNSNSPKSPTAGGC

[0166] AAAAA variant:

[0167] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVAAGATGETAVFTGNSNSPASPTAGGC

[0168] AAAAK 변이체:

[0169] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVAAGATGETAVFTGNSNSPASPTKGGC

[0170] AAKAA 변이체:

[0171] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVAAGATGETKVFTGNSNSPASPTAGGC

[0172] AAKAK 변이체:

[0173] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVAAGATGETKVFTGNSNSPASPTKGGC

[0174] KAAAA 변이체:

[0175] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKAGATGETAVFTGNSNSPASPTAGGC

[0176] E1347A 변이체:

[0177] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVAGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNNSNSPKSPTKGGC

[0178] SSVG variants:

[0179] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGVIPVKRGATGETKVFTGNSNGPKSPTKGGC

[0180] GSAG mutants:

[0181] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLEGKVFSSGGTPPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNGPKSPTKGGC

[0182] GSVS variants:

[0183] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLEGKVFSSGGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGVIPVKRGATGETKVFTGNNSNSPKSPTKGGC

[0184] B-2. Other cytosine deaminases

[0185] In some embodiments, the cytosine deaminase that can be used in the present invention is APOBEC1. This enzyme typically acts on single-stranded DNA, but can be useful for base editing when fused with an enzyme domain having nickase function.

[0186] The above “APOBEC1” is an abbreviation for Apolipoprotein B mRNA Editing Catalytic Polypeptide 1, and refers to an enzyme that has the activity of deaminating cytosine to uracil. APOBEC1 is originally known as an enzyme involved in editing apolipoprotein B mRNA in mammalian hepatocytes, etc., but this enzyme can exhibit the activity of deaminating cytosine bases within a specific sequence of DNA or RNA. The APOBEC1 that can be used in the present invention is not necessarily limited to the naturally occurring wild type, and also includes functional mutants, orthologs, or artificially improved proteins that exhibit similar activity.

[0187] In some embodiments, the cytosine deaminase is APOBEC1 or an ortholog thereof having the following amino acid sequence:

[0188] APOBEC1:

[0189] MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSIWRHTSQNTNKHVEVNFIEKFTTERYFCPNTRCSITWFLSWSPCGECSRAITEFLSRYPHVTLFI YIARLYHHADPRNRQGLRDLISSGVTIQIMTEQESGYCWRNFVNYSPSNEAHWPRYPHLWVRLYVLELYCIILGLPPCLNILRRKQPQLTFFTIALQSCHYQRLPPHILWATGLK

[0190] In some embodiments, the cytosine deaminase that can be used in the present invention is AID.

[0191] The above “AID” ​​stands for Activation-Induced Cytidine Deaminase, which refers to a cytosine deaminase that plays an essential role in the generation of antibody diversity in the immune system. The AID that can be used in some embodiments of the present invention is not necessarily limited to the naturally occurring wild type, but also includes functional variants thereof, orthologs, or artificially improved proteins that exhibit similar activity.

[0192] In some embodiments, the cytosine deaminase is AID or an ortholog thereof having the amino acid sequence:

[0193] AID:

[0194] MDSLLMNRRKFLYQFKNVRWAKGRRETYLCYVVKRRDSATSFSLDFGYLRNKNGCHVELLFLRYISDWDLDPGRCYRVTWFTSWSPCYDCARHVADFLRGNPNLSLRIFTARLYFCEDRKAEPEGLRRLHRAGVQIAIMTFKDYFYCWNTFVENHERTFKAWEGLHENSVRLSRQLRRILLPLYEVDDLRDAFRTLGL

[0195] In some embodiments, the cytosine deaminase that can be used in the present invention is a variant of TadA8e that has been mutated to have cytosine deaminase activity. Such variants may include, for example, one or more of amino acid residues 6, 26, 27, 28, 46, 48, 49, 61, 74, 76, 77, 82, 96, 107, 108, 112, 114, 115, 119, 122, 127, 142, 143, 151, 154, and 158 of the amino acid sequence of SEQ ID NO: 8 (TadA8e) in which one or more of the amino acid residues is mutated to a different amino acid. With regard to the composition of such cytosine deaminase, reference may be made to WO 2022 / 060185, WO 2023 / 086953, etc., which are incorporated by reference in their entirety into this application, and which were already known prior to the present application.

[0196] The DNA base editor of the present invention comprises cytosine deaminase. Cytosine deaminase is an enzyme that converts cytosine bases in DNA to uracil through a deamination reaction, and plays a key role in correcting the C:G base pair at the corresponding site to T:A.

[0197] The cytosine deaminase used in the present invention may be a full-length protein or may be split into two parts to maintain enzymatic activity. The split form may be incorporated into two fusion proteins, each of which can be utilized as a dimeric base editor.

[0198] Thus, the cytosine deaminase used in the present invention is not limited to the name or origin of the enzyme, and includes any protein, a homologue thereof, or a variant thereof having biological activity capable of inducing base correction by selectively converting a cytosine base on a target DNA to uracil.

[0199] C. dUDG

[0200] In some embodiments, a DNA base editor according to the present invention comprises dUDG.

[0201] With regard to dUDG included in the base editor of the present invention, the contents described under the “1. dUDG” section of this specification are quoted as is.

[0202] D. Additional polypeptide elements

[0203] The DNA base editor according to the present invention may further comprise additional polypeptides or protein components in addition to (i) a DNA binding protein, (ii) a cytosine deaminase, and (iii) dUDG.

[0204] In addition to the DNA binding protein, cytosine deaminase, and dUDG, the DNA base editor of the present invention may additionally include various additional protein or peptide sequences. These additional components are used to enhance base editing efficiency or to control intracellular delivery to the target sequence and localization within organelles.

[0205] In some embodiments, the DNA base editor of the present invention may comprise an NLS.

[0206] The above "NLS (nuclear localization signal)" refers to an amino acid sequence motif required for protein translocation from the cytoplasm to the nucleus. NLSs are generally composed of short sequences rich in basic amino acids such as lysine or arginine, and mediate protein entry into the nucleus through interaction with nuclear transport receptors (e.g., importins).

[0207] The NLSs that can be used in some embodiments of the present invention are not limited to a specific sequence, structure, or origin, and various forms of NLSs can be used as long as they maintain their functional characteristics. Information regarding the function and sequence of NLSs is widely known through prior literature and publicly available materials, and those skilled in the art can select or modify an NLS sequence suitable for the intended purpose based on this information.

[0208] In some embodiments, the DNA base editor of the present invention may comprise a NES.

[0209] The above "NES (nuclear export signal)" refers to a peptide sequence or a functional variant thereof capable of inducing protein movement into the cytoplasm. NES generally has the function of binding to a nuclear export receptor (exportin) to transport a protein from the nucleus to the cytoplasm, and typically has a structural characteristic in which 4 to 5 hydrophobic amino acids are arranged at specific intervals. The NES used in some embodiments of the present invention is not limited to a specific sequence, structure, or origin, and may include various forms of nuclear export sequences as long as the functional characteristics are maintained.

[0210] The NES used in some embodiments of the present invention may be derived from various proteins existing in nature, and artificially designed sequences may also be utilized. For example, the NES derived from the HIV-1 Rev protein (e.g., LQLPPLERLTL), the NES derived from the PKI protein (e.g., LALKLAGLDI), or the NES derived from the NS2 protein of the mouse minivirus (MVM) (e.g., VDEMTKKFGTLTIHDTEK) are widely known as representative sequences, and various variants having similar nuclear export functions are also well known to those skilled in the art.

[0211] In some embodiments, the DNA base editor of the present invention may comprise MTS.

[0212] The above "MTS (mitochondrial targeting sequence)" refers to an amino acid sequence that induces the transport of a protein to the mitochondria after translation. MTS is generally located at the N-terminus and forms a unique α-helical structure with a repeating arrangement of basic and hydrophobic amino acids, which interacts with the mitochondrial inner membrane transport complex to achieve transport. The MTS that can be used in some embodiments of the present invention is not limited to a specific sequence, length, structure, or origin, and is understood to include any functional sequence that can effectively direct a protein to the mitochondria.

[0213] Information on MTS has been identified from various biological proteins, and relevant sequences are widely available through public databases such as UniProt, NCBI, and MitoCarta. Using this publicly available sequence information, those skilled in the art can select or combine MTS sequences appropriate for the desired protein to design it.

[0214] The MTS used in some embodiments of the present invention may be derived from various mitochondrial proteins existing in the natural world, and a sequence artificially designed to have a specific mitochondrial transport function may also be utilized. For example, MTS derived from human superoxide dismutase 2 (SOD2) protein (e.g., MALSRAVCGTSRQLAPVLGYLGSRQKHSLPD), MTS derived from cytochrome c oxidase subunit 8A (COX8A) (e.g., MASVLTPLLLRGLTGSARRLPVPRAKIHSL), or MTS derived from human mitochondrial ATP synthase F1β subunit (e.g., MLGFVGRVAAAPASGALRRLTPSASLPPAQLLLRAAPTAVHPVRDYAAQ) are widely known as representative sequences, and similarly, various variants having a protein transport function to mitochondria are well known to those skilled in the art.

[0215] In some embodiments, the DNA base editor of the present invention may comprise a CTS.

[0216] The above “chloroplast targeting sequence (CTS)” refers to an amino acid sequence that induces transport of a protein translated in the cytoplasm to the chloroplast. It is generally located at the N-terminus of the protein, and forms a structure in which positively charged amino acids and hydrophobic amino acids are repeatedly arranged to interact with a transport complex that passes through the outer and inner membranes of the chloroplast, thereby mediating transport to the chloroplast stroma. The CTS in the present specification is not limited to a specific sequence, length, structure, or biological origin, and is understood to include any functional sequence that can effectively direct a protein to the chloroplast.

[0217] CTS sequences have been identified from various plant-derived proteins (e.g., RuBisCO subunits, ferredoxin, plastocyanin, etc.) and are widely available in public databases such as UniProt, TAIR, and NCBI. Those skilled in the art can design a CTS suitable for a target protein by referencing or combining these published sequences. Therefore, the CTS that can be used in some embodiments of the present invention is not limited to a specific CTS sequence, but may include various homologs, functional analogs, or variants that can functionally induce chloroplast transport.

[0218] Additional polypeptide elements such as those described above may be present in the form of a fusion protein (comprising at least a DNA binding protein and a cytosine deaminase) or may be present separately from such a fusion protein.

[0219] When present in the form of a fusion protein, localization signal sequences such as NLS, NES, MTS, and CTS are preferably located at the N-terminus of the fusion protein, and can be appropriately designed depending on the cellular organelle in which the proofreading system is to function.

[0220] In some embodiments, the DNA base editor of the present invention may include additional sequences for biotechnology techniques, such as tags.

[0221] In some embodiments, the fusion protein and / or polypeptide comprised in the base editor according to the present invention may include a protein tag sequence, such as a His-tag, FLAG-tag, HA-tag, or Myc-tag, to facilitate purification or detection.

[0222] These tag sequences are additionally located at the N- or C-terminus of the fusion protein or polypeptide, providing experimental or manufacturing convenience. For example, the His-tag can be utilized for protein purification using metal columns, while the FLAG-tag or HA-tag are useful for antibody-based Western blotting or immunostaining.

[0223] The above tag can be designed and positioned within a range that does not affect protein function, and can also be configured with a removable cleavage sequence if necessary.

[0224] E. fusion protein

[0225] All or part of the components constituting the DNA base editor according to the present invention may be present in the form of a fusion protein.

[0226] E-1. Unit fusion protein A

[0227] In some embodiments, a DNA base editor according to the present invention may comprise (i) one or more DNA binding proteins (e.g., zinc finger proteins, TALE proteins, etc.) and (ii) a cytosine deaminase in the form of a single fusion protein.

[0228] In some preferred embodiments, the sequence of the DNA binding protein and cytosine deaminase within the fusion protein is as follows. The sequence below is intended to indicate the relative positions of the indicated components, and other components that may be present, including linkers, are omitted.

[0229] (N-terminal) [DNA binding protein] - [cytosine deaminase] (C-terminal)

[0230] Various additional polypeptide components described in the “D. Additional Polypeptide Components” section above may be added to the arrangement described above. These may also be directly linked to other protein components or linked via one or more linkers.

[0231] The protein components mentioned in the above arrangement may be directly connected to each other or connected via one or more linkers.

[0232] E-2. Unit fusion protein B

[0233] In some embodiments, a DNA base editor according to the present invention may comprise (i) one or more DNA binding proteins (e.g., zinc finger proteins, TALE proteins, etc.), (ii) cytosine deaminase, and (iii) dUDG in the form of a single fusion protein.

[0234] The sequence of DNA binding protein, cytosine deaminase, and dUDG included in the fusion protein can vary.

[0235] In some preferred embodiments, the sequence of the DNA binding protein, cytosine deaminase, and dUDG within the fusion protein is as follows. The sequence below is intended to indicate the relative positions of the indicated components, and other components that may be present, including linkers, are omitted.

[0236] (N-terminal) [DNA binding protein] - [cytosine deaminase] - [dUDG] (C-terminal)

[0237] (N-terminal) [DNA binding protein] - [dUDG] - [cytosine deaminase] (C-terminal)

[0238] Various additional polypeptide components described in the “Additional Polypeptide Components” section above may be added to the arrangement described above. These may also be directly linked to other protein components or linked via one or more linkers.

[0239] The protein components mentioned in the above arrangement may be directly connected to each other or connected via one or more linkers.

[0240] E-3. Linker

[0241] The term "linker" as used herein refers to an amino acid linker, which is an amino acid sequence that covalently connects two or more functional protein domains, peptides, or other biological molecular elements. Such linkers serve to provide sufficient flexibility, length, or spatial separation so that each connected component can maintain its own structural or functional activity, and may sometimes be designed to include a specific secondary structure (e.g., an α-helix) or recognition sequence. For example, a repeating sequence based on glycine (G) and serine (S) (e.g., GGGGS)n is known as a representative example that confers high flexibility and water solubility.

[0242] The linker that can be used in the DNA editing editor according to the present invention is not limited to a specific amino acid sequence, length, or structure. It can be a naturally occurring sequence or an artificially designed sequence, and any amino acid sequence having a variety of lengths, sequence combinations, or structural characteristics is encompassed, as long as the function of each component connected via the linker is substantially maintained. Those skilled in the art can select or design an appropriate linker based on the characteristics of the target domain and the intended application.

[0243] In some embodiments, the DNA editing editor according to the present invention may comprise one or more linkers selected from the following linkers:

[0244] 2a.a. Linker: GS

[0245] 4a.a. Linker: GSGS

[0246] 5a.a. Linker: PGSGS

[0247] 5a.a. Linker: TGEKQ

[0248] 10a.a. Linker: SGAQGSTLDF

[0249] 13a.a. Linker: AAEFGIRIPGEKP

[0250] 14a.a. Linker: AAEFGIHGVPAAMG

[0251] 16a.a. Linker: SGSETPGTSESATPES

[0252] 24a.a. Linker: SGTPHEVGVYTLSGTPHEVGVYTL

[0253] 32a.a. Linker: GSGGSSGGSSGSETPGTSESATPESSGGSSGGS

[0254] E-4. Monomeric Base Editor

[0255] In some embodiments, a DNA base editor according to the present invention comprises a DNA binding protein and a cytosine deaminase as a single fusion protein (monomeric base editor).

[0256] In some embodiments where the base editor according to the present invention is a monomeric base editor, DddA is used as the cytosine deaminase. tox When used, the cytosine deaminase is non-toxic full-length DddA tox It is desirable. For example, GSVG can be used.

[0257] In some embodiments where the DNA base editor according to the present invention is a monomeric base editor, the base editor may comprise a monomeric fusion protein A or B.

[0258] In some embodiments of a base editor comprising a single fusion protein A, dUDG is present separately from the single fusion protein.

[0259] E-5. Dimeric Base Editor

[0260] In some embodiments, the DNA base editor according to the present invention comprises a DNA binding protein and a cytosine deaminase as two fusion proteins (dimeric base editor).

[0261] In certain embodiments where the DNA base editor according to the present invention is a dimeric base editor, the cytosine deaminase exists in two fragments, each of which can be divided into a first fusion protein and a second fusion protein. An example of a fragment is DddA. tox 1333N / 1333C or DddA tox Includes 1397N / 1397C.

[0262] In some embodiments where the DNA base editor according to the present invention is a dimeric base editor, the two fusion proteins included in the base editor may be as follows.

[0263] Dimeric base editor composition 1st fusion protein 2nd fusion protein 1 unit fusion protein A unit fusion protein A 2 unit fusion protein A unit fusion protein B 3 unit fusion protein B unit fusion protein A 4 unit fusion protein B unit fusion protein B

[0264] In some embodiments where the base editor takes the dimeric base editor configuration 1, the dUDG is present separately from the two fusion proteins.

[0265] In some embodiments where the DNA base editor according to the present invention is a dimeric base editor, DNA editing occurs in the region between the DNA sequences to which the two DNA binding proteins included in the two fusion proteins each bind, and this region is referred to as a “spacer.”

[0266] F. C-to-T base correction efficacy

[0267] The base editor of the present invention is useful for correcting cytosine (C) bases to thymine (T) bases. The inventors of the present invention have confirmed that even when using dUDG instead of UGI in a C-to-T base correction system, a C-to-T correction efficiency equivalent to or higher than that achieved when using UGI can be achieved.

[0268] The above C-to-T correction is a C-to-T correction in nuclear or organelle DNA.

[0269] In some embodiments, the base editor of the present invention is suitable for C-to-T correction in nuclear DNA.

[0270] In some embodiments, the base editor of the present invention is suitable for C-to-T correction in chloroplast DNA.

[0271] In some embodiments, the base editor of the present invention is suitable for C-to-T correction in mitochondrial DNA.

[0272] In some embodiments, a base editor according to the present invention is characterized by having a superior efficiency in correcting cytosine (C) bases to thymine (T) bases compared to a base editor of the same composition except that it does not include dUDG.

[0273] In some embodiments, the base editor according to the present invention is characterized by having a superior efficiency in correcting cytosine (C) bases to thymine (T) bases in chloroplast DNA compared to a base editor of the same composition except that it does not include dUDG.

[0274] In some embodiments, a base editor according to the present invention is characterized by having a superior efficiency in correcting cytosine (C) bases to thymine (T) bases compared to a base editor of the same configuration except that UGI is used instead of dUDG.

[0275] In some embodiments, the base editor according to the present invention is characterized by having a superior efficiency in correcting cytosine (C) bases to thymine (T) bases in chloroplasts compared to a base editor of the same configuration except that UGI is used instead of dUDG.

[0276] In some embodiments, the base editor according to the present invention has a C-to-T base correction efficiency of 50% or more, 60% or more, 70% or more, 80% or more, 90% or more, or 95% or more. Such high efficiency is very useful for in vivo and ex vivo gene correction therapy, crop trait improvement, etc.

[0277] In some embodiments, the base editor according to the present invention has a C-to-T base correction efficiency of 50% or more, 60% or more, 70% or more, 80% or more, 90% or more, or 95% or more in chloroplast DNA. Such high efficiency is very useful for improving crop traits, etc.

[0278] 3. Uses of dUDG

[0279] One aspect of the present invention relates to a dUDG for use in C-to-T base correction.

[0280] UDG, an enzyme with uracil removal catalytic activity, has been recognized as a factor that interferes with C-to-T base proofreading, and therefore its use in C-to-T base proofreading has never been considered. Meanwhile, for the purpose of analyzing the structure and function of UDG, an inactive form of UDG (dUDG) with the uracil removal activity removed has been produced. However, the possibility of utilizing such dUDG in a C-to-T base proofreading system has not been suggested at all.

[0281] The above C-to-T base correction is C-to-T base correction in nuclear or organelle DNA.

[0282] In some embodiments, the dUDG according to the present invention is suitable for C-to-T base correction in nuclear DNA.

[0283] In some embodiments, the dUDG according to the present invention is suitable for C-to-T base correction in chloroplast DNA.

[0284] In some embodiments, the dUDG according to the present invention is suitable for C-to-T base correction in mitochondrial DNA.

[0285] The above dUDG is a UDG variant that can bind to DNA containing uracil but does not have the activity of removing uracil, and can be derived from any UDG.

[0286] In some embodiments, the dUDG is a variant having an amino acid mutation in motif A or motif B that is not present in wild-type UDG. This includes having an amino acid mutation in both motif A and motif B that is not present in wild-type UDG.

[0287] In some embodiments, the dUDG is a variant having an amino acid mutation in an amino acid sequence corresponding to the amino acid sequence KTVIIGQDPYHGPGQAMGLSF included in UDG of Arabidopsis thaliana or an amino acid sequence corresponding to HPSGLSA included in UDG of Arabidopsis thaliana. This includes variants having amino acid mutations in both the amino acid sequence corresponding to the amino acid sequence KTVIIGQDPYHGPGQAMGLSF included in UDG of Arabidopsis thaliana and the amino acid sequence corresponding to HPSGLSA included in UDG of Arabidopsis thaliana.

[0288] In some embodiments, the dUDG is a variant comprising an amino acid corresponding to an aspartic acid (D), tyrosine (Y), phenylalanine (F), or glutamine (Q) residue in the amino acid sequence KTVIIGQDPYHGPGQAMGLSF included in Arabidopsis thaliana UDG, which is mutated to another amino acid, or an amino acid corresponding to a histidine (H) or leucine (L) residue in the amino acid sequence HPSGLSA included in Arabidopsis thaliana UDG, which is mutated to another amino acid. This includes a variant in which an amino acid corresponding to an aspartic acid (D), tyrosine (Y), phenylalanine (F) or glutamine (Q) residue in the amino acid sequence KTVIIGQDPYHGPGQAMGLSF contained in Arabidopsis thaliana UDG is mutated to another amino acid, and an amino acid corresponding to a histidine (H) or leucine (L) residue in the amino acid sequence HPSGLSA contained in Arabidopsis thaliana UDG is also mutated to another amino acid.

[0289] In some embodiments, the dUDG is a variant comprising an amino acid corresponding to an aspartic acid (D) residue in the amino acid sequence KTVIIGQDPYHGPGQAMGLSF included in Arabidopsis thaliana UDG, which is mutated to another amino acid, or an amino acid corresponding to a histidine (H) residue in the amino acid sequence HPSGLSA included in Arabidopsis thaliana UDG, which is mutated to another amino acid. This includes variants comprising an amino acid corresponding to an aspartic acid (D) residue in the amino acid sequence KTVIIGQDPYHGPGQAMGLSF included in Arabidopsis thaliana UDG, which is mutated to another amino acid, and an amino acid corresponding to a histidine (H) residue in the amino acid sequence HPSGLSA included in Arabidopsis thaliana UDG, which is also mutated to another amino acid.

[0290] In some embodiments, the dUDG is a variant comprising an amino acid corresponding to an aspartic acid (D) residue in the amino acid sequence KTVIIGQDPYHGPGQAMGLSF included in Arabidopsis thaliana UDG, which is mutated to asparagine (N), alanine (A), or a conservative amino acid substitution thereof, or an amino acid corresponding to a histidine (H) residue in the amino acid sequence HPSGLSA included in Arabidopsis thaliana UDG, which is mutated to alanine (A) or a conservative substitution thereof. This includes a variant in which an amino acid corresponding to an aspartic acid (D) residue in the amino acid sequence KTVIIGQDPYHGPGQAMGLSF contained in Arabidopsis thaliana UDG is mutated to asparagine (N), alanine (A), or a conservative amino acid substitution thereof, and an amino acid corresponding to a histidine (H) residue in the amino acid sequence HPSGLSA contained in Arabidopsis thaliana UDG is also mutated to alanine (A) or a conservative substitution thereof.

[0291] In some embodiments, the dUDG is derived from UDG of Arabidopsis thaliana.

[0292] In some embodiments, the dUDG has the amino acid sequence of UDG of Arabidopsis thaliana, but has D173N or H295A, or both of these amino acid mutations.

[0293] In some embodiments, the dUDG comprises the amino acid sequence of SEQ ID NO: 1 or an amino acid sequence having at least 80% sequence identity thereto.

[0294] In some embodiments, the dUDG has the amino acid sequence of human UNG2, but has D154N or H277A, or both of these amino acid mutations.

[0295] In some embodiments, the dUDG has the amino acid sequence of E. coli UNG, but has D64N or H187A, or both of these amino acid mutations.

[0296] As for the amino acid sequence of dUDG and sequence alignment of different species or protein families, orthologous proteins, the contents described under the section “1. dUDG” in this specification are cited as is.

[0297] 4. Polynucleotide

[0298] One aspect of the present invention relates to a polynucleotide encoding the DNA base editor of the present invention. With respect to the "DNA base editor of the present invention," the content described in the "2. DNA base editor" section of this specification is incorporated herein by reference.

[0299] The polynucleotide according to the present invention is a polynucleotide encoding a protein component (fusion protein and / or polypeptide) constituting the DNA base editor of the present invention as described above, and the polynucleotide may be DNA or RNA. The DNA or RNA includes sequences contained within mRNA, cDNA, synthetic DNA, plasmid DNA, linear DNA, or viral vectors.

[0300] The polynucleotide according to the present invention may be a single polynucleotide encoding one fusion protein constituting the monomeric base editor as described above, or a plurality of polynucleotides encoding one fusion protein constituting the monomeric base editor and polypeptide(s) that may exist separately therefrom, or two polynucleotides constituting a dimeric base editor, or a plurality of nucleotides encoding two fusion proteins constituting the dimeric base editor and polypeptide(s) that may exist separately therefrom, and the two or multiple polynucleotides may be produced in the form of one polynucleotide through a known biotechnology technique such as the Golden Gate technique.

[0301] A person skilled in the art can easily obtain the amino acid sequence of each constituent protein by referring to the contents disclosed in the present application specification and the published protein sequences registered in databases such as NCBI GenBank and UniProt, and produce a polynucleotide according to the present invention. Based on the obtained amino acid sequences, a nucleotide sequence can be designed considering the codon usage frequency suitable for the host organism, and further, it is obvious to a person skilled in the art or within the scope of routine experimental techniques to produce a polynucleotide containing the sequence using commercial services for custom production of synthetic genes (e.g., IDT, GenScript, etc.) and vector cloning, PCR amplification, and DNA assembly technologies (Gibson assembly, Golden Gate, etc.).

[0302] The polynucleotide(s) may include expression control sequences (e.g., promoter) in addition to the coding region so as to sufficiently express the function of the encoded protein(s).

[0303] In some embodiments of the present invention, the polynucleotide may be codon optimized to suit the type of expression system (e.g., bacteria, plants, mammalian cells, etc.). Codon optimization is a general technique for increasing gene expression efficiency, and can increase protein expression by adjusting the nucleotide sequence according to the tRNA utilization frequency of the target organism.

[0304] 5. Base correction composition

[0305] One aspect of the present invention relates to a base correction composition comprising the DNA base editor of the present invention or a polynucleotide encoding the same. With respect to the "DNA base editor of the present invention," the content described in the "2. DNA base editor" section of this specification is incorporated herein by reference.

[0306] The base correction composition according to the present invention is for correcting cytosine (C) base to thymine (T) in DNA of a nucleus or cell organelle.

[0307] In some embodiments, the base correction composition according to the present invention is suitable for correcting a cytosine (C) base to thymine (T) in DNA of a nucleus.

[0308] In some embodiments, the base correction composition according to the present invention is suitable for correcting a cytosine (C) base to thymine (T) in the DNA of a chloroplast.

[0309] In some embodiments, the base correction composition according to the present invention is suitable for correcting a cytosine (C) base to thymine (T) in mitochondrial DNA.

[0310] The base correction composition according to the present invention may comprise the DNA base editor of the present invention or a polynucleotide encoding the same as described above and a biocompatible carrier.

[0311] The above "biocompatible carrier" refers to a material that can be included in the composition according to the present invention, and which does not significantly inhibit physiological functions when in contact with a biological system—e.g., cells, tissues, organs, or entire organisms of humans, animals, or plants—and does not induce toxic or immune reactions, etc. Such a carrier can be selected in various ways depending on pharmaceutical, biological, or agricultural use, and can play a role in improving the stability, permeability, and delivery efficiency of a base correction enzyme or a protein, nucleic acid, or auxiliary molecule related thereto.

[0312] The biocompatible carrier that can be used in the base correction composition according to the present invention is not limited to a specific chemical structure, physical form or origin, and may include, for example, a buffer solution, a surfactant, a liposome, a nanoparticle, a hydrogel, a polymeric material (e.g., PEG, PVA, PLA, PLGA), a natural or synthetic polysaccharide (e.g., dextran, hyaluronic acid, chitosan), a plant- or microbial-derived polymer, a liposome, a micelle, a biopolymer, or a combination thereof. In addition, the biocompatible carrier may be appropriately selected or combined depending on a specific delivery route or application target (e.g., human tissue, animal tissue, plant tissue, etc.), and may also include a biologically acceptable solvent, preservative, stabilizer, buffer, surfactant, reducing agent, etc.

[0313] A person skilled in the art can easily select and combine a carrier suitable for a given application purpose, delivery route, or target organism based on information already widely known through various literature and public databases regarding the types, properties, and application methods of the biocompatible carriers.

[0314] The base correction composition according to the present invention may be provided in various physical forms. The physical form may be selected based on the intended application, route of administration, stability, storage conditions, or manufacturing process, and falls within the general pharmaceutical design criteria for enhancing the efficacy and ease of use of the composition.

[0315] In some embodiments, dUDG may be included in one of the fusion proteins or may exist independently of the fusion protein. Existing dUDG independently of the fusion protein means that dUDG is expressed in a separate protein form from the DNA binding protein and cytosine deaminase.

[0316] In some embodiments, the base correction composition according to the present invention may be in the form of a liquid, suspension, gel, powder, lyophilisate, tablet, capsule, or injectable composition. It may also be formulated as a liposome, nanoparticle, lipid nanoparticle (LNP), or other delivery particle.

[0317] The base correction composition aspect of the present invention is not limited to the above physical form, and all formulation modifications that can be appropriately selected and manufactured by a person skilled in the art according to the purpose are included in the scope of the present invention.

[0318] The base correction composition according to the present invention can be applied in various ways to ensure effective delivery to the target cell, tissue, or organism. The application method may vary depending on the target species, cell type, delivery route, or formulation characteristics, and is selected based on the stability, efficacy, and biological compatibility of the composition.

[0319] The base correction composition of the present invention can be applied to plants or plant cells. Methods for applying to plants may include agroinfiltration, Agrobacterium-mediated delivery, gene gun delivery, electroporation, or direct intratissue injection. The composition to be applied may be prepared in the form of protein, DNA, mRNA, or ribonucleoprotein (RNP), and may be used with a delivery vehicle capable of penetrating plant cell walls (e.g., cell-penetrating peptide, non-targeting nanoparticle, etc.) depending on the purpose.

[0320] The base correction composition of the present invention can be flexibly applied to various biological subjects, and any in vivo, ex vivo, or in vitro delivery method that can be selected by a person skilled in the art to achieve the desired effect is included in the scope of application of the present invention.

[0321] The base correction composition of the present invention can be used to directly manipulate cells in an extracellular environment (in vitro), and can also be used for the purpose of directly inducing base correction in an intracellular environment (in vivo).

[0322] The base correction composition according to the present invention can be usefully applied to base correction technology that enables precise manipulation of genetic sequences by selectively converting specific DNA bases in vivo or in vitro. In particular, the composition can be used to precisely correct a desired target base sequence through C-to-T correction, which converts cytosine (C) to thymine (T).

[0323] Because this correction reaction occurs without DNA double-strand breaks, it has the advantage of being less mutagenic and enhancing genome stability compared to existing gene editing technologies. Therefore, the composition of the present invention can be widely utilized in various fields, including plant variety improvement, industrial microorganism improvement, and biotechnology research.

[0324] 6. Transmitter

[0325] One aspect of the present invention relates to a delivery system comprising the DNA base editor of the present invention or a polynucleotide encoding the same. With respect to the "DNA base editor of the present invention," the content described in the "2. DNA base editor" section of this specification is incorporated herein by reference.

[0326] The delivery vehicle according to the present invention refers to a means for effectively delivering the DNA base editor of the present invention or one or more polynucleotides (e.g., DNA, mRNA, etc.) encoding the same to a target site in a cell or a living body. Such a delivery vehicle may include various components to enhance the cell penetration efficiency, intracellular stability, organelle targeting ability, or in vivo distribution characteristics of the editor component, and preferably has acceptable properties such as biocompatibility, biodegradability, and non-immunogenicity. The delivery vehicle of the present invention aims to increase the efficiency and specificity of gene correction, while minimizing cytotoxicity and reducing the possibility of affecting non-target tissues.

[0327] The term "vector" or "delivery vehicle" as used herein refers to a biological or non-biological means capable of effectively delivering the DNA base editor of the present invention or one or more polynucleotides encoding it into cells. These delivery vehicles can be categorized into various types based on their structure, origin, mechanism of action, etc., and can be appropriately selected depending on the intended application target (e.g., plant cells, bacteria, etc.) and administration method.

[0328] In some embodiments, the vector may be a viral vector, including but not limited to adeno-associated virus (AAV), lentivirus, adenovirus, retrovirus, bacteriophage-based vector, and the like.

[0329] In other embodiments, the carrier may be a non-viral carrier, including, for example, lipid nanoparticles (LNPs), polymeric nanoparticles, cationic liposomes, lipofectins, peptide-based carriers, electroporation, or nanoneedle-based systems.

[0330] Additionally, in some embodiments for application to plants, the carrier may comprise a physical delivery means implemented by Agrobacterium tumefaciens strains, plant virus-based vectors, protoplast delivery systems, or gene gun technology.

[0331] A person skilled in the art can select and combine appropriate carriers based on known techniques, depending on the characteristics of a specific base editor system, the delivery route, and the type of cell or organism to which it is applied.

[0332] The delivery system according to the present invention is applicable to various cells and organisms and can be selectively adjusted according to the purpose. Specifically, the delivery target includes a eukaryotic cell or a prokaryotic cell, and in some embodiments, the delivery target may be a plant tissue, cell, or embryo.

[0333] 7. Base correction method

[0334] One aspect of the present invention relates to a method for correcting a cytosine (C) base to a thymine (T) base, comprising introducing the DNA base editor of the present invention as described above into a cell containing target DNA for base correction, or expressing the DNA base editor within a cell containing target DNA for base correction. With respect to the "DNA base editor of the present invention," the contents described in the "2. DNA base editor" section of the present specification are incorporated herein by reference.

[0335] In some embodiments, the method comprises introducing a DNA base editor, a base correction composition (e.g., comprising a polynucleotide(s) encoding a DNA base editor of the present invention), and a delivery vehicle (e.g., comprising a polynucleotide(s) encoding a DNA base editor of the present invention) into a cell containing target DNA for base correction, as described above.

[0336] The above method is a method of correcting cytosine (C) bases in nuclear or organelle DNA to thymine (T) bases.

[0337] In some embodiments, the target DNA is nuclear DNA, and the method is a method of correcting a cytosine (C) base in nuclear DNA to a thymine (T) base.

[0338] In some embodiments, the target DNA is chloroplast DNA, and the method is a method of correcting a cytosine (C) base in chloroplast DNA to a thymine (T) base.

[0339] In some embodiments, the target DNA is mitochondrial DNA, and the method is a method of correcting a cytosine (C) base in mitochondrial DNA to a thymine (T) base.

[0340] The above base correction method can be performed in vitro, ex vivo, or in vivo, and can be designed according to various application purposes such as research purposes and trait improvement purposes. In particular, the DNA base editor of the present invention can efficiently correct cytosine bases in target DNA to thymine (C:G → T:A), and thus has wide applicability such as introduction of agriculturally useful traits and exploration of functional genes.

[0341] The components used in the base correction method of the present invention may include one or more of the DNA base editor described above, a polynucleotide encoding the same (e.g., mRNA or DNA), and a carrier (e.g., adeno-associated virus vector, lipid nanoparticle, etc.) containing the components. The components may be introduced into cells singly or in combination, and when present in the form of a fusion protein, stable expression and base correction activity can be provided through optimized binding between the component proteins. In addition, the components may additionally include an organelle targeting sequence (MTS, CTS, etc.), NES, or NLS, as needed.

[0342] The base correction method of the present invention is characterized by selectively correcting a specific base on DNA with another base, and the target of correction may be various, such as a mutation causing a genetic disease, an abnormal expression control region, or an artificial mutation for imparting a specific trait.

[0343] The method for expressing the DNA base editor or its components for implementing the base correction method of the present invention is not particularly limited and can be performed using technical means widely known to those skilled in the art. For example, a polynucleotide encoding a desired protein can be cloned into an appropriate expression vector, then introduced into a cell to induce transcription and / or translation, thereby causing expression. The expression can include both transient or stable expression within the cell, and can be implemented in various ways depending on the type of expression vector (e.g., plasmid, viral vector, etc.), the selection of promoter, the cell type, the introduction method, etc. Those skilled in the art can select and apply an appropriate expression system and conditions considering the desired cell type and base correction efficiency.

[0344] In the base correction method of the present invention, the DNA base editor or a composition comprising the same can be introduced into cells through various physical or chemical methods. For example, lipofection, electroporation, microinjection, viral vector delivery, nanoparticle delivery, Agrobacterium-mediated delivery, etc. can be used, and an appropriate delivery method can be selected depending on the type of cell being introduced, the target gene location, the target organism species, etc. In addition, the introduction conditions (e.g., pH, temperature, incubation time, introduction amount, etc.) can be easily optimized by a person skilled in the art, taking into account the desired base correction efficiency and cell viability.

[0345] The base correction method of the present invention is useful for correcting cytosine (C) bases to thymine (T) bases. The inventors of the present invention have confirmed that even when using dUDG instead of UGI in a C-to-T base correction system, correction efficiency equivalent to or higher than that achieved when using UGI can be achieved.

[0346] In some embodiments, the base correction method according to the present invention has a C-to-T base correction efficiency of 50% or more, 60% or more, 70% or more, 80% or more, 90% or more, or 95% or more. Such high efficiency is very useful for in vivo and ex vivo gene correction therapy, crop trait improvement, etc.

[0347] In some embodiments, the base correction method according to the present invention has a C-to-T base correction efficiency of 50% or more, 60% or more, 70% or more, 80% or more, 90% or more, or 95% or more in chloroplast DNA. Such high efficiency is very useful for improving crop traits, etc.

[0348] 8. Corrective organisms

[0349] The DNA base correction composition of the present invention can be applied to the genomes of various eukaryotic cells, including plant cells. The target for correction includes not only nuclear DNA but also the genomes of cellular organelles such as mitochondrial DNA and chloroplast DNA.

[0350] The present invention is useful for regulating gene expression through substitution of specific bases in plant cells or protoplasts, or for producing plants having specific genotypes.

[0351] The DNA base editor or correction composition of the present invention can be introduced into plant cells or plant-derived protoplasts to selectively correct specific bases on a target base sequence. The base editor or correction composition can be directly introduced or can be delivered to plant cells, tissues, or protoplasts using techniques such as Agrobacterium-mediated delivery, PEG-mediated delivery, or gene gun, thereby inducing base correction of nuclear DNA, chloroplast DNA, or mitochondrial DNA. Target DNA may include nuclear DNA, mitochondrial DNA, and chloroplast DNA. In particular, when targeting chloroplast DNA, a chloroplast transit signal (CTS) and / or a nuclear export signal (NES) can be included in the fusion protein or transporter to induce movement into the chloroplast. Similarly, when targeting mitochondrial DNA, a mitochondrial targeting signal (MTS) can be added to facilitate target-specific organelle delivery.

[0352] Plant cells or protoplasts that have undergone base correction can be regenerated into whole plants under appropriate regeneration conditions. The regenerated plants contain the corrected base sequence, and the base correction is genetically stable and can be passed on to future generations. Progeny plants can be obtained through self-pollination or crossbreeding of the regenerated plants, and include F1, F2, and subsequent generations with the corrected genotype. These corrected plants and their progeny can be utilized in various fields, such as breeding, functional crop development, and molecular farming.

[0353] In some embodiments, the DNA base correction composition of the present invention can induce C-to-T base correction at a specific site within the plant cell genome. For example, a C present at a specific position in a wild-type sequence is converted to a T by the action of cytosine deaminase, which can serve as a useful tool for restoring abnormal gene sequences, regulating the expression of specific genes, or improving specific traits. The correction site can be located in the target DNA within the nucleus or various cellular organelles such as chloroplasts and mitochondria.

[0354] Plants obtained through the above DNA base correction technology can form seeds with corrected genotypes, and these seeds can stably transmit the corrected genotypes to future generations. The present invention encompasses such seeds, plants germinated from such seeds, and even tissues, cells, or biological byproducts obtained therefrom. Plants derived from such seeds maintain the corrected genotypes and can be continuously propagated or propagated into future plants with the base-corrected traits in the same manner.

[0355] The base editing technology according to the present invention can be utilized as a powerful molecular biological tool for improving plant traits, modifying specific gene functions, or regulating target gene expression. For example, it can be applied to the development of functional crops or transgenic plants by targeting genes involved in the photosynthetic pathway, disease resistance genes, salt tolerance, or drought tolerance genes. In particular, chloroplast gene editing, which is inherited through cytoplasmic inheritance, offers advantageous characteristics in terms of trait stability and prevention of environmental transmission.

[0356] In some embodiments, plant cells, protoplasts, or plants containing a base sequence in which cytosine (C) in the wild-type target DNA is replaced with thymine (T) can be generated. Such corrected cells or plants can be utilized in various plant biotechnology fields, such as functional analysis, metabolic regulation, crop improvement, and optimization of biosynthetic pathways.

[0357] In some embodiments, the base-correcting organism according to the present invention is a part of a plant selected from the group consisting of cereals, legumes, root vegetables, vegetables, fruits, and forage crops.

[0358] In some embodiments, the base-correcting organism according to the present invention is an algae or a part thereof.

[0359] In some embodiments, the base-correcting organism according to the present invention is a descendant or clone of a plant or algae, or a portion thereof.

[0360] The above “part of a progeny or clone” may include (i) a part of a cell or tissue derived from the progeny or clone, (ii) a part of a cell lineage or organ within the progeny or clone, or (iii) a part of an individual among the progeny or clone population having a corrected genotype. For example, when base correction is reflected only in some leaf tissue or flower tissue among plant tissues, or in some embodiments having a corrected genotype, the base correction organism according to the present invention is a seed obtained from a progeny or clone of the plant or algae.

[0361] In some embodiments, the base-correcting organism according to the present invention is a seed obtained from a progeny or clone of said plant or algae.

[0362] The present invention can be explained with specific embodiments exemplified below based on the above, but is not limited thereto.

[0363] 1. A DNA base editor for correcting a cytosine (C) base to a thymine (T) base, comprising a DNA binding protein, a cytosine deaminase, and a dead uracil DNA glycosylase (dUDG), wherein the cytosine deaminase exists in a full-length form or in the form of two split bodies, and when the cytosine deaminase exists in the form of two split bodies, the cytosine deaminase activity is exhibited through dimerization thereof.

[0364] 2. A DNA base editor according to the above-mentioned first paragraph, wherein the dUDG can bind to DNA containing uracil but does not have the activity of removing uracil.

[0365] 3. A DNA base editor according to the above-mentioned first or second clause, wherein the dUDG has an amino acid mutation in motif A or motif B that does not exist in the wild-type UDG.

[0366] 4. A base editor according to any one of the above-mentioned claims 1 to 3, wherein the dUDG has an amino acid mutation in an amino acid sequence corresponding to the amino acid sequence KTVIIGQDPYHGPGQAMGLSF included in the UDG of Arabidopsis thaliana or an amino acid sequence corresponding to HPSGLSA included in the UDG of Arabidopsis thaliana.

[0367] 5. A base editor according to any one of the above-mentioned claims 1 to 4, wherein the amino acid mutation is a mutation of an amino acid corresponding to an aspartic acid (D), tyrosine (Y), phenylalanine (F), or glutamine (Q) residue in the amino acid sequence KTVIIGQDPYHGPGQAMGLSF included in Arabidopsis thaliana UDG into another amino acid, or an amino acid corresponding to a histidine (H) or leucine (L) residue in the amino acid sequence HPSGLSA included in Arabidopsis thaliana UDG into another amino acid.

[0368] 6. A base editor according to any one of the above-mentioned claims 1 to 5, wherein the amino acid mutation comprises a mutation of an amino acid corresponding to an aspartic acid (D) residue in the amino acid sequence KTVIIGQDPYHGPGQAMGLSF included in Arabidopsis thaliana UDG to another amino acid, or a mutation of an amino acid corresponding to a histidine (H) residue in the amino acid sequence HPSGLSA included in Arabidopsis thaliana UDG to another amino acid.

[0369] 7. A base editor according to any one of the above-mentioned claims 1 to 6, wherein the DNA binding protein is independently selected from the group consisting of a zinc finger protein, a TALE (transcription-activator-like effector) protein, and a CRISPR-associated nuclease.

[0370] 8. A base editor according to any one of the above-mentioned claims 1 to 7, wherein the cytosine deaminase is APOBEC (apolipoprotein B editing complex), AID (activation-induced deaminase), double-strand DNA-specific cytosine deaminase, or TadA (tRNA-specific adenosine deaminase) having C-to-T proofreading activity.

[0371] 9. A DNA base editor comprising a DNA binding protein, cytosine deaminase, and dead uracil DNA glycosylase (dUDG).

[0372] Containing the above DNA binding protein and cytosine deaminase in the form of two fusion proteins,

[0373] The above cytosine deaminase is a double-stranded DNA-specific cytosine deaminase, which exists in the form of two split entities and exerts cytosine deaminase activity through their dimerization.

[0374] A base editor for correcting a cytosine (C) base to a thymine (T) base, wherein the two fusion proteins each independently contain one DNA binding protein and divide the two fragments.

[0375] 10. A base editor for correcting a cytosine (C) base to a thymine (T) base, wherein dUDG is included in at least one of the two fusion proteins in the above-mentioned paragraph 9.

[0376] 11. A base editor according to any one of the above-mentioned claims 1 to 10, wherein the DNA is chloroplast DNA and optionally further comprises a chloroplast transit signal (CTS) and / or NES.

[0377] 12. A base editor according to any one of the above-mentioned claims 1 to 10, wherein the DNA is mitochondrial DNA and optionally further comprises a mitochondrial transfer signal (MTS) and / or a nuclear export signal (NES).

[0378] 13. A base editor according to any one of the above-mentioned items 1 to 10, wherein the DNA is nuclear DNA and optionally further comprises a nuclear localization signal (NLS).

[0379] 14. A base editor according to any one of the above-mentioned clauses 1 to 13, wherein the DNA is DNA of a plant cell.

[0380] 15. A DNA base editor characterized in that the efficiency of correcting cytosine (C) bases to thymine (T) bases is superior to that of a base editor of the same composition, except that it does not include dUDG, according to any one of the above-mentioned claims 1 to 14.

[0381] 16. A DNA base editor for C-to-T base correction in chloroplast DNA, characterized in that the efficiency of correction of cytosine (C) base to thymine (T) base is superior to that of a base editor of the same composition except that it does not include dUDG in the above-mentioned 15th paragraph.

[0382] 17. A DNA base editor characterized in that the efficiency of correcting cytosine (C) bases to thymine (T) bases is superior to that of a base editor of the same composition, except that a uracil glycosylase inhibitor (UGI) is used instead of dUDG according to any one of the above-mentioned claims 1 to 14.

[0383] 18. A DNA base editor for C-to-T base correction in chloroplast DNA, characterized in that the efficiency of correcting cytosine (C) bases to thymine (T) bases is superior to that of a base editor of the same composition, except that a uracil glycosylase inhibitor (UGI) is used instead of dUDG in the above-mentioned 17th paragraph.

[0384] 19. A DNA base editor according to any one of the above-mentioned claims 1 to 18, wherein the efficiency of correcting a cytosine (C) base to a thymine (T) base is 50% or more, 60% or more, 70% or more, 80% or more, 90% or more, or 95% or more.

[0385] 20. A DNA base editor according to the above-mentioned item 19, wherein the efficiency of correcting cytosine (C) base to thymine (T) base in chloroplast DNA is 50% or more, 60% or more, 70% or more, 80% or more, 90% or more, or 95% or more.

[0386] 21. A polynucleotide(s) encoding a base editor as described in any one of claims 1 to 20 mentioned above.

[0387] 22. A DNA base correction composition comprising a base editor or polynucleotide(s) encoding the same as described in any one of claims 1 to 20 mentioned above.

[0388] 23. A method for correcting a cytosine (C) base to a thymine (T) base, comprising introducing a base editor as described in any one of the above-mentioned items 1 to 20 or a base correction composition as described in the above-mentioned item 22 into a cell containing target DNA for base correction.

[0389] 24. A plant cell or protoplast in which C-to-T base correction has been performed by the base correction method described in the above-mentioned clause 23.

[0390] 25. A plant or part thereof grown or cultured from a plant cell or protoplast according to the above-mentioned Article 24.

[0391] 26. A plant or part thereof which is a descendant or clone of a plant according to the above-mentioned paragraph 25.

[0392] 27. Seeds obtained from the plant according to the above-mentioned item 25 or 26.

[0393] 28. A plant or its offspring, or a part thereof, grown from a seed according to the above-mentioned clause 27.

[0394] 29. A plant cell or protoplast according to the above-mentioned item 24, a plant or a part thereof according to any one of the above-mentioned items 25, 26 and 28, or a seed according to the above-mentioned item 26, wherein the cytosine (C) base in the wild-type target DNA sequence is corrected to the thymine (T) base.

[0395] 30. A UDG variant for use in C-to-T base correction that can bind to DNA containing uracil but does not have the activity to remove uracil.

[0396] 31. A UDG variant according to the above-mentioned clause 30, which has an amino acid mutation in motif A or motif B that is not present in the wild-type UDG.

[0397] 32. A UDG variant according to the above-mentioned claim 30 or 31, wherein the UDG variant has an amino acid mutation in an amino acid sequence corresponding to the amino acid sequence KTVIIGQDPYHGPGQAMGLSF included in UDG of Arabidopsis thaliana or an amino acid sequence corresponding to HPSGLSA included in UDG of Arabidopsis thaliana.

[0398] 33. A UDG variant according to any one of the above-mentioned items 30 to 32, wherein the amino acid mutation comprises a mutation of an amino acid corresponding to an aspartic acid (D), tyrosine (Y), phenylalanine (F), or glutamine (Q) residue in the amino acid sequence KTVIIGQDPYHGPGQAMGLSF included in Arabidopsis thaliana UDG to another amino acid, or a mutation of an amino acid corresponding to a histidine (H) or leucine (L) residue in the amino acid sequence HPSGLSA included in Arabidopsis thaliana UDG to another amino acid.

[0399] 34. A UDG variant according to any one of the above-mentioned items 30 to 33, wherein the amino acid mutation comprises a mutation of an amino acid corresponding to an aspartic acid (D) residue in the amino acid sequence KTVIIGQDPYHGPGQAMGLSF contained in Arabidopsis thaliana UDG to another amino acid, or a mutation of an amino acid corresponding to a histidine (H) residue in the amino acid sequence HPSGLSA contained in Arabidopsis thaliana UDG to another amino acid.

[0400] 35. A UDG variant comprising the amino acid sequence of SEQ ID NO: 1 or an amino acid sequence having at least 80% sequence identity thereto, according to any one of claims 30 to 34 mentioned above.

[0401] In another aspect, the present invention can be described with specific embodiments exemplified below based on the above contents, but is not limited thereto.

[0402] 1. A method for correcting a cytosine (C) base to a thymine (T) base, comprising introducing a DNA base editor into a cell containing target DNA for base correction or expressing the DNA base editor within a cell containing target DNA for base correction.

[0403] The DNA base editor comprises a DNA binding protein, a cytosine deaminase, and a dead uracil DNA glycosylase (dUDG), wherein the cytosine deaminase exists in a full-length form or in the form of two splits, and when it exists in the form of two splits, it exhibits cytosine deaminase activity through dimerization thereof.

[0404] Base editing method.

[0405] 2. A base correction method according to the above-mentioned first paragraph, wherein the dUDG can bind to DNA containing uracil but does not have the activity of removing uracil.

[0406] 3. A base correction method according to the above-mentioned first or second clause, wherein the dUDG has an amino acid mutation in motif A or motif B that does not exist in the wild-type UDG.

[0407] 4. A base correction method according to any one of the above-mentioned claims 1 to 3, wherein the dUDG has an amino acid mutation in an amino acid sequence corresponding to the amino acid sequence KTVIIGQDPYHGPGQAMGLSF included in the UDG of Arabidopsis thaliana or an amino acid sequence corresponding to HPSGLSA included in the UDG of Arabidopsis thaliana.

[0408] 5. A base correction method, wherein among the above-mentioned clause 4, the amino acid mutation is a mutation in which an amino acid corresponding to an aspartic acid (D), tyrosine (Y), phenylalanine (F), or glutamine (Q) residue in the amino acid sequence KTVIIGQDPYHGPGQAMGLSF included in Arabidopsis thaliana UDG is mutated to another amino acid, or an amino acid corresponding to a histidine (H) or leucine (L) residue in the amino acid sequence HPSGLSA included in Arabidopsis thaliana UDG is mutated to another amino acid.

[0409] 6. A base correction method according to the above-mentioned claim 4 or 5, wherein the amino acid mutation comprises a mutation of an amino acid corresponding to an aspartic acid (D) residue in the amino acid sequence KTVIIGQDPYHGPGQAMGLSF included in Arabidopsis thaliana UDG to another amino acid, or a mutation of an amino acid corresponding to a histidine (H) residue in the amino acid sequence HPSGLSA included in Arabidopsis thaliana UDG to another amino acid.

[0410] 7. A base correction method according to any one of the above-mentioned claims 1 to 6, wherein the DNA binding protein is independently selected from the group consisting of a zinc finger protein, a TALE (transcription-activator-like effector) protein, and a CRISPR-associated nuclease.

[0411] 8. A base correction method according to any one of the above-mentioned clauses 1 to 7, wherein the cytosine deaminase is APOBEC (apolipoprotein B editing complex), AID (activation-induced deaminase), double-strand DNA-specific cytosine deaminase, or TadA (tRNA-specific adenosine deaminase) having C-to-T proofreading activity.

[0412] 9. In any one of the above-mentioned clauses 1 to 8, the DNA base editor comprises a DNA binding protein and a cytosine deaminase in the form of two fusion proteins,

[0413] The above cytosine deaminase is a double-stranded DNA-specific cytosine deaminase, which exists in the form of two split entities and exerts cytosine deaminase activity through their dimerization.

[0414] The two fusion proteins each independently contain one DNA binding protein, and are divided into two fragments.

[0415] Base editing method.

[0416] 10. A base correction method according to the above-mentioned 9th paragraph, wherein dUDG is included in at least one of the two fusion proteins.

[0417] 11. A base correction method according to any one of the above-mentioned claims 1 to 10, wherein the target DNA is chloroplast DNA and the DNA base editor optionally additionally includes a chloroplast transit signal (CTS) and / or NES.

[0418] 12. A base correction method according to any one of the above-mentioned claims 1 to 10, wherein the target DNA is mitochondrial DNA, and the DNA base editor optionally additionally includes a mitochondrial transfer signal (MTS) and / or a nuclear export signal (NES).

[0419] 13. A base correction method according to any one of the above-mentioned claims 1 to 10, wherein the target DNA is nuclear DNA and the DNA base editor optionally additionally includes a nuclear localization signal (NLS).

[0420] 14. A base correction method according to any one of the above-mentioned clauses 1 to 13, wherein the target DNA is DNA of a plant cell.

[0421] 15. A DNA base editing method according to any one of the above-mentioned claims 1 to 14, characterized in that the DNA base editor has a superior efficiency in correcting cytosine (C) bases to thymine (T) bases compared to an editor of the same configuration except that the DNA base editor does not include dUDG.

[0422] 16. A DNA base correction method in which the DNA base editor has a superior correction efficiency of cytosine (C) base to thymine (T) base compared to a base editor of the same configuration except that the DNA base editor does not contain dUDG, and the target DNA is chloroplast DNA.

[0423] 17. A DNA base editing method according to any one of the above-mentioned claims 1 to 14, characterized in that the DNA base editor has a superior efficiency in correcting cytosine (C) bases to thymine (T) bases compared to a base editor of the same configuration, except that the DNA base editor uses a uracil glycosylase inhibitor (UGI) instead of dUDG.

[0424] 18. A DNA base correction method in which the DNA base editor has a superior correction efficiency of a cytosine (C) base to a thymine (T) base compared to a base editor of the same configuration, except that the DNA base editor uses a uracil glycosylase inhibitor (UGI) instead of dUDG, and the target DNA is chloroplast DNA.

[0425] 19. A DNA base correction method according to any one of the above-mentioned clauses 1 to 18, wherein the correction efficiency of a cytosine (C) base to a thymine (T) base is 50% or more, 60% or more, 70% or more, 80% or more, 90% or more, or 95% or more.

[0426] 20. A DNA base correction method according to the above-mentioned item 19, wherein the correction efficiency of cytosine (C) base to thymine (T) base in chloroplast DNA is 50% or more, 60% or more, 70% or more, 80% or more, 90% or more, or 95% or more.

[0427] 21. A plant cell or protoplast in which C-to-T base correction has been performed by the base correction method described in any one of the above-mentioned items 1 to 20.

[0428] 22. A plant or part thereof grown or cultured from a plant cell or protoplast according to the above-mentioned Article 21.

[0429] 23. A plant or part thereof which is a descendant or clone of a plant according to the above-mentioned paragraph 22.

[0430] 24. Seeds obtained from the plant according to the above-mentioned item 22 or 23.

[0431] 25. A plant or its offspring, or a part thereof, grown from a seed according to the above-mentioned clause 24.

[0432] 26. A plant cell or protoplast according to the above-mentioned item 25, a plant or a part thereof according to any one of the above-mentioned items 22, 23 and 25, or a seed according to the above-mentioned item 24, wherein the cytosine (C) base in the wild-type target DNA sequence is corrected to the thymine (T) base.

[0433] In another aspect, the present invention can be described with specific embodiments exemplified below based on the above, but is not limited thereto.

[0434] 1. Use of dead uracil DNA glycosylase (dUDG) for C-to-T base correction.

[0435] 2. The use of the above-mentioned first paragraph, wherein dUDG can bind to DNA containing uracil but does not have the activity of removing uracil.

[0436] 3. The use according to the above-mentioned first or second clause, wherein the dUDG has an amino acid mutation in motif A or motif B that does not exist in the wild-type UDG.

[0437] 4. The use according to any one of the above-mentioned claims 1 to 3, wherein the dUDG has an amino acid mutation in an amino acid sequence corresponding to the amino acid sequence KTVIIGQDPYHGPGQAMGLSF included in the UDG of Arabidopsis thaliana or an amino acid sequence corresponding to HPSGLSA included in the UDG of Arabidopsis thaliana.

[0438] 5. A use according to any one of the above-mentioned claims 1 to 4, wherein the amino acid mutation of the dUDG comprises a mutation of an amino acid corresponding to an aspartic acid (D), tyrosine (Y), phenylalanine (F), or glutamine (Q) residue in the amino acid sequence KTVIIGQDPYHGPGQAMGLSF included in Arabidopsis thaliana UDG to another amino acid, or a mutation of an amino acid corresponding to a histidine (H) or leucine (L) residue in the amino acid sequence HPSGLSA included in Arabidopsis thaliana UDG to another amino acid.

[0439] 6. A use according to any one of the above-mentioned claims 1 to 5, wherein the amino acid mutation of the dUDG comprises a mutation of an amino acid corresponding to an aspartic acid (D) residue in the amino acid sequence KTVIIGQDPYHGPGQAMGLSF included in Arabidopsis thaliana UDG to another amino acid, or a mutation of an amino acid corresponding to a histidine (H) residue in the amino acid sequence HPSGLSA included in Arabidopsis thaliana UDG to another amino acid.

[0440] 7. The use according to any one of the above-mentioned claims 1 to 6, wherein the dUDG comprises an amino acid sequence of SEQ ID NO: 1 or an amino acid sequence having at least 80% sequence identity therewith.

[0441] 8. A use according to any one of the above-mentioned claims 1 to 7, wherein C-to-T base correction occurs by a base editor comprising a DNA binding protein, cytosine deaminase, and dUDG, wherein the cytosine deaminase exists in a full-length form or in the form of two splits, and when it exists in the form of two splits, it exhibits cytosine deaminase activity through dimerization thereof.

[0442] 9. The use according to the above-mentioned paragraph 8, wherein the DNA binding protein is independently selected from the group consisting of a zinc finger protein, a TALE (transcription-activator-like effector) protein, and a CRISPR-associated nuclease.

[0443] 10. The use according to claim 8 or 9, wherein the cytosine deaminase is APOBEC (apolipoprotein B editing complex), AID (activation-induced deaminase), double-strand DNA-specific cytosine deaminase, or TadA (tRNA-specific adenosine deaminase) having C-to-T proofreading activity.

[0444] 11. In any one of the above-mentioned clauses 8 to 10, the DNA base editor comprises a DNA binding protein and a cytosine deaminase in the form of two fusion proteins,

[0445] The above cytosine deaminase is a double-stranded DNA-specific cytosine deaminase, which exists in the form of two split entities and exerts cytosine deaminase activity through their dimerization.

[0446] The two fusion proteins each independently contain one DNA binding protein, and are divided into two fragments.

[0447] use.

[0448] 12. The use according to the above-mentioned clause 11, wherein dUDG is included in at least one of the two fusion proteins.

[0449] Hereinafter, the present invention will be described in detail with reference to the following examples. However, the following examples are provided only to illustrate the present invention and the present invention is not limited thereto.

[0450] Example 1: Chloroplast psaA gene base correction

[0451] 1.1. Cloning of base editor and production of transformed plants

[0452] DNA encoding nucleotide editors targeting the psaA gene in chloroplasts of Arabidopsis (Arabidopsis thaliana) was cloned and transformed into transgenic plants using Agrobacterium. The composition of the nucleotide editors used is as follows, and the DNA sequences that bind via TALE proteins and the spacer regions where nucleotide editing takes place are as shown in Figure 3.

[0453] Number Fusion protein composition1CTS + 3xFlag + AtpsaA Left TALE + 1397N + UGI2CTS + 3xFlag + AtpsaA Right TALE + 1397C + UGI3CTS + 3xFlag + AtpsaA Left TALE + 1397N + dUDG4CTS + 3xFlag + AtpsaA Right TALE + 1397C + dUDG5CTS + 3xFlag + AtpsaA Left TALE + 1397C + UGI6CTS + 3xFlag + AtpsaA Right TALE + 1397N + UGI7CTS + 3xFlag + AtpsaA Left TALE + 1397C + dUDG8CTS + 3xFlag + AtpsaA Right TALE + 1397N + dUDG9CTS + 3xFlag + AtpsaA Right TALE + GSVG + UGI10CTS + 3xFlag + AtpsaA Right TALE + GSVG + dUDG

[0454] A linker is used between the protein components of the above fusion protein, and is indicated in the amino acid sequence of the fusion protein described herein.

[0455] Specifically, genetic constructs encoding base editors (fusion proteins) as shown in the table above were designed to be positioned between the RPS5A promoter and the 35S terminator, and cloned using Gibson assembly, Golden Gate, restriction enzymes, etc. into a vector suitable for transforming Agrobacterium tumefaciens, a strain capable of delivering genetic constructs to the Arabidopsis nuclear genome using T-DNA.

[0456] Specifically, plasmids corresponding to target DNAs were selected from a set of TALE subarray plasmids consisting of a total of 424 (6 × 64 tripartite plasmids + 2 × 16 bipartite plasmids + 2 × 4 monopartite plasmids). These subarray plasmids encode repeat units required for TALE proteins to recognize specific DNA sequences and were designed to contain a BsaI restriction enzyme recognition site for Golden Gate assembly. Each repeat unit has specificity for a specific nucleotide (e.g., NI for A, HD for C, NN for G, NG for T) (Kim, Y. et al., 2013). The selected TALE subarray plasmids are cleaved with Bsa I restriction enzyme and then ligated to form complementary sticky ends. This process results in the synthesis of a TALE array between the N-terminal and C-terminal domains (see Figure 19).

[0457] The assembled gene sequence was finally inserted into the destination vector to generate a base editor plasmid targeting a specific sequence (see Fig. 19). At this time, the destination vector was prepared in advance through DNA synthesis and Gibson assembly method, and it contained the RPS5A promoter, CTS sequence, 3XFlag tag, N-terminal domain, TALE repeat array insertion site, C-terminal domain, and DddA. tox It can include components such as a fragment and dUDG (deactivated uracil DNA glycosylase). This modular assembly method allows for the rapid and efficient construction of customized base-editing tools for various DNA target sequences.

[0458] DddA toxTo avoid intracellular toxicity, two plasmids containing each TALE sequence were synthesized so that each inactive fragment can bind to the target DNA, and these two plasmids were joined into one plasmid using the Golden Gate method (see Fig. 20). The Golden Gate assembly method cleaves the SapI recognition site and allows for the ligation with complementary sticky ends. This allows the gene fragments to be efficiently inserted into the destination vector in the correct direction and order. The central plasmid in Fig. 8B shows an intermediate insertion step, and successful insertion is indicated when GFP is removed. The final completed plasmid (bottom plasmid) has a spectinomycin resistance gene as a selectable marker and contains the right arm (RB; Right Border) and left arm (LB; Left Border) sequences of T-DNA, which are essential for Agrobacterium-mediated plant transformation. This final plasmid has a structure that allows the expression of a complete base editor under the control of the RPS5A promoter. It is designed to perform base editing functions by targeting specific DNA sequences within plant cells.

[0459] The finally obtained recombinant plasmid was transformed into Agrobacterium strain GV3101, and then Arabidopsis thaliana Colombia (Col-0) plants were transformed using the floral dipping technique according to the published method (Zhang et al., Nat. Protoc. 1, 641-646, 2006).

[0460] 1.2. Confirmation of base correction through sequencing and measurement of base correction efficiency

[0461] DNA was extracted from untransformed wild-type Col-0 plants and plants grown from first-generation seeds of transformed Arabidopsis, and sequences were analyzed using targeted deep sequencing. Base correction efficiency (frequency) was calculated as the percentage of sequencing reads that reflected the desired base correction among all sequencing reads.

[0462] As shown in Fig. 4, the experimental results showed that the C-to-T base correction efficiency was significantly higher when dUDG was connected than when UGI was connected. These results are consistent with the results obtained by split DddA. tox In the case of the dimeric base editor (left side of Fig. 4), as well as the monomeric DddA tox The same was observed in the case of the monomeric base editor (right side of Fig. 4).

[0463]

[0464] Example 2: Chloroplast 16S rRNA gene (site 1) base correction

[0465] 2.1. Cloning of base editors and production of transformed plants

[0466] DNA encoding base editors targeting the 16S rRNA gene (site 1) in the chloroplast of Arabidopsis (Arabidopsis thaliana) was cloned and transformed into transgenic plants using Agrobacterium. The composition of the base editors used is as follows, and the DNA sequence binding via the TALE protein and the spacer region where base editing takes place are as shown in Figure 5.

[0467] Number fusion protein composition 1 CTS + 3xFlag + 16S rRNA Left TALE + GSVG + UGI 2 CTS + 3xFlag + 16S rRNA Left TALE + GSVG + dUDG

[0468] A linker is used between the protein components of the above fusion protein, and is indicated in the amino acid sequence of the fusion protein described herein.

[0469] The TALE production and base editor cloning process is as described in Example 1.

[0470] 2.2. Confirmation of base correction through sequencing and measurement of base correction efficiency

[0471] DNA was extracted from untransformed wild-type Col-0 plants and plants grown from first-generation seeds of transformed Arabidopsis, and sequences were analyzed using targeted deep sequencing. Base correction efficiency (frequency) was calculated as the percentage of sequencing reads that reflected the desired base correction among all sequencing reads.

[0472] As shown in Figs. 6 to 9, the experimental results showed that the C-to-T base correction efficiency was significantly higher when dUDG was connected than when UGI was connected.

[0473]

[0474] Example 3: Chloroplast 16S rRNA gene (site 2) base correction

[0475] 3.1. Cloning of base editors and production of transformed plants

[0476] DNA encoding base editors targeting the 16S rRNA gene (site 2) in the chloroplast of Arabidopsis (Arabidopsis thaliana) was cloned and transformed into transgenic plants using Agrobacterium. The composition of the base editors used is as follows, and the DNA sequence binding via the TALE protein and the spacer region where base editing takes place are as shown in Figure 10.

[0477] Number Fusion protein composition1CTS + 3xFlag + 16S rRNA Site 2 Left TALE + 1397N + UGI2CTS + 3xFlag + 16S rRNA Site 2 Right TALE + 1397C + UGI3CTS + 3xFlag + 16S rRNA Site 2 Left TALE + 1397C + UGI4CTS + 3xFlag + 16S rRNA Site 2 Right TALE + 1397N + UGI5CTS + 3xFlag + 16S rRNA Site 2 Left TALE + 1397N + dUDG6CTS + 3xFlag + 16S rRNA Site 2 Right TALE + 1397C + dUDG7CTS + 3xFlag + 16S rRNA Site 2 Left TALE + 1397C + dUDG8CTS + 3xFlag + 16S rRNA Site 2 Right TALE + 1397N + dUDG9CTS + 3xFlag + 16S rRNA Site 2 Right TALE + GSVG + UGI10CTS + 3xFlag + 16S rRNA Site 2 Right TALE + GSVG + dUDG

[0478] A linker is used between the protein components of the above fusion protein, and is indicated in the amino acid sequence of the fusion protein described herein.

[0479] The TALE production and base editor cloning process is as described in Example 1.

[0480] 3.2. Confirmation of base correction through sequencing and measurement of base correction efficiency

[0481] DNA was extracted from untransformed wild-type Col-0 plants and plants grown from first-generation seeds of transformed Arabidopsis, and sequences were analyzed using targeted deep sequencing. Base correction efficiency (frequency) was calculated as the percentage of sequencing reads that reflected the desired base correction among all sequencing reads.

[0482] As shown in Fig. 11, the C-to-T base correction efficiency was significantly higher when dUDG was connected than when UGI was connected. This trend was observed in split DddA. tox As well as the dimeric base editor (Fig. 11 left), monomeric DddA tox The same was observed in the monomeric base editor (Fig. 11, right).

[0483]

[0484] Example 4: Chloroplast atpB gene base correction

[0485] 4.1. Cloning of base editors and production of transformed plants

[0486] Transgenic plants were created by cloning DNA encoding nucleotide editors targeting the atpB gene in chloroplasts of Arabidopsis (Arabidopsis thaliana) and transformation using Agrobacterium. The composition of the nucleotide editors used is as follows, and the DNA sequences that bind via TALE proteins and the spacer regions where nucleotide editing takes place are as shown in Figure 12.

[0487] Number fusion protein composition 1CTS + 3xFlag + atpB L TALE + GSVG + UGI 2CTS + 3xFlag + atpB L TALE + GSVG + dUDG

[0488] A linker is used between the protein components of the above fusion protein, and is indicated in the amino acid sequence of the fusion protein described herein.

[0489] The TALE production and base editor cloning process is as described in Example 1.

[0490] 4.2. Confirmation of base correction through sequencing and measurement of base correction efficiency

[0491] DNA was extracted from untransformed wild-type Col-0 plants and plants grown from first-generation seeds of transformed Arabidopsis, and sequences were analyzed using targeted deep sequencing. Base correction efficiency (frequency) was calculated as the percentage of sequencing reads that reflected the desired base correction among all sequencing reads.

[0492] As shown in Fig. 13, the C-to-T base correction efficiency was significantly higher when dUDG was connected than when UGI was connected.

[0493]

[0494] Example 5: Mitochondrial atp1 gene base correction

[0495] 5.1. Cloning of base editors and production of transformed plants

[0496] Transgenic plants were created by cloning DNA encoding nucleotide editors targeting the atp1 gene in Arabidopsis (Arabidopsis thaliana) mitochondria and transformation using Agrobacterium. The composition of the nucleotide editors used is as follows, and the DNA sequences that bind via the TALE protein and the spacer regions where nucleotide editing takes place are as shown in Figure 14.

[0497] Number of fusion proteins composition 1MTS + 3xFlag + atp1 Left TALE + 1397N + dUDG 2MTS + 3xFlag + atp1 Right TALE + 1397C + dUDG 3MTS + 3xFlag + atp1 Left TALE + 1397C + dUDG 4MTS + 3xFlag + atp1 Right TALE + 1397N + dUDG

[0498] A linker is used between the protein components of the above fusion protein, and is indicated in the amino acid sequence of the fusion protein described herein.

[0499] The TALE production and base editor cloning process is as described in Example 1.

[0500] 5.2. Confirmation of base correction through sequencing and measurement of base correction efficiency

[0501] DNA was extracted from untransformed wild-type Col-0 plants and plants grown from first-generation seeds of transformed Arabidopsis, and sequences were analyzed using targeted deep sequencing. Base correction efficiency (frequency) was calculated as the percentage of sequencing reads that reflected the desired base correction among all sequencing reads.

[0502] As shown in Figs. 15 and 16, the experimental results showed that the C-to-T base correction efficiency was significantly enhanced when dUDG was connected.

[0503]

[0504] Example 6: Nuclear gene base editing

[0505] 6.1. Cloning of base editors and production of transformed plants

[0506] DNA encoding nucleotide editors targeting the PDS3 and CESA3 genes in the Arabidopsis (Arabidopsis thaliana) nucleus was cloned and transformed into transgenic plants using Agrobacterium. The composition of the nucleotide editors used is as follows.

[0507] Number gRNA and fusion protein constructs 1 AtU6 promoter-PDS3_Q15 sgRNA-RPA5A promoter-rApobec-nCas9 (D10A)-dUDG-NLS-HA 2 AtU6 promoter-CESA3_Q15 sgRNA-RPA5A promoter-rApobec-nCas9 (D10A)-dUDG-NLS-HA

[0508] A linker is used between the protein components of the above fusion protein, and is indicated in the amino acid sequence of the fusion protein described herein.

[0509] Specifically, genetic constructs encoding base editors (fusion proteins) as shown in the table above were designed to be positioned between the RPS5A promoter and the 35S terminator, and the sgRNA was positioned downstream of the U6 promoter, and cloned into a vector suitable for transforming Agrobacterium tumefaciens, a strain capable of delivering genetic constructs to the Arabidopsis nuclear genome using T-DNA, using Gibson assembly, Golden Gate, restriction enzymes, etc. Finally, the obtained recombinant plasmid was transformed into the Agrobacterium strain GV3101, and then Arabidopsis thaliana Columbia (Col-0) plants were transformed using the floral dipping technique according to the known method (Zhang et al., Nat. Protoc. 1, 641-646, 2006).

[0510] 6.2. Confirmation of base correction through sequencing and measurement of base correction efficiency

[0511] DNA was extracted from untransformed wild-type Col-0 plants and plants grown from first-generation seeds of transformed Arabidopsis, and sequences were analyzed using targeted deep sequencing. Base correction efficiency (frequency) was calculated as the percentage of sequencing reads that reflected the desired base correction among all sequencing reads.

[0512] As shown in Figs. 17 and 18, the experimental results showed that the C-to-T base correction efficiency was significantly enhanced when dUDG was connected.

[0513]

[0514] order

[0515] The amino acid sequences of the proteins used in the above examples are as follows.

[0516] CTS: MDSQLVLSLKLNPSFTPLSPLFPFTPCSSFSPSLRFSSCYSRRLYSPVTVYAAK

[0517] MTS: MFKQASRLSRSVAAAASSKSVTTRAFSTELPSTLDS

[0518] NLS: PKKKRKV

[0519] 3xFlag: DYKDHDGDYKDHDIDYKDDDDK

[0520] HA: YPYDVPDYA

[0521] 1397N: GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGTPPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEG

[0522] 1397C: AIPVKRGATGETKVFTGNNSNSPKSPTKGGC

[0523] GSVG: GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLEGKVFSGSGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGVIPVKRGATGETKVFTGNSNGPKSPTKGGC

[0524] AD: SEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN

[0525] UGI: TNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML

[0526] UDG: MASSTPKTLMDFFQPAKRLKASPSSSSFPAVSVAGGSRDLGSVANSPPRVTVTTSVADDSSGLTPEQIARAEFNKFVAKSKRNLAVCSERVTKAKSEGNCYVPLSELLVEESWLKALPGEFHKPYAKSLSDFLEREIITDSKPLIYPPQHLIFNALNTTPFDRV KTVIIGQDPYHGPGQAMGLSFSVPEGKLPSSLLNIFKELHKDVGCSIPRHGNLQKWAVQGVLLLNAVLTVRSKQPNSHAKKGWEQFTDAVIQSISQQKEGVVFLLWGRYAQEKSKLIDATKHHILTAAHPSGLSANRGFFDCRHFSRANQLLEEMGIPPIDWQL

[0527] dUDG: MASSTPKTLMDFFQPAKRLKASPSSSSFPAVSVAGGSRDLGSVANSPPRVTVTTSVADDSSGLTPEQIARAEFNKFVAKSKRNAVCSERVTKAKSEGNCYVPLSELLVEESWLKALPGEFHKPYAKSLSDFLEREIITDSKSPLIYPPQHLIFNALNTTPFDRVKTVIIGQNPYHGPGQAMGLSFSVPEGEKLPSSLLNIFKELHKDVGCSIPRHGNLQKWAVQGVLLNAVLTVRSKQPNSHAKKGWEQFTDAVIQSISQQKEGVVFLLWGRYAQEKSKLIDATKHHILTAAHPSGLSANRGFFDCRHFSRANQLLEEMGIPIDWQL

[0528] AtpsaA Left TALE: DLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLG

[0529] AtpsaA Right TALE: DLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLG

[0530] 16S rRNA site 1 Left TALE: DLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLG

[0531] rApobec: MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSIWRHTSQNTNKHVEVNFIEKFTTERYFCPNTRCSITWFLSWSPCGECSRAITEFLSRYPHVTLFIYIARLYHHADPRNRQGLRDLISSGVTIQIMTEQESGYCWRNFVNYSPSNEAHWPRYPHLWVRLYVLELYCIILGLPPCLNILRRKQPQLTFFTIALQSCHYQRLPPHILWATGLKSGSETPGTSESATPES

[0532] nCas9 (D10A): IAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRLKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD

[0533] 16S rRNA site 1 Right SPEECH:

[0534] DLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFFTAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQ LDTGQLLKIAKRGGVTAAVHAWRNALTGAPLNLTQVVAIASNNGGKQALETVQRLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDH GLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQD HGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQ AHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLC QDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVL CQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTQVVAIASNNGGKQALETVQRLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPV LCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVQLSRPDPALALTNDHLVALACLGGRPALDAVKKGLG

[0535] 16S rRNA Site 2 Left TALE: DLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLG

[0536] 16S rRNA Site 2 Right TALE: DLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLG

[0537] atpB Left AGAIN: DLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFFTAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQ LDTGQLLKIAKRGGVTAAVHAWRNALTGAPLNLTPEQVVAIASNGGGKQALETVQRLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDH GLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQA HGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQ AHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLC QDHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLPVLCQDHGLTPEQVVAIASNGGGKQALETVQRLLPVL CQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLPV LCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNGGGKQALESIVQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLG

[0538] atp 1 Left TALE: DLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLG

[0539] atp 1 Right TALE:

[0540] DLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLG

[0541] CTS + 3xFlag + AtpsaA Left TALE + 1397N + UGI:

[0542]

[0543] CTS + 3xFlag + AtpsaA Right TALE + 1397C + UGI:

[0544]

[0545] CTS + 3xFlag + AtpsaA Left TALE + 1397N + dUDG:

[0546]

[0547] CTS + 3xFlag + AtpsaA Right TALE + 1397C + dUDG:

[0548]

[0549] CTS + 3xFlag + AtpsaA Left TALE + 1397C + UGI:

[0550]

[0551] CTS + 3xFlag + AtpsaA Right TALE + 1397N + UGI:

[0552]

[0553] CTS + 3xFlag + AtpsaA Left TALE + 1397C + dUDG:

[0554]

[0555] CTS + 3xFlag + AtpsaA Right TALE + 1397N + dUDG:

[0556]

[0557] CTS + 3xFlag + AtpsaA Right TALE + GSVG + UGI:

[0558]

[0559] CTS + 3xFlag + AtpsaA Right TALE + GSVG + dUDG:

[0560]

[0561] CTS + 3xFlag + 16S rRNA site 1 Left TALE + GSVG + UGI:

[0562] MDSQLVLSLKLNPSFTPLSPLFPFTPCSSFSPSLRFSSCYSRRLYSPVTVYAAKDYKDHDGDYKDHDIDYKDDDDKPGSGSDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLEGKVFSSGGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGVIPVKRGATGETKVFTGNSNGPKSPTKGGCGSSGTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML (밑줄친 부분은 링커에 해당함)

[0563] CTS + 3xFlag + 16S rRNA site 1 Left TALE + GSVG + dUDG:

[0564]

[0565] CTS + 3xFlag + 16S rRNA Site 2 Left TALE + 1397N + UGI:

[0566]

[0567] CTS + 3xFlag + 16S rRNA Site 2 Right TALE + 1397C + UGI:

[0568]

[0569] CTS + 3xFlag + 16S rRNA Site 2 Left TALE + 1397C + UGI:

[0570]

[0571] CTS + 3xFlag + 16S rRNA Site 2 Right TALE + 1397N + UGI:

[0572]

[0573] CTS + 3xFlag + 16S rRNA Site 2 Left TALE + 1397N + dUDG:

[0574]

[0575] CTS + 3xFlag + 16S rRNA Site 2 Right TALE + 1397C + dUDG:

[0576]

[0577] CTS + 3xFlag + 16S rRNA Site 2 Left TALE + 1397C + dUDG:

[0578]

[0579] CTS + 3xFlag + 16S rRNA Site 2 Right TALE + 1397N + dUDG:

[0580]

[0581] CTS + 3xFlag + 16S rRNA Site 2 Right TALE + GSVG + UGI:

[0582]

[0583] CTS + 3xFlag + 16S rRNA Site 2 Right TALE + GSVG + dUDG:

[0584]

[0585] CTS + 3xFlag + atpB L TALE + GSVG + UGI:

[0586]

[0587]

[0588] MTS + 3xFlag + atp 1 Left TALE + 1397N + dUDG:

[0589]

[0590]

[0591]

[0592]

[0593] rApobec-nCas9 (D10A)-dUDG-NLS-HA:

[0594] MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSIWRHTSQNTNKHVEVNFIEKFTTERYFCPNTRCSITWFLSWSPCGECSRAITEFLSRYPHVTLFIYIARLYHHADPRNRQGLRDLISSGVTIQIMTEQESGYCWRNFVNYSPSNEAHWPRYPHLWVRLYVLELYCIILGLPPCLNILRRKQPQLTFFTIALQSCHYQRLPPHILWATGLKSGSETPGTSESATPESSGSETPGTSESATPESIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGDSGSETPGTSESATPESMASSTPKTLMDFFQPAKRLKASPSSSSFPAVSVAGGSRDLGSVANSPPRVTVTTSVADDSSGLTPEQIARAEFNKFVAKSKRNLAVCSERVTKAKSEGNCYVPLSELLVEESWLKALPGEFHKPYAKSLSDFLEREIITDSKSPLIYPPQHLIFNALNTTPFDRVKTVIIGQNPYHGPGQAMGLSFSVPEGEKLPSSLLNIFKELHKDVGCSIPRHGNLQKWAVQGVLLLNAVLTVRSKQPNSHAKKGWEQFTDAVIQSISQQKEGVVFLLWGRYAQEKSKLIDATKHHILTAAHPSGLSANRGFFDCRHFSRANQLLEEMGIPPIDWQLGGSGPPKKKRKVYPYDVPDYA (밑줄친 부분은 링커에 해당함)

[0595] The sgRNA sequences used in the above examples are as follows.

[0596] PDS3_Q15: GCCTTATCAAAACGGGTTTT

[0597] CESA3_S983: GTCTCTTATGCTATCAACAG

Claims

1. A method for correcting a cytosine (C) base to a thymine (T) base, comprising introducing a DNA base editor into a cell containing target DNA for base correction or expressing the DNA base editor within a cell containing target DNA for base correction. The DNA base editor comprises a DNA binding protein, a cytosine deaminase, and a dead uracil DNA glycosylase (dUDG), wherein the cytosine deaminase exists in a full-length form or in the form of two splits, and when it exists in the form of two splits, it exhibits cytosine deaminase activity through dimerization thereof. Base editing method.

2. A base correction method in the first paragraph, wherein the dUDG can bind to DNA containing uracil but does not have the activity of removing uracil.

3. A base correction method in claim 1, wherein the dUDG has an amino acid mutation in motif A or motif B that does not exist in wild-type UDG.

4. A base correction method in claim 1, wherein the dUDG has an amino acid mutation in an amino acid sequence corresponding to the amino acid sequence KTVIIGQDPYHGPGQAMGLSF included in UDG of Arabidopsis thaliana or an amino acid sequence corresponding to HPSGLSA included in UDG of Arabidopsis thaliana.

5. A base correction method according to claim 4, wherein the amino acid mutation is a mutation in which an amino acid corresponding to an aspartic acid (D), tyrosine (Y), phenylalanine (F), or glutamine (Q) residue in the amino acid sequence KTVIIGQDPYHGPGQAMGLSF included in Arabidopsis thaliana UDG is mutated to another amino acid, or an amino acid corresponding to a histidine (H) or leucine (L) residue in the amino acid sequence HPSGLSA included in Arabidopsis thaliana UDG is mutated to another amino acid.

6. A base correction method according to claim 4, wherein the amino acid mutation includes a mutation of an amino acid corresponding to an aspartic acid (D) residue in the amino acid sequence KTVIIGQDPYHGPGQAMGLSF included in Arabidopsis thaliana UDG to another amino acid, or a mutation of an amino acid corresponding to a histidine (H) residue in the amino acid sequence HPSGLSA included in Arabidopsis thaliana UDG to another amino acid.

7. A base correction method according to claim 1, wherein the DNA binding protein is independently selected from the group consisting of a zinc finger protein, a TALE (transcription-activator-like effector) protein, and a CRISPR-associated nuclease.

8. A base correction method according to claim 1, wherein the base correction method is a double strand DNA-specific cytosine deaminase.

9. In the first paragraph, the DNA base editor comprises a DNA binding protein and a cytosine deaminase in the form of two fusion proteins, The above cytosine deaminase is a double-stranded DNA-specific cytosine deaminase, which exists in the form of two split entities and exerts cytosine deaminase activity through their dimerization. The two fusion proteins each independently contain one DNA binding protein, and are divided into two fragments. Base editing method.

10. A base correction method according to claim 9, wherein dUDG is included in at least one of the two fusion proteins.

11. A base correction method according to claim 1, wherein the target DNA is chloroplast DNA and the DNA base editor optionally additionally includes a chloroplast transit signal (CTS) and / or NES.

12. A base correction method according to claim 1, wherein the target DNA is mitochondrial DNA, and the DNA base editor optionally additionally includes a mitochondrial transfer signal (MTS) and / or a nuclear export signal (NES).

13. A base correction method according to claim 1, wherein the target DNA is nuclear DNA and the DNA base editor optionally additionally includes a nuclear localization signal (NLS).

14. A base correction method according to claim 1, wherein the target DNA is DNA of a plant cell.

15. A DNA base editing method characterized in that, in the first paragraph, the DNA base editor has a superior efficiency in correcting cytosine (C) bases to thymine (T) bases compared to an editor of the same configuration, except that the DNA base editor does not include dUDG.

16. A DNA base correction method in which the DNA base editor has a superior correction efficiency of a cytosine (C) base to a thymine (T) base compared to a base editor of the same configuration except that the DNA base editor does not contain dUDG, and the target DNA is chloroplast DNA.

17. A DNA base editing method characterized in that, in the first paragraph, the DNA base editor has a superior efficiency in correcting a cytosine (C) base to a thymine (T) base compared to a base editor of the same configuration, except that the DNA base editor uses a uracil glycosylase inhibitor (UGI) instead of dUDG.

18. A DNA base correction method in which the DNA base editor has a superior correction efficiency of a cytosine (C) base to a thymine (T) base compared to a base editor of the same configuration, except that the DNA base editor uses a uracil glycosylase inhibitor (UGI) instead of dUDG in the 17th paragraph, and the target DNA is chloroplast DNA.

19. A DNA base correction method in paragraph 1, wherein the correction efficiency of cytosine (C) base to thymine (T) base is 50% or more, 60% or more, 70% or more, 80% or more, 90% or more, or 95% or more.

20. A DNA base correction method according to claim 19, wherein the correction efficiency of cytosine (C) base to thymine (T) base in chloroplast DNA is 50% or more, 60% or more, 70% or more, 80% or more, 90% or more, or 95% or more.

21. A plant cell or protoplast in which C-to-T base correction has been performed by the base correction method described in paragraph 1.

22. A plant or part thereof grown, cultured or propagated from a plant cell or protoplast according to Article 21.

23. Seeds obtained from the plant according to Article 22.

24. A plant or part thereof according to paragraph 22, or a seed according to paragraph 23, wherein the cytosine (C) base in the wild-type target DNA sequence is corrected to the thymine (T) base, or a plant or part thereof, or a seed.

Citation Information

Patent Citations

  • Continuous waste supply device for waste pyrolysis device

    KR102608575B1

  • Exhaust structure of mold

    KR102756855B1

  • High efficiency base editors comprising gam

    WO2019139645A2

  • Base-editing systems

    WO2021087246A1

  • KR20190127797A