Adenine base correction system

A base correction composition with a DNA-binding protein and full-length cytosine deaminase addresses the challenge of selectively correcting adenine bases in plant organelles, providing efficient and targeted genetic modifications.

JP2026507910APending Publication Date: 2026-03-06GREENGENE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025552156
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-03-06
Filing Date
2024-03-06
Publication Date
2026-03-06

AI Technical Summary

Technical Problem

Conventional gene correction tools are ineffective for selectively correcting adenine bases in plant organelles like mitochondria and chloroplasts without affecting cytosine bases, and existing cytosine base editors are toxic and require large vectors.

Method used

A base correction composition comprising a DNA-binding protein, full-length cytosine deaminase, and adenine deaminase, with optional uracil-DNA glycosylase, selectively corrects adenine bases to guanine without affecting cytosine, usable in both nuclear and organelle DNA.

Benefits of technology

Achieves selective adenine-to-guanine correction in plant and animal cells, including organelles, without the limitations of conventional tools, enabling targeted genetic modifications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026507910000001_ABST
    Figure 2026507910000001_ABST
Patent Text Reader

Abstract

The present invention relates to a base correction composition having the activity of correcting adenine A bases in DNA to guanine G bases, and a method for base correction using the same. The present invention is useful for correcting bases in nuclear DNA or organelle DNA, particularly in organelle DNA such as chloroplasts and mitochondria. According to one aspect of the present invention, the base correction composition comprises a DNA-binding protein, cytosine deaminase, adenine deaminase, and uracil-DNA glycosylase (UDG), and such a base correction composition has the activity of selectively correcting only adenine bases, substantially without correcting cytosine C bases. The present invention also provides a system for correcting adenine bases in plant cell DNA to guanine bases using a single fusion protein comprising a DNA-binding protein, cytosine deaminase, and adenine deaminase.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a base correction composition having the activity of correcting adenine A bases in DNA to guanine G bases, and a base correction method using the same.

[0002] The present invention is useful for correcting bases in nuclear DNA or organelle DNA, particularly in organelle DNA such as chloroplasts and mitochondria. According to one aspect of the present invention, a base correction composition comprises a DNA-binding protein, cytosine deaminase, adenine deaminase, and uracil-DNA glycosylase (UDG), and such a base correction composition has the activity of selectively correcting only adenine bases without substantially correcting cytosine C-bases. The present invention also provides a system for correcting adenine bases in plant cell DNA to guanine bases using a single fusion protein comprising a DNA-binding protein, cytosine deaminase, and adenine deaminase. [Background technology]

[0003] Fusion proteins combining a DNA-binding protein and a deaminase allow for DNA mutagenesis by substituting nucleotides in a gene and / or correcting bases and / or correcting point mutations that induce genetic disorders and / or transducing single nucleotides in a targeted manner to introduce desired single nucleotide mutations in prokaryotic cells and eukaryotic cells such as humans, without generating DNA double-strand breaks (DSBs).

[0004] Programmable gene construction tools, such as zinc-finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), clustered regularly interspaced short palindromic repeat (CRISPR) systems, and base editors composed of CRISPR-associated protein 9 (Cas9) variants and nucleotide deaminase proteins, have the potential to be used for plant genetic research and crop trait improvement through base sequence alterations. However, these conventional gene correction tools are not suitable for correcting DNA bases in plant organelles, including mitochondria and chloroplasts, particularly because they cannot deliver the guide RNAs required to activate the most widely used CRISPR systems. Plant organelles encode various essential genes required for photosynthesis and cellular respiration. Methods and tools for correcting genes in these organelles are critically needed for functional studies of these genes and for crop productivity and trait improvement.

[0005] The bacterial toxin DddAtox is an enzymatic portion of a bacterial toxin derived from Burkholderia cenocepacia that can deaminate cytosines in double-stranded DNA. Because DddAtox is toxic to cells, it is split into two inactive splits to avoid toxicity in host cells. Each split (or half) can be linked to a DNA-binding protein designed to bind to DNA and used as a pair to form a functional cytosine base editor (DdCBE, DddA-derived cytosine base editor).

[0006] In principle, this deamination enzymatic reaction is activated by the proximity of two inactivated halves on the target DNA by DNA-binding proteins, and a cytosine-to-thymine (C-to-T) base conversion occurs between the binding sites of the two DNA-binding proteins. When a TALE (transcription activator-like effector) protein is used as the DNA-binding protein, the cytosine-to-thymine base conversion is induced in a region of 14–18 bases between the two TALE binding sites. Therefore, the efficiency of base correction may vary depending on the DNA-binding specificity of the TALE protein.

[0007] On the other hand, by linking DddAtox cytosine deaminase and adenine deaminase, which can correct adenine to guanine (A-to-G), to a DNA-binding protein, the adenine base can be corrected. Using a TALE (transcription activator-like effector) protein as the DNA-binding protein, a base editor linking cytosine deaminase and adenine deaminase (TALED, TALE-linked deaminase) can introduce a diverse spectrum of mutations, unlike cytosine base editors that only correct C-to-T. Summary of the Invention [Problem to be solved by the invention]

[0008] Conventional TALEDs contain both adenine deaminase and cytosine deaminase, resulting in the correction of not only adenine bases but also cytosine bases present in target DNA. Experimental results have confirmed that this is particularly pronounced in plant chloroplasts. Because there is a need to selectively correct only adenine bases while maintaining the wild-type sequence of cytosine bases, the inventors of the present invention aimed to develop a method that can selectively correct only adenine bases to guanine bases without causing the correction of undesired cytosine bases.

[0009] On the other hand, DddAtox cytosine deaminase is used in the form of two inactivated segments to avoid toxicity. Therefore, two TALE modules must be linked to each segment so that each segment is located close to the target DNA to be corrected. However, the use of two TALE modules in this manner has disadvantages, such as limitations on the selection of target sites for DNA correction and an increase in the size of the vector required to express the base editor. Therefore, the inventors of the present invention provide a method for correcting adenine bases in plant cell DNA using a monomeric TALED linked to a full-length DddAtox cytosine deaminase from which toxicity has been removed. Another object of the present invention is to provide a method for selectively correcting only adenine bases without causing correction of undesired cytosine bases using such a monomeric TALED linked to a full-length DddAtox cytosine deaminase. [Means for solving the problem]

[0010] The present invention provides a base correction composition for correcting adenine bases to guanine, comprising a DNA-binding protein, cytosine deaminase, and adenine deaminase, and additionally containing uracil-DNA glycosylase (UDG). The base correction composition selectively corrects only adenine bases, correcting the corresponding adenine base to guanine, without substantially causing correction of cytosine bases. Using this base correction composition, a base correction system capable of selectively correcting only adenine bases can be provided, which can be used to correct not only nuclear DNA but also organelle DNA such as chloroplasts and mitochondria. The base correction composition of the present invention can be used in both plant and animal cells. In particular, by transforming plant cells with the base correction composition of the present invention, plants or their seeds in which only the adenine base of a desired target DNA is selectively corrected can be obtained.

[0011] The present invention also provides a base correction composition for correcting adenine bases in plant cells to guanine, which comprises a DNA-binding protein, cytosine deaminase, and adenine deaminase, wherein the cytosine deaminase is contained in a full-length form rather than in a split form. Use of the base correction composition eliminates the need for a fusion protein containing two or more DNA-binding proteins (preferably TALE proteins) and deaminase, allowing for unlimited selection of target positions for DNA correction. [Brief explanation of the drawings]

[0012] [Figure 1] Schematic diagram of a base editor targeting the plant chloroplast gene psaA using a DddAtox fragment, showing the absence (A) and presence (B) of UDG. CTS stands for chloroplast transit signal, AD for adenine deaminase, NTD for N-terminal domain, CTC for C-terminal domain, and UDG for uracil DNA glycosylase. "Right repeats" and "Left TALE repeats" each represent a TALE array, and the underlined sequences indicate the binding sites of TALE proteins (consisting of an N-terminal domain, TALE array, and C-terminal domain). Base correction by deaminase occurs between the DNA sequences bound by the two TALE proteins. [Figure 2] Schematic diagrams of a base editor targeting the plant chloroplast gene psaA using full-length DddAtox GSVG, showing the absence (A) and presence (B) of UDG. GSVG refers to a mutant (SEQ ID NO: 9) in which the cytosine deaminase contains the amino acid sequence of SEQ ID NO: 2, with S at position 37, G at position 59, A at position 109, and S at position 129 replaced with G, S, V, and G, respectively. [Figure 3]This figure shows the base correction efficiency of C2, the position where cytosine base correction occurs in the psaA target DNA in a first-generation plant transformed with the base editor of the present invention. C2 refers to the second base, cytosine, in the DNA sequence to be corrected, located between the first TALE protein (left TALE: SEQ ID NO: 10) and the second TALE protein (right TALE: SEQ ID NO: 11) of the base correction composition used as shown in Figures 1 and 2. Col-0 refers to an untransformed wild-type plant. Figure A shows the case where a pair of TALEs linked to DddAtox fragments was used, and Figure B shows the case where a monomeric TALED linked to full-length DddAtox was used. Col-0 refers to an untransformed wild-type plant. "L" indicates the left TALE having the amino acid sequence of SEQ ID NO: 10, and "R" indicates the right TALE having the amino acid sequence of SEQ ID NO: 11. "AD" indicates adenine deaminase having the amino acid sequence of SEQ ID NO: 1. "GSVG" indicates full-length DddAtox GSVG having the amino acid sequence of SEQ ID NO: 9. [Figure 4] The base correction rate at the psaA target position in three wild-type plants (Col-0) is shown. Since the theoretical value should be zero, the displayed value is interpreted as being due to sequencing errors, and this can be considered as a reference value for determining whether actual base correction has occurred. [Figure 5] This figure shows the base correction efficiency of psaA target DNA in 24 first-generation plants transformed with a base editor that uses a DddAtox fragment and a TALE pair that does not use UDG. L-1397N indicates the left TALE linked to the N-terminal DddAtox fragment, and R-1397C-AD indicates the right TALE linked to the C-terminal DddAtox fragment and adenine deaminase. [Figure 6]The DddAtox fragment was used, and the base correction efficiency of psaA target DNA was shown in 24 first-generation plants transformed with a base editor using a TALE pair in which UDG was linked to the left TALE and adenine deaminase was linked to the right TALE. [Figure 7] The DddAtox fragment was used, and the efficiency of base correction of psaA target DNA was shown in 24 first-generation plants transformed with a base editor using a TALE pair in which adenine deaminase and UDG were linked to the right TALE. [Figure 8] The DddAtox fragment was used, and the base correction efficiency of psaA target DNA was shown in 24 first-generation plants transformed with a base editor using a TALE pair in which UDG was linked to the left TALE and adenine deaminase and UDG were linked to the right TALE. [Figure 9] This shows the base correction efficiency of psaA target DNA in 17 first-generation plants transformed with a base editor using a TALE pair in which adenine deaminase was linked to the left TALE, using a DddAtox fragment but not UDG. [Figure 10] The DddAtox fragment was used, and the efficiency of base correction of psaA target DNA was shown in 14 first-generation plants transformed with a base editor using a TALE pair in which adenine deaminase and UDG were linked to the left TALE. [Figure 11] The DddAtox fragment was used, and the base correction efficiency of psaA target DNA was shown in 24 first-generation plants transformed with a base editor using a TALE pair in which adenine deaminase was linked to the left TALE and UDG was linked to the right TALE. [Figure 12]The DddAtox fragment was used, and the base correction efficiency of psaA target DNA was shown in 16 first-generation plants transformed with a base editor using a TALE pair in which adenine deaminase and UDG were linked to the left TALE and UDG was linked to the right TALE. [Figure 13] The full-length DddAtox mutant was used, and the base correction efficiency of the psaA target was shown in 20 first-generation plants transformed with the UDG-free monomeric TALED base editor. [Figure 14] The full-length DddAtox mutant was used, and the base correction efficiency of psaA target DNA was shown in 29 first-generation plants transformed with a UDG-linked monomeric TALED base editor. [Figure 15] Schematic diagram of a base editor targeting the plant chloroplast gene psbA using DddAtox fragments or full-length DddAtox (SEQ ID NO: 9). CTS represents the chloroplast transit signal, NTD represents the N-terminal domain of the TALE protein, and CTD represents the C-terminal domain of the TALE protein. "Left TALE repeats" and "Right TALE repeats" represent the TALE array, respectively. The underlined sequences represent the binding sites of the TALE protein (composed of the N-terminal domain, TALE array, and C-terminal domain). Base correction by deaminase occurs between the DNA sequences bound by the two TALE proteins. Figure 15A shows the results using DddAtox fragments (1397N and 1397C), while Figures 15B and 15C show the results using full-length DddAtox ("GSVG"; SEQ ID NO: 9) with only the Left TALE repeats and / or only the Right TALE repeats, respectively. [Figure 16]The results of measuring the base correction efficiency for each individual that survived atrazine treatment of first-generation plants transformed with the base editors shown in Figure 15 are shown. Col-0 is an untransformed wild-type individual. "Left1" and "L1" have the same meaning and refer to the left TALE protein of SEQ ID NO: 14. "L2" refers to the left TALE sequence of SEQ ID NO: 15. "R1" refers to the right TALE protein of SEQ ID NO: 16. "Right 2" and "R2" have the same meaning and refer to the right TALE protein of SEQ ID NO: 17. DETAILED DESCRIPTION OF THE INVENTION

[0013] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one skilled in the art to which this invention belongs. Generally, the terms used herein are well known and commonly used in the art.

[0014] As used herein, the terms "correction," "editing," and "editing" are used interchangeably and refer to a method of altering a nucleic acid sequence by selective mutation of a specific genomic target, including but not limited to a gene, promoter, open reading frame, or any nucleic acid sequence.

[0015] As used herein, the terms "base editor," "base editing system," and "base correction system" are used interchangeably and refer to a substance that has the activity of altering a nucleic acid sequence by selectively mutating a genomic target, and include a combination of one or more different base editors. Depending on the context, the terms "base editor," "base editing system," or "base correction system" may be in the form of a polypeptide (which may be a fusion protein) or a polynucleotide, or a combination thereof, or may refer to a composition comprising one or more polypeptides (which may be fusion proteins) or polynucleotides, or a combination thereof. Therefore, as used herein, the term "base correction system" or "base correction composition" may include a single base editor or a combination of two or more different base editors, where the different base editors can be used simultaneously or separately.

[0016] As used herein, the term "fusion protein" refers to a polypeptide formed by the linkage of two or more different polypeptides via peptide bonds. Fusion proteins used in the present invention include DNA-binding proteins and deaminase enzymes, or fragments thereof, and may contain additional sequences such as UDG, NLS, NES, MTS, and CTS, or may contain additional sequences for biotechnology techniques such as tags. These polypeptides may be directly linked and / or linked via a linker. A linker refers to any molecule that links two different molecules, and any linker known in the biotechnology field to be useful for providing fusion proteins or protein conjugates may be used. For example, a peptide linker containing 1 to 100 amino acid residues may be used.

[0017] The method for designing and constructing a fusion protein or a polynucleotide encoding the same can be any method known in the art, and the polynucleotide can be inserted into a vector, which can then be introduced into a cell. The individual proteins constituting the fusion protein of the present invention are typically cloned into a single polynucleotide and expressed as a single polypeptide (fusion protein). However, one or more of the individual proteins can also be cloned into separate polynucleotides and expressed as two or more separate polynucleotides, and such cases also fall within the scope of the present invention.

[0018] As used herein, the terms "target," "target," "target site," or "target region" refer to a pre-defined nucleic acid sequence of any composition and / or length. Such a target region includes, but is not limited to, a gene, a promoter, an open reading frame, or any nucleic acid sequence.

[0019] The present invention provides a base correction composition having the activity of correcting adenine bases in DNA to guanine, the base correction composition comprising one or more fusion proteins, each of which comprises a DNA binding protein and cytosine deaminase, at least one of which comprises adenine deaminase, and at least one of which comprises UDG, wherein the cytosine deaminase can exist in the form of a full-length or two fragments.

[0020] The base correction composition according to the present invention is characterized by containing both cytosine deaminase and adenine deaminase. In particular, a TALE protein is used as a DNA-binding protein, and a base editor in which both cytosine deaminase and adenine deaminase are linked is called a TALED. Unlike previous DdCBE base editors, which only allowed for cytosine base correction of specific motifs in organelle DNA, TALED allows for A-to-G base correction. In the present specification, the term "TALED" refers to a base editor in which both cytosine deaminase and adenine deaminase are linked to a TALE protein. The TALE protein can be used in the form of a monomer (e.g., when cytosine deaminase is used in its full-length form) or in the form of a paired dimer (e.g., when cytosine deaminase is used in its split form). Adenine deaminase can be linked to all or only one of the TALE proteins.

[0021] The base-correction composition of the present invention can be used to correct not only nuclear DNA but also organelle DNA such as chloroplasts and mitochondria. The base-correction composition of the present invention can be used in both plant and animal cells, and in particular, by transforming plant cells with the base-correction composition of the present invention, it can be used to obtain plants or their seeds in which only the adenine base of the desired target DNA is selectively corrected.

[0022] The cytosine deaminase used in the present invention may be in the form of a first segment and a second segment, and the first segment and the second segment each have a form that can be bound to a DNA-binding protein.

[0023] A cytosine deaminase that can be used in a base editor according to the present invention refers to an amino-deaminase that has the activity of converting cytosine to uridine, and can be derived from and / or mutated (e.g., engineered and / or evolved) any organism (e.g., eukaryote or prokaryote), including, but not limited to, algae, bacteria, fungi, plants, invertebrates, and mammals. For example, the cytosine deaminase can be derived from and / or mutated by APOBEC (apolipoprotein B editing complex), AID (activation-induced deaminase), the bacterial adenine deaminase TadA (tRNA-specific adenosine deaminase), or an ortholog thereof, or a cytosine deaminase or fragment thereof derived from and / or mutated by the bacterial cytosine deaminase DddA or an ortholog thereof. The cytosine deaminase mutated from TadA mentioned above may be, for example, a polypeptide in which one or more of the amino acid residues at positions 6, 26, 27, 28, 46, 48, 49, 61, 74, 76, 77, 82, 96, 107, 108, 112, 114, 115, 119, 122, 127, 142, 143, 151, 154, and 158 of the amino acid sequence of SEQ ID NO: 1 are mutated to another amino acid. For example, the polypeptide may be a polypeptide in which the amino acid at position 27 of the amino acid sequence of SEQ ID NO: 1 is mutated to lysine, the amino acid at position 28 to alanine, the amino acid at position 61 to isoleucine, and the amino acid at position 96 to asparagine. With regard to the configuration of cytosine deaminase that can be used in the present invention, reference may be made to content that was already known prior to the filing of this application, including International Patent Application Publications WO2022 / 060185, WO2023 / 086953, etc., which are incorporated by reference in their entirety in this application.

[0024] When a base editor according to the present invention includes a cytosine deaminase, the cytosine deaminase may be in the form of a first segment and a second segment, or may be in the form of a full-length segment. When the cytosine deaminase has the form of a first segment and a second segment, the first segment and the second segment each have a form that allows them to be linked to a DNA-binding protein.

[0025] As used herein, when two proteins are "linked," they may be directly linked or indirectly linked via a linker or another protein (or proteins).

[0026] The cytosine deaminase used in the present invention may be DddAtox, which is a part of a bacterial toxin derived from Burkholderia cenocepacia that exhibits enzymatic function and can deaminate cytosine in double-stranded DNA. DddAtox may comprise the amino acid sequence of SEQ ID NO: 2.

[0027] (SEQ ID NO: 2) wild-type DddAtox GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC

[0028] Since DddAtox is toxic to cells, it can be used in the form of two inactivated split bodies, i.e., the first split body and the second split body, to avoid toxicity in host cells. When the cytosine deaminase used in the present invention is used in the form of the first split body and the second split body, each of the first split body and the second split body has no deamination activity.

[0029] The first segment of the DddAtox cytosine deaminase can comprise the sequence from the N-terminus to G33, G44, A54, N68, G82, N98, or G108 in the amino acid sequence of SEQ ID NO: 2. The second segment can comprise the sequence from G34, P45, G55, N69, T83, A99, or A109 in the amino acid sequence of SEQ ID NO: 2 to the C-terminus.

[0030] Preferably, the first segment of DddAtox cytosine deaminase comprises the sequence from the N-terminus to G44 of the amino acid sequence of SEQ ID NO: 2 (SEQ ID NO: 3 below), and the second segment comprises the sequence from P45 to the C-terminus (SEQ ID NO: 4 below).

[0031] (SEQ ID NO: 3) wild-type DddAtox G1333-N GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGG

[0032] (SEQ ID NO: 4) wild-type DddAtox G1333-C PTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC

[0033] Preferably, the first segment of DddAtox cytosine deaminase comprises the sequence from the N-terminus to G108 of the amino acid sequence of SEQ ID NO: 2 (SEQ ID NO: 5 below), and the second segment may comprise the sequence from A109 to the C-terminus (SEQ ID NO: 6 below).

[0034] (SEQ ID NO: 5) wild-type DddAtox G1397-N GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEG

[0035] (SEQ ID NO: 6) wild-type DddAtox G1397-C AIPVKRGATGETKVFTGNSNSPKSPTKGGC

[0036] When the first and second segments of DddAtox are used as cytosine deaminase, one or more amino acids located on the surface where the first and second segments of cytosine deaminase bind to each other may be substituted with other amino acids. For example, the first and second segments of DddAtox may have the amino acid sequences of SEQ ID NO: 3 (G1333-N) and SEQ ID NO: 4 (G1333-C), respectively. In this case, one or more amino acids selected from the group consisting of positions 3, 5, 10, 11, 13, 14, 15, 16, 17, 18, 19, 28, 30, and 31 of SEQ ID NO: 3, or one or more amino acids selected from the group consisting of positions 13, 16, 17, 20, 21, 28, 29, 30, 31, 32, 33, 56, 57, 58, and 60 of SEQ ID NO: 4 may be substituted with other amino acids, but are not limited thereto. As another example, the first fragment of DddAtox may comprise the amino acid sequences of SEQ ID NO: 5 (G1397-N) and SEQ ID NO: 6 (G1397-C), in which case one or more amino acids selected from the group consisting of positions 87, 88, 91, 92, 95, 100, 101, 102, and 103 of SEQ ID NO: 5, or one or more amino acids selected from the group consisting of positions 13, 14, 15, and 16 of SEQ ID NO: 6 may be substituted with other amino acids, but is not limited thereto. The "other amino acids" refer to alanine, isoleucine, leucine, methionine, phenylalanine, proline, tryptophan, valine, aspartic acid, cysteine, glutamine, glycine, serine, threonine, tyrosine, aspartate, glutamic acid, arginine, histidine, lysine, and all known variants of these amino acids, excluding the amino acids present in the wild-type protein at the original mutated positions. By using such mutants, highly efficient and precise C-to-T correction can be achieved without causing undesired non-targeted C-to-T correction due to the inability of two DddAtox fragment pairs, each linked to a DNA-binding protein, to function properly when not bound to DNA.

[0037] As used herein, the terms "G1333-N," "G1333N," or "1333N" can refer to the first fragment of wild-type DddAtox having the amino acid sequence of SEQ ID NO: 3 or an amino acid variant thereof, and the terms "G1333-C," "G1333C," or "1333C" can refer to the second fragment of wild-type DddAtox having the amino acid sequence of SEQ ID NO: 4 or an amino acid variant thereof.

[0038] As used herein, the terms "G1397-N," "G1397N," or "1397N" can refer to the first fragment of wild-type DddAtox having the amino acid sequence of SEQ ID NO: 5 or an amino acid variant thereof, and the terms "G1397-C," "G1397C," or "1397C" can refer to the second fragment of wild-type DddAtox having the amino acid sequence of SEQ ID NO: 6 or an amino acid variant thereof.

[0039] The cytosine deaminase used in the present invention can be in its full-length form. In this case, the full-length cytosine deaminase (e.g., DddAtox) has its amino acid sequence modified to be non-toxic and / or have only low toxicity. Positively charged amino acids are specifically concentrated at the C-terminus of DddAtox. Because DNA is negatively charged, it binds to positively charged amino acids in proteins. Substitution of these positively charged amino acids weakens the binding ability of DddAtox to DNA, thereby reducing and / or eliminating intracellular toxicity. In other words, if the positively charged amino acids are substituted to eliminate toxicity, cloning can be performed in E. coli, allowing full-length DddAtox to be obtained. Such a non-toxic full-length cytosine deaminase can be provided by substituting one or more, two or more, three or more, four or more, or five or more amino acids in the wild-type amino acid sequence of SEQ ID NO: 2 with other amino acids. The "other amino acid" refers to an amino acid selected from alanine, isoleucine, leucine, methionine, phenylalanine, proline, tryptophan, valine, aspartic acid, cysteine, glutamine, glycine, serine, threonine, tyrosine, aspartate, glutamic acid, arginine, histidine, lysine, and all known variants of the amino acids, excluding the amino acid that the wild-type protein has at the original mutation position. For example, the other amino acid may be alanine.

[0040] The non-toxic full-length DddAtox may comprise an amino acid sequence selected from the group consisting of the following amino acid sequences:

[0041] A1341D KRKKA mutation GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYDNAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTAGGC

[0042] AAAAA mutant GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVAAGATGETAVFTGNSNSPASPTAGGC

[0043] AAAAK mutant GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVAAGATGETAVFTGNSNSPASPTKGGC

[0044] AAKAA mutation GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVAAGATGETKVFTGNSNSPASPTAGGC

[0045] AAKAK mutant GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVAAGATGETKVFTGNSNSPASPTKGGC

[0046] KAAAA mutant GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKAGATGETAVFTGNSNSPASPTAGGC

[0047] E1347A mutant GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVAGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC

[0048] Preferably, the full-length cytosine deaminase mutant that can be used in the present invention may have one or more amino acid substitutions selected from the group consisting of S to G at position 37, G to S at position 59, A to V at position 109, and S to G at position 129 of the amino acid sequence of SEQ ID NO: 2.

[0049] More preferably, the full-length cytosine deaminase mutant that can be used in the present invention may have all of the following substitutions in the amino acid sequence of SEQ ID NO: 2: S to G at position 37, G to S at position 59, A to V at position 109, and S to G at position 129. In this case, the sequence is as follows:

[0050] GSVG mutants GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLEGKVFSSGGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGVIPVKRGATGETKVFTGNSNGPKSPTKGGC

[0051] As another example, a full-length cytosine deaminase mutant that can be used in the present invention can comprise the following sequence:

[0052] SSVG mutant GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSGSGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGVIPVKRGATGETKVFTGNSNGPKSPTKGGC

[0053] GSAG variant GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLEGKVFSGSGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNGPKSPTKGGC

[0054] GSVS variant GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLEGKVFSGSGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGVIPVKRGATGETKVFTGNSNSPKSPTKGGC

[0055] An adenine deaminase that can be used in a base editor according to the present invention refers to an amino-detaching enzyme that converts adenine base to hypoxanthine (inosine as a nucleotide) and can be derived from any organism (e.g., a eukaryote or a prokaryote), including, but not limited to, algae, bacteria, fungi, plants, invertebrates, and mammals, such as E. coli, S. aureus, S. typhi, S. putrefaciens, H. influenzae, or C. crescentus, and / or can be mutated (e.g., engineered and / or evolved). Such an adenine deaminase can be, for example, APOBEC, AID, or TadA, or a mutant thereof. The aforementioned TadA can be, for example, TadA8e (SEQ ID NO: 1) or a truncated or mutant thereof (e.g., a mutant that has been improved and / or evolved to be applicable to deoxynucleotides). The TadA8e mutants mentioned above may be those in which one or more of the amino acid residues at positions 23, 28, 30, 36, 46, 48, 49, 51, 76, 82, 84, 106, 108, 110, 111, 146, 147, 152, 154, 155, 156, and 157 of SEQ ID NO: 1 have been mutated to other amino acids. Regarding the construction of adenine deaminase that can be used in the present invention, reference may be made to International Patent Application Publications WO2022 / 060185, WO2023 / 086953, etc., which are incorporated herein by reference in their entirety. The adenine deaminase that can be used in the present invention may be a base-correcting composition comprising the amino acid sequence of SEQ ID NO: 1 or a conservative amino acid substitution thereof.

[0056] When cytosine deaminase is used in the form of split bodies in a base editor according to the present invention, adenine deaminase can be linked to the N-terminus or C-terminus of the first split body of cytosine deaminase and / or the N-terminus or C-terminus of the second split body.

[0057] With regard to the configuration of deaminase that can be used in the present invention, reference may be made to content that was already known prior to the filing of this application, including International Patent Application Publication No. WO2022 / 060185, the entire contents of which are incorporated herein by reference.

[0058] The DNA-binding protein according to the present invention can be a zinc finger protein, a TALE protein, or a CRISPR-linked nuclease, and / or a combination thereof. Regarding the configuration of zinc finger proteins, TALE proteins, and CRISPR-linked nucleases, reference may be made to International Patent Application Publication WO2022 / 060185, the entire contents of which are incorporated herein by reference, as well as other content already known prior to the filing of this application.

[0059] The DNA-binding protein according to the present invention may be a TALE protein. The TALE protein of the present invention refers to a protein that binds to nucleotides in a sequence-specific manner through one or more TALE-repeat modules. The TALE protein comprises at least one TALE-repeat module, preferably, but not limited to, 1 to 30 TALE-repeat modules. As used herein, a TALE-repeat module may also be referred to as a "TALE array," and the term "TALE protein" refers to a configuration comprising an N-terminal domain and a C-terminal domain (which may include half domains) on both sides of the TALE array. As used herein, the term "TALE" may refer only to a "TALE array" or to a "TALE protein" depending on the context.

[0060] When a TALE protein is used as the DNA-binding protein of the base editor used in the present invention, a single-module TALE array or a multi-module TALE array (e.g., a dual-module TALE array consisting of a first TALE array and a second TALE array) can be used. For example, a first TALE protein (left TALE) can be bound to a first segment of cytosine deaminase (first fusion), and a second TALE protein (right TALE) can be bound to a second segment of cytosine deaminase (second fusion). In this case, the first TALE protein (left TALE) is linked to the first segment of cytosine deaminase in the NC direction, and the second TALE protein (right TALE) is linked to the second segment of cytosine deaminase, and the first TALE protein (left TALE) and the second TALE protein (right TALE) are each linked to a structure consisting of an N-terminal domain, a TALE array, and a C-terminal domain (which may include a half domain). The first splitter can be an N-terminal splitter and / or a C-terminal splitter of full-length cytosine deaminase, and the second splitter can also be an N-terminal splitter and / or a C-terminal splitter of full-length cytosine deaminase. Even when cytosine deaminase is full-length, a single-module TALE array or a multi-module TALE array can be used. When a single-module TALE array is used, a single TALE domain and cytosine deaminase are linked in the NC direction. When a dual-module TALE array is used, a first TALE protein (left or right TALE) is linked to full-length cytosine deaminase in the NC direction, and a second TALE protein (left or right TALE) can be included separately.

[0061] When cytosine deaminase is used in the form of a first segment and a second segment, and the DNA-binding protein is a TALE protein, the base editor of the present invention can have the form of a composition of a first fusion in which a first TALE is linked to the first segment of cytosine deaminase, and a second fusion in which a second TALE is linked to the second segment of cytosine deaminase. The first fusion and second fusion have the structures N'-TALE-first segment (cytosine deaminase)-C' and N'-TALE-second segment (cytosine deaminase)-C', respectively. Adenine deaminase can be linked to the first fusion, the second fusion, or both. Specifically, it can be linked to the N-terminus or C-terminus of the first segment of cytosine deaminase in the first fusion, and / or the N-terminus or C-terminus of the second segment of cytosine deaminase in the second fusion.

[0062] When cytosine deaminase is used in its full-length form and the DNA-binding protein has the form of a single TALE module as a TALE protein, the single TALE module and cytosine deaminase are included in the N-terminal direction, and in this case, adenine deaminase is linked toward the C-terminal direction of the single TALE module and can be linked to the N-terminal or C-terminal of cytosine deaminase. In the specification of this application, the term "monomeric TALED" refers to a form in which a single TALE module is used as the DNA-binding protein and cytosine deaminase and adenine deaminase are linked.

[0063] When cytosine deaminase is used in its full-length form and the DNA-binding protein is a TALE protein in the form of a dual TALE module, the base editor of the present invention can have the form of a composition of a first fusion in which a first TALE module and cytosine deaminase are linked in the N-C direction, and a second fusion containing adenine deaminase and a second TALE. The first fusion and second fusion have the structures N'-TALE-cytosine deaminase-C' and N'-TALE-adenine deaminase-C', respectively, and adenine deaminase can be bound to the N-terminus or C-terminus of the TALE.

[0064] According to one embodiment of the present invention, the base correction composition of the present invention is characterized by containing uracil DNA glycosylase (UDG). Specifically, one or more of the fusion proteins contained in the base correction composition of the present invention contains UDG. UDG is known to recognize naturally damaged DNA and selectively remove only uracil bases from DNA. The A-to-G base editor of the present invention contains a cytosine deaminase such as DddAtox, which deaminates cytosine bases to uracil. Therefore, co-expression of UDG removes the converted uracil base, restoring the original DNA sequence and preventing undesired cytosine base correction. This technique may be particularly useful for selectively correcting only adenine bases in environments where the natural expression rate of UDG is relatively low, particularly in organelles of plant cells. UDGs of any species, such as Arabidopsis, human, mouse, tobacco, and rice, can be used, and the UDGs can contain various lengths, such as 2, 5, 10, 16, 24, or 32 amino acids. For example, the UDG used in the present invention contains the amino acid sequence of SEQ ID NO: 7, but is not limited thereto.

[0065] UDG is preferably used in the form of a fusion protein together with the DNA-binding protein, cytosine deaminase, and adenine deaminase used in the present invention, but it can also be delivered to target DNA in a form (protein or polynucleotide) separate from the above components. When used in the form of a fusion protein, UDG can be linked to the C-terminus of cytosine deaminase (or its fragment) or the C-terminus of adenine deaminase, but is not necessarily limited thereto.

[0066] According to one embodiment of the present invention, a base editor according to the present invention containing UDG can selectively correct only adenine bases in DNA without substantially correcting cytosine bases to thymine bases. For example, a UDG-linked TALED according to the present invention can correct cytosine bases to thymine at a frequency of less than 20%, less than 10%, less than 5%, or less than 1%. The base correction frequency (or efficiency) can be measured as the base correction frequency (or efficiency) calculated when sequencing target DNA from a cell transformed to express a base editor according to the present invention, and can be expressed, for example, as the percentage of sequencing reads that reflect the desired base correction results among all sequencing reads for DNA obtained from a base-corrected cell or individual.

[0067] One or more of the fusion proteins included in the base correction composition according to the present invention may further include a nuclear export signal (NES). Attaching an NES to a base correction protein can result in more efficient base correction. The NES sequence may be any signal sequence (e.g., VDEMTKKFGTLTIHDTEK) that confers the ability to enable nuclear export. A natural NES or an artificially synthesized NES may be used. For example, it may be derived from the mirute virus of mice (MVM), but is not limited thereto.

[0068] One or more of the fusion proteins contained in the base correction composition according to the present invention may further contain a mitochondrial targeting sequence (MTS). The MTS that can be used in the present invention is any signal sequence capable of transporting into mitochondria, and may be a natural MTS present at the N-terminus of various mitochondrial proteins, or an artificially synthesized MTS. When an MTS is used in the base correction composition according to the present invention, its position may vary, for example, it may be linked directly or indirectly (e.g., via a linker and / or other protein components) to the N-terminus of a DNA-binding protein or the N-terminus of an NES, but is not limited thereto.

[0069] One or more of the fusion proteins contained in the base correction composition according to the present invention may further contain a chloroplast transit signal (CTS), and such a base correction composition is useful for correcting chloroplast, plastid, or leucoplast DNA in plant cells. The CTS that can be used in the present invention is any signal sequence capable of transporting DNA into chloroplasts, including natural CTSs present at the N-terminus of various chloroplast proteins, as well as artificially synthesized CTSs. When a CTS is used in the base correction composition according to the present invention, its location may vary, including, but not limited to, direct or indirect (e.g., via a linker and / or other protein components) linkage to the N-terminus of a DNA-binding protein or the N-terminus of an NES.

[0070] One or more of the fusion proteins contained in the base correction composition according to the present invention may contain a nuclear localization signal (NLS). The NLS that can be used in the present invention may be any signal sequence capable of transporting the protein into the nucleus, including natural NLSs present at the N-terminus of various nuclear proteins and / or artificially synthesized NLSs. When an NLS is used in the base correction composition according to the present invention, its location may vary. For example, it may be linked directly or indirectly (e.g., via a linker and / or other protein components) to the N-terminus of a DNA-binding protein, but this is not intended to be limiting.

[0071] The base editor according to the present invention has the form of a fusion protein. Specifically, the fusion protein includes a DNA-binding protein and a deaminase, or fragments thereof, and may include additional sequences such as UDG, NLS, NES, MTS, and CTS. It may also include additional sequences for bioengineering techniques, such as tags. These polypeptides may be linked directly and / or via a linker. Such fusion proteins (or polynucleotides encoding the fusion proteins) can be designed and constructed using any method known in the field of bioengineering.

[0072] In the present invention, base editors can be in the form of polynucleotides encoding the fusion proteins described herein. Such polynucleotides can be inserted into vectors, which can then be introduced into cells.

[0073] The present invention also provides plant cells transformed with the vectors, plants grown therefrom, plants that are progeny and / or clones of the plants, and seeds obtained from such plants.

[0074] In the plant cells, plants, and seeds according to the present invention, adenine bases in wild-type DNA are corrected to guanine. In particular, plant cells, plants, and seeds transformed with a base editor containing UDG (or a polynucleotide encoding it) may selectively correct only adenine bases without correcting cytosine bases in wild-type DNA.

[0075] The present invention also provides a method for correcting adenine bases in DNA to guanine, comprising expressing the base correction composition described herein or a polynucleotide encoding the same in an animal or animal cell, or a plant or plant cell of interest, where the DNA can be nuclear and / or organelle DNA, preferably plant organelle DNA.

[0076] Based on the above, the present invention relates to the following (1) to (51), but is not limited thereto.

[0077] (1) A base correction composition having the activity of correcting adenine A bases in DNA to guanine G bases, the base correction composition comprising one or more fusion proteins, each of which comprises a DNA-binding protein and cytosine deaminase, at least one of the one or more fusion proteins comprising adenine deaminase, at least one of the one or more fusion proteins comprising uracil DNA glycosylase (UDG), and the cytosine deaminase being present in the form of a full-length or two fragments.

[0078] (2) A base correction composition according to (1), wherein the DNA is nuclear DNA or organelle DNA.

[0079] (3) In (1) or (2), the base correcting composition, wherein the cytosine deaminase is APOBEC (apolipoprotein B editing complex), AID (activation-induced deaminase), DddAtox, or a mutant thereof.

[0080] (4) A base correcting composition according to any one of (1) to (3), wherein the cytosine deaminase is contained in a full-length form comprising the amino acid sequence of SEQ ID NO: 2, and one or more amino acids selected from the group consisting of 37, 59, 109, and 129 of the amino acid sequence of SEQ ID NO: 2 are substituted with other amino acids.

[0081] (5) In (4), the base correcting composition is one in which the cytosine deaminase has the amino acid sequence of SEQ ID NO: 2 and has one or more amino acid substitutions selected from the group consisting of S to G at position 37, G to S at position 59, A to V at position 109, and S to G at position 129.

[0082] (6) In any one of (1) to (3), a base correction composition is provided, in which cytosine deaminase is contained in the form of a first segment and a second segment, one fusion protein contains the first segment, and another fusion protein contains the second segment, in which one or more amino acids located on the dimerization surface of the first segment and the second segment are replaced with other amino acids.

[0083] (7) In (6), a base correcting composition in which the first segment of cytosine deaminase contains the sequence from the N-terminus to G at position 33, G at position 44, A at position 54, N at position 68, G at position 82, N at position 98, or G at position 108 in the amino acid sequence of SEQ ID NO: 2, and the second segment of cytosine deaminase contains the sequence from G at position 34, P at position 45, G at position 55, N at position 69, T at position 83, A at position 99, or A at position 109 in the amino acid sequence of SEQ ID NO: 2 to the C-terminus.

[0084] (8) In (6), the base correcting composition is characterized in that the first segment of cytosine deaminase contains the amino acid sequence of SEQ ID NO: 5, and the second segment of cytosine deaminase contains the amino acid sequence of SEQ ID NO: 6, in which one or more amino acids selected from the group consisting of positions 87, 88, 91, 92, 95, 100, 101, 102, and 103 of SEQ ID NO: 5 or one or more amino acids selected from the group consisting of positions 13, 14, 15, and 16 of SEQ ID NO: 6 are substituted with other amino acids.

[0085] (9) In any one of (1) to (8), the adenine deaminase is APOBEC (apolipoprotein B editing complex), AID (activation-induced deaminase), TadA (tRNA-specific adenosine deaminase), or a mutant thereof.

[0086] (10) In any one of (1) to (9), the base correction composition, wherein the one or more fusion proteins each comprise one or more DNA binding proteins selected from the group consisting of zinc finger proteins, TALE (transcriptional activator-like effector) proteins, and CRISPR-associated nucleases.

[0087] (11) In any one of (1) to (10), the base correction composition, wherein the DNA is organelle DNA and the one or more fusion proteins each contain a nuclear export signal (NES).

[0088] (12) In any one of (1) to (11), the base correction composition, wherein the DNA is mitochondrial DNA, and the one or more fusion proteins each contain a mitochondrial targeting sequence (MTS).

[0089] (13) In any one of (1) to (11), the base correction composition, wherein the DNA is chloroplast, plastid, or leucoplast DNA, and one or more fusion proteins each contain a chloroplast transit signal (CTS).

[0090] (14) A base correction composition according to any one of (1) to (10), wherein the DNA is nuclear DNA and one or more fusion proteins contain a nuclear localization signal (NLS).

[0091] (15) A base correction composition according to any one of (1) to (14), which does not substantially cause correction of cytosine C base to thymine T base.

[0092] (16) A base correction composition according to any one of (1) to (14), which corrects cytosine C bases to thymine T bases at a frequency of less than 20%.

[0093] (17) A base correction composition according to any one of (1) to (14), which corrects cytosine C bases to thymine T bases at a frequency of less than 10%.

[0094] (18) A base correction composition according to any one of (1) to (14), which corrects cytosine C bases to thymine T bases at a frequency of less than 5%.

[0095] (19) A base correction composition according to any one of (1) to (14), which corrects a cytosine C base to a thymine T base at a frequency of less than 1%.

[0096] (20) A base correcting composition having the activity of correcting an adenine A base in the DNA of a plant cell to a guanine G base, the base correcting composition comprising a DNA binding protein, a cytosine deaminase (preferably, a full-length form of the cytosine deaminase), and a fusion protein containing adenine deaminase (preferably, the full-length form of the cytosine deaminase comprises the amino acid sequence of SEQ ID NO: 9).

[0097] (21)(20) The base corrector composition, wherein the DNA is nuclear DNA or organelle DNA.

[0098] (22) The base correcting composition according to (20) or (21), wherein the cytosine deaminase is APOBEC (apolipoprotein B editing complex), AID (activation-induced deaminase), DddAtox, or a mutant thereof.

[0099] (23) In any one of (20) to (22), the base correction composition is characterized in that the cytosine deaminase is substituted with another amino acid at one or more amino acids selected from the group consisting of positions 37, 59, 109, and 129 of the amino acid sequence of SEQ ID NO: 2.

[0100] (24) In (23), the cytosine deaminase is a base correcting composition in which S at position 37 of the amino acid sequence of SEQ ID NO: 2 is replaced with G, and / or G at position 59 is replaced with S, and / or A at position 109 is replaced with V, and / or S at position 129 is replaced with G in the amino acid sequence of SEQ ID NO: 2.

[0101] (25) A base correction composition according to any one of (20) to (24), wherein the adenine deaminase is APOBEC (apolipoprotein B editing complex), AID (activation-induced deaminase), TadA (tRNA-specific adenosine deaminase), or a mutant thereof.

[0102] (26) In any one of (20) to (25), the base correction composition, wherein the fusion protein comprises one or more DNA binding proteins selected from the group consisting of zinc finger proteins, TALE proteins, and CRISPR-associated nucleases.

[0103] (27) The base correction composition according to any one of (20) to (26), wherein the DNA is organelle DNA and the fusion protein contains a nuclear export signal (NES).

[0104] (28) A base correction composition according to any one of (20) to (27), wherein the DNA is mitochondrial DNA and the fusion protein contains a mitochondrial targeting sequence (MTS).

[0105] (29) A base correction composition according to any one of (20) to (27), wherein the DNA is chloroplast, chloroplast, or leucoplast DNA, and the fusion protein contains a chloroplast transit signal (CTS).

[0106] (30) The base correction composition according to any one of (20) to (26), wherein the DNA is nuclear DNA and the fusion protein contains a nuclear localization signal (NLS).

[0107] (31) A base correcting composition according to any one of (20) to (30), wherein the fusion protein comprises uracil DNA glycosylase (UDG).

[0108] (32) A base correcting composition according to (31), which does not substantially correct a cytosine C base to a thymine T base.

[0109] (33) A base correcting composition according to (31), which corrects a cytosine C base to a thymine T base at a frequency of less than 20%.

[0110] (34) A base correcting composition according to (31), which corrects a cytosine C base to a thymine T base at a frequency of less than 10%.

[0111] (35) A base correcting composition according to (31), which corrects a cytosine C base to a thymine T base at a frequency of less than 5%.

[0112] (36) A base correcting composition according to (31), which corrects a cytosine C base to a thymine T base at a frequency of less than 1%.

[0113] (37) A polynucleotide encoding any one of one or more fusion proteins contained in the base correction composition according to any one of (1) to (36), or a combination of two or more of the above polynucleotides.

[0114] (38) A vector comprising a polynucleotide or a combination of polynucleotides according to (37).

[0115] (39) A base correction composition comprising the vector according to (38).

[0116] (40) A plant cell transformed with a vector according to (38) or a plant cell containing a polynucleotide or a combination of polynucleotides according to (37).

[0117] (41) Plants grown from plant cells by (40).

[0118] (42) Plants that are descendants or clones of plants according to (41)

[0119] (43) Seeds obtained from plants according to (41) or (42).

[0120] (44) A plant cell, plant, or seed according to (40), a plant according to (41) or (42), or a seed according to (43), in which the adenine A base in the wild-type DNA has been corrected to guanine G.

[0121] (45) In (44), the cytosine C base in the wild-type DNA is not corrected in a plant cell, plant, or seed.

[0122] (46) A method for correcting adenine A bases in plant organelle DNA to guanine G, comprising expressing a base correction composition according to any one of (1) to (36) or a polynucleotide or combination of polynucleotides according to (37) in a plant or plant cell of interest.

[0123] (47) The method according to (46), wherein correction of a cytosine C base to a thymine T base does not substantially occur.

[0124] (48) A method according to (46), in which the correction of a cytosine C base to a thymine T base occurs at a frequency of less than 20%.

[0125] (49) A method according to (46), in which the correction of a cytosine C base to a thymine T base occurs at a frequency of less than 10%.

[0126] (50)(46) A method in which the correction of a cytosine C base to a thymine T base occurs at a frequency of less than 5%.

[0127] The present invention will be described in detail below with reference to the following examples, but the following examples are for illustrative purposes only and are not intended to limit the scope of the present invention.

[0128] Example 1: Transformation of Arabidopsis with the psaA base editor

[0129] We cloned DNA encoding the following base editor targeting the psbA gene in the chloroplasts of Arabidopsis thaliana, and created transgenic plants through Agrobacterium-mediated transformation. The gene sequence used is as follows:

[0130] [Table 1]

[0131] The sequences of the CTS, Left TALE, Right TALE, ABE8.0, GSVG, and UDG used above are as follows:

[0132] CTS (chloroplast transport signal) MDSQLVLSLKLNPSFTPLSPLFPFTPCSSFSPSLRFSSCYSRRLYSPVTVYAAK (SEQ ID NO: 8)

[0133] ABE8.0 (adenine deaminase) SEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN (SEQ ID NO: 1)

[0134] 1397N (DddAtox division body) GSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEG (SEQ ID NO: 5)

[0135] 1397C (DddAtox fragment): GSAIPVKRGATGETKVFTGNSNSPKSPTKGGC (SEQ ID NO: 6)

[0136] GSVG (full-length DddAtox mutant): GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLEGKVFSSGGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGVIPVKRGATGETKVFTGNSNGPKSPTKGGC (SEQ ID NO: 9)

[0137] UDG (uracil DNA glycosylase): MASSTPKTLMDFFQPAKRLKASPSSSSFPAVSVAGGSRDLGSVANSPPRVTVTTSVADDSSGLTPEQIARAEFNKFVAKSKRNLAVCSERVTKAKSEGNCYVPLSELLVEESWLKALPGEFHKPYAKSLSDFLEREIITDSKSPLIYPPQHLIFNALNTTPFDRVKTVIIGQDPYHGPGQAMGLSFSVPEGEKLPSSLLNIFKELHKDVGCSIPRHGNLQKWAVQGVLLLNAVLTVRSKQPNSHAKKGWEQFTDAVIQSISQQKEGVVFLLWGRYAQEKSKLIDATKHHILTAAHPSGLSANRGFFDCRHFSRANQLLEEMGIPPIDWQL (SEQ ID NO: 7)

[0138] Left TALE 1 (including N-terminal domain and C-terminal domain): (SEQ ID NO: 10)

[0139] Right TALE (including N-terminal domain and C-terminal domain): DLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHERAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDH GLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQ AHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVL CQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLG (Sequence number 11)

[0140] The sequence of the linker used is as follows: Linker 1: GS Linker 2: SGSETPGTSESATPES (SEQ ID NO: 12) Linker 3: LVGS

[0141] Specifically, to correct the psaA gene present in Arabidopsis chloroplast genomes using the gene construct, the construct was first placed between the RPS5A promoter, which induces expression during embryogenesis, and the 35S terminator, designed to induce correction from an early stage of development. The construct was then cloned using Gibson assembly, Golden Gate, restriction enzymes, etc., into a vector suitable for transforming Agrobacterium tumefaciens, a strain capable of transferring the gene construct to Arabidopsis nuclear genomes using T-DNA. The Agrobacterium strain GV3101 transformed with the vector was then used to transform Arabidopsis thaliana Columbia-0 (Col-0) plants by floral dipping according to a known method (Zhang et al., Nat. Protoc. 1, 641-646 (2006)).

[0142] Example 2: Confirmation of base correction through sequencing and measurement of correction efficiency

[0143] DNA was extracted from untransformed wild-type Col-0 plants and plants obtained by growing seeds from the first generation of transformed Arabidopsis, and the sequences were analyzed using targeted deep sequencing. The base correction efficiency (frequency) was calculated as the percentage of sequencing reads that reflected the desired base correction results among all sequencing reads (Figures 3 to 14).

[0144] As a result of the experiments, all of the base editors used in the present invention corrected the adenine base at the target site of the psaA gene with high efficiency, as shown in Figures 1 and 2. In particular, it was confirmed for the first time that the adenine base in the chloroplast DNA of plant cells was effectively corrected when a monomeric TALED linked to the full-length, non-toxic form of DddAtox cytosine deaminase (GSVG mutant) was used.

[0145] In addition, to determine the extent to which unwanted C-to-T correction can be avoided when using UDG-linked TALEDs, the average frequency at which correction occurred at the C2 base among the correction target sites in the psaA gene was summarized for each base editor used in the table below.

[0146] [Table 2]

[0147] As can be seen in the table above, when UDG was additionally ligated, C-to-T correction was significantly reduced in all base editors used, and no C-to-T correction was observed, almost to the same extent as in wild-type Col-0 (control group), in which no base correction occurs.

[0148] The experimental results also showed that by using the UDG-linked TALED of the present invention, plant individuals could be obtained in which cytosine bases were not corrected and only adenine bases were corrected to guanine (e.g., individuals #4, #10, and #14 in Figure 8).

[0149] Example 3: Transformation of Arabidopsis thaliana with the psbA base editor

[0150] In order to correct the base sequence of 5'-AGT-3', which encodes serine 264, in the psbA gene found in the chloroplasts of Arabidopsis thaliana, to 5'-GGT-3', which encodes glycine, a fusion protein was constructed as shown in Table 3. Correcting serine 264 to glycine confers resistance to the atrazine herbicide. To express the fusion protein, the following DNA was cloned and transformed using Agrobacterium to create a transgenic plant.

[0151] [Table 3]

[0152] The sequences of the components of the fusion protein used are as follows:

[0153] CTS (chloroplast transport signal) MDSQLVLSLKLNPSFTPLSPLFPFTPCSSFSPSLRFSSCYSRRLYSPVTVYAAK (SEQ ID NO: 8)

[0154] 1397N (DddAtox division) GSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEG (SEQ ID NO: 5)

[0155] 1397C (DddAtox division) GSAIPVKRGATGETKVFTGNSNSPKSPTKGGC (SEQ ID NO: 6)

[0156] GSVG (full-length DddAtox mutant) GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLEGKVFSSGGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGVIPVKRGATGETKVFTGNSNGPKSPTKGGC (SEQ ID NO: 9)

[0157] AD(TadA8e) SEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN (SEQ ID NO: 1)

[0158] UDG (uracil DNA glycosylase): MASSTPKTLMDFFQPAKRLKASPSSSSFPAVSVAGGSRDLGSVANSPPRVTVTTSVADDSSGLTPEQIARAEFNKFVAKSKRNLAVCSERVTKAKSEGNCYVPLSELLVEESWLKALPGEFHKPYAKSLSDFLEREIITDSKSPLIYPPQHLIFNALNTTPFDRVKTVIIGQDPYHGPGQAMGLSFSVPEGEKLPSSLLNIFKELHKDVGCSIPRHGNLQKWAVQGVLLLNAVLTVRSKQPNSHAKKGWEQFTDAVIQSISQQKEGVVFLLWGRYAQEKSKLIDATKHHILTAAHPSGLSANRGFFDCRHFSRANQLLEEMGIPPIDWQL (SEQ ID NO: 7)

[0159] Linker 11 GSSGSETPGTSESATPES (SEQ ID NO: 13) Linker 12 LVGS Linker 13 SGSETPGTSESATPES (SEQ ID NO: 18) Linker 14 GS

[0160] psbA Left 1 TALE protein (SEQ ID NO: 14: Includes the N-terminal domain and the C-terminal domain.)

[0161] psbA Left 2 TALE protein (SEQ ID NO: 15: Includes the N-terminal domain and the C-terminal domain.)

[0162] psbA Right 1 TALE protein (SEQ ID NO: 16: Includes the N-terminal domain and the C-terminal domain.)

[0163] psbA Right 2 TALE protein (SEQ ID NO: 17: Includes the N-terminal domain and the C-terminal domain.)

[0164] Specifically, the gene construct was used to amend the 5'-AGT-3' sequence encoding serine 264 in the psbA gene, which is inserted into Arabidopsis thaliana chloroplast genomes, to 5'-GGT-3' encoding glycine. The gene construct was first engineered to induce the amendment early in development between the RPS5A promoter, which induces expression during embryogenesis, and the 35S terminator. A suitable vector for transforming Agrobacterium tumefaciens, a strain capable of transferring gene constructs to Arabidopsis nuclear genomes using T-DNA, was cloned using Gibson assembly, Golden Gate, and restriction enzymes. The Agrobacterium strain GV3101 transformed with the vector was then used to transform Arabidopsis thaliana Columbia-0 (Col-0) plants via floral dipping according to a previously published method (Zhang et al., Nat. Protoc. 1, 641-646 (2006)).

[0165] Approximately 15,000 to 20,000 seeds of each of the 12 fusion proteins were introduced into Arabidopsis plants using the floral dipping method and sown in soil. Seven and 14 days later, the plants were treated with atrazine (40 g per hectare). The surviving first-generation transgenic plants were then analyzed for gene correction efficiency. The results confirmed that A-to-G base correction was achieved when a fusion protein base editor containing UDG was used according to the present invention (Figure 16).

Claims

1. A base correction composition having an activity of correcting adenine A bases in DNA to guanine G bases, the base correction composition comprising one or more fusion proteins, each of which comprises a DNA binding protein and cytosine deaminase, at least one of the one or more fusion proteins comprising adenine deaminase, at least one of the one or more fusion proteins comprising uracil DNA glycosylase (UDG), and the cytosine deaminase being present in the form of a full-length or two segments.

2. In claim 1, A base corrector composition, wherein the DNA is nuclear DNA or organelle DNA.

3. In claim 1 or claim 2, A base correcting composition, wherein the cytosine deaminase is APOBEC (apolipoprotein B editing complex), AID (activation-induced deaminase), or DddAtox, or a mutant thereof.

4. In any one of claims 1 to 3, A base correcting composition, wherein cytosine deaminase is contained in its full-length form, comprising the amino acid sequence of SEQ ID NO: 2, and one or more amino acids selected from the group consisting of 37, 59, 109, and 129 of the amino acid sequence of SEQ ID NO: 2 are substituted with other amino acids.

5. In claim 4, A base correcting composition, wherein the cytosine deaminase has the amino acid sequence of SEQ ID NO: 2 and has one or more amino acid substitutions selected from the group consisting of a substitution of S for G at position 37, a substitution of G for S at position 59, a substitution of A for V at position 109, and a substitution of S for G at position 129.

6. In any one of claims 1 to 3, A base correcting composition comprising cytosine deaminase in the form of a first segment and a second segment, one fusion protein comprising the first segment and another fusion protein comprising the second segment, wherein one or more amino acids located on the dimerization surface of the first segment and the second segment are substituted with other amino acids.

7. In claim 6, A base correcting composition, wherein the first segment of cytosine deaminase comprises the sequence from the N-terminus to G at position 33, G at position 44, A at position 54, N at position 68, G at position 82, N at position 98, or G at position 108 in the amino acid sequence of SEQ ID NO: 2, and the second segment of cytosine deaminase comprises the sequence from G at position 34, P at position 45, G at position 55, N at position 69, T at position 83, A at position 99, or A at position 109 in the amino acid sequence of SEQ ID NO: 2 to the C-terminus.

8. The base correcting composition according to claim 6, wherein the first segment of cytosine deaminase comprises the amino acid sequence of SEQ ID NO: 5, and the second segment of cytosine deaminase comprises the amino acid sequence of SEQ ID NO: 6, wherein one or more amino acids selected from the group consisting of positions 87, 88, 91, 92, 95, 100, 101, 102, and 103 of SEQ ID NO: 5 or one or more amino acids selected from the group consisting of positions 13, 14, 15, and 16 of SEQ ID NO: 6 are substituted with other amino acids.

9. In any one of claims 1 to 8, A base correcting composition, wherein the adenine deaminase is APOBEC (apolipoprotein B editing complex), AID (activation-induced deaminase), or TadA (tRNA-specific adenosine deaminase), or a mutant thereof.

10. In any one of claims 1 to 9, A base correction composition, wherein the one or more fusion proteins each comprise one or more DNA binding proteins selected from the group consisting of zinc finger proteins, TALE (transcriptional activator-like effector) proteins, and CRISPR-associated nucleases.

11. In any one of claims 1 to 10, A base corrected composition, wherein the DNA is organelle DNA and the one or more fusion proteins each contain a nuclear export signal (NES).

12. In any one of claims 1 to 11, A base-correcting composition, wherein the DNA is mitochondrial DNA, and the one or more fusion proteins each contain a mitochondrial targeting sequence (MTS).

13. In any one of claims 1 to 11, A base-corrected composition, wherein the DNA is chloroplast, plastid, or leucoplast DNA, and the one or more fusion proteins each contain a chloroplast transit signal (CTS).

14. In any one of claims 1 to 10, A base corrected composition, wherein the DNA is nuclear DNA and one or more of the fusion proteins comprises a nuclear localization signal (NLS).

15. In any one of claims 1 to 14, A base correcting composition that does not substantially cause correction of cytosine C bases to thymine T bases.

16. In any one of claims 1 to 14, A base correcting composition that corrects cytosine C bases to thymine T bases at a frequency of less than 20%, less than 10%, less than 5%, or less than 1%.

17. A base correcting composition having the activity of correcting an adenine A base in the DNA of a plant cell to a guanine G base, the base correcting composition comprising a fusion protein containing a DNA binding protein, cytosine deaminase, and adenine deaminase.

18. In claim 17, A base corrector composition, wherein the DNA is nuclear DNA or organelle DNA.

19. In claim 17 or claim 18, A base correcting composition, wherein the cytosine deaminase is APOBEC (apolipoprotein B editing complex), AID (activation-induced deaminase), TadA (tRNA-specific adenosine deaminase), or DddAtox, or a mutant thereof.

20. In any one of claims 17 to 19, A base correcting composition, wherein the cytosine deaminase is such that one or more amino acids selected from the group consisting of positions 37, 59, 109, and 129 of the amino acid sequence of SEQ ID NO: 2 are substituted with other amino acids.

21. In claim 20, A base correcting composition, wherein the cytosine deaminase is an amino acid sequence of SEQ ID NO: 2 in which S at position 37 is replaced with G, and / or G at position 59 is replaced with S, and / or A at position 109 is replaced with V, and / or S at position 129 is replaced with G.

22. In any one of claims 17 to 21, A base correcting composition, wherein the adenine deaminase is APOBEC (apolipoprotein B editing complex), AID (activation-induced deaminase), TadA (tRNA-specific adenosine deaminase), or a mutant thereof.

23. In any one of claims 17 to 22, A base correcting composition, wherein the fusion protein comprises one or more DNA binding proteins selected from the group consisting of zinc finger proteins, TALE proteins, and CRISPR-associated nucleases.

24. In any one of claims 17 to 23, A base corrected composition, wherein the DNA is organelle DNA and the fusion protein comprises a nuclear export signal (NES).

25. In any one of claims 17 to 24, A base correcting composition, wherein the DNA is mitochondrial DNA and the fusion protein comprises a mitochondrial targeting sequence (MTS).

26. In any one of claims 17 to 24, A base-corrected composition, wherein the DNA is chloroplast, variegated, or leucoplast DNA, and the fusion protein comprises a chloroplast transit signal (CTS).

27. In any one of claims 17 to 23, A base corrected composition, wherein the DNA is nuclear DNA and the fusion protein comprises a nuclear localization signal (NLS).

28. In any one of claims 17 to 27, A base correcting composition, wherein the fusion protein comprises uracil DNA glycosylase (UDG).

29. 29. In claim 28, A base correcting composition that does not substantially cause correction of cytosine C bases to thymine T bases.

30. 29. In claim 28, A base correcting composition that corrects cytosine C bases to thymine T bases at a frequency of less than 20%, less than 10%, less than 5%, or less than 1%.

31. A polynucleotide encoding any one of one or more fusion proteins contained in the base correction composition according to any one of claims 17 to 30, or a combination of two or more of the polynucleotides.

32. A vector comprising a polynucleotide or a combination of polynucleotides according to claim 31.

33. A base correcting composition comprising a vector according to claim 32.

34. 33. A plant cell transformed with a vector according to claim 32 and / or containing a polynucleotide or combination of polynucleotides according to claim 31.

35. 35. A plant grown from a plant cell according to claim 34.

36. 36. A plant that is a descendant or clone of a plant according to claim 35.

37. A seed obtained from a plant according to claim 35 or claim 36.

38. 38. A plant cell, plant or seed in which the adenine A base of the wild-type DNA has been corrected to guanine G, such as a plant cell according to claim 34, a plant according to claim 35 or claim 36, or a seed according to claim 37.

39. 39. A plant cell, plant, or seed according to claim 38, wherein the cytosine C base of the wild-type DNA is not corrected.

40. A method for correcting adenine A bases in plant organelle DNA to guanine G, comprising expressing a base correction composition according to any one of claims 1 to 30, or a polynucleotide or combination of polynucleotides according to claim 31 in a plant or plant cell of interest.

41. 41. The method of claim 40, wherein substantially no correction of cytosine C bases to thymine T bases occurs.

42. 41. The method of claim 40, wherein correction of cytosine C bases to thymine T bases occurs at a frequency of less than 20%, less than 10%, less than 5%, or less than 1%.