Adenine base editing system

By incorporating UDG and full-length DddAtox cytosine deaminase into the TALED tool, the problems of non-selectivity and target site limitation in adenine base editing in existing technologies have been solved, enabling selective editing of adenine bases in nuclear DNA and organelle DNA.

CN120826468APending Publication Date: 2025-10-21GREENGENE INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202480015386.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-03-06
Filing Date
2024-03-06
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

Existing TALED tools cannot selectively edit only adenine bases when editing plant chloroplast DNA, and the use of DddAtox cytosine deaminase in the form of a split variant results in limited target site selection and a large vector size.

Method used

A base editing composition containing uracil DNA glycosylase (UDG) is used, which combines full-length DddAtox cytosine deaminase and adenine deaminase to form the monomer TALED, avoiding cytosine base editing and achieving selective editing of adenine bases.

Benefits of technology

It enables selective editing of adenine bases in nuclear DNA and organelle DNA (such as chloroplasts and mitochondria), avoiding unnecessary editing of cytosine bases, and expanding the selectivity of editing target sites and the flexibility of vectors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120826468A_ABST
    Figure CN120826468A_ABST
Patent Text Reader

Abstract

The present invention relates to a base editing composition having an activity of editing an adenine (A) base of DNA into a guanine (G) base, and a base editing method using the same. The present invention has bases for editing nuclear DNA or organelle DNA, in particular bases for editing organelle DNA such as chloroplast or mitochondria. According to one aspect of the present invention, a base editing composition comprises a DNA binding protein, a cytosine deaminase, an adenine deaminase, and a uracil DNA glycosylase (UDG), and has an activity of selectively editing only an adenine base without substantially editing a cytosine (C) base. The present invention also provides a system for editing an adenine base of a plant cell DNA into a guanine base using a single fusion protein comprising a DNA binding protein, a cytosine deaminase and an adenine deaminase.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a base editing composition having the activity of editing the adenine (A) base of DNA to a guanine (G) base and a base editing method using the same. The present invention is useful for editing bases of nuclear DNA or organelle DNA, and is particularly useful for editing bases of organelle DNA such as chloroplasts or mitochondria. According to one aspect of the present invention, the base editing composition comprises a DNA binding protein, cytosine deaminase, adenine deaminase, and uracil DNA glycosylase (UDG), and the base editing composition has the activity of selectively editing only adenine bases without substantially editing cytosine (C) bases. The present invention also provides a system for editing adenine bases of plant cell DNA to guanine bases using a single fusion protein comprising a DNA binding protein, cytosine deaminase, and adenine deaminase. Background Art

[0002] The fusion protein linked to the DNA binding protein and the deaminase can induce DNA mutations (e.g., converting single nucleotides) in a targeted manner without generating double-strand breaks (DSBs) in the genome, thereby replacing nucleotides or editing bases, or editing point mutations that induce genetic disorders, or introducing desired single nucleotide mutations in prokaryotic cells and eukaryotic cells such as humans.

[0003] Programmable genome editing tools such as zinc-finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), clustered regularly interspaced short palindromic repeats (CRISPR) systems, and base editors composed of CRISPR-associated protein 9 (Cas9) variants and base deaminase proteins have the potential to be used for plant genetic research and crop trait improvement by changing base sequences. However, these existing genome editing tools are not suitable for editing DNA bases in plant organelles, including mitochondria and chloroplasts, especially because the guide RNA required to run the most widely used CRISPR system cannot be delivered to the organelles. Plant organelles encode many essential genes required for photosynthesis and cellular respiration. Methods or tools for editing the genes of these organelles are very necessary for studying the functions of these genes or improving crop yield and traits.

[0004] Bacterial toxin DddAtox It is derived from Burkholderia cenocepacia and plays the role of a bacterial toxin enzyme that deaminates cytosine in double-stranded DNA. tox It is toxic to cells. In order to avoid toxicity to host cells, it is usually divided into two inactivated splits for use. Each split (or half) can be connected to a DNA-binding protein designed to bind to DNA and can be used as a cytosine base editor (DdCBE, DddA-derived cytosine base editor) with pairing function.

[0005] In principle, this deamination enzymatic reaction requires two inactive halves to be brought into proximity on the target DNA by DNA-binding proteins in order to be activated, with cytosine to thymine (C-to-T) base editing occurring between the two DNA-binding protein binding sites. When transcription activator-like effector (TALE) proteins are used as DNA-binding proteins, the cytosine to thymine base conversion is induced in a region of 14 to 18 bases between the two TALE binding sites. Therefore, the efficiency of base editing may be altered by the DNA binding specificity of the TALE protein.

[0006] On the other hand, if DddA tox By linking cytosine deaminase and adenine deaminase, which can trigger adenine to guanine (A-to-G) editing, to DNA-binding proteins, editing of adenine bases can be achieved. Using transcription activator-like effector (TALE) proteins as DNA-binding proteins, TALE-linked deaminase (TALED) base editors linking cytosine deaminase and adenine deaminase differ from cytosine base editors that only trigger C to T editing and can achieve a variety of mutation spectra. Summary of the Invention

[0007] Problems to be solved by the invention

[0008] Existing TALEDs, because they possess both adenine and cytosine deaminases, can edit not only adenine bases present in target DNA but also cytosine bases. Experimental results confirm that this is particularly pronounced in plant chloroplasts. Because cytosine bases must be selectively edited to adenine bases while maintaining the wild-type sequence, the inventors of the present invention aimed to develop a method that selectively edits adenine bases to guanine bases without triggering the editing of undesired cytosine bases.

[0009] On the other hand, due to DddA tox In order to avoid toxicity, cytosine deaminase is divided into two inactive segments and used. Therefore, two TALE modules are required to connect the two segments so that the two segments can be located close to each other on the target DNA to be edited. In this case, there are the following disadvantages: if two TALE modules are used, the target position selection for DNA editing is limited, and the size of the vector required for expressing the base editor must be larger. Therefore, the inventors of the present invention aim to provide a method of using a full-length DddA connected to a detoxified form. tox In addition, another object of the present invention is to provide a method for editing adenine bases in plant cell DNA using a monomeric TALED of cytosine deaminase. tox The monomeric TALED of cytosine deaminase selectively edits only adenine bases without triggering the editing of cytosine bases.

[0010] Means used to solve problems

[0011] The present invention provides a base editing composition, which edits adenine bases to guanine base editing compositions, and the base editing composition further comprises uracil DNA glycosylase (UDG) on the basis of DNA binding protein, cytosine deaminase and adenine deaminase. The base editing composition does not substantially trigger the editing of cytosine bases, but only selectively edits adenine bases, thereby having the activity of editing the adenine bases to guanine. By utilizing this base editing composition, a base editing system that can only selectively edit adenine bases can be provided, and this system can be used not only for editing nuclear DNA, but also for editing organelle DNA such as chloroplasts or mitochondria. The base editing compositions of the present invention can be used in both plant cells and animal cells, especially by converting plant cells into the base editing compositions of the present invention, so as to obtain plants or their seeds that selectively edit only the adenine bases of the desired target DNA.

[0012] The present invention also provides a base editing composition that edits the adenine base of a plant cell to guanine, wherein the base editing composition comprises a DNA binding protein, a cytosine deaminase, and an adenine deaminase, and the cytosine deaminase is included in a full-length form rather than a segmented form. When using the base editing composition, there is no need to use a fusion protein comprising two or more DNA binding proteins (preferably TALE proteins) and deaminases, thereby allowing for unrestricted selection of target positions for DNA editing. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 To utilize DddA tox Schematic diagram of the base editor of the segmented body targeting the plant chloroplast gene psaA, showing the case without UDG (A) and with UDG (B). CTS represents chloroplast transit signal, AD represents adenine deaminase, NTD represents N-terminal domain, CTC represents C-terminal domain, and UDG represents uracil DNA glycosylase. "Right repeats" and "Left TALE repeats" represent TALE arrays, respectively, and the base sequences marked in blue represent the sites to which the TALE protein (composed of the N-terminal domain, TALE array, and C-terminal domain) binds. Deaminase-based base editing occurs between the DNA base sequences bound by the two TALE proteins.

[0014] Figure 2 Is the use of full-length DddA tox Schematic diagram of the base editor GSVG targeting the plant chloroplast gene psaA, showing case A without UDG and case B with UDG. GSVG refers to a variant of cytosine deaminase comprising the amino acid sequence of SEQ ID NO. 2, in which serine (S) at position 37, glycine (G) at position 59, alanine (A) at position 109, and serine (S) at position 129 are substituted with G, S, V, and G, respectively (SEQ ID NO. 9).

[0015] Figure 3 The base editing efficiency of the position C2 where cytosine base editing occurs in the psaA target DNA in the first generation plants transformed by the base editor of the present invention is shown. C2 refers to the position C2 where cytosine base editing occurs. Figure 1 and Figure 2The second base in the DNA sequence to be edited between the first TALE protein (TALE on the left; SEQ ID NO. 10) and the second TALE protein (TALE on the right; SEQ ID NO. 11) of the base editing composition used is cytosine. Col-0 refers to a wild-type plant individual that has not been transformed. Figure A is a plant that uses a DddA-linked tox In the case of a split TALE pair, Figure B shows the use of a TALE with full-length DddA attached. tox Col-0 is a wild-type individual that has not been transformed. "L" represents the left TALE with the amino acid sequence of SEQ ID NO.10, and "R" represents the right TALE with the amino acid sequence of SEQ ID NO.11. "AD" represents adenine deaminase with the amino acid sequence of SEQ ID NO.1. "GSVG" represents the full-length DddA with the amino acid sequence of SEQ ID NO.9. tox GSVG.

[0016] Figure 4 The base editing rate at the psaA target site in three wild-type plants (Col-0) is shown. Since this value should theoretically be zero, it is interpreted as due to sequencing error and serves as a benchmark for determining whether base editing has actually occurred.

[0017] Figure 5 Shown by using DddA tox Base editing efficiency of psaA target DNA in 24 first-generation plants transformed with a TALE pair that splits and does not use UDG. L-1397N indicates a base editing site with DddA attached to the N-terminus. tox The left TALE of the split body, R-1397C-AD, indicates the C-terminal DddA is connected tox The right TALE of the segmentosome and adenine deaminase.

[0018] Figure 6 Use DddA tox The segmented figure shows the base editing efficiency of the psaA target DNA in 24 first-generation plants transformed with a base editor of a TALE pair in which the left TALE is linked to UDG and the right TALE is linked to adenine deaminase.

[0019] Figure 7 Use DddA tox The segmented figure shows the base editing efficiency of the psaA target DNA in 24 first-generation plants transformed with a base editor in which the right TALE is linked to adenine deaminase and UDG.

[0020] Figure 8 Use DddA tox The segmented figure shows the base editing efficiency of the psaA target DNA in 24 first-generation plants transformed with a base editor of a TALE pair in which the left TALE is linked to UDG and the right TALE is linked to adenine deaminase and UDG.

[0021] Figure 9 Use DddA tox The figure of the split body without using UDG is a graph showing the base editing efficiency of the psaA target DNA in 17 first-generation plants transformed with a TALE pair in which the TALE on the left is linked to adenine deaminase.

[0022] Figure 10 Use DddA tox The segmented figure shows the base editing efficiency of the psaA target DNA in 14 first-generation plants transformed with a base editor in which the TALE pair on the left is linked to adenine deaminase and UDG.

[0023] Figure 11 Use DddA tox The segmented figure shows the base editing efficiency of the psaA target DNA in 24 first-generation plants transformed with a base editor of a TALE pair in which the left TALE is linked to adenine deaminase and the right TALE is linked to UDG.

[0024] Figure 12 Use DddA tox The segmented figure shows the base editing efficiency of the psaA target DNA in 16 first-generation plants transformed with a base editor of a TALE pair in which the left TALE is linked to adenine deaminase and UDG, and the right TALE is linked to UDG.

[0025] Figure 13 Is the use of full-length DddA tox The figure of the variant is a graph showing the base editing efficiency of the psaA target DNA in 20 first-generation plants transformed with the monomeric TALED base editor that does not contain UDG.

[0026] Figure 14 Is the use of full-length DddA tox The figure of the variant is a graph showing the base editing efficiency of the psaA target DNA in 29 first-generation plants transformed with the monomeric TALED base editor linked to UDG.

[0027] Figure 15 Is the use of DddA toxSplit or full-length DddA tox (SEQ ID NO.9) Schematic diagram of a base editor targeting the plant chloroplast gene psbA. CTS represents the chloroplast transit signal, NTD represents the N-terminal domain of the TALE protein, and CTD represents the C-terminal domain of the TALE protein. "Left TALE repeats (left TALErepeats)" and "right TALErepeats (right TALErepeats)" represent the TALE arrays, respectively, and the base sequences underlined in blue represent the sites to which the TALE protein (composed of the N-terminal domain, TALE array, and C-terminal domain) binds. Deaminase-based base editing occurs between the DNA base sequences bound by the two TALE proteins. Figure 15 The A part is to use DddA tox Figure of the split body (1397N and 1397C), Figure 15 Part B and Figure 15 The C part is to use the full length DddA tox ("GSVG"; SEQ ID NO. 9), and plots using only the left TALE repeat sequence (left TALErepeats) or only the right TALE repeat sequence (right TALErepeats), respectively.

[0028] Figure 16 Show use Figure 15 The results of base editing efficiency determination were performed on the individuals that survived the atrazine treatment in the first generation plants transformed with the base editor shown. Col-0 is a wild-type individual that was not transformed. "Left side 1 (left side 1)" and "L1" have the same meaning, both representing the left TALE protein of SEQ ID NO.14. "L2" represents the left TALE sequence of SEQ ID NO.15. "R1" represents the right TALE protein of SEQ ID NO.16. "Right side 2 (right side 2)" and "R2" have the same meaning, both representing the right TALE protein of SEQ ID NO.17. DETAILED DESCRIPTION

[0029] Unless defined otherwise, all technical and scientific terms used in this specification have the same meaning as commonly understood by those skilled in the art to which the present invention pertains. Generally, the terms used in this specification are well known and commonly used in the art.

[0030] As used herein, the terms "correction," "modification," and "editing" are used interchangeably to refer to methods for altering a nucleic acid sequence by selectively mutagenizing a specific genomic target. Such specific genomic targets include, but are not limited to, genes, promoters, open reading frames, or any nucleic acid sequence.

[0031] The terms "base editor", "base editing system" and "base correction system" used in this specification are used interchangeably to refer to substances that have the activity of changing nucleic acid sequences through selective mutation of genomic targets, and include combinations of more than one different base editors. The terms "base editor", "base editing system" or "base correction system" used in this specification may be a polypeptide (which may be a fusion protein) or a polynucleotide or a combination thereof, or a composition comprising one or more polypeptides (which may be fusion proteins) or polynucleotides or a combination thereof, depending on the context. Therefore, the terms "base editing system" or "base editing composition" used in this specification may include one base editor or a combination of two or more different base editors, wherein the different base editors may be used simultaneously or individually.

[0032] The term "fusion protein" as used in this specification refers to a polypeptide formed by two or more different polypeptides bound by peptide bonds. The fusion protein used in the present invention comprises a DNA binding protein and a deaminase or a segment thereof, and may include additional sequences such as UDG, NLS, NES, MTS, CTS, and may include additional sequences for biotechnology methods such as tags, and these individual polypeptides may be directly connected or connected through a linker. "Linker" refers to any molecule that connects two different molecules. In the field of biotechnology, any linker known to provide a fusion protein or protein conjugate can be used, for example, a peptide linker comprising 1 to 100 amino acid residues can be used.

[0033] The design and construction of a fusion protein or a polynucleotide encoding the same can be performed using any method known in the art. The polynucleotide can be inserted into a vector, and the vector can be introduced into a cell. The individual proteins constituting the fusion protein according to the present invention are typically cloned as a single polynucleotide and expressed as a single polypeptide (fusion protein). However, one or more of the individual proteins can also be cloned as (multiple) separate polynucleotides and expressed as two or more separate polypeptides, and this scenario also falls within the scope of the present invention.

[0034] As used herein, the terms "target," "targeting," "target site," or "targeting portion" refer to a predetermined nucleic acid sequence of any composition and / or length. Such target sites include, but are not limited to, genes, promoters, open reading frames, or any nucleic acid sequence.

[0035] The present invention provides a base editing composition that has the activity of editing adenine bases in DNA to guanine. The base editing composition comprises one or more fusion proteins, each of which comprises a DNA-binding protein and a cytosine deaminase. At least one of the fusion proteins comprises adenine deaminase, and at least one of the fusion proteins comprises UDG. The cytosine deaminase can exist in full-length or split form.

[0036] The base editing composition according to the present invention is characterized in that it contains cytosine deaminase and adenine deaminase. In particular, TALE protein is used as a DNA binding protein, and the base editors to which cytosine deaminase and adenine deaminase are connected are called TALED. Unlike the DdCBE base editor that could only perform cytosine base editing in a specific motif in organelle DNA, TALED can achieve base editing from A to G. In this application specification, the term "TALED" refers to a base editor to which cytosine deaminase and adenine deaminase are connected to TALE proteins, and the TALE protein can be used as a monomer (for example, when cytosine deaminase is used in full-length form), or as a paired dimer form (for example, when cytosine deaminase is used in segmented form), and adenine deaminase can also be connected to all of the transcription activator-like effector proteins, or only to one of them.

[0037] The base editing composition according to the present invention can be used not only to edit nuclear DNA, but also to edit organelle DNA such as chloroplasts or mitochondria. The base editing composition of the present invention can be used in both plant cells and animal cells. In particular, by converting plant cells into the base editing composition of the present invention, plants or their seeds can be obtained that selectively edit only the adenine bases of the desired target DNA.

[0038] The cytosine deaminase used in the present invention may be contained in the form of a first segment and a second segment, and the first segment and the second segment each have a form bound to a DNA-binding protein.

[0039] The cytosine deaminase that can be used in the base editor according to the present invention refers to a deaminase having the activity of converting cytosine base to uridine, which can be derived from or mutated from (e.g., engineered or evolved) any organism (e.g., eukaryotic or prokaryotic), wherein the organism includes algae, bacteria, fungi, plants, invertebrates and mammals, but is not limited thereto. For example, it can be: APOBEC (apolipoprotein B editing complex); AID (Activation-induced cytidine deaminase); cytosine deaminase derived from or mutated from TadA (tRNA-specific adenosine deaminase) as a bacterial adenine deaminase or its ortholog; or cytosine deaminase or its fragment derived from or mutated from DddA as a bacterial cytosine deaminase or its ortholog. The cytosine deaminase mutated from the above-mentioned TadA may be, for example, a polypeptide in which one or more of the amino acid residues at positions 6, 26, 27, 28, 46, 48, 49, 61, 74, 76, 77, 82, 96, 107, 108, 112, 114, 115, 119, 122, 127, 142, 143, 151, 154, and 158 in the amino acid sequence of SEQ ID NO. 1 are mutated to another amino acid. For example, the polypeptide may be a polypeptide in which the amino acid residue at position 27 is mutated to lysine, the amino acid residue at position 28 is mutated to alanine, the amino acid residue at position 61 is mutated to isoleucine, and the amino acid residue at position 96 is mutated to asparagine in the amino acid sequence of SEQ ID NO. 1. Regarding the composition of cytosine deaminase that can be used in the present invention, reference can be made to the contents known before the date of this application, in particular, all technical contents described in International Patent Application Publication Nos. WO2022 / 060185 and WO2023 / 086953, which are incorporated into the present application by reference.

[0040] When the base editor according to the present invention includes cytosine deaminase, the cytosine deaminase can be included in the form of a first segment and a second segment, or in a full-length form. When the first segment and the second segment are in the form of a first segment and a second segment, the first segment and the second segment are each connected to a DNA binding protein.

[0041] In this specification, when it is mentioned that two proteins are "linked", they may be directly linked or indirectly linked via a linker or other protein(s).

[0042] The cytosine deaminase used in the present invention can be DddA tox ,DddAtox It is the enzymatic part of the bacterial toxin from Burkholderia cenocepacia that deaminates cytosine in double-stranded DNA. tox It may comprise the amino acid sequence of SEQ ID NO.2.

[0043] SEQ ID NO.2 shows: (wild-type DddA tox )

[0044] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC

[0045] Due to DddA tox Toxic to cells, in order to avoid toxicity in host cells, it can be used in the form of two inactive splits (i.e., a first split and a second split). When the cytosine deaminase used in the present invention is used in the form of a first split and a second split, each of the first split and the second split has no deamination activity.

[0046] The DddA tox The first segment of cytosine deaminase may comprise a sequence from the N-terminus to G33, G44, A54, N68, G82, N98, or G108 in the amino acid sequence of SEQ ID NO.2, and the second segment may comprise a sequence from G34, P45, G55, N69, T83, A99, or A109 to the C-terminus in the amino acid sequence of SEQ ID NO.2.

[0047] Preferably, DddA tox The first segment of cytosine deaminase may include the sequence from the N-terminus to G44 in the amino acid sequence of SEQ ID NO. 2 (SEQ ID NO. 3 below), and the second segment may include the sequence from P45 to the C-terminus (SEQ ID NO. 4 below).

[0048] SEQ ID NO.3: (wild type DddA tox G1333-N)

[0049] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGG

[0050] SEQ ID NO.4: (wild type DddA tox G1333-C)

[0051] PTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVN MTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC

[0052] Preferably, DddA tox The first segment of cytosine deaminase may further comprise the sequence from the N-terminus to G108 in the amino acid sequence of SEQ ID NO. 2 (SEQ ID NO. 5 below), and the second segment may further comprise the sequence from A109 to the C-terminus (SEQ ID NO. 6 below).

[0053] SEQ ID NO.5: (wild type DddA tox G1397-N)

[0054] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEG

[0055] SEQ ID NO.6: (wild type DddA tox G1397-C)

[0056] AIPVKRGATGETKVFTGNSNSPKSPTKGGC

[0057] When DddA tox When the first segment and the second segment of the cytosine deaminase are used as cytosine deaminase, one or more amino acids on the surface where the first segment and the second segment of the cytosine deaminase bind to each other may be substituted with other amino acids. toxThe first segment and the second segment may respectively comprise the amino acid sequences of SEQ ID NO.3 (G1333-N) and SEQ ID NO.4 (G1333-C). In this case, one or more amino acids selected from the group consisting of positions 3, 5, 10, 11, 13, 14, 15, 16, 17, 18, 19, 28, 30 and 31 of SEQ ID NO.3, or one or more amino acids selected from the group consisting of positions 13, 16, 17, 20, 21, 28, 29, 30, 31, 32, 33, 56, 57, 58 and 60 of SEQ ID NO.4 may be substituted with other amino acids, but are not limited thereto. As another example, the first segment of DddAtox may comprise the amino acid sequences of SEQ ID NO.5 (G1397-N) and SEQ ID NO.6 (G1397-C). In this case, one or more amino acids selected from the group consisting of positions 87, 88, 91, 92, 95, 100, 101, 102, and 103 of SEQ ID NO.5 or one or more amino acids selected from the group consisting of positions 13, 14, 15, and 16 of SEQ ID NO.6 may be substituted with other amino acids, but are not limited thereto. The “other amino acids” refer to amino acids selected from alanine, isoleucine, leucine, methionine, phenylalanine, proline, tryptophan, valine, asparagine, cysteine, glutamine, glycine, serine, threonine, tyrosine, aspartic acid, glutamic acid, arginine, histidine, lysine, and all known variants of the amino acids, excluding the amino acids originally present in the wild-type protein at the mutation site. When such variants are used, the two DddA tox When a pair of split bodies cannot bind to DNA, they cannot function properly, allowing for very efficient and precise cytosine to thymine editing (C-to-T editing) without causing unwanted off-target C-to-T editing.

[0058] In this specification, the term "G1333-N", "G1333N" or "1333N" may refer to the wild-type DddA having the amino acid sequence of SEQ ID NO. 3. tox The first segment or amino acid variant thereof, the term "G1333-C", "G1333C" or "1333C" may refer to the wild-type DddA having the amino acid sequence of SEQ ID NO.4 tox or an amino acid variant thereof.

[0059] In this specification, the term "G1397-N", "G1397N" or "1397N" may refer to the wild-type DddA having the amino acid sequence of SEQ ID NO.5.tox The first segment or amino acid variant thereof, the term "G1397-C", "G1397C" or "1397C" may refer to the wild-type DddA having the amino acid sequence of SEQ ID NO.6 tox or an amino acid variant thereof.

[0060] The cytosine deaminase used in the present invention can be used in full-length form. In this case, the full-length cytosine deaminase used (e.g., DddA tox ) has an amino acid sequence modified to be non-toxic or only have low toxicity. tox The C-terminus of DddA specifically aggregates positively charged amino acids. Since DNA is negatively charged, it binds to positively charged amino acids in proteins. By replacing these positively charged amino acids, the binding of DddA to DddA can be weakened. tox The binding force to DNA can be improved, thereby reducing or eliminating the toxicity in cells. In other words, if the toxicity is eliminated by replacing positively charged amino acids, cloning can be performed in E. coli, thereby ensuring the full-length DddA. tox Such non-toxic full-length cytosine deaminase can be provided by replacing one or more, two or more, three or more, four or more, or five or more amino acids in the wild-type amino acid sequence of SEQ ID NO.2 with other amino acids. The "other amino acids" refer to amino acids selected from alanine, isoleucine, leucine, methionine, phenylalanine, proline, tryptophan, valine, asparagine, cysteine, glutamine, glycine, serine, threonine, tyrosine, aspartic acid, glutamic acid, arginine, histidine, lysine and all known variants of the amino acids, excluding the amino acids originally present in the wild-type protein at the mutation site. For example, the other amino acid can be alanine.

[0061] The non-toxic full-length DddA tox It may comprise an amino acid sequence selected from the group consisting of the following amino acid sequences.

[0062] A1341D KRKKA variant

[0063] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYDNAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTAGGC

[0064] AAAAA variant

[0065] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVAAGATGETAVFTGNSNSPASPTAGGC

[0066] AAAAK variant

[0067] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVAAGATGETAVFTGNSNSPASPTKGGC

[0068] AAKAA variant

[0069] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVAAGATGETKVFTGNSNSPASPTAGGC

[0070] AAKAK variant

[0071] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVAAGATGETKVFTGNSNSPASPTKGGC

[0072] KAAAA variant

[0073] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKAGATGETAVFTGNSNSPASPTAGGC

[0074] E1347A variant

[0075] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVAGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC

[0076] Preferably, the full-length cytosine deaminase variant useful in the present invention may have one or more amino acid substitutions selected from the following group: a group consisting of a substitution of S at position 37 to G, a substitution of G at position 59 to S, a substitution of A at position 109 to V, and a substitution of S at position 129 to G in the amino acid sequence of SEQ ID NO. 2.

[0077] More preferably, the full-length cytosine deaminase variant that can be used in the present invention may have all of the amino acid sequence of SEQ ID NO. 2, wherein S at position 37 is substituted with G, G at position 59 is substituted with S, A at position 109 is substituted with V, and S at position 129 is substituted with G. In this case, its sequence is as follows.

[0078] GSVG variants

[0079] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLEGKVFSSGGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGVIPVKRGATGETKVFTGNSNGPKSPTKGGC

[0080] As yet another example, a full-length cytosine deaminase variant useful in the present invention may comprise the following sequence.

[0081] SSVG variants

[0082] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGVIPVKRGATGETKVFTGNSNGPKSPTKGGC

[0083] GSAG variants

[0084] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLEGKVFSSGGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNGPKSPTKGGC

[0085] GSVS variant

[0086] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLEGKVFSSGGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGVIPVKRGATGETKVFTGNSNSPKSPTKGGC

[0087] The adenine deaminase that can be used in the base editor of the present invention refers to a deaminase having the activity of converting an adenine base into hypoxanthine (inosine as a nucleoside), which can be derived from or mutated from (e.g., engineered or evolved) any organism (e.g., a eukaryote or a prokaryote), for example, it can be derived from or mutated from Escherichia coli (Escherichiacoli, E. coli), Staphylococcus aureus (S. aureus), Salmonella Typhi (S. Typhi), Shewanella putrefaciens (S. putrefaciens), Haemophilus influenzae (H. influenzae) or Caulobacter crescentus (C. crescentus), wherein the organism includes algae, bacteria, fungi, plants, invertebrates and mammals, but is not limited thereto. Such adenine deaminase may be, for example, APOBEC, AID, or TadA, or a variant thereof. The TadA may be, for example, TadA8e (SEQ ID NO.1) or a cleaved form or variant thereof (e.g., a variant improved or evolved to be suitable for deoxynucleotides). The variant of TadA8e may be, for example, one or more of amino acid residues 23, 28, 30, 36, 46, 48, 49, 51, 76, 82, 82, 84, 106, 108, 110, 111, 146, 147, 152, 154, 155, 156, and 157 of SEQ ID NO.1 mutated to other amino acids. With respect to the composition of the adenine deaminase that can be used in the present invention, reference may be made to the contents known before the date of this application, in particular, all technical contents described in International Patent Application Publication Nos. WO2022 / 060185 and WO2023 / 086953, which are incorporated herein by reference. The base editing composition may include adenine deaminase that can be used in the present invention and may also include the amino acid sequence of SEQ ID NO.1 or a conservative amino acid substitution thereof.

[0088] In the base editor according to the present invention, when cytosine deaminase is used in the form of a segment, adenine deaminase can be connected to the N-terminus or C-terminus of the first segment of cytosine deaminase, and / or the N-terminus or C-terminus of the second segment.

[0089] Regarding the composition of the deaminase that can be used in the present invention, in addition to International Patent Application Publication No. WO2022 / 060185, which is incorporated by reference in its entirety herein, known contents before the present application can also be cited.

[0090] The DNA-binding protein according to the present invention can be a zinc finger protein, a TALE protein, or a CRISPR-associated nuclease, or a combination thereof. Regarding the composition of zinc finger proteins, TALE proteins, and CRISPR-associated nucleases, reference can be made to the publicly known content prior to the date of this application, in particular, all technical contents described in International Patent Application Publication No. WO2022 / 060185, etc., which are incorporated herein by reference.

[0091] The DNA binding protein according to the present invention can be a TALE protein. The TALE protein of the present invention refers to a protein that binds to nucleotides in a sequence-specific manner through one or more TALE-repeat modules. The TALE protein comprises at least one TALE-repeat module, preferably, 1 to 30 TALE-repeat modules, but is not limited thereto. As used in this specification, the TALE-repeat module can be referred to as a "TALE array", and the expression "TALE protein" refers to a structure comprising an N-terminal domain and a C-terminal domain (which may include a half-domain) on both sides of the TALE array. Depending on the context, the term "TALE" used in this specification may mean only a "TALE array" or a "TALE protein".

[0092] When using TALE proteins as DNA binding proteins used in base editors according to the present invention, single-module TALE arrays or multi-module TALE arrays (e.g., a dual-module TALE array of a first TALE array and a second TALE array) can also be used. For example, the first segment of cytosine deaminase can be combined with a first TALE protein (left TALE) (first fusion), and the second segment of the cytosine deaminase can be combined with a second TALE protein (right TALE) (second fusion). In this case, along the NC direction, the first TALE protein (left TALE) is connected to the first segment of cytosine deaminase, the second TALE protein (right TALE) is connected to the second segment of cytosine deaminase, and the first TALE protein (left TALE) and the second TALE protein (right TALE) are respectively connected by the structures of the N-terminal domain, the TALE array and the C-terminal domain (which may include a half domain). The first segment can be an N-terminal side segment or a C-terminal side segment of a full-length cytosine deaminase, and the second segment can also be an N-terminal side segment or a C-terminal side segment of a full-length cytosine deaminase. Even if the cytosine deaminase is full-length, a single-module TALE array or a multi-module TALE array can be used. When a single-module TALE array is used, a single TALE domain and a cytosine deaminase are connected along the NC direction. When a double-module TALE is used, along the NC direction, the first TALE protein (left TALE or right TALE) is connected to the full-length cytosine deaminase, and the second TALE protein (left or right TALE) can be included separately.

[0093] When cytosine deaminase is used in the form of a first segment and a second segment, and the DNA binding protein is a TALE protein, the base editor of the present invention may have the following composition form: the first segment of the cytosine deaminase is connected to the first fusion of the first TALE, and the second segment of the cytosine deaminase is connected to the second fusion of the second TALE. The first fusion and the second fusion have N'-TALE-first segment (cytosine deaminase)-C' and N'-TALE-second segment (cytosine deaminase)-C' structures, respectively. Adenine deaminase can be connected to the first fusion or the second fusion, or to both, specifically, it can be bound to the N-terminus or C-terminus of the first segment of the cytosine deaminase in the first fusion, and / or the N-terminus or C-terminus of the second segment of the cytosine deaminase in the second fusion.

[0094] When the cytosine deaminase is used in full-length form and the DNA-binding protein is a TALE protein having a single TALE module, the single TALE module and the cytosine deaminase are included along the C-terminus of the single TALE module. In this case, the adenine deaminase is linked along the C-terminus of the single TALE module and can be linked to the N-terminus or C-terminus of the cytosine deaminase. In this specification, the term "monomeric TALED" refers to a form in which a single TALE module is used as the DNA-binding protein, and the cytosine deaminase and adenine deaminase are linked.

[0095] When cytosine deaminase is used in full-length form and the DNA binding protein is a TALE protein having a dual TALE module, the base editor of the present invention may have the following composition form: a first fusion of a first TALE module and a cytosine deaminase connected in the NC direction, and a second fusion comprising an adenine deaminase and a second TALE. The first fusion and the second fusion have the structures of N'-TALE-cytosine deaminase-C' and N'-TALE-adenine deaminase-C', respectively, and the adenine deaminase can be bound to the N-terminus or C-terminus of the TALE.

[0096] According to one embodiment of the present invention, the base editing composition of the present invention is characterized in that it contains uracil DNA glycosylase (UDG). Specifically, one or more of the fusion proteins contained in the base editing composition according to the present invention contains UDG. It is known that UDG recognizes damaged DNA in the natural state and has the activity of selectively removing only uracil bases in DNA. The A to G base editor according to the present invention includes DddA tox The cytosine deaminase is expressed, so the cytosine base is deaminated and converted into uracil. Like this, when UDG is expressed at the same time, the converted uracil base can be removed and the original DNA sequence can be repaired, thereby avoiding unwanted cytosine base editing. This technology is particularly useful for selectively editing only adenine bases in environments where the natural expression level of UDG is low, especially in the organelles of plant cells. UDG of any species can be used, for example, UDG of Arabidopsis, humans, mice, tobacco, rice, etc., and UDG can have different lengths, for example, comprising 2, 5, 10, 16, 24 or 32 amino acids. For example, the UDG used in the present invention may comprise the amino acid sequence of SEQ ID NO.7, but is not limited thereto.

[0097] UDG is preferably used as a fusion protein with the DNA-binding protein, cytosine deaminase, and adenine deaminase used in the present invention, but can also be delivered to the target DNA in a form separated from these components (protein or polynucleotide). When used as a fusion protein, UDG can be linked to the C-terminus of cytosine deaminase (or its fragment) or the C-terminus of adenine deaminase, but is not necessarily limited thereto.

[0098] According to one embodiment of the present invention, a base editor according to the present invention comprising UDG is capable of selectively editing only adenine bases of DNA without substantially triggering editing of cytosine bases to thymine bases. For example, in a UDG-linked TALED according to the present invention, cytosine bases can be edited to thymine at a frequency of less than 20%, less than 10%, less than 5% or less than 1%. The base editing frequency (or efficiency) can be determined by sequencing the base editing frequency (or efficiency) calculated for target DNA in cells transformed by expressing the base editor according to the present invention, for example, as the percentage of sequencing reads reflecting the desired base editing result out of all sequencing reads of DNA obtained from base-edited cells or individuals.

[0099] One or more of the fusion proteins included in the base editing composition according to the present invention may further include a nuclear export signal (NES). When the NES is attached to the base editing protein, base editing can be performed with higher efficiency. The NES sequence can be any signal sequence that confers nuclear export capability (e.g., VDEMTKKFGTLTIHDTEK), and a natural nuclear export signal (Nuclear Export Signal, NES) or an artificially synthesized NES can be used. For example, it can be derived from mirute virus of mice (MVM), but is not limited thereto.

[0100] One or more of the fusion proteins included in the base editing composition according to the present invention may also include a mitochondrial target sequence (MTS). The MTS that can be used in the present invention may be any signal sequence that has the ability to translocate into mitochondria, and may be a natural MTS present at the N-terminus of a variety of mitochondrial proteins. In addition, artificially synthesized MTS may also be used. When using MTS in the base editing composition according to the present invention, its position may be diversified, for example, it may be directly or indirectly (for example, through a linker and / or other protein components) connected to the N-terminus of the DNA binding protein or the N-terminus of the NES, but is not limited thereto.

[0101] One or more of the fusion proteins included in the base editing composition according to the present invention may further include a chloroplast transit signal (CTS), and the base editing composition may be used to edit chloroplast, chromoplast or leucoplast DNA in plant cells. The CTS that can be used in the present invention can be any signal sequence that has the ability to translocate into the chloroplast, and can be a natural CTS present at the N-terminus of a variety of chloroplast proteins. In addition, an artificially synthesized CTS can also be used. When using a CTS in the base editing composition according to the present invention, its position can be diversified, for example, it can be directly or indirectly (for example, through a linker and / or other protein components) connected to the N-terminus of a DNA binding protein or the N-terminus of an NES, but is not limited thereto.

[0102] One or more of the fusion proteins included in the base editing composition according to the present invention may include a nuclear localization signal (NLS). The NLS that can be used in the present invention can be any signal sequence that has the ability to translocate into the nucleus, or a natural NLS present at the N-terminus of a variety of nucleoproteins, or an artificially synthesized NLS. When using an NLS in the base editing composition according to the present invention, its position can be diversified, for example, it can be directly or indirectly (for example, through a linker and / or other protein components) connected to the N-terminus of the DNA binding protein, but is not limited thereto.

[0103] The base editor according to the present invention has the form of a fusion protein. Specifically, the fusion protein comprises a DNA binding protein and a deaminase or a segment thereof, and may include additional sequences such as UDG, NLS, NES, MTS, CTS, and may include additional sequences for biotechnology methods such as tags, and these polypeptides may be directly connected or connected through a linker. The fusion protein (or polynucleotide encoding the fusion protein) can be designed and constructed by any method known in the field of biotechnology.

[0104] In the present invention, the base editor can be in the form of a polynucleotide encoding a fusion protein described in the specification of this application. The polynucleotide can be inserted into a vector, and the vector can be introduced into a cell.

[0105] The present invention also provides plant cells transformed by the vector, plants cultured from the plant cells, progeny or cloned plants of the plant, and seeds obtained from the plant.

[0106] In the plant cells, plants, and seeds according to the present invention, the adenine base of the wild-type DNA is edited to guanine. In particular, in plant cells, plants, and seeds transformed with a base editor comprising UDG (or a polynucleotide encoding the base editor), while the cytosine base of the wild-type DNA is not edited, only the adenine base can be selectively edited.

[0107] The present invention also provides a method for editing adenine bases in DNA to guanine, comprising the following steps: expressing the base editing composition described in this application or a polynucleotide encoding the same in an animal or animal cell, or a plant or plant cell of interest. The DNA may be nuclear DNA or organelle DNA, preferably plant organelle DNA.

[0108] The present invention can be directed to the following embodiments (1) to (51) based on the above contents, but is not limited thereto.

[0109] Embodiment (1): A base editing composition having the activity of editing the adenine (A) base of DNA into the guanine (G) base, wherein the base editing composition comprises one or more fusion proteins, each fusion protein comprises a DNA binding protein and a cytosine deaminase, and at least one of the one or more fusion proteins comprises adenine deaminase, at least one of the one or more fusion proteins comprises uracil DNA glycosylase (UDG), and the cytosine deaminase exists in the form of full length or two split bodies.

[0110] Embodiment (2): The base editing composition according to embodiment (1), wherein the DNA is nuclear DNA or organelle DNA.

[0111] Embodiment (3): The base editing composition according to embodiment (1) or embodiment (2), wherein the cytosine deaminase is apolipoprotein B editing complex (APOBEC), activation-induced cytidine deaminase (AID) or DddA tox , or variations thereof.

[0112] Embodiment (4): A base editing composition according to any one of embodiments (1) to (3), wherein the cytosine deaminase comprises the amino acid sequence of SEQ ID NO.2 and is contained in full-length form, and one or more amino acids selected from the group consisting of positions 37, 59, 109 and 129 in the amino acid sequence of SEQ ID NO.2 are replaced by different amino acids.

[0113] Embodiment (5): A base editing composition according to embodiment (4), wherein the cytosine deaminase comprises the amino acid sequence of SEQ ID NO. 2 and has one or more amino acid substitutions selected from the group consisting of serine (S) at position 37 being replaced by glycine (G), glycine (G) at position 59 being replaced by serine (S), alanine (A) at position 109 being replaced by valine (V), and serine (S) at position 129 being replaced by glycine (G).

[0114] Embodiment (6): A base editing composition according to any one of embodiments (1) to (3), wherein cytosine deaminase is contained in the form of a first segment and a second segment, one fusion protein contains the first segment, and another fusion protein contains the second segment, and one or more amino acids located on the dimerization surface of the first segment and the second segment are replaced by different amino acids.

[0115] Embodiment (7): A base editing composition according to embodiment (6), wherein the first segment of cytosine deaminase comprises a sequence from the N-terminus to glycine (G) at position 33, glycine (G) at position 44, alanine (A) at position 54, asparagine (N) at position 68, glycine (G) at position 82, asparagine (N) at position 98, or glycine (G) at position 108 in the amino acid sequence of SEQ ID NO.2, and the second segment of cytosine deaminase comprises a sequence from glycine (G) at position 34, proline (P) at position 45, glycine (G) at position 55, asparagine (N) at position 69, threonine (T) at position 83, alanine (A) at position 99 or alanine (A) at position 109 to the C-terminus in the amino acid sequence of SEQ ID NO.2.

[0116] Embodiment (8): A base editing composition according to embodiment (6), wherein the first segment of cytosine deaminase comprises the amino acid sequence of SEQ ID NO.5, the second segment of cytosine deaminase comprises the amino acid sequence of SEQ ID NO.6, and one or more amino acids selected from the group consisting of positions 87, 88, 91, 92, 95, 100, 101, 102 and 103 of SEQ ID NO.5 or one or more amino acids selected from the group consisting of positions 13, 14, 15 and 16 of SEQ ID NO.6 are replaced by different amino acids.

[0117] Embodiment (9): A base editing composition according to any one of embodiments (1) to (8), wherein the adenine deaminase is apolipoprotein B editing complex (APOBEC), activation-induced cytidine deaminase (AID) or tRNA-specific adenosine deaminase (TadA), or a variant thereof.

[0118] Embodiment (10): A base editing composition according to any one of embodiments (1) to (9), wherein the one or more fusion proteins each contain one or more DNA-binding proteins in the group consisting of zinc finger protein, transcription activator-like effector (TALE) protein and CRISPR-associated nuclease.

[0119] Embodiment (11): A base editing composition according to any one of embodiments (1) to (10), wherein the DNA is organellar DNA, and the one or more fusion proteins each contain a nuclear export signal (NES).

[0120] Embodiment (12): A base editing composition according to any one of embodiments (1) to (11), wherein the DNA is mitochondrial DNA, and the one or more fusion proteins each contain a mitochondrial targeting sequence (MTS).

[0121] Embodiment (13): A base editing composition according to any one of embodiments (1) to (11), wherein the DNA is chloroplast DNA, chromoplast DNA or leucoplast DNA, and the one or more fusion proteins each contain a chloroplast transit signal (CTS).

[0122] Embodiment (14): A base editing composition according to any one of embodiments (1) to (10), wherein the DNA is nuclear DNA and the one or more fusion proteins contain a nuclear localization signal (NLS).

[0123] Embodiment (15): A base editing composition according to any one of embodiments (1) to (14), wherein editing of cytosine (C) bases to thymine (T) bases is not substantially triggered.

[0124] Embodiment (16): A base editing composition according to any one of embodiments (1) to (14), wherein cytosine (C) bases are edited to thymine (T) bases at a frequency of less than 20%.

[0125] Embodiment (17): The base editing composition according to any one of (1) to (14), wherein cytosine (C) bases are edited to thymine (T) bases at a frequency of less than 10%.

[0126] Embodiment (18): A base editing composition according to any one of embodiments (1) to (14), wherein cytosine (C) bases are edited to thymine (T) bases at a frequency of less than 5%.

[0127] Embodiment (19): A base editing composition according to any one of embodiments (1) to (14), wherein cytosine (C) bases are edited to thymine (T) bases at a frequency of less than 1%.

[0128] Embodiment (20): A base editing composition having the activity of editing the adenine (A) base of the DNA of a plant cell into a guanine (G) base, wherein the base editing composition (the full-length cytosine deaminase preferably comprises the amino acid sequence of SEQ ID NO.9) comprises a fusion protein comprising a DNA binding protein, a cytosine deaminase (preferably a full-length cytosine deaminase) and an adenine deaminase.

[0129] Embodiment (21): A base editing composition according to embodiment (20), wherein the DNA is nuclear DNA or organelle DNA.

[0130] Embodiment (22): The base editing composition according to embodiment (20) or (21), wherein the cytosine deaminase is apolipoprotein Bediting complex (APOBEC), activation-induced cytidine deaminase (AID) or DddA tox , or variations thereof.

[0131] Embodiment (23): A base editing composition according to any one of embodiments (20) to (22), wherein the cytosine deaminase is an enzyme in which one or more amino acids selected from the group consisting of positions 37, 59, 109 and 129 in the amino acid sequence of SEQ ID NO.2 are replaced by different amino acids.

[0132] Embodiment (24): A base editing composition according to embodiment (23), wherein, in the amino acid sequence of SEQ ID NO.2 of cytosine deaminase, serine (S) at position 37 is replaced by glycine (G), and / or glycine (G) at position 59 is replaced by serine (S), and / or alanine (A) at position 109 is replaced by valine (V), and / or serine (S) at position 129 is replaced by glycine (G).

[0133] Embodiment (25): A base editing composition according to any one of embodiments (20) to (24), wherein the adenine deaminase is apolipoprotein B editing complex (APOBEC), activation-induced deaminase (AID) or tRNA-specific adenosine deaminase (TadA), or a variant thereof.

[0134] Embodiment (26): A base editing composition according to any one of embodiments (20) to (25), wherein the fusion protein is selected from one or more DNA binding proteins in the group consisting of zinc finger proteins, transcription activator-like effector proteins and CRISPR-associated nucleases.

[0135] Embodiment (27): A base editing composition according to any one of embodiments (20) to (26), wherein the DNA is organellar DNA and the fusion protein comprises a nuclear export signal.

[0136] Embodiment (28): A base editing composition according to any one of embodiments (20) to (27), wherein the DNA is mitochondrial DNA and the fusion protein comprises a mitochondrial targeting sequence (MTS).

[0137] Embodiment (29): A base editing composition according to any one of embodiments (20) to (27), wherein the DNA is chloroplast, chromatin or leucoplast DNA, and the fusion protein comprises a chloroplast transport signal (CTS).

[0138] Embodiment (30): A base editing composition according to any one of embodiments (20) to (26), wherein the DNA is nuclear DNA and the fusion protein comprises a nuclear localization signal (NLS).

[0139] Embodiment (31): A base editing composition according to any one of embodiments (20) to (30), wherein the fusion protein comprises uracil DNA glycosylase (UDG).

[0140] Embodiment (32): A base editing composition according to embodiment (31), wherein editing of cytosine (C) bases to thymine (T) bases is not substantially triggered.

[0141] Embodiment (33): The base editing composition according to embodiment (31), wherein cytosine (C) bases are edited to thymine (T) bases at a frequency of less than 20%.

[0142] Embodiment (34): The base editing composition according to embodiment (31), wherein cytosine (C) bases are edited to thymine (T) bases at a frequency of less than 10%.

[0143] Embodiment (35): The base editing composition according to embodiment (31), wherein cytosine (C) bases are edited to thymine (T) bases at a frequency of less than 5%.

[0144] Embodiment (36): The base editing composition of embodiment (31), wherein cytosine (C) bases are edited to thymine (T) bases at a frequency of less than 1%.

[0145] Embodiment (37): A polynucleotide or a combination of polynucleotides encoding any one of the more than one fusion proteins contained in the base editing composition according to any one of embodiments (1) to (36).

[0146] Embodiment (38): A vector comprising the polynucleotide or combination of polynucleotides according to (37).

[0147] Embodiment (39): A base editing composition comprising the vector according to (38).

[0148] Embodiment (40): A plant cell transformed with the vector according to (38) or comprising the polynucleotide or combination of polynucleotides according to (37).

[0149] Embodiment (41): A plant cultured from the plant cell according to embodiment (40).

[0150] Embodiment (42): A plant which is a descendant or clone of the plant according to embodiment (41).

[0151] Embodiment (43): A seed obtained from the plant according to embodiment (41) or (42).

[0152] Embodiment (44): A plant cell according to embodiment (40), a plant according to embodiment (41) or (42), or a seed according to embodiment (43), wherein the adenine (A) base of the wild-type DNA is edited to guanine (G).

[0153] Embodiment (45): The plant cell, plant or seed according to embodiment (44), wherein the cytosine (C) base of the wild-type DNA is not edited.

[0154] Embodiment (46): A method for editing the adenine (A) base of the organelle DNA of a plant to guanine (G), wherein the method comprises the following steps: expressing the base editing composition according to any one of embodiments (1) to (36) or the polynucleotide or combination of polynucleotides according to embodiment (37) in the plant or plant cell of interest.

[0155] Embodiment (47): A method according to embodiment (46), wherein editing of cytosine (C) bases to thymine (T) bases is not substantially triggered.

[0156] Embodiment (48): A method according to embodiment (46), wherein editing of cytosine (C) bases to thymine (T) bases occurs at a frequency of less than 20%.

[0157] Embodiment (49): A method according to embodiment (46), wherein editing of cytosine (C) bases to thymine (T) bases occurs at a frequency of less than 10%.

[0158] Embodiment (50): A method according to embodiment (46), wherein cytosine (C) bases are edited to thymine (T) bases at a frequency of less than 5%.

[0159] Hereinafter, the present invention will be described in detail by the following examples. However, the following examples are only for illustrating the present invention and the present invention is not limited to these examples.

[0160] Example 1: Arabidopsis psaA base editor transformation

[0161] DNA encoding the following base editor targeting the Arabidopsis thaliana chloroplast gene psaA was cloned and transformed with Agrobacterium to produce plants. The gene sequence used is as follows.

[0162] [Table 1]

[0163]

[0164] The sequences of the CTS, left TALE, right TALE, ABE8.0, GSVG and UDG used are as follows.

[0165] CTS (chloroplast transit signal):

[0166] MDSQLVLSLKLNPSFTPLSPLFPFTPCSSSFSPSLRFSSCYSRRLYSPVT VYAAK(SEQ ID NO.8)

[0167] ABE8.0 (adenine deaminase):

[0168] SEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN(SEQ ID NO.1)

[0169] 1397N(DddA tox Split body):

[0170] GSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGG PTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTET LLPENAKMTVVPPEG(SEQ ID NO.5)

[0171] 1397C(DddA tox Split body):

[0172] GSAIPVKRGATGETKVFTGNSNSPKSPTKGGC(SEQ ID NO.6)

[0173] GSVG(Full-length DddA tox variant):

[0174] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLEGKVFSSGGPT PYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLL PENAKMTVVPPEGVIPVKRGATGETKVFTGNSNGPKSPTKGGC(SEQID NO.9)

[0175] UDG(Uracil DNA glycosylase):

[0176] MASSTPKTLMDFFQPAKRLKASPSSSSFPAVSVAGGSRDLGSVANSPPRVTVTTSVADDSSGLTPEQIARAEFNKFVAKSKRNLAVCSERVTKAKSEGNCYVPLSELLVEESWLKALPGEFHKPYAKSLSDFLEREIITDSKSPLIYPPQHLIFNALNTTPFDRVKTVIIGQDPYHGPGQAMGLSFSVPEGEKLPSSLLNIFKELHKDVGCSIPRHGNLQKWAVQGVLLLNAVLTVRSKQPNSHAKKGWEQFTDAVIQSISQQKEGVVFLLWGRYAQEKSKLIDATKHHILTAAHPSGLSANRGFFDCRHFSRANQLLEEMGIPPIDWQL(SEQ ID NO.7)

[0177] Left TALE 1(containing N-terminal domain and C-terminal domain):

[0178] DLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLG(SEQ ID NO.10)

[0179] Right TALE (including N-terminal domain and C-terminal domain):

[0180] DLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLG(SEQ ID NO.11)

[0181] The sequences of the linkers used are as follows.

[0182] Linker 1:

[0183] GS

[0184] Linker 2:

[0185] SGSETPGTSESATPES(SEQ ID NO.12)

[0186] Connector 3:

[0187] LVGS

[0188] Specifically, in order to use the gene construct to edit the psaA gene present in the chloroplast genome of Arabidopsis, it was designed as follows: first, it was located between the RPS5A promoter and the 35S terminator, whose expression can be induced during embryogenesis, so that editing can be triggered from the early developmental stage. A vector suitable for transforming Agrobacterium tumefaciens (the Agrobacterium tumefaciens can use T-DNA to deliver gene constructs to the nuclear genome of Arabidopsis) was cloned using Gibson Assembly, Golden Gate cloning, and restriction endonucleases. Subsequently, the Agrobacterium strain GV3101 transformed with the vector was used to transform Arabidopsis Columbia ecotype-0 (Col-0) plants by the floral dip method according to known methods (Zhang et al., Nat. Protoc. 1, 641-646 (2006)).

[0189] Example 2: Confirming base editing and determining editing efficiency by sequencing

[0190] DNA was extracted from untransformed wild-type Col-0 individuals and individuals cultured from first-generation transformed Arabidopsis seeds, and then analyzed using targeted deep sequencing. Base editing efficiency (frequency) was calculated based on the percentage of sequencing reads reflecting the desired base editing result in the total sequencing reads ( Figures 3 to 14 ).

[0191] The experimental results confirmed that the base editors used in the present invention can edit efficiently, such as Figure 1 and Figure 2 The adenine base in the target site of the psaA gene shown is particularly effective when using a detoxified full-length form of DddA linked to the psaA gene. tox In the case of a monomeric TALED of cytosine deaminase (GSVG variant), it was also confirmed for the first time that adenine bases in chloroplast DNA of plant cells were efficiently edited.

[0192] In addition, to determine the extent to which unwanted C to T editing can be avoided when using UDG-linked TALEDs, the average frequency of C2 base editing in the editing target site of the psaA gene is summarized in the table below for each base editor used.

[0193] [Table 2]

[0194] base editors C2 base editing efficiency of the psaA gene (average) Control group (Col-0) 0.17% L-1397N+R-1397C-AD 46.14% L-1397N+R-1397C-AD-UDG 11.72% L-1397N-UDG+R-1397C-AD 4.53% L-1397N-UDG+R-1397C-AD-UDG 1.54% L-1397C-AD+R-1397N 26.38% L-1397C-AD-UDG+R-1397N 3.21% L-1397C-AD+R-1397N-UDG 1.47% L-1397C-AD-UDG+R-1397N-UDG 0.7% R-AD-GSVG 23.96% R-AD-GSVG-UDG 1.37%

[0195] As shown in the table above, it was confirmed that in all base editors used, when UDG was further linked, C to T editing was significantly reduced, and no C to T editing was confirmed at a level almost equivalent to that of the wild-type Col-0 (control group) in which base editing did not occur.

[0196] The experimental results also show that when the UDG-linked TALED according to the present invention is used, it can ensure that only the adenine base is edited to guanine in the plant individual (e.g., Figure 8 #4, #10, and #14 individuals).

[0197] Example 3: Arabidopsis psbA base editor transformation

[0198] The fusion proteins shown in Table 3 were created by base-editing the 5'-AGT-3' base sequence encoding serine at position 264 in the psbA gene, which is present in the chloroplasts of Arabidopsis thaliana, to 5'-GGT-3', which encodes glycine. Editing serine at position 264 to glycine conferred resistance to the herbicide atrazine. To express the fusion protein, the following DNA was cloned and transformed with Agrobacterium to create transformed plants.

[0199] [Table 3]

[0200]

[0201]

[0202] The sequences of the fusion protein components used are as follows.

[0203] CTS (chloroplast transit signal)

[0204] MDSQLVLSLKLNPSFTPLSPLFPFTPCSSSFSPSLRFSSCYSRRLYSPVTVYAAK(SEQ ID NO.8)

[0205] 1397N(DddA tox Split body)

[0206] GSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEG(SEQ ID NO.5)

[0207] 1397C(DddA tox Split body)

[0208] GSAIPVKRGATGETKVFTGNSNSPKSPTKGGC(SEQ ID NO.6)

[0209] GSVG(Full-length DddA tox variant)

[0210] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLEGKVFSSGGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGVIPVKRGATGETKVFTGNSNGPKSPTKGGC(SEQ IDNO.9)

[0211] AD(TadA8e)

[0212] SEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN(SEQ ID NO.1)

[0213] UDG(Uracil DNA glycosylase)

[0214] MASSTPKTLMDFFQPAKRLKASPSSSSFPAVSVAGGSRDLGSVANSPPRVTVTTSVADDSSGLTPEQIARAEFNKFVAKSKRNLAVCSERVTKAKSEGNCYVPLSELLVEESWLKALPGEFHKPYAKSLSDFLEREIITDSKSPLIYPPQHLIFNALNTTPFDRVKTVIIGQDPYHGPGQAMGLSFSVPEGEKLPSSLLNIFKELHKDVGCSIPRHGNLQKWAVQGVLLLNAVLTVRSKQPNSHAKKGWEQFTDAVIQSISQQKEGVVFLLWGRYAQEKSKLIDATKHHILTAAHPSGLSANRGFFDCRHFSRANQLLEEMGIPPIDWQL(SEQ ID NO.7)

[0215] Adapter 11

[0216] GSSGSETPGTSESATPES(SEQ ID NO.13)

[0217] Adapter 12

[0218] LVGS

[0219] Adapter 13

[0220] SGSETPGTSESATPES(SEQ ID NO.18)

[0221] Adapter 14

[0222] GS

[0223] psbA left 1 TALE protein

[0224] DLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLG (SEQ ID NO.14; containing N-terminal domain and C-terminal domain)

[0225] 2 TALE protein on the left side of psbA

[0226] DLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLG (SEQ ID NO.15; containing N-terminal domain and C-terminal domain)

[0227] 1 TALE protein to the right of psbA

[0228] DLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLG (SEQ ID NO.16; containing N-terminal domain and C-terminal domain)

[0229] 2 TALE protein to the right of psbA

[0230] DLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLG(Including SEQ ID NO.17; including N-terminal domain and C-terminal domain)

[0231] Specifically, in order to use the gene construct to edit the base sequence of 5'-AGT-3' encoding serine 264 in the gene psbA present in the chloroplast genome of Arabidopsis thaliana to 5'-GGT-3' encoding glycine, the gene construct was designed as follows: first, it was located between the RPS5A promoter and the 35S terminator, which can be induced to express during embryogenesis, so that editing can be triggered from the early developmental stage; Gibson assembly, Golden Gate cloning, restriction enzymes, etc. were used to clone a vector suitable for transforming Agrobacterium tumefaciens (the Agrobacterium tumefaciens can use T-DNA to deliver gene constructs to the nuclear genome of Arabidopsis thaliana); then, using the Agrobacterium strain GV3101 transformed with the vector, floral infusion was performed according to a known method (Zhang et al., Nat. Protoc. 1, 641-646 (2006)). Arabidopsis thaliana Columbia ecotype-0 (Col-0) plants were transformed using the dip method.

[0232] About 15,000 to 20,000 seeds of Arabidopsis thaliana introduced with the 12 fusion proteins were sown in the soil by the floral-dip method. After 7 and 14 days, the seeds were treated with atrazine (40 g / hectare) to ensure the survival of the first generation of transformed plants. The gene editing efficiency was analyzed. The results confirmed that when the fusion protein base editor containing UDG according to the present invention was used, A to G base editing was achieved ( Figure 16 ).

Claims

1. A base editing composition having the activity of editing adenine bases of DNA into guanine bases, characterized in that The base editing composition comprises one or more fusion proteins, each fusion protein comprises a DNA binding protein and a cytosine deaminase, and at least one of the one or more fusion proteins comprises adenine deaminase, and at least one of the one or more fusion proteins comprises uracil DNA glycosylase, and the cytosine deaminase exists in full-length or two split forms.

2. The base editing composition according to claim 1, characterized in that The DNA is nuclear DNA or organellar DNA.

3. The base editing composition according to claim 1 or 2, characterized in that The cytosine deaminase is apolipoprotein B mRNA editing complex, activation-induced cytidine deaminase, tRNA-specific adenosine deaminase, or DddA tox , or variations thereof.

4. The base editing composition according to any one of claims 1 to 3, characterized in that The cytosine deaminase comprises the amino acid sequence of SEQ ID NO. 2 in full-length form, and one or more amino acids selected from the group consisting of positions 37, 59, 109 and 129 in the amino acid sequence of SEQ ID NO. 2 are substituted with different amino acids.

5. The base editing composition according to claim 4, characterized in that The cytosine deaminase comprises the amino acid sequence of SEQ ID NO. 2 and has one or more amino acid substitutions selected from the group consisting of substitution of serine at position 37 with glycine, substitution of glycine at position 59 with serine, substitution of alanine at position 109 with valine, and substitution of serine at position 129 with glycine.

6. The base editing composition according to any one of claims 1 to 3, characterized in that The cytosine deaminase is contained in the form of a first segment and a second segment, and one fusion protein contains the first segment, and another fusion protein contains the second segment, and one or more amino acids located on the dimerization surface of the first segment and the second segment are substituted with different amino acids.

7. The base editing composition according to claim 6, characterized in that The first segmented form of cytosine deaminase comprises a sequence from the N-terminus to glycine at position 33, glycine at position 44, alanine at position 54, asparagine at position 68, glycine at position 82, asparagine at position 98, or glycine at position 108 in the amino acid sequence of SEQ ID NO. 2, and the second segmented form of cytosine deaminase comprises a sequence from glycine at position 34, proline at position 45, glycine at position 55, asparagine at position 69, threonine at position 83, alanine at position 99, or alanine at position 109 to the C-terminus in the amino acid sequence of SEQ ID NO.

2.

8. The base editing composition according to claim 6, characterized in that The first segmented form of cytosine deaminase comprises the amino acid sequence of SEQ ID NO.5, and the second segmented form of cytosine deaminase comprises the amino acid sequence of SEQ ID NO.6, and one or more amino acids selected from the group consisting of positions 87, 88, 91, 92, 95, 100, 101, 102 and 103 of SEQ ID NO.5 or one or more amino acids selected from the group consisting of positions 13, 14, 15 and 16 of SEQ ID NO.6 are substituted by different amino acids.

9. The base editing composition according to any one of claims 1 to 8, characterized in that The adenine deaminase is an apolipoprotein B mRNA editing complex, an activation-induced cytidine deaminase, or a tRNA-specific adenosine deaminase, or a variant thereof.

10. The base editing composition according to any one of claims 1 to 9, characterized in that The one or more fusion proteins each comprise one or more DNA-binding proteins selected from the group consisting of zinc finger proteins, transcription activator-like effector proteins, and CRISPR-associated nucleases.

11. The base editing composition according to any one of claims 1 to 10, characterized in that The DNA is organellar DNA, and the one or more fusion proteins each comprise a nuclear export signal.

12. The base editing composition according to any one of claims 1 to 11, characterized in that The DNA is mitochondrial DNA, and the one or more fusion proteins each contain a mitochondrial target sequence.

13. The base editing composition according to any one of claims 1 to 11, characterized in that The DNA is chloroplast DNA, chromoplast DNA or leucoplast DNA, and the one or more fusion proteins each contain a chloroplast transport signal.

14. The base editing composition according to any one of claims 1 to 10, characterized in that The DNA is nuclear DNA, and the one or more fusion proteins contain a nuclear localization signal.

15. The base editing composition according to any one of claims 1 to 14, characterized in that The editing of cytosine bases to thymine bases is essentially not triggered.

16. The base editing composition according to any one of claims 1 to 14, characterized in that Cytosine bases are edited to thymine bases at a frequency of less than 20%, less than 10%, less than 5%, or less than 1%.

17. A base editing composition having the activity of editing adenine bases in the DNA of plant cells into guanine bases, characterized in that: The base editing composition includes a fusion protein comprising a DNA binding protein, a cytosine deaminase, and an adenine deaminase.

18. The base editing composition according to claim 17, characterized in that The DNA is nuclear DNA or organellar DNA.

19. The base editing composition according to claim 17 or 18, characterized in that The cytosine deaminase is apolipoprotein B mRNA editing complex, activation-induced cytidine deaminase, tRNA-specific adenosine deaminase, or DddA tox , or variations thereof.

20. The base editing composition according to any one of claims 17 to 19, wherein The cytosine deaminase is an enzyme in which one or more amino acids selected from the group consisting of positions 37, 59, 109, and 129 in the amino acid sequence of SEQ ID NO. 2 are substituted with different amino acids.

21. The base editing composition according to claim 20, characterized in that In the amino acid sequence of SEQ ID NO. 2 of cytosine deaminase, serine at position 37 is substituted by glycine, and / or glycine at position 59 is substituted by serine, and / or alanine at position 109 is substituted by valine, and / or serine at position 129 is substituted by glycine.

22. The base editing composition according to any one of claims 17 to 21, wherein The adenine deaminase is an apolipoprotein B mRNA editing complex, an activation-induced cytidine deaminase, or a tRNA-specific adenosine deaminase, or a variant thereof.

23. The base editing composition according to any one of claims 17 to 22, wherein The fusion protein comprises one or more DNA binding proteins selected from the group consisting of zinc finger proteins, transcription activator-like effector proteins and CRISPR-associated nucleases.

24. The base editing composition according to any one of claims 17 to 23, wherein The DNA is organellar DNA, and the fusion protein contains a nuclear export signal.

25. The base editing composition according to any one of claims 17 to 24, wherein The DNA is mitochondrial DNA, and the fusion protein comprises a mitochondrial target sequence.

26. The base editing composition according to any one of claims 17 to 24, wherein The DNA is chloroplast DNA, variegated DNA, or leucoplast DNA, and the fusion protein comprises a chloroplast transport signal.

27. The base editing composition according to any one of claims 17 to 23, wherein The DNA is nuclear DNA, and the fusion protein contains a nuclear localization signal.

28. The base editing composition according to any one of claims 17 to 27, wherein The fusion protein comprises uracil DNA glycosylase.

29. The base editing composition according to claim 28, characterized in that The editing of cytosine bases to thymine bases is essentially not triggered.

30. The base editing composition according to claim 28, wherein Cytosine bases are edited to thymine bases at a frequency of less than 20%, less than 10%, less than 5%, or less than 1%.

31. A polynucleotide or a combination of two or more polynucleotides, characterized in that: The polynucleotide or a combination of two or more polynucleotides encodes any one of the one or more fusion proteins contained in the base editing composition according to any one of claims 1 to 30.

32. A carrier, characterized in that , Comprising the polynucleotide or combination of polynucleotides according to claim 31.

33. A base editing composition, characterized in that Comprising the vector according to claim 32.

34. A plant cell, characterized in that Transformed by the vector according to claim 32 or comprising the polynucleotide or combination of polynucleotides according to claim 31 .

35. A plant, characterized in that Culturing the plant cell according to claim 34.

36. A plant, characterized in that The plant is a descendant or clone of the plant according to claim 35.

37. A seed, characterized in that Obtained from the plant according to claim 35 or 36.

38. A plant cell, plant or seed, characterized in that The plant cell is the plant cell according to claim 34, the plant is the plant according to claim 35 or 36, or the seed is the seed according to claim 37, and the adenine base of the wild-type DNA of the plant cell, plant or seed is edited to guanine.

39. The plant cell, plant or seed according to claim 38, characterized in that The cytosine bases of wild-type DNA are not edited.

40. A method for editing adenine bases in plant organelle DNA into guanine, characterized in that: The following steps are involved: Expressing the base editing composition of any one of claims 1 to 30 or the polynucleotide or combination of polynucleotides according to claim 31 in a plant or plant cell of interest.

41. The method according to claim 40, wherein The editing of cytosine bases to thymine bases is essentially not triggered.

42. The method according to claim 40, wherein Cytosine bases are edited to thymine bases at a frequency of less than 20%, less than 10%, less than 5%, or less than 1%.

Citation Information

Patent Citations

  • Targeted deaminase and base editing using same

    WO2022060185A1

  • Compositions and methods for the treatment of hereditary angioedema (HAE)

    WO2023086953A1