Nicking enzyme mediated plant DNA base editing

By expressing a base editing composition containing nicking enzymes and deaminases in plant organelles, the problem of editing plant organelle DNA using existing tools has been solved, enabling highly efficient editing of adenine to guanine and cytosine to thymine, improving editing efficiency and avoiding cytotoxicity.

CN121399264APending Publication Date: 2026-01-23GREENGENE INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202480042853.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-06-28
Filing Date
2024-06-27
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

Existing genome editing tools are difficult to effectively edit the DNA of plant organelles such as mitochondria and chloroplasts, mainly because they cannot deliver guide RNA from the CRISPR system into the organelles or express multiple compounds simultaneously. Furthermore, DddAtox is cytotoxic and has low efficiency when used in segments.

Method used

By linking a nicking enzyme to a DNA-binding protein, and combining it with adenine deaminase and/or cytosine deaminase, targeted editing of plant organelle DNA can be achieved, avoiding the use of DddAtox. The nicking enzyme removes adenine and cytosine amino groups from single-stranded DNA, thus achieving base editing.

Benefits of technology

This technology enables highly efficient targeted editing of adenine to guanine and cytosine to thymine in plant organelle DNA, improving base editing efficiency, especially for cytosine editing in the 5'-GC-3' sequence, while avoiding the cytotoxicity issues of DddAtox.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121399264A_ABST
    Figure CN121399264A_ABST
Patent Text Reader

Abstract

The present invention relates to a base editing method for plant DNA, the method comprising a step of expressing a base editing composition in a target plant, plant cell or protoplast, the base editing composition comprising one or more DNA binding proteins and one or more zymoproteins, or comprising a polynucleotide encoding said proteins. The invention is used for editing bases in nuclear DNA or organelle DNA of plants, especially for editing bases in organelle DNA of plants such as chloroplast or mitochondria. The base editing composition can be used for realizing targeted editing from adenine (A) base to guanine (G), targeted editing from cytosine (C) base to thymine (T) base or a method for realizing the two editing methods at the same time in plant DNA (Deoxyribose Nucleic Acid). In addition, the present invention relates to a plant cell, plant or seed implementing such base editing.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present invention relates to a method of base editing of plant DNA, comprising a step of expressing a base editing composition comprising one or more DNA binding proteins and one or more enzyme proteins, or polynucleotides encoding the proteins, in a target plant, plant cell or protoplast. The present invention is useful for editing bases in the nuclear DNA or organelle DNA of a plant, and in particular, for editing bases in the organelle DNA of a plant such as chloroplast or mitochondria. Such a base editing composition can be used in a method for achieving targeted editing of adenine (A) bases to guanine (G) bases, targeted editing of cytosine (C) bases to thymine (T) bases, or both in plant DNA. In addition, the present invention relates to a plant cell, plant or seed that achieves such base editing. BACKGROUND

[0002] Fusion proteins of DNA binding proteins and deaminases are capable of inducing DNA variation (e.g., single nucleotide conversion) in a targeted manner to substitute nucleotides or edit bases in the genome without generating DNA double-strand breaks (DSBs), or to edit point mutations that cause genetic disorders, or to introduce desired single nucleotide variations in prokaryotic and eukaryotic cells (e.g., humans).

[0003] Programmable genome editing tools such as zinc-finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), clustered regularly interspaced short palindromic repeats (CRISPR) systems, and base editors composed of a CRISPR-associated protein 9 (Cas9) variant lacking nucleic acid degradation efficiency and a base deaminase protein, have the potential to be used for plant genetic research and crop trait improvement through base sequence alteration. However, these existing genome editing tools are not suitable for editing DNA sequences of plant organelles including mitochondria and chloroplasts, and the main reason is that it is not possible to deliver guide RNA required for the most widely used CRISPR system into the organelle, or it is difficult to express two compounds simultaneously in the organelle.

[0004] Furthermore, mitochondria and chloroplasts in plants are organelles with DNA of different morphology than that of animal nuclei or mitochondria. For example, human mitochondrial DNA has 16,569 base pairs, while Arabidopsis thaliana mitochondrial DNA has 366,924 base pairs. On the other hand, Arabidopsis thaliana chloroplast DNA has 154,478 base pairs. As mentioned above, the length and composition of plant chloroplast and mitochondrial DNA differ from those of animal nuclei or mitochondria, and the types of genes they contain are also completely different. In particular, plant chloroplasts contain proteins related to DNA replication and repair, which are different from those found in animal nuclei or mitochondria. Therefore, the DNA repair patterns exhibited when changes occur in plant chloroplast and mitochondrial DNA differ from those in animal nuclei or mitochondria. Thus, methods effective in editing animal nuclear or mitochondrial DNA may not be effective in editing plant organelle DNA. Plant organelles encode many essential genes required for photosynthesis and respiration. Methods or tools for editing the genes of these organelles are crucial for studying the function of these genes or improving crop productivity and traits.

[0005] Prior to this invention, base editing of organelle DNA could be performed using a strain derived from Burkholderia neonsis (…). Burkholderia cenocepacia DddA, a bacterial toxin tox To achieve this. DddA tox It originates from Burkholderia neonsis ( Burkholderia cenocepacia This part, possessing the enzymatic function of bacterial toxins, can deaminate cytosine in double-stranded DNA. Due to DddA... tox It is cytotoxic; to avoid toxicity to host cells, DddA... tox It is divided into two inactivating splits for use, and each split can be linked to a DNA-binding protein designed to bind to DNA, thus serving as a pairwise cytosine base editor (DddA-derived cytosine base editor, DdCBE). On the other hand, if DddA... tox By linking cytosine deaminase and adenine deaminase capable of causing adenine-to-guanine (A-to-G) editing to DNA-binding proteins, editing of adenine bases can be achieved (see International Patent Application Publication WO2022 / 060185A).

[0006] Existing technical documents Patent documents Patent Document 1: International Patent Publication WO2022 / 060185A1 Summary of the Invention The problem the invention aims to solve This invention attempts to achieve this without using DddA tox Editing the bases of plant organelle DNA is a technique used in this context. Specifically, base editing of plant organelle DNA is attempted by linking nickases and deaminases to DNA-binding proteins. Deaminases induce base editing by removing the amino groups of cytosine and adenine from single-stranded DNA (not double-stranded DNA). When organelle DNA, which is a double helix, is induced into single-stranded DNA by nickases, the deaminases remove the amino groups of cytosine and adenine, thereby inducing base editing.

[0007] means for solving problems This invention relates to a DNA base editing composition comprising a nicking enzyme and a method for base editing plant DNA. The method includes the step of expressing the base editing composition in a target plant, plant cell, or protoplast. The base editing composition comprises one or more DNA-binding proteins and one or more enzyme proteins, or comprises polynucleotides encoding the proteins. The one or more enzyme proteins comprise a nicking enzyme and further comprise adenine deaminase and / or cytosine deaminase.

[0008] In another aspect of the invention, the base editing composition comprises one or more fusion proteins or one or more polynucleotides encoding one or more fusion proteins, each of the one or more fusion proteins independently comprising a DNA-binding protein, at least one of the one or more fusion proteins comprising a nicking enzyme, and at least one of the one or more fusion proteins comprising adenine deaminase and / or cytosine deaminase.

[0009] The cytosine deaminases that can be used in this invention preferably have deaminase activity for single-stranded DNA.

[0010] The DNA base editing composition of the present invention can be used to provide a method for achieving targeted editing of adenine (A) bases to guanine (G) bases, targeted editing of cytosine (C) bases to thymine (T) bases, or both, in plant DNA, preferably plant organelle DNA, and also provides plant cells, plants, or seeds for achieving such base editing.

[0011] Invention Effects This invention provides a method for base editing of plant DNA, comprising the step of expressing a DNA base editing composition in a target plant, plant cell, or protoplast. The plant cell, protoplast, plant, and seed according to this base editing method can edit cytosine bases to thymine or adenine bases to guanine in a wild-type target DNA sequence, or both types of editing. This invention can be achieved by using a nicking enzyme and by using DddA. toxThe existing base editor achieves the same effect as targeted C-to-T and / or A-to-G editing, without using DddA. tox In particular, the present invention provides a base editing method that can perform both C-to-T and A-to-G editing, even when using a nicking enzyme and a deaminase. Furthermore, unlike existing cytosine base editors, the base editing method of the present invention can edit not only cytosine bases present in the 5'-TC-3' sequence of the target DNA sequence, but also cytosine bases present in the 5'-AC-3', 5'-CC-3', or 5'-GC-3' sequences to thymine. In particular, it can edit cytosine bases present in the 5'-GC-3' sequence to thymine, thereby improving base editing efficiency. Attached Figure Description

[0012] Figure 1 This is a schematic diagram of targeting the plant (lettuce) chloroplast gene psaA using a base editor containing a nicking enzyme (MutH) and a deaminase. CTS represents the chloroplast transit signal, NTD represents the N-terminal domain of the TALE protein, and CTD represents the C-terminal domain of the TALE protein. Single-underlined base sequences represent the base sequences that bind to the TALE protein in the psaA gene. The left side shows the base sequences that bind to the protein at "LspsaA site 1 Left TALE," and the right side shows the base sequences that bind to the proteins at "LspsaA site 1 Right 1 TALE" and "LspsaA site Right 2 TALE." Double-underlined base sequences (5'-GATC-3') represent the base sequences that MutH recognizes in the psaA gene. The base positions located between the two TALE proteins are indicated by subscript numbers, and the first base after the DNA site where the left TALE binds is considered number 1 and the counting begins thereafter. Unless otherwise stated, this base numbering is the same in subsequent figures. UGI stands for uracil glycosylase inhibitor.

[0013] Figure 2Part a shows the C-to-T base editing efficiency in the plant (lettuce) chloroplast gene psaA obtained using nickase (MutH) and cytosine deaminases (TadA-CDd, rApobec, TadA8eE27R N46L). "L" refers to the "LspsaA site 1 Left TALE", and "R1" and "R2" refer to the "LspsaA site 1 Right 1 TALE" and "LspsaA site 2 Right 2 TALE" proteins, respectively. "2aa lin" refers to the linker (using 2 amino acids) between the TALE protein and the deaminase. Figure 2 In part b, the bases actually edited by base editing are indicated by square brackets ([ ]), and the frequencies of wild-type alleles and edited alleles obtained from the editing results are expressed as percentages.

[0014] Figure 3 Part a shows the C-to-T base editing efficiency in the plant (lettuce) chloroplast gene psaA obtained using nickase (MutH) and cytosine deaminases (TadA-CDd, rApobec, TadA8eE27R N46L), using an 18-amino acid linker. "L" refers to "LspsaA site 1 Left TALE", and "R1" and "R2" refer to the proteins "LspsaA site 1 Right 1 TALE" and "LspsaA site 2 Right TALE", respectively. "18aa lin" refers to the linker between the TALE protein and the deaminase (using 18 amino acids). Figure 3 Part b shows the frequencies (in percentage) of wild-type alleles and edited alleles obtained from the editing results.

[0015] Figure 4 Part a shows the A-to-G base editing efficiency in the plant (lettuce) chloroplast gene psaA obtained using the nicking enzyme (MutH) and adenine deaminase (TadA8e). "L" refers to the "LspsaA site 1 Left TALE" protein, and "R1" refers to the "LspsaA site 1 Right 1 TALE" protein. Figure 4Part b shows the frequencies (in percentage) of wild-type alleles and edited alleles obtained from the editing results.

[0016] Figure 5 This relates to the base editor based on the present invention using nicking enzymes and cytosine deaminases (TadA-CDd, rApobec) versus the existing known DddA... tox A comparison of the C-to-T base editing efficiency in the psaA gene of plant (lettuce) chloroplasts using a base editor. (This is a comparison using DddA.) tox The base editor uses DddA tox Segmentation (1397N and 1397C) or using full-length DddA tox (GSVG) DdCBE. In all experimental groups, the TALE protein used was the "LspsaA site 1 Left TALE" protein (denoted as "L") and the "LspsaA site 1 Right 1 TALE" protein (denoted as "R1"). C4 is the C base located in the 5'-GC-3' sequence, and G5 and C8 are the C bases located in the 5'-TC-3' sequence. Figure 5 The results show that when using the base editor of the present invention, compared with using the existing DddA... tox Unlike other base editors, the C base located in the 5'-GC-3' sequence motif can be edited to G.

[0017] Figure 6 The present invention utilizes a base editor employing nickase (Nt.BspD6I(C)) and adenine deaminase (TadA8e) in conjunction with the existing known DddA... tox A comparison of the A-to-G base editing efficiency of the psaA gene in plant (lettuce) chloroplasts using a base editor. The target gene for editing the psaA gene is... Figures 1 to 5 Different psaA gene sites were used in the experiments, therefore the TALE proteins used were also different. In all experimental groups, the TALE proteins used were "LspsaA site 2 Left TALE" (represented as "L" in this figure) and "LspsaA site 2 Right TALE" (represented as "R" in this figure).

[0018] Figure 7This is a schematic diagram of targeting the plant (lettuce) chloroplast gene psbA using Nt.BspD6I(C) as a nicking enzyme and a base editor containing a deaminase. The single-underlined base sequences represent the base sequences bound to the TALE protein in the psbA gene. The left side shows the base sequences bound to the proteins “LspsbA Left 1 TALE” and “LspsbA Left 2”, while the right side shows the base sequences bound to the proteins “LspsbA Right 1 TALE”, “LspsbA Right 2 TALE”, and “LspsbA Right 3 TALE”.

[0019] Figure 8 Part a shows the A-to-G base editing efficiency in the plant (lettuce) chloroplast gene psbA obtained using Nt.BspD6I(C) as a nicking enzyme and adenine deaminase (TadA8e). “L1” and “L2” refer to the proteins “LspsbA Left 1 TALE” and “LspsbA Left 2 TALE”, respectively. “R1”, “R2”, and “R3” refer to the proteins “LspsbA Right 1 TALE”, “LspsbA Right 2 TALE”, and “LspsbA Right 3 TALE”, respectively. “2aa lin” refers to the linker (using 2 amino acids) between the TALE protein and the deaminase. Figure 8 Part b shows the frequencies (in percentage) of wild-type alleles and edited alleles obtained from the editing results.

[0020] Figure 9The diagram shows the A-to-G base editing efficiency in the plant (lettuce) chloroplast gene psbA using Nt.BspD6I(C) as a nicking enzyme and adenine deaminase (TadA8e), with results obtained using an 18-amino acid linker. "L1" and "L2" refer to the proteins "LspsbA Left 1 TALE" and "LspsbA Left 2 TALE" respectively, while "R1", "R2", and "R3" refer to the proteins "LspsbA Right 1 TALE", "LspsbA Right 2 TALE", and "LspsbA Right 3 TALE" respectively. "18aa lin" refers to the linker between the TALE protein and the deaminase (using 18 amino acids).

[0021] Figure 10 Part a shows the C-to-T base editing efficiency in the plant (lettuce) chloroplast gene psbA obtained using Nt.BspD6I(C) as a nicking enzyme and cytosine deaminase (TadA-CDd). "L1" refers to the "LspsbA Left 1 TALE" protein, and "R3" refers to the "LspsbA Right 3 TALE" protein. "2aa lin" refers to the linker (using 2 amino acids) between the TALE protein and the deaminase. Figure 10 Part b shows the frequencies (in percentage) of wild-type alleles and edited alleles obtained from the editing results.

[0022] Figure 11 Part a shows the results of simultaneous C-to-T and A-to-G base editing in the plant (lettuce) chloroplast gene psbA when used in conjunction with the nicking enzyme (Nt-BspD6I(C)) and cytosine deaminase (TadA-CDd) and adenine deaminase (TadA8e). "L1" refers to the "LspsbA Left 1 TALE" protein, and "R3" refers to the "LspsbA Right 3 TALE" protein. "2aa lin" refers to the linker (using 2 amino acids) between the TALE protein and the deaminase. The two base editors used in the experiment differed in the linking order of the nicking enzyme and adenine deaminase. Figure 11 Part b shows the frequencies (in percentage) of wild-type alleles and edited alleles obtained from the editing results.

[0023] Figure 12 Part a shows the results of simultaneous C-to-T and A-to-G base editing in the plant (lettuce) chloroplast gene psaA when MutH is used as the nicking enzyme and TadA-dual is used as the deaminase. "L" refers to the "LspsaA Left site 1 TALE" protein, and "R1" refers to the "LspsaA site 1 Right 1 TALE" protein. "2aa lin" refers to the linker (using 2 amino acids) between the TALE protein and the deaminase. Figure 12 Part b shows the frequencies (in percentage) of wild-type alleles and edited alleles obtained from the editing results.

[0024] Figure 13 This diagram illustrates the C-to-T base editing efficiency in the Arabidopsis chloroplast gene psaA obtained using MutH as a nicking enzyme and cytosine deaminases (TadA-CDd or rApobec) as deaminases. It also shows the C-to-T base editing efficiency values, average values, and error ranges at the C8 site of the psaA gene obtained from a single T1 plant derived from transformed Arabidopsis. The average base editing efficiency was approximately 74.2% when using TadA-CDd and approximately 77.8% when using rApobec. The base sequences shown represent the bases located between two TALE proteins, and "AtpsaA Left TALE (AtpsaALeft TALE)" and "AtpsaA Right TALE (AtpsaA Right TALE)" were used as TALE proteins. "Col-0" is an untransformed wild-type individual.

[0025] Figure 14 This shows the base editing efficiency obtained from a single T1 individual ( Figure 13 The average values ​​shown are based on the results obtained from individuals using TadA-CDd as a deaminase.

[0026] Figure 15 This shows the base editing efficiency obtained from a single T1 individual ( Figure 13 The average values ​​shown are based on the results obtained from individuals using rAPOBEC as a deaminase.

[0027] Figure 16This paper illustrates the C-to-T base editing efficiency in the Arabidopsis chloroplast gene psaA obtained using a heterodimer of two FokI variants (Quasi FokI) as a nicking enzyme and a cytosine deaminase (TadA-CDd) as a deaminase. It also shows the C-to-T base editing efficiency values, average values, and error ranges for G5 and C8 of the psaA gene obtained from a single T1 plant derived from transformed Arabidopsis. C-to-T editing in G5 refers to editing the C (cytosine) bases present on the opposite DNA strand of G5. The average base editing efficiency in G5 is approximately 56.1%, and in C8, it is approximately 26.3%. The base sequences shown represent bases located between two TALE proteins, and "AtpsaA Left TALE" and "AtpsaA Right TALE" are used as TALE proteins. “Col-0” is an untransformed wild-type individual. “40aa lin” refers to a linker consisting of 40 amino acids connecting a FokI variant with amino acid variations of E490K and I538K and a FokI variant with amino acid variations of D450A, Q486E, and I499L.

[0028] Figure 17 This shows the base editing efficiency obtained from a single T1 individual ( Figure 16 The basis for the average value shown).

[0029] Figure 18 The efficiency of A-to-G base editing in the Arabidopsis chloroplast gene psbA obtained using Nt-BspD6I(C) as a nicking enzyme and TadA8e as a deaminase is shown, and the results obtained from a single T1 individual obtained from transformed Arabidopsis are presented.

[0030] Figure 19 The efficiency of A-to-G base editing in the Arabidopsis chloroplast gene psbA, obtained using Nt-BspD6I(C) as a nicking enzyme and TadA8e as a deaminase, is shown, along with results obtained from a single T2 individual from transformed Arabidopsis. "AtpsbA Left TALE" and "AtpsbA Right TALE" were used as TALE proteins. The psbA base editor used produced the following result through A-to-G base editing: changing the serine residue at position 264 (encoded by 5'-AGT-3') of the D1 protein encoded by the psbA gene to a glycine residue (encoded by 5'-GGT-3'), which confers resistance to the atrazine herbicide. Figure 19The eight individuals shown are all individuals that survived even after treatment with atrazine.

[0031] Figure 20 It is from Figure 19 The T3 individuals derived from seeds of individuals #2-1, #2-3, #2-4, and the wild-type individuals are shown to exhibit growth when treated with culture medium (1 / 2 MS) alone, or further treated with Basta (PPT: glufosinate) or atrazine herbicides. The growth of T3 individuals derived from #2-1, #2-3, and #2-4 was not inhibited even with atrazine treatment. The exogenous gene conferring Basta resistance was introduced into #2-1, which survived Basta treatment.

[0032] Figure 21 yes Figure 20 The sequencing analysis results of T3 individuals with atrazine resistance shown indicate that they all exhibited more than 99% homogeneity.

[0033] Figure 22 Yes Figure 21 The results of PCR analysis on T3 individuals with atrazine resistance using primers for the Basta resistance gene and the psbA gene were analyzed. Figure 20 The atrazine-resistant #2-1 derivative #2-1-1 showed a foreign gene, while the Basta-resistant #2-3 and #2-4 derivatives #2-3-1 and #2-4-1 showed no foreign gene introduction. Detailed Implementation

[0034] Unless otherwise defined, all technical and scientific terms used in this specification have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. Generally, the terms used in this specification are those well-known and commonly used in the art.

[0035] The terms “correction,” “modification,” and “editing” used in this specification are used interchangeably and refer to methods for altering nucleic acid sequences through selective mutations of specific genomic targets. These specific genomic targets include, but are not limited to, genes, promoters, open reading frames, or any nucleic acid sequence.

[0036] The terms "base editor," "base editing system," and "base correction system" used in this specification are used interchangeably to refer to a substance having the activity of altering nucleic acid sequences through selective mutation of genomic targets, and include a combination of one or more different base editors. The terms "base editor," "base editing system," or "base correction system" used in this specification, depending on the context, can be a polypeptide (which may be a fusion protein) or a polynucleotide or a combination thereof, or a composition containing one or more polypeptides (which may be fusion proteins) or polynucleotides or combinations thereof. Therefore, the terms "base editing system" or "base editing composition" used in this specification can contain one base editor or a combination of two or more different base editors, wherein the different base editors can be used simultaneously or individually.

[0037] As used in this specification, the term "fusion protein" refers to a polypeptide composed of two or more different polypeptides linked by peptide bonds. The fusion proteins used in this invention include DNA-binding proteins and enzyme proteins. In addition to the DNA-binding proteins and enzyme proteins, they may contain signal sequences such as NLS, NES, MTS, and CTS, and may include additional sequences for biotechnological methods such as tagging. These individual polypeptides can be directly linked or linked via linkers. A "linker" is any molecule that links two different molecules. Any linker known in the field of biotechnology for providing fusion proteins or protein conjugates can be used; for example, peptide linkers containing 1 to 100 amino acid residues can be used. Such fusion proteins (or polynucleotides encoding fusion proteins) can be designed and prepared using any method known in the field of biotechnology.

[0038] In this specification, when referring to the “linkage” of two proteins, it can be a direct link or an indirect link through a linker or other proteins (multiple).

[0039] The design and construction methods of the fusion protein or the polynucleotide encoding it of the present invention can employ any method known in the art. The polynucleotide can be inserted into a vector, and the vector can be introduced into cells. The individual proteins constituting the fusion protein of the present invention are typically cloned as a single polynucleotide and expressed as a single polypeptide (fusion protein). However, one or more of the individual proteins can also be cloned as (multiple) separate polynucleotides and expressed as two or more separate polypeptides, and this is also within the scope of the present invention.

[0040] As used in this specification, the terms "target" and "target site" refer to a predetermined nucleic acid sequence of any composition and / or length. Such target sites include, but are not limited to, genes, promoters, or any nucleic acid sequence.

[0041] The terms “gene delivery body,” “delivery composition,” or “vector” used in this specification are used interchangeably to refer to a medium or substance used to deliver or introduce the base composition of the present invention into a cell, tissue, or organism containing a target DNA for base editing, or to express the target DNA.

[0042] The present invention provides a base editing composition having the activity of editing adenine (A) bases of DNA to guanine (G) and / or editing cytosine (C) bases to thymine (T), wherein the base editing composition comprises one or more DNA-binding proteins and one or more enzyme proteins, or comprises multiple polynucleotides encoding said proteins, wherein said proteins may exist in the form of two fusion proteins.

[0043] This invention provides a DNA base editing composition and a DNA base editing method using the base editing composition. The composition comprises one or more DNA-binding proteins and one or more enzyme proteins, or comprises a polynucleotide encoding the protein. The invention also provides a base editing composition and a DNA base editing method using the base editing composition, wherein the one or more enzyme proteins comprise a nicking enzyme, and further comprise adenine deaminase and / or cytosine deaminase.

[0044] The base editing compositions of the present invention and the DNA base editing methods using said compositions may contain one or more DNA-binding proteins and one or more enzyme proteins in the form of one or more fusion proteins, preferably in the form of two fusion proteins. In this case, each of the two fusion proteins independently comprises a DNA-binding protein, at least one of the two fusion proteins contains a nicking enzyme, and at least one of the two fusion proteins independently comprises adenine deaminase and / or cytosine deaminase. In this respect, the DNA base editing compositions of the present invention may contain two polynucleotides encoding the two fusion proteins as described above.

[0045] When the base editing composition of the present invention comprises two fusion proteins, each of the two fusion proteins independently comprises a DNA-binding protein, and one or both of the two fusion proteins may contain a nicking enzyme. Additionally, one or both of the two fusion proteins may independently contain adenine deaminase or cytosine deaminase. When the base editing composition of the present invention comprises both adenine deaminase and cytosine deaminase, the adenine deaminase and cytosine deaminase may be contained in the same fusion protein or in different fusion proteins. When the base editing composition of the present invention comprises two fusion proteins, one of the two fusion proteins may contain adenine deaminase and / or cytosine deaminase, and the other may contain a nicking enzyme; or, of the two fusion proteins, (1) one may contain adenine deaminase, and the other may contain a nicking enzyme and cytosine deaminase; or (2) one may contain cytosine deaminase, and the other may contain a nicking enzyme and adenine deaminase. For example, one of the two fusion proteins may contain cytosine deaminase, and the other may contain a nicking enzyme. For example, one of the two fusion proteins may contain adenine deaminase, and the other may contain a nicking enzyme. Alternatively, one of the two fusion proteins may contain cytosine deaminase, and the other may contain both a nicking enzyme and adenine deaminase. In this respect, the base editing composition of the present invention may contain two polynucleotides encoding the two fusion proteins as described above.

[0046] The DNA base editing composition or the DNA base editing method encoding the polynucleotide of the present invention can be used not only to edit nuclear DNA but also to edit organelle DNA. Furthermore, it provides a DNA base editing method comprising the step of expression in a plant, plant cell, or protoplast. The organelles include mitochondria, chloroplasts, chromoplasts, and leucoplasts. The base editing composition of the present invention can be used in plant cells and animal cells, preferably in plants, plant cells, or protoplasts. The DNA can be plant nuclear DNA or organelle DNA, preferably plant chloroplast DNA.

[0047] The DNA-binding protein used in the base editing method of the present invention may be selected from the group consisting of zinc finger protein, transcription activator-like effector (TALE) protein, CRISPR-associated nuclease, or combinations thereof. Regarding the composition of zinc finger protein, TALE protein, and CRISPR-associated nuclease, reference can be made to all information known prior to the date of this application, particularly the entire technical content described in International Patent Application Publication WO2022 / 060185, which is incorporated herein by reference.

[0048] In one embodiment of the present invention, the DNA-binding protein may be a zinc finger protein. A "zinc finger protein," referred to as "ZF" or "ZFP," is a protein or a portion of a larger protein that binds to DNA in a sequence-specific manner via one or more zinc finger motifs. A zinc finger protein may contain 3 to 6 zinc finger motifs, but the zinc finger protein used in this invention is not limited to this. A zinc finger motif is a small polypeptide domain consisting of approximately 20 to 40 amino acid residues, with four amino acid residues as cysteine ​​and / or histidine that may be suitably positioned and coordinated to a zinc ion. For example, depending on the type of residues that coordinate to a zinc ion, zinc finger motifs may be classified as C2H2 (Cys2-His2), C3H (Cys2-CysHis), C4 (Cys2-Cys2), etc. The zinc finger motifs used in the zinc finger proteins used in this invention may be naturally occurring zinc fingers or variants thereof in any eukaryote, including humans, or artificially manufactured.

[0049] In one embodiment of the present invention, the DNA-binding protein may be a TALE protein. The TALE protein of the present invention refers to a protein that binds to nucleotides in a sequence-specific manner via one or more TALE-repeat modules. The TALE protein comprises at least one TALE-repeat module, preferably 1 to 30 TALE-repeat modules, but is not limited thereto. As used herein, a TALE-repeat module may be referred to as a "TALE array," and the expression "TALE protein" refers to a configuration comprising an N-terminal domain and a C-terminal domain (which may include a half domain) on each side of the TALE array. Depending on the context, the term "TALE" as used herein may refer to either "TALE array" or "TALE protein."

[0050] When two or more DNA-binding proteins are used in the base editing method of the present invention, the two or more DNA-binding proteins may be of the same or different types. For example, when two DNA-binding proteins are used in the base editing composition of the present invention (e.g., when two fusion proteins are used), the two DNA-binding proteins may both be zinc finger proteins, both be TALE proteins, or both be CRISPR-associated nucleases. Alternatively, one of the two DNA-binding proteins may be a zinc finger protein, and the other may be a TALE protein or a CRISPR-associated nuclease, or one may be a TALE protein, and the other may be a zinc finger protein or a CRISPR-associated nuclease. Preferably, when the DNA-binding protein used in the present invention is a TALE protein and is used in the form of two fusion proteins, the two fusion proteins each contain a TALE protein. In this respect, the base editing composition of the present invention may contain two or more polynucleotides encoding two or more DNA-binding proteins, which may be the same or different as described above.

[0051] When two DNA-binding proteins are used in the base editing system of the present invention, one can be designated by the modifier "first" or "left," and the other by the modifier "second" or "right." For example, when both DNA-binding proteins are TALE, one TALE protein can be called the "first TALE protein" or "left TALE protein," and the other TALE protein can be called the "second TALE protein" or "right TALE protein."

[0052] When the TALE protein is used as the DNA-binding protein of the base editing composition of the present invention, a single-module TALE array or a multi-module TALE array (e.g., a dual-module TALE array consisting of a first TALE (or left TALE) array and a second TALE (or right TALE) array) can also be used. When the base editing composition of the present invention comprises two fusion proteins, the two fusion proteins each have a first TALE protein and a second TALE protein. For example, one of the first TALE protein and the second TALE protein may be linked to adenine deaminase, and the other TALE protein may be linked to a nicking enzyme. Alternatively, for example, one of the first TALE protein and the second TALE protein may be linked to cytosine deaminase, and the other TALE protein may be linked to a nicking enzyme. The adenine deaminase, cytosine deaminase, and nicking enzyme may each be linked to the C-terminal domain of the TALE protein. When the base editing composition of the present invention comprises one fusion protein, the fusion protein has a single-module TALE protein. For example, the single-module TALE protein may be linked to both adenine deaminase and a nicking enzyme. Additionally, for example, monomodule TALE proteins can be linked to cytosine deaminases and nicking enzymes. When linked to a monomodule TALE protein, the nicking enzyme and deaminase (cytosine deaminase or adenine deaminase) can be attached to the C-terminal domain of the TALE protein in the order of nicking enzyme and deaminase, or in the order of deaminase (cytosine deaminase or adenine deaminase) and nicking enzyme.

[0053] The base editing composition of the present invention comprises both cytosine deaminase and adenine deaminase, enabling simultaneous targeted editing of cytosine (C) bases to thymine (T) bases and targeted editing of adenine (A) bases to guanine (G) bases. When using this dual-base editor configuration, a single-module TALE array or a multi-module TALE array (e.g., a dual-module TALE array consisting of a first TALE (or left TALE) array and a second TALE (or right TALE) array) can be used. When the dual-base editor of the present invention comprises two fusion proteins, the two fusion proteins each have a first TALE protein and a second TALE protein. For example, one of the first TALE protein and the second TALE protein may be linked to both adenine deaminase and cytosine deaminase, and the other TALE protein may be linked to a nicking enzyme. Alternatively, for example, one of the first TALE protein and the second TALE protein may be linked to cytosine deaminase, and the other TALE protein may be linked to both adenine deaminase and a nicking enzyme. Additionally, for example, one of the first and second TALE proteins may be linked to an adenine deaminase, and the other TALE protein may be linked to a cytosine deaminase and a nicking enzyme. The adenine deaminase, cytosine deaminase, and nicking enzyme may each be linked to the C-terminal domain of the TALE protein. When two or more of the adenine deaminase, cytosine deaminase, and nicking enzyme are linked to a TALE protein, they may be linked to the TALE protein in any order. When the dual-base editor of the present invention comprises a fusion protein, the fusion protein has a monomodal TALE protein. For example, the monomodal TALE protein may be linked to an adenine deaminase, a cytosine deaminase, and a nicking enzyme. The adenine deaminase, cytosine deaminase, and nicking enzyme may be linked to the TALE protein in any order (e.g., via the C-terminal domain).

[0054] The enzyme protein used in the base editing method of the present invention is characterized by comprising a nick enzyme. A nick enzyme is any enzyme that produces single-strand breaks (also called "nicks") in double-stranded DNA, that is, it has the activity of cutting one strand of the DNA double helix but not the other strand. The nick enzyme can be selected from MutH, BspD6I, FokI (including homodimers or heterodimers), BsaI, BsmBI, BsmAI, BsrDI, CviPII, BspQI, AlwI and I-TevI, fragments thereof and variants thereof, or their conserved amino acid substitutions. For example, MutH (SEQ ID NO: 6), Nt.BspD6I(C) (the catalytically active fragment of Nt.BspD6I; SEQ ID NO: 8), or a heterodimer of the catalytically active domain of FokI (SEQ ID NO: 32) can be used as the nick enzyme. MutH can be used in the form of a variant or fragment thereof, wherein one or more amino acids selected from the group consisting of residues 48, 91, 94, 184, and 212 of the amino acid sequence of SEQ ID NO: 6 are mutated to other amino acids. For example, it may have substitutions of one or more amino acids selected from the group consisting of K48A, E91A, F94A, R184A, and Y212S. FokI can be used in the form of a variant or fragment thereof, wherein one or more amino acids selected from the group consisting of residues 469, 483, 486, 487, 490, 496, 499, 537, and 538 of the amino acid sequence of SEQ ID NO: 71 are mutated to other amino acids, and can be used in monomeric or dimer form, wherein the dimer may be a homodimer or a heterodimer. For example, it can be used in the form of a heterodimer as shown in SEQ ID NO: 32 (a FokI fragment (domain) with amino acid substitutions of E490K and I538K and a FokI fragment (domain) with amino acid substitutions of D450A, Q486E and I499L linked by 40 amino acid linkers). The amino acid substitution positions (numbers) mentioned above are based on SEQ ID NO: 71. It is a common practice in the biotechnology field to describe specific amino acid substitutions by using the positional number of the unsubstituted amino acid residues and the sequence number of the substituted amino acid.

[0055] The nicking enzyme can be used as a fusion protein linked to a DNA-binding protein and / or other proteins, or it can be expressed as a standalone protein. When more than one fusion protein is used in the base editing method of the present invention, only one of the more than one fusion protein contains the nicking enzyme, or multiple fusion proteins may contain the same or different nicking enzymes. Additionally, when multiple fusion proteins are used in the base editing system of the present invention, the nicking enzyme may be included in the fusion protein, which contains a DNA-binding protein that binds to DNA located upstream of the 5' or 3' of the target DNA site for base editing. In this respect, the base editing composition of the present invention may contain a polynucleotide encoding the nicking enzyme as described above, which may be linked to a polynucleotide encoding other proteins, or may exist alone.

[0056] The enzyme protein used in the base editing method of the present invention may comprise cytosine deaminase and / or adenine deaminase. The cytosine deaminase and / or adenine deaminase may be used as a fusion protein linked to a DNA-binding protein and / or other proteins, or may be expressed as a separate protein. When more than one fusion protein is used in the base editing method of the present invention, only one of the more than one fusion protein contains cytosine deaminase and / or adenine deaminase, or multiple fusion proteins may contain cytosine deaminase and / or adenine deaminase. When multiple fusion proteins may contain cytosine deaminase and / or adenine deaminase, the multiple fusion proteins may contain mutually different cytosine deaminases and / or adenine deaminases. Additionally, when multiple fusion proteins are used in the base editing method of the present invention, the cytosine deaminase and / or adenine deaminase may be independently contained in the fusion protein, which comprises a DNA-binding protein that binds to DNA located 5' upstream or 3' of the target DNA site for base editing. When the cytosine deaminase and / or adenine deaminase are used, their positions are independent of the nicking enzyme. That is, when the cytosine deaminase and / or adenine deaminase are used, they may be contained in the same fusion protein as the nicking enzyme, or they may be contained in different fusion proteins, or both may coexist. Furthermore, when contained in the same fusion protein, the cytosine deaminase and / or adenine deaminase may be directly or indirectly (e.g., via linkers and / or other protein components) linked to the N-terminus or C-terminus of the nicking enzyme. In this respect, the base editing composition of the present invention may contain polynucleotides encoding the cytosine deaminase and / or adenine deaminase as described above, which may be independently linked to polynucleotides encoding other proteins, or exist alone.

[0057] The adenine deaminase used in the base editing method of the present invention refers to any deaminase having the activity of converting adenine bases into hypoxanthine (inosine as a nucleoside). In this specification, the term "adenine deaminase" includes enzymes that possess both adenine deaminase and cytosine deaminase activities. The adenine deaminase used in the present invention can be derived from or mutated (e.g., engineered or evolved) any organism (e.g., eukaryotes or prokaryotes), for example, it can be derived from or mutated from *Escherichia coli* (…). Escherichia coli , E. coli Staphylococcus aureus ( Staphylococcus aureus , S. aureus ), enteric Salmonella ( Salmonella Typhi , S. Typhi Shewanella putrefactive bacteria ( Shewanella putrefaciens , S. putrefaciens Haemophilus influenzae ( ) Haemophilus influenzae , H. influenzae ) or Crescentella ( Caulobacter crescentus , C. crescentus The organisms mentioned include, but are not limited to, algae, bacteria, fungi, plants, invertebrates, and mammals. The adenine deaminase may be, for example, tRNA-specific adenosine deaminase (TadA) or a variant thereof.

[0058] The TadA mentioned can be, for example, TadA8e (SEQ ID NO: 28) or its cleavage or variants (e.g., variants modified or evolved in a manner applicable to deoxynucleotides). Variants of the mentioned TadA8e can be, for example, mutations of one or more amino acids from the group consisting of positions 26, 28, 30, 46, 48, 49, 73, 82, 84, 96, 106, 108, 110, and 111 of SEQ ID NO: 28, to other amino acids, or conserved amino acid substitutions thereof. For example, the amino acid variant may contain substitutions of one or more amino acids selected from the group consisting of V28Q, V28R, A48W, F84M, V106A, K110S, K110T, K110V, R111F, R111Q, R111S, R111T, and R111Y. For example, variants of TadA8e may include substitutions of one or more amino acids selected from the group consisting of R26, V28, A48, Y73, and H96, and may have, for example, the amino acid sequence of SEQ ID NO: 30. The composition of the adenine deaminases that can be used in this invention is based on information known prior to the date of this application, particularly the complete technical content described in the International Patent Application Publications WO2022 / 060185, WO2023 / 086953, etc., which are incorporated herein by reference.

[0059] The cytosine deaminase that can be used in the base editing compositions and base editing methods of the present invention refers to any deaminase having the activity of converting cytosine bases into uridine. In this specification, the term "cytosine deaminase" includes enzymes that possess both adenine deaminase and cytosine deaminase activities. The cytosine deaminases that can be used in the present invention can be derived from or mutated (e.g., engineered or evolved) any organism (e.g., eukaryotes or prokaryotes), wherein said organisms include, but are not limited to, algae, bacteria, fungi, plants, invertebrates, and mammals. For example, it could be: apolipoprotein B editing complex (APOBEC), activation-induced cytidine deaminase (AID), a tRNA-specific adenosine deaminase (TadA) or its ortholog, derived from or mutated from bacterial adenosine deaminase, or a variant thereof. The cytidine deaminase could be, for example, APOBEC (apolipoprotein B editing complex), AID (activation-induced deaminase), TadA (tRNA-specific adenosine deaminase), or a variant thereof. The cytosine deaminase with the TadA8e mutation mentioned above can be, for example, one or more amino acid residues at positions 6, 26, 27, 28, 46, 48, 49, 61, 73, 74, 76, 77, 82, 96, 107, 108, 112, 114, 115, 119, 122, 127, 142, 143, 151, 154, and 158 in the amino acid sequence of SEQ ID NO: 28, especially one or more amino acid residues at positions 26, 27, 28, 48, 61, 73, 76, 96, 151, 154, and 158 mutated to another amino acid, for example, it can have the amino acid sequence of SEQ ID NO: 22, SEQ ID NO: 26, or SEQ ID NO: 30. The composition of the cytosine deaminase that can be used in this invention may be referenced from the contents known prior to the date of this application, in particular all the technical contents recorded in the international patent application publications WO2022 / 060185, WO2023 / 086953, etc., which are incorporated herein by reference.

[0060] Preferably, when cytosine deaminase is used in the base editing method of the present invention, the cytosine deaminase is not the bacterial cytosine deaminase DddA. tox When cytosine deaminase is used in this invention, the cytosine deaminase preferably has deaminase activity for single-stranded DNA.

[0061] The polypeptides used in the base editing method of the present invention, in particular, one or more of adenine deaminase, cytosine deaminase and nicking enzyme, can be expressed from polynucleotides that have been codon-optimized for expression in plants.

[0062] One or more of the fusion proteins used in the base editing compositions of the present invention and the base editing methods using said base editing compositions, as signal proteins, may contain a nuclear export signal (NES) as part of the fusion protein. When an NES is attached to a base editing protein, base editing can be performed with higher efficiency. The NES sequence can be any signal sequence conferring nuclear export capability (e.g., VDEMTKKFGTLTIHDTEK), and can use natural nuclear export signals (NES) or artificially synthesized NES. For example, it may be derived from mouse parvovirus (MVM), but is not limited thereto. When the NES is used in the base editing system of the present invention, its position can be diverse, for example, it can be directly or indirectly (e.g., through linkers and / or other protein components) linked to the N-terminus of a DNA-binding protein, but is not limited thereto. In this regard, the base editing compositions of the present invention may contain polynucleotides encoding NES as described above.

[0063] One or more of the fusion proteins used in the base editing compositions of the present invention and the base editing methods using said base editing compositions, as signal proteins, may include a mitochondrial targeting sequence (MTS) as part of the fusion protein. The MTS used in the present invention can be any signal sequence capable of translocating into mitochondria, and can be a natural MTS present at the N-terminus of various mitochondrial proteins; in addition, artificially synthesized MTS may also be used. When the MTS is used in the base editing compositions of the present invention, its location can be diverse, for example, it can be directly or indirectly (e.g., through linkers and / or other protein components) linked to the N-terminus of a DNA-binding protein or the N-terminus of a NES, but is not limited thereto. In this respect, the base editing compositions of the present invention may contain a polynucleotide encoding the MTS as described above.

[0064] The base editing compositions of the present invention and the base editing methods using said base editing compositions may contain one or more fusion proteins as signal proteins, including a chloroplast transit signal (CTS) as part of the fusion protein. Such base editing compositions may be used to edit chloroplast, chromoplast, or leucoplast DNA in plant cells. The CTS used in the present invention can be any signal sequence capable of translocating into chloroplasts, and may be a naturally occurring CTS present at the N-terminus of various chloroplast proteins; additionally, artificially synthesized CTS may also be used. When a CTS is used in the base editing compositions of the present invention, its location can be diverse, for example, it may be directly or indirectly (e.g., through linkers and / or other protein components) linked to the N-terminus of a DNA-binding protein or the N-terminus of a NES, but is not limited thereto. In this respect, the base editing compositions of the present invention may contain polynucleotides encoding the CTS as described above.

[0065] One or more of the fusion proteins used in the base editing compositions of the present invention and the base editing methods using said base editing compositions, as signal proteins, may contain a nuclear localization signal (NLS). The NLS used in the present invention can be any signal sequence capable of translocating into the nucleus, or it can be a natural NLS present at the N-terminus of various nuclear proteins, or a synthetically produced NLS. When an NLS is used in the base editing compositions of the present invention, its location can be diverse; for example, it can be directly or indirectly (e.g., through linkers and / or other protein components) attached to the N-terminus of a DNA-binding protein, but is not limited thereto. In this respect, the base editing compositions of the present invention may contain a polynucleotide encoding an NLS as described above.

[0066] The base editing compositions used in this invention may also contain a uracil glycosylase inhibitor (UGI) as an enzyme protein. UGI is preferably included in a base editor fusion protein containing cytosine deaminase because the UGI can increase C-to-T base editing efficiency by inhibiting the activity of uracil DNA glycosylase (UDG), an enzyme that catalyzes the removal of uracil (U) from DNA to repair mutant DNA. When using UGI in the base editing compositions of this invention, it can be used as a fusion protein linked to other proteins or as a standalone protein expression. When used as a fusion protein, the location of UGI can be varied; preferably, it can be directly or indirectly (e.g., via linkers and / or other protein components) linked to the C-terminus of cytosine deaminase. In this respect, the base editing compositions of this invention may contain a polynucleotide encoding the UGI as described above, which may be linked to a polynucleotide encoding other proteins or exist alone.

[0067] When the protein used in the base editing method of the present invention is in the form of a fusion protein (multiple), the fusion protein comprises a DNA-binding protein and an enzyme protein. In addition to the DNA-binding protein and the enzyme protein, it may contain signal sequences such as NLS, NES, MTS, and CTS, and may contain additional sequences for biotechnological methods such as tagging. These individual polypeptides can be directly linked or linked via linkers. A "linker" refers to any molecule that connects two different molecules. In the field of biotechnology, any linker known to be used to provide fusion proteins or protein conjugates can be used; for example, peptide linkers containing 1 to 100 amino acid residues can be used. In this respect, the base editing composition of the present invention may comprise multiple polynucleotides encoding the fusion proteins (multiple) described above. Such fusion proteins (or polynucleotides encoding fusion proteins) can be designed and prepared by any method known in the field of biotechnology.

[0068] In this invention, the base editing composition may be in the form of multiple polynucleotides encoding multiple proteins or fusion proteins as described in this application specification. Such polynucleotides may be inserted into a gene delivery vehicle (e.g., a vector), and said vector may be introduced into cells. The polynucleotides may be RNA, such as mRNA. When the base editing composition of this invention is delivered in mRNA form, compared to delivery in the form of a vector using DNA, the transcription to mRNA process is not required, thus allowing for faster initiation of gene editing, a higher likelihood of achieving transient protein expression, and reduced off-target editing effects.

[0069] One aspect of the present invention relates to a delivery composition comprising the base editing composition of the present invention, said delivery composition having any of the following forms: delivery of a base editing composition comprising a plurality of proteins or fusion proteins or polynucleotides encoding thereof, as described throughout this specification, to a cell, tissue, or organism containing target DNA for base editing. Delivery of the base editing proteins (a plurality of), fusion proteins (a plurality of), or polynucleotides (a plurality of) of the present invention into cells by means of electroporation or the like, or modification thereof to enable good delivery to the nucleus or organelles, or forms with attached individual portions, are also included in the delivery composition of the present invention.

[0070] When the base editing composition is in the form of a polynucleotide, the delivery composition of the present invention may also be referred to as a "gene delivery body" or a "vector". When the base editing composition is in the form of a polynucleotide, the delivery composition of the present invention may be a viral delivery body or a non-viral delivery body, preferably a viral delivery body. For example, the viral delivery body may contain adenovirus, adeno-associated virus (AAV), retrovirus, lentivirus, poxvirus, vaccinia virus, and herpesvirus, but is not limited thereto. For example, the non-viral gene delivery body may contain plasmids, metal nanoparticles, lipid nanoparticles, or polymer nanoparticles, but is not limited thereto. The base editing composition contained in the "gene delivery body" or "delivery composition" may contain multiple base editing proteins or fusion proteins or multiple polynucleotides encoding them as described in the specification of this application, specifically, it may be in the form of more than one polynucleotide.

[0071] This invention also provides a DNA base editing method, comprising the steps of expressing or introducing into a target animal or animal cell, or plant or plant cell, the base editing protein (multiple) or fusion protein (multiple) or polynucleotide (multiple) encoding the protein described in this application. The DNA base editing method can achieve targeted editing of adenine (A) bases to guanine (G) bases or cytosine (C) bases to thymine (T) bases, or both, in nuclear DNA or organelle DNA. Simultaneous A-to-G and C-to-T editing can be achieved in two ways: by using a structure containing both adenine deaminase and cytosine deaminase, or by using a deaminase structure having both adenine deaminase and cytosine deaminase activity alone. The C-to-T editing achieved by the base editing method of the present invention can not only edit the cytosine (C) bases present in the 5'-TC-3' sequence of the target DNA sequence, but also edit the cytosine (C) bases in the C, CC or GC base sequences of the a portion of the target DNA sequence into thymine (T) bases. In particular, it can edit the cytosine (C) bases in the GC base sequence of the target DNA sequence into thymine (T) bases.

[0072] One aspect of the present invention provides a plant cell or protoplast that selectively edits cytosine and / or adenine bases of a desired target DNA (especially chloroplast DNA) using the base editing method of the present invention.

[0073] Another aspect of the present invention provides a plant or a portion thereof grown or cultured from plant cells or protoplasts edited from the bases described above, or a progeny or clone of said plant or a portion thereof.

[0074] Another aspect of the invention provides a seed obtained from the plant, its offspring or clone, or a portion thereof as described above.

[0075] Another aspect of the present invention provides a plant or its offspring, or a portion thereof, grown from the seeds described above.

[0076] When the base editing composition of the present invention is in the form of a polynucleotide, the transformation of plants by such polynucleotide can be achieved using transformation techniques well known to those skilled in the art of biotechnology. For example, transformation methods using Agrobacterium (e.g., Agrobacterium tumefaciens, Agrobacterium rhizogene, etc.), microprojectile bombardment, electroporation, PEG-mediated fusion, microinjection, liposome-mediated method, inplanta transformation, vacuum infiltration method, floral meristem dipping method, and Agrobacterium spraying method can be used. Regarding Agrobacterium transformation, a binary vector system utilizing two replicons can be used.

[0077] When the base editing composition of the present invention is in the form of a polynucleotide, a viral vector may be used, for example, including but not limited to: Geminivirus, Tobacco Rattle Virus (TRV), Tomato Mosaic Virus (ToMV), Foxtail Mosaic Virus (FoMV), Barley Yellow Striate Mosaic Virus (BYSMV), Sonchus Yellow Net Rhabdovirus (SYNV), etc.

[0078] When the base editing composition of the present invention is present in the form of mRNA, the mRNA can be delivered directly or via a carrier. Methods for delivering mRNA molecules into cells, including methods for delivering mRNA into cells in vivo or in vitro, are considered. For example, mRNA molecules can be delivered into cells by comprising lipids (e.g., liposomes, micelles, etc.), nanoparticles or nanotubes, or cationic compounds (e.g., polyethyleneimine or PEI). Depending on the circumstances, a bolistic method (e.g., a gene gun or a bio-ballistic particle delivery system) can be used to deliver mRNA into cells. The carrier may comprise, but is not limited to, cell-penetrating peptides (CPPs), nanoparticles, or polymers.

[0079] The plant cells, protoplasts, plants, and seeds of this invention can edit cytosine bases in wild-type target DNA sequences to thymine or adenine bases to guanine. One of the advantages of this invention is that, unlike existing cytosine base editors, it can edit not only cytosine bases present in the 5'-TC-3' sequence of the target DNA sequence, but also cytosine bases present in the 5'-AC-3', 5'-CC-3', or 5'-GC-3' sequences to thymine, particularly cytosine bases present in the 5'-GC-3' sequence. Furthermore, even without the application of DddA… tox It can also be implemented using DddA-based... tox Existing base editors are used to enable targeted C-to-T and / or A-to-G base editing.

[0080] The inventors of this invention can successfully obtain individuals in a homologous state, achieved by base editing only in the desired target bases of chloroplast DNA in T2 and subsequent generations of individuals transformed with the base editing composition of this invention. Homologousity refers to the identical replication of DNA in chloroplasts or mitochondria, forming a genetically stable state, which is a very important aspect of improving plant traits through gene editing.

[0081] In technologies that manipulate plant genes and provide improved plants through gene editing, commercialization is difficult due to regulations related to genetically modified plants (GMOs) where the foreign gene introduced for gene editing has not been knocked out. The inventors of this invention have demonstrated that by using the base editing composition of this invention, herbicide-resistant plants can be stably obtained even without the presence of a foreign gene conferring herbicide resistance.

[0082] The present invention may be based on the above content and may involve the following embodiments (1) to (34), but is not limited thereto.

[0083] Implementation Scheme (1): A method for base editing of plant DNA, the method comprising the step of expressing a DNA base editing composition in a target plant, plant cell or protoplast, wherein the base editing composition comprises one or more DNA binding proteins and one or more enzyme proteins, or comprises a polynucleotide encoding the protein, wherein the one or more enzyme proteins comprises a nickase and further comprises adenine deaminase and / or cytosine deaminase.

[0084] Implementation Scheme (2): According to the base editing method of Implementation Scheme (1), wherein the DNA base editing composition comprises one or more fusion proteins or one or more polynucleotides encoding one or more fusion proteins, each of the one or more fusion proteins independently comprises a DNA binding protein, at least one of the one or more fusion proteins comprises a nicking enzyme, and at least one of the one or more fusion proteins comprises adenine deaminase and / or cytosine deaminase.

[0085] Implementation Scheme (3): According to the base editing method described in Implementation Scheme (2), at least one of the more than one fusion protein contains a cytosine deaminase, wherein the cytosine deaminase is an apolipoprotein B editing complex (APOBEC), an activation-induced deaminase (AID), or a tRNA-specific adenosine deaminase (TadA), or a variant thereof.

[0086] Implementation scheme (4): According to the base editing method described in implementation scheme (3), wherein cytosine deaminase has deaminase activity for single-stranded DNA.

[0087] Implementation scheme (5): The base editing method according to implementation scheme (3), wherein the cytosine deaminase is tRNA-specific adenosine deaminase (TadA) or a variant thereof.

[0088] Implementation scheme (6): According to the base editing method described in implementation scheme (3), wherein the cytosine deaminase comprises an amino acid sequence selected from the group consisting of the amino acid sequences of SEQ ID NO: 22, 24, 26 and 30 or their conserved amino acid substitutes.

[0089] Implementation Scheme (7): According to the base editing method described in Implementation Scheme (2), wherein at least one of the more than one fusion protein contains an adenine deaminase, the adenine deaminase being tRNA-specific adenosine deaminase (TadA) or a variant thereof.

[0090] Implementation scheme (8): According to the base editing method described in implementation scheme (7), wherein the adenine deaminase comprises the amino acid sequence of SEQ ID NO: 28 or SEQ ID NO: 30 or its conserved amino acid substitutes.

[0091] Implementation scheme (9): a base editing method according to any one of implementation schemes (1) to (8), wherein the nicking enzyme is selected from the group consisting of MutH, BspD6I, FokI (including FokI-FokI homodimer or heterodimer), BsaI, BsmBI, BsmAI, BsrDI, CviPII, BspQI, AlwI and I-TevI, fragments thereof and variants thereof.

[0092] Implementation scheme (10): According to the base editing method described in implementation scheme (9), wherein the nicking enzyme comprises an amino acid sequence selected from the group consisting of SEQ ID NO: 6, SEQ ID NO: 8 and SEQ ID NO: 32 or a conserved amino acid substitute thereof.

[0093] Implementation scheme (11): According to the base editing method described in implementation scheme (1) or (2), wherein each of the more than one DNA binding proteins is independently selected from the group consisting of zinc finger protein, transcription activator-like effector (TALE) protein and CRISPR-related nuclease.

[0094] Implementation scheme (12): According to the base editing method described in implementation scheme (1) or (2), wherein one or more of adenine deaminase, cytosine deaminase and nicking enzyme are expressed by a polynucleotide expression that has been codon-optimized for expression in plants.

[0095] Implementation scheme (13): The base editing method according to any one of implementation schemes (1) to (12), wherein the DNA is plant organelle DNA.

[0096] Implementation scheme (14): The base editing method according to any one of implementation schemes (3) to (6), wherein the fusion protein containing cytosine deaminase also contains uracil glycosylase inhibitor (UGI).

[0097] Implementation scheme (15): According to the base editing method described in implementation scheme (2), wherein the one or more fusion proteins further include a nuclear export signal (NES).

[0098] Implementation scheme (16): According to the base editing method described in implementation scheme (2), wherein the one or more fusion proteins further include a mitochondrial targeting sequence (MTS).

[0099] Implementation scheme (17): According to the base editing method described in implementation scheme (2), wherein the one or more fusion proteins further include a chloroplast transit signal (CTS).

[0100] Implementation scheme (18): According to the base editing method described in implementation scheme (2), wherein one or more fusion proteins each contain a nuclear localization signal (NLS).

[0101] Implementation Scheme (19): According to the base editing method of Implementation Scheme (2), wherein the DNA base editing composition comprises two fusion proteins or two polynucleotides encoding two fusion proteins, one of the two fusion proteins comprising adenine deaminase and / or cytosine deaminase, and the other comprising a nicking enzyme.

[0102] Implementation scheme (20): According to the base editing method of implementation scheme (2), wherein the DNA base editing composition comprises two fusion proteins or two polynucleotides encoding two fusion proteins, wherein, of the two fusion proteins, (1) one contains adenine deaminase and the other contains nicking enzyme and cytosine deaminase, and (2) one contains cytosine deaminase and the other contains nicking enzyme and adenine deaminase.

[0103] Implementation Scheme (21): The base editing method according to Implementation Scheme (2), wherein the DNA base editing composition comprises a fusion protein, the fusion protein comprising a nicking enzyme and comprising adenine deaminase and / or cytosine deaminase.

[0104] Implementation scheme (22): The base editing method according to any one of implementation schemes (1) to (21) wherein targeted editing of cytosine (C) base to thymine (T) base and targeted editing of adenine (A) base to guanine (G) base can be achieved simultaneously.

[0105] Implementation scheme (23): According to any one of the implementation schemes (1) to (21), the base editing method is used to edit the cytosine (C) base (preferably the cytosine (C) base in the GC base sequence) into the thymine (T) base in the TC, AC, CC or GC base sequence of the target DNA sequence.

[0106] Implementation scheme (24): A plant cell or protoplast that performs DNA base editing by a base editing method according to any one of implementation schemes (1) to (23).

[0107] Implementation scheme (25): A plant or part of a plant that is grown or cultured from plant cells or protoplasts as described in implementation scheme (24).

[0108] Implementation scheme (26): a plant or part of a plant, which is a progeny or clone of the plant described in implementation scheme (25).

[0109] Implementation scheme (27): A seed obtained from a plant according to implementation scheme (25) or (26).

[0110] Implementation scheme (28): a plant or a progeny of a plant, or a part of a plant, which grows from a seed according to implementation scheme (27).

[0111] Implementation scheme (29): A plant cell or protoplast, plant or part of a plant, or seed, wherein the plant cell or protoplast described in implementation scheme (24), the plant or part thereof described in implementation scheme (25), (26) or (28), or the seed described in implementation scheme (27), wherein in the wild-type target DNA sequence, cytosine (C) bases are edited to thymine (T) bases or adenine (A) bases are edited to guanine (G) bases.

[0112] Implementation scheme (30): According to the plant cell or protoplast, plant or part of plant or seed described in implementation scheme (29), wherein in the TC, AC, CC or GC sequence of the target DNA sequence, the cytosine (C) base, preferably the cytosine (C) base in the GC base sequence is edited to a thymine (T) base.

[0113] Implementation Scheme (31): A base editing composition comprising one or more DNA-binding proteins and one or more enzyme proteins, or one or more polynucleotides encoding them, said one or more enzyme proteins comprising a nicking enzyme, and further comprising adenine deaminase and / or cytosine deaminase.

[0114] Implementation Scheme (32): The base editing composition according to Implementation Scheme (31), wherein the one or more DNA-binding proteins and the one or more enzyme proteins exist in the form of one or more fusion proteins, each of the one or more fusion proteins independently contains a DNA-binding protein, at least one of the one or more fusion proteins contains a nicking enzyme, and at least one of the one or more fusion proteins contains adenine deaminase and / or cytosine deaminase.

[0115] Implementation scheme (33): A delivery composition comprising one or more carriers, said carriers comprising the base editing composition according to implementation scheme (31) or (32).

[0116] Implementation scheme (34): a plant cell or protoplast, or a plant or part of a plant containing them, or a seed, which is transformed from one or more carriers according to implementation scheme (33).

[0117] The present invention will be described in detail below through the following embodiments. However, the following embodiments are only for illustrating the present invention, and the present invention is not limited to these embodiments.

[0118] Example 1: Plasmid Construction The DNA encoding a base editor fusion protein was cloned, said base editor fusion protein being grown on lettuce (…). Lactuca sativa ) and Arabidopsis thaliana ( Arabidopsis thaliana Base editing was performed on the psaA and psbA genes of the target gene.

[0119] As shown in Table 1 below, the DNA encoding the psaA and psbA gene sequences of TALE proteins that recognize lettuce (denoted as Ls) and Arabidopsis thaliana (denoted as At) was constructed using the known Golden Gate cloning technique and cloned in a manner that linked the DNA sequences of other proteins that constitute the base editor fusion protein.

[0120] Table 1

[0121] Specifically, in order to transform Arabidopsis thaliana ( Arabidopsis thalianaUsing the Golden Gate cloning technique, the TALE protein-coding base sequence was assembled at the desired position in a master vector containing the RPS5A promoter and including the CTS and tag sequences. Using PrimeSTAR® GXL DNA polymerase (TAKARA), PCR products encoding the desired fusion protein were generated and assembled using NEBuilder. These were then chemically transformed into *E. coli* DH5α, and antibiotic-resistant colonies were screened. Sanger sequencing was then used to analyze the surviving colonies. Each construct, containing deaminase and cleavage enzyme, was digested using AatII and PmeI restriction endonucleases, and then incorporated into the monolithic construct using T4 DNA ligase.

[0122] To transform lettuce protoplasts, the base sequence encoding the TALE protein was assembled using the aforementioned master vector. A PCR product encoding the desired fusion protein was generated using PrimeSTAR® GXL DNA polymerase (TAKARA), and assembled into a vector [pCaMV35S PPDK-CTS-3xFlag-Nos terminator] cleaved into BamHI and XhoI using NEBuilder. The assembled construct was chemically transformed into DH5α, and antibiotic-resistant colonies were screened. The plasmid, analyzed by Sanger sequencing, was purified using the ZymoPURE® II plasmid midiprep kit for transfection.

[0123] The cloned DNA sequence is structured as follows: For protoplast transformation in lettuce: p35SPPDK-CTS-3Xflag-TALE-2aa lin-TadA-CDd-4aa lin 2-UGI-4aa lin 2-UGI-Nos ter.

[0124] p35SPPDK-CTS-3Xflag-TALE-2aa lin-rApobec 1-4aa lin 2-UGI-4aa lin 2-UGI-Nos ter.

[0125] p35SPPDK-CTS-3Xflag-TALE-2aa lin-TadA E27R N46L-4aa lin 2-UGI-4aalin 2-UGI-Nos ter.

[0126] p35SPPDK-CTS-3Xflag- TALE-18aa lin- TadA-CDd -4aa lin 2-UGI-4aa lin2-UGI -Nos ter。

[0127] p35SPPDK-CTS-3Xflag- TALE-18aa lin-rApobec 1-4aa lin 2-UGI-4aa lin 2-UGI -Nos ter。

[0128] p35SPPDK-CTS-3Xflag- TALE-18aa lin- TadA E27R N46L-4aa lin 2-UGI-4aalin 2-UGI -Nos ter。

[0129] p35SPPDK-CTS-3Xflag- TALE-2aa lin-TadA8e-Nos ter。

[0130] p35SPPDK-CTS-3Xflag- TALE-18aa lin-TadA8e-Nos ter。

[0131] p35SPPDK-CTS-3Xflag- TALE-4aa lin 1- MutH -Nos ter。

[0132] p35SPPDK-CTS-3Xflag- TALE-4aa lin 1- Nt.BspD6I(C) -Nos ter。

[0133] p35SPPDK-CTS-3Xflag- TALE-2aa lin-1397N-4aa lin 2-UGI-Nos ter。

[0134] p35SPPDK-CTS-3Xflag- TALE-2aa lin-1397C-4aa lin 2-UGI-Nos ter。

[0135] p35SPPDK-CTS-3Xflag- TALE-2aa lin-GSVG-4aa lin 2-UGI-Nos ter。

[0136] p35SPPDK-CTS-3Xflag- TALE-2aa lin-1397N-Nos ter。

[0137] p35SPPDK-CTS-3Xflag-TALE-2aa lin-1397C-16aa lin-TadA8e-Nos ter.

[0138] p35SPPDK-CTS-3Xflag-TALE-18aa lin-TadA8e-4aa lin 3-GSVG-Nos ter.

[0139] Arabidopsis thaliana transformation: pRPS5A-CTS-3Xflag-TALE-4aalin 1-MutH-35Ster.

[0140] pRPS5A-CTS-3Xflag- TALE-4aa lin 1- Nt.BspD6I(C)-35S ter.

[0141] pRPS5A-CTS-3Xflag- TALE-17aa lin- FokI E490K I538K-40aa- FokI D450AQ486E I499L-35Ster.

[0142] pRPS5A-CTS-3Xflag- TALE-4aa lin 1- 2aa lin-TadA-CDd-4aa lin 2-UGI-4aalin 2-UGI -35S ter.

[0143] pRPS5A-CTS-3Xflag- TALE-4aa lin 1- 2aa lin- rApobec 1-4aa lin 2-UGI-4aa lin 2-UGI -35S ter.

[0144] pRPS5A-CTS-3Xflag- TALE-4aa lin 1- 2aa lin-TadA8e-35S ter.

[0145] The DNA used for protoplast transformation in lettuce is located between the 35SPPDK promoter and the Nos terminator, while the DNA used for plant transformation in Arabidopsis thaliana is located between the RPS5A promoter and the 35S terminator.

[0146] The DNA sequences used are as follows: CTS (SEQ ID NO: 1): ATGGATTCACAGCTAGTCTTGTCTCTGAAGCTGAATCCAAGCTTCACTCCTCTTTCTCCTCTCTTCCCTTTCACTCCATGTTCTTCTTTTTCGCCGTCGCTCCGGTTTTCTTCTTGCTACTCCCGCCGCCTCTATTCTCCGGTTACCGTCTACGCCGCGAAG。

[0147] 3xflag(SEQ ID NO:3): GACTACAAAGACCATGACGGTGATTATAAAGATCATGACATCGATTACAAGGATGACGATGACAAG。

[0148] MutH(SEQ ID NO:5): ATGTCACAGCCTAGACCACTACTAAGTCCTCCCGAAACAGAAGAGCAATTACTAGCACAGGCTCAACAATTGAGCGGTTATACCTTAGGGGAGCTGGCAGCTTTGGCTGGACTGGTTACTCCAGAGAATTTAAAGAGAGACAAGGGTTGGATCGGCGTACTATTAGAGATATGGCTTGGCGCATCTGCTGGCAGTAAACCGGAACAGGACTTCGCCGCCTTGGGTGTGGAACTAAAAACGATACCCGTGGACTCTCTGGGTCGTCCACTCGAAACGACCTTTGTATGTGTCGCACCTTTAACAGGTAATTCCGGGGTAACATGGGAAACATCCCATGTGAGGCATAAGCTGAAGCGTGTCCTATGGATTCCGGTCGAAGGCGAAAGGTCTATACCCTTAGCTAAGAGGCGTGTCGGATCACCGCTCTTGTGGTCCCCTAATGAAGAAGAAGATAGACAGTTACGAGAGGATTGGGAGGAGCTTATGGATATGATCGTATTAGGTCAGATCGAGAGAATTACCGCTAGGCACGGCGAGTACTTACAAATAAGACCGAAGGCTGCTAATGCGAAGGCTTTAACCGAAGCCATCGGCGCACGAGGTGAACGTATATTAACACTGCCTCGTGGTTTCTATCTGAAGAAAAACTTTACGTCCGCTCTATTAGCGAGACATTTCTTAATACAATGA。

[0149] Nt.BspD6I(C) (SEQ ID NO:7): ATGCGACAATTAGAGGAAGTTATCGACTTACTAGAAGTCTACCATGAGAAGAAAAATGTGATTGAGGAAAAGATAAAGGCACGATTTATTGCCAATAAGAACACGGTTTTCGAGTGGCTTACGTGGAACGGGTTTATAATTCTGGGGAACGCCCTGGAGTATAAAAACAATTTTGTAATCGACGAAGAATTACAACCAGTAACCCACGCTGCCGGTAACCAACCAGATATGGAGATCATATATGAAGATTTCATCGTCTTAGGCGAAGTGACCACGTCTAAAGGAGCCACCCAATTTAAAATGGAGAGCGAGCCCGTGACTCGTCACTACTTGAATAAGAAGAAAGAACTTGAAAAGCAAGGCGTTGAGAAGGAACTGTATTGTCTTTTCATAGCACCAGAGATAAATAAAAACACGTTTGAAGAGTTCATGAAGTATAATATAGTACAGAACACTAGAATAATTCCCCTGAGTTTGAAGCAATTCAATATGCTACTTATGGTACAAAAGAAACTCATCGAGAAGGGAAGGAGACTCTCTAGTTATGATATTAAGAACTTGATGGTCAGCTTATATAGGACGACGATCGAGTGCGAGAGAAAGTATACCCAGATAAAGGCGGGACTGGAAGAAACCCTTAATAACTGGGTCGTTGATAAGGAGGTCAGGTTCTGA。

[0150] 2 aa linker (Linker): GGATCC。

[0151] 4 aa linker (Linker) 1 (SEQ ID NO: 9): GGATCCGGCAGT。

[0152] 4 aa linker (Linker) 2 (SEQ ID NO: 11): AGCGGCGGGAGC。

[0153] 4 aa linker (Linker) 3 (SEQ ID NO: 13): CTAGTCGGTTCC。

[0154] 16aa linker (Linker) (SEQ ID NO: 15): TCAGGAAGCGAAACTCCTGGTACCTCAGAGTCCGCTACTCCCGAATCC.

[0155] 17aa linker (Linker) (SEQ ID NO: 17): TCCGGCGGAGGAGGTAGCGGCGGCGGTGGCTCAGGCGGTGGGGGTAGTAGT.

[0156] 18aa linker (Linker) (SEQ ID NO: 19): GGATCCTCAGGAAGCGAAACTCCTGGTACCTCAGAGTCCGCTACTCCCGAATCC.

[0157] TadA-CDd (SEQ ID NO: 21): ATGTCTGAGGTGGAGTTTTCCCACGAGTACTGGATGAGACATGCCCTGACCCTGGCCAAGAGGGCACGGGATGAGAGGAAGGCACCTGTGGGAGCCGTGCTGGTGCTGAACAATAGAGTGATCGGCGAGGGCTGGAACAGAGCCATCGGCCTGCACGACCCAACAGCCCATGCCGAAATTATCGCCCTGAGACAGGGCGGCCTGGTCATGCAGAACTACAGACTGATTGACGCCACCCTGTACGTGACATTCGAGCCTTGCGTGATGTGCGCCGGCGCCATGATCAACTCTAGGATCGGCCGCGTGGTGTTTGGCGTGAGGAACTCAAAAAGAGGCGCCGCAGGCTCCCTGATGAACGTGCTGAACTACCCCGGCATGAATCACCGCGTCGAAATTACCGAGGGAATCCTGGCAGATGAATGTGCCGCCCTGCTGTGCGATTTCTATCGGATGCCTAGACAGGTGTTCAATGCTCAGAAGAAGGCCCAGAGCTCCATCAAC.

[0158] rApobec 1 (SEQ ID NO: 23): AGCAGTGAGACTGGACCCGTTGCAGTCGATCCCACTTTGCGTCGTCGAATAGAGCCGCACGAGTTTGAAGTTTTCTTCGACCCACGAGAACTTCGAAAAGAAACTTGCTTGTTGTATGAGATTAACTGGGGGGGAAGGCATAGTATATGGAGGCATACATCCCAGAATACTAACAAGCATGTAGAGGTGAACTTTATTGAAAAGTTCACGACCGAGAGGTACTTCTGTCCTAACACGAGGTGCAGTATCACATGGTTCTTATCATGGTCACCTTGCGGCGAGTGCAGTAGAGCAATTACTGAATTTCTGTCTAGATACCCCCATGTTACCCTATTCATATACATCGCAAGACTTTACCACCACGCGGACCCAAGGAATCGTCAAGGACTTAGAGACTTGATATCTTCCGGGGTGACAATTCAGATTATGACGGAACAAGAGAGCGGTTACTGTTGGAGGAACTTCGTCAATTATTCCCCTTCTAACGAAGCTCACTGGCCTCGATACCCCCACCTATGGGTACGACTCTACGTCCTTGAACTCTACTGCATTATACTAGGGCTACCACCATGTTTGAACATTCTGAGACGTAAGCAACCACAGCTAACTTTCTTCACAATTGCCCTGCAATCTTGCCATTACCAGAGGTTACCTCCTCACATCTTGTGGGCAACCGGTTTGAAAAGCGGGTCCGAAACACCGGGTACCTCCGAATCAGCAACCCCCGAGAGT。

[0159] TadA E27R N46L(SEQ ID NO:25): TCTGAGGTGGAGTTTTCCCACGAGTACTGGATGAGACATGCCCTGACCCTGGCCAAGAGGGCACGGGATGAGAGGCGTGTGCCTGTGGGAGCCGTGCTGGTGCTGAACAATAGAGTGATCGGCGAGGGCTGGCTTAGAGCCATCGGCCTGCACGACCCAACAGCCCATGCCGAAATTATGGCCCTGAGACAGGGCGGCCTGGTCATGCAGAACTACAGACTGATTGACGCCACCCTGTACGTGACATTCGAGCCTTGCGTGATGTGCGCCGGCGCCATGATCCACTCTAGGATCGGCCGCGTGGTGTTTGGCGTGAGGAACTCAAAAAGAGGCGCCGCAGGCTCCCTGATGAACGTGCTGAACTACCCCGGCATGAATCACCGCGTCGAAATTACCGAGGGAATCCTGGCAGATGAATGTGCCGCCCTGCTGTGCGATTTCTATCGGATGCCTAGACAGGTGTTCAATGCTCAGAAGAAGGCCCAGAGCTCCATCAAC。

[0160] TadA8e(SEQ ID NO:27): TCTGAGGTGGAGTTTTCCCACGAGTACTGGATGAGACATGCCCTGACCCTGGCCAAGAGGGCACGGGATGAGAGGGAGGTGCCTGTGGGAGCCGTGCTGGTGCTGAACAATAGAGTGATCGGCGAGGGCTGGAACAGAGCCATCGGCCTGCACGACCCAACAGCCCATGCCGAAATTATGGCCCTGAGACAGGGCGGCCTGGTCATGCAGAACTACAGACTGATTGACGCCACCCTGTACGTGACATTCGAGCCTTGCGTGATGTGCGCCGGCGCCATGATCCACTCTAGGATCGGCCGCGTGGTGTTTGGCGTGAGGAACTCAAAAAGAGGCGCCGCAGGCTCCCTGATGAACGTGCTGAACTACCCCGGCATGAATCACCGCGTCGAAATTACCGAGGGAATCCTGGCAGATGAATGTGCCGCCCTGCTGTGCGATTTCTATCGGATGCCTAGACAGGTGTTCAATGCTCAGAAGAAGGCCCAGAGCTCCATCAAC。

[0161] TadA - dual (SEQ ID NO: 29): ATGTCTGAGGTGGAGTTTTCCCACGAGTACTGGATGAGACATGCCCTGACCCTGGCCAAGAGGGCACGGGATGAGGGAGAGGCACCTGTGGGAGCCGTGCTGGTGCTGAACAATAGAGTGATCGGCGAGGGCTGGAACAGAAGAATCGGCCTGCACGACCCAACAGCCCATGCCGAAATTATCGCCCTGAGACAGGGCGGCCTGGTCATGCAGAACTCTAGACTGATTGACGCCACCCTGTACGTGACATTCGAGCCTTGCGTGATGTGCGCCGGCGCCATGATCAACTCTAGGATCGGCCGCGTGGTGTTTGGCGTGAGGAACTCAAAAAGAGGCGCCGCAGGCTCCCTGATGAACGTGCTGAACTACCCCGGCATGAATCACCGCGTCGAAATTACCGAGGGAATCCTGGCAGATGAATGTGCCGCCCTGCTGTGCGATTTCTATCGGATGCCTAGACAGGTGTTCAATGCTCAGAAGAAGGCCCAGAGCTCCATCAAC。

[0162] FokI E490K I538K-40aa- FokI D450A Q486E I499L(SEQ ID NO:31): AAGAGAACCAGACTAGAAACAAGCATTTAAACCCGAACGAGTGGTGGAAAGTTTATCCCAGTTCAGTGACAGAGTTTAAGTTCCTGTTCGTCAGTGGACATTTTAAAGGGAACTACAAGGCTCAACTCACGAGGCTCAACCACATAACCAACTGCAACGGAGCTGTTTTATCCGTTGAAGAGTTGTTGATCGGAGGTGAAATGATCAAGGCGGGCACATTGACCCTCGAGGAAGTAAGGAGGAAGTTCAATAACGGGGAGATTAATTTT。

[0163] DddA tox 1397N(SEQ ID NO:33): GGATCAGGTAGCTACGCACTTGGTCCTTACCAGATTAGCGCACCCCAACTCCCCGCCTATAATGGTCAAACCGTCGGGACCTTTTACTACGTAAACGATGCTGGTGGGCTGGAATCCAAAGTATTCTCCTCAGGGGGCCCTACACCCTACCCCAACTACGCCAATGCTGGTCATGTAGAAGGGCAGTCAGCACTGTTTATGCGCGATAATGGTATAAGCGAGGGGTTGGTCTTCCATAACAACCCAGAGGGTACTTGTGGCTTCTGTGTGAATATGACTGAAACCCTTCTGCCCGAAAATGCCAAGATGACTGTCGTCCCACCTGAAGGC。

[0164] DddA tox 1397C(SEQ ID NO:35): GGCAGTGCCATACCTGTGAAGCGGGGAGCAACAGGGGAGACAAAGGTGTTCACAGGCAACTCTAACAGTCCAAAGAGCCCCACCAAAGGCGGGTGT。

[0165] GSVG(SEQ ID NO:37): GGCAGCTACGCCCTGGGTCCGTATCAGATTAGCGCCCCGCAGCTGCCAGCATACAATGGTCAGACCGTGGGTACCTTCTACTATGTGAACGACGCTGGCGGTCTGGAGGGCAAGGTGTTTAGCAGCGGCGGTCCAACCCCGTACCCAAACTATGCCAATGCCGGTCATGTGGAGAGTCAGAGCGCCCTGTTCATGCGTGATAACGGCATCAGCGAGGGTCTGGTGTTCCACAACAACCCGGAAGGCACCTGCGGTTTTTGCGTGAACATGACCGAGACCCTGCTGCCGGAAAACGCGAAAATGACCGTGGTGCCGCCGGAAGGTGTCATTCCAGTGAAGCGCGGCGCTACCGGTGAAACCAAAGTGTTTACCGGTAACAGCAACGGCCCGAAGAGCCCGACCAAAGGCGGTTGC。

[0166] 2x UGI(SEQ ID NO:39): ACCAATCTCTCCGACATCATTGAAAGAAACAGGAAAGCAGCTGGTGATACAAGAATCAATATTGCTTCCAGAAGAAGTGGAGGAGGTCATTGGAAACAAGCCGGAGTCAGACATTCTTGTAC ACACTGCTTATGATGAAAGTACTGATGAGAATGTTATGTTGTTAACATCTGATGCACCTGAATAATAAACCATGGGCATTATTCAAGATTCAAATGGAGAAAACAAATCAAGATGCTAAGCGGC GGGAGCACCAATCTCTCCGACATCATTTGAAAGAAACAGGAAAGCAGCTGGTGATACAAGAATCAATATTGCTTCCAGAAGAAGTGGAGGAGGTCATTGGAAACAAGCCGGAGTCAGACATTCT TGTACACACTGCTTATGATGAAAGTACTGATGAGAATGTTATGTTGTTAACATCTGATGCACCTGAATAATAAACCATGGGCATTAGTTATTCAAGATTCAAATGGAGAAAACAAATCAAGATGCTA。

[0167] UGI (SEQ ID NO:41): ACCAATCTCTCCGACATCATTGAAAGAAACAGGAAAGCAGCTGGTGATACAAGAATCAATATTGCTTCCAGAAGAAGTGGAGGAGGTCATTGGAAACAAGCCGGAGTCAGACATTCTTGT ACACACTGCTTATGATGAAAGTACTGATGAGAATGTTATGTTGTTAACATCTGATGCACCTGAATAATAAACCATGGGCATTAGTTATTCAAGATTCAAATGGAGAAAACAAATCAAGATGCTA。

[0168] The TALE is a DNA link.

[0169] LspsaA site 1 Left TALE (SEQ ID NO:43): CCCTGCACAGGTGGTTGCTATTGCCTCCAACGGTGGAGGCAAACAGGCCCTTGAGACTGTGCAACGCCTCCTTCCCGTGCTGTGCCAAGATCATGGGCTGACACCAGCCCAAGTCGTCGCTATCGCAAGTAATATCGGCGGGAAGCAAGCTTTGGAAACAGTGCAACGCTTGTTGCCAGTTCTGTGCCAAGCCCACGGTCTGACTCCGGAACAGGTGGTGGCGATTGCAAGCAACGGCGGCGGCAAACAGGCTCTAGAGAGCATTGTTGCCCAGCTCTCCAGACCTGATCCGGCGCTAGCCGCGTTGACAAATGATCATCTTGTTGCTCTTGCTTGTCTTGGCGGACGACCTGCCCTGGATGCTGTGAAAAAGGGGTTGGGG。

[0170] LspsaA site 1 Right 1 TALE (SEQ ID NO: 45): ACCCGATCAAGTTGTTGCTATCGCTAGCCATGATGGAGGGAAACAAGCCCTTGAGACTGTGCAACGGCTGCTTCCAGTGTTGTGCCAAGCTCATGGACTTACTCCCGATCAGGTCGTGGCTATTGCATCAAATGGTGGTGGCAAACAAGCACTGGAAACCGTTCAAAGGTTGCTTCCTGTTCTGTGTCAGGACCACGGTCTGACTCCGGAACAGGTGGTGGCGATTGCAAGCAACGGCGGCGGCAAACAGGCTCTAGAGAGCATTGTTGCCCAGCTCTCCAGACCTGATCCGGCGCTAGCCGCGTTGACAAATGATCATCTTGTTGCTCTTGCTTGTCTTGGCGGACGACCTGCCCTGGATGCTGTGAAAAAGGGGTTGGGG。

[0171] LspsaA site 1 Right 2 TALE (SEQ ID NO: 47): TCCCGAACAAGTTGTTGCTATTGCCTCTAATGGAGGGGGAAACAAGCTGGAAACCGTTCAAAGGTTGCTCCCTGTCCTCTGCCAGGCACATGGTCTGACCCCAGCCCAGGTCGTGGCTATTGCCAGTAATGGGGGCGGCAAGCAGGCCCTTGAGACTGTCCAGCCAGCTGTTGTTGTTGTT ATCACGGTCTGACTCCGGAACAGGTGGTGGCGATTGCAAGCAACGGCGGCGGCAAACAGGCTCTAGAGAGCATTGTTGCCCAGCTCCAGACCTGATCCGGCGCTAGCCGCGTTGACAAATGATCATCTTGTTGCTCTTGTCTTGGCGGACGACCTGCCCTGAACAGGGTGGGAGGGGAGTT

[0172] LspsaA (LspsaA site 2 Left TALE) (SEQ ID NO:49): CCCCGATCAGGTCGTCCGCAATCGCATCTAACATCGGGGGCAAGCAAGCACTGGAAACAGTGCAGAGGCTCTTGCCCGTTCTCTGTCAAGACCACGGACTTACCCCCGACCAGGTGGTCGCAATCGCCTCCAACATAGGTGGAAAACAGGCTCTCGAGACAGTTCAAAGGTTGCTCCCTGTTGTTGTTAG CTCACGGTCTGACTCCGGAACAGGTGGTGGCGATTGCAAGCAACGGCGGCGGCAAACAGGCTCTAGAGAGCATTGTTGCCCAGCTCCAGACCTGATCCGGCGCTAGCCGCGTTGACAAATGATCATCTGTTGCTCTTGCTTGGCTTGGCGGACGACCTGCCCTGAACAGGTGGTGGGAGGGATT

[0173] LspsaA site 2 Right TALE (SEQ ID NO:51): ACCAGCCCAAGTCGTTGCCATCGCCAGCAACAATGGGGGCAAGCAAGCTCTCGAAACCGTTCAGAGACTTCTGCCCGTGTTGTGCCAAGATCATGGGTTGACCCCAGCTCAGGTGGTCGCAATTGCTTCAAATAACGGGGGCAAGCAAGCACTGGAGACTGTTCAACGCCTCCTGCCCGTCCTCTGCCAAGCACACGGTCTGACTCCGGAACAGGTGGTGGCGATTGCAAGCAACGGCGGCGGCAAACAGGCTCTAGAGAGCATTGTTGCCCAGCTCTCCAGACCTGATCCGGCGCTAGCCGCGTTGACAAATGATCATCTTGTTGCTCTTGCTTGTCTTGGCGGACGACCTGCCCTGGATGCTGTGAAAAAGGGGTTGGGG。

[0174] LspsbA Left 1 TALE (SEQ ID NO: 53): ACCTGAGCAGGTGGTTGCAATTGCCAGCAATGGTGGAGGCAAACAAGCTCTGGAGACAGTGCAGAGACTTTTGCCTGTCCTTTGCCAGGCCCACGGATTGACCCCAGACCAGGTTGTCGCTATTGCATCACATGACGGTGGCAAGCAAGCTCTCGAAACTGTCCAGAGATTGCTCCCTGTCTTGTGTCAAGCACACGGTCTGACTCCGGAACAGGTGGTGGCGATTGCAAGCAACGGCGGCGGCAAACAGGCTCTAGAGAGCATTGTTGCCCAGCTCTCCAGACCTGATCCGGCGCTAGCCGCGTTGACAAATGATCATCTTGTTGCTCTTGCTTGTCTTGGCGGACGACCTGCCCTGGATGCTGTGAAAAAGGGGTTGGGG。

[0175] LspsbA Left 2 TALE (SEQ ID NO: 55): CCCAGAGCAAGTGGTGGCCATCGCTTCAAACATAGGAGGAAAGCAAGCCCTCGAGACTGTTCAAAGACTGCTTCCCGTTCTCTGTCAAGCACATGGGTTGACTCCAGCACAGGTTGTTGCCATTGCTAGTAATATTGGTGGTAAACAGGCATTGGAAACTGTGCAAAGACTCCTTCCTGTCCTGTGTCAGGACCACGGTCTGACTCCGGAACAGGTGGTGGCGATTGCAAGCAACGGCGGCGGCAAACAGGCTCTAGAGAGCATTGTTGCCCAGCTCTCCAGACCTGATCCGGCGCTAGCCGCGTTGACAAATGATCATCTTGTTGCTCTTGCTTGTCTTGGCGGACGACCTGCCCTGGATGCTGTGAAAAAGGGGTTGGGG。

[0176] LspsbA Right 1 TALE (SEQ ID NO: 57): TCCAGACCAGGTCGTTGCTATTGCCAGCAATAACGGAGGTAAGCAAGCATTGGAGACTGTTCAGCGGCTCCTCCCTGTCCTCTGTCAGGCCCACGGGCTGACTCCTGCACAGGTGGTGGCCATCGCTTCAAACGGGGGGGGGAAGCAAGCCCTCGAAACTGTTCAACGGTTGTTGCCTGTTTTGTGTCAGGATCACGGTCTGACTCCGGAACAGGTGGTGGCGATTGCAAGCAACGGCGGCGGCAAACAGGCTCTAGAGAGCATTGTTGCCCAGCTCTCCAGACCTGATCCGGCGCTAGCCGCGTTGACAAATGATCATCTTGTTGCTCTTGCTTGTCTTGGCGGACGACCTGCCCTGGATGCTGTGAAAAAGGGGTTGGGG。

[0177] LspsbA Right 2 TALE (SEQ ID NO: 59): CCCTGCTCAAGTTGTTGCAATCGCTTCCAATAACGGCGGTAAACAAGCCCTTGAGACAGTGCAAAGGCTCTTGCCAGTGCTCTGTCAAGACCACGGTCTGACTCCGGAACAGGTGGTGGCGATTGCAAGCAACGGCGGCGGCAAACAGGCTCTAGAGAGCATTGTTGCCCAGCTCTCCAGACCTGATCCGGCGCTAGCCGCGTTGACAAATGATCATCTTGTTGCTCTTGCTTGTCTTGGCGGACGACCTGCCCTGGATGCTGTGAAAAAGGGGTTGGGG。

[0178] LspsbA Right 3 TALE (SEQ ID NO: 61): CCCCGATCAGGTCGTCGCAATCGCATCTAACATCGGGGGCAAGCAAGCACTGGAAACAGTGCAGAGGCTCTTGCCCGTTCTCTGTCAAGACCACGGACTTACCCCCGACCAGGTGGTCGCAATCGCCTCCAACATAGGTGGAAAACAGGCTCTCGAGACAGTTCAAAGGTTGCTCCCTGTGTTGTGCCAAGCTCACGGTCTGACTCCGGAACAGGTGGTGGCGATTGCAAGCAACGGCGGCGGCAAACAGGCTCTAGAGAGCATTGTTGCCCAGCTCTCCAGACCTGATCCGGCGCTAGCCGCGTTGACAAATGATCATCTTGTTGCTCTTGCTTGTCTTGGCGGACGACCTGCCCTGGATGCTGTGAAAAAGGGGTTGGGG。<​​​CCCTGCACAGGTGGTTGCTATTGCCTCCAACGGTGGAGGCAAACAGGCCCTTGAGACTGTGCAACGCCTCCTTCCCGTGCTGTGCCAAGATCATGGGCTGACACCAGCCCAAGTCGTCGCTATCGCAAGTAATATCGGCGGGAAGCAAGCTTTGGAAACAGTGCAACGCTTGTTGCCAGTTCTGTGCCAAGCCCACGGTCTGACTCCGGAACAGGTGGTGGCGATTGCAAGCAACGGCGGCGGCAAACAGGCTCTAGAGAGCATTGTTGCCCAGCTCTCCAGACCTGATCCGGCGCTAGCCGCGTTGACAAATGATCATCTTGTTGCTCTTGCTTGTCTTGGCGGACGACCTGCCCTGGATGCTGTGAAAAAGGGGTTGGGG。

[0180] AtpsaA Right TALE (SEQ ID NO: 65): TCCTGCTCAGGTTGTGGCCATTGCCAGCCACGATGGGGGTAAGCAAGCACTTGAAACAGTTCAAAGACTGCTTCCCGTGCTTTGTCAGGCACACGGGCTGACTCCCGCACAAGTCGTCGCCATCGCCTCACATGACGGAGGCAAACAAGCACTGGAGACAGTTCAACGCCTCCTCCCTGTCTTGTGCCAAGACCACGGTCTGACTCCGGAACAGGTGGTGGCGATTGCAAGCAACGGCGGCGGCAAACAGGCTCTAGAGAGCATTGTTGCCCAGCTCTCCAGACCTGATCCGGCGCTAGCCGCGTTGACAAATGATCATCTTGTTGCTCTTGCTTGTCTTGGCGGACGACCTGCCCTGGATGCTGTGAAAAAGGGGTTGGGG。

[0181] AtpsbA Left TALE (SEQ ID NO: 67): CCCAGAGCAAGTGGTGGCCATCGCTTCAAACATAGGAGGAAAGCAAGCCCTCGAGACTGTTCAAAGACTGCTTCCCGTTCTCTGTCAAGCACATGGGTTGACTCCAGCACAGGTTGTTGCCATTGCTAGTAATATTGGTGGTAAACAGGCATTGGAAACTGTGCAAAGACTCCTTCCTGTCCTGTGTCAGGACCACGGTCTGACTCCGGAACAGGTGGTGGCGATTGCAAGCAACGGCGGCGGCAAACAGGCTCTAGAGAGCATTGTTGCCCAGCTCTCCAGACCTGATCCGGCGCTAGCCGCGTTGACAAATGATCATCTTGTTGCTCTTGCTTGTCTTGGCGGACGACCTGCCCTGGATGCTGTGAAAAAGGGGTTGGGG。

[0182] AtpsbA Right TALE (SEQ ID NO: 69): CCCCGATCAGGTCGTCGCAATCGCATCTAACATCGGGGGCAAGCAAGCACTGGAAACAGTGCAGAGGCTCTTGCCCGTTCTCTGTCAAGACCACGGACTTACCCCCGACCAGGTGGTCGCAATCGCCTCCAACATAGGTGGAAAACAGGCTCTCGAGACAGTTCAAAGGTTGCTCCCTGTGTTGTGCCAAGCTCACGGTCTGACTCCGGAACAGGTGGTGGCGATTGCAAGCAACGGCGGCGGCAAACAGGCTCTAGAGAGCATTGTTGCCCAGCTCTCCAGACCTGATCCGGCGCTAGCCGCGTTGACAAATGATCATCTTGTTGCTCTTGCTTGTCTTGGCGGACGACCTGCCCTGGATGCTGTGAAAAAGGGGTTGGGG。

[0183] Example 2: Plant Transformation and Growth Conditions The *Agrobacterium tumefaciens* strain GV3101 was transformed using the vector cloned in Example 1, and *Arabidopsis thaliana* plants were transformed using the floral dip method according to known methods (Zhang et al., *Nat. Protoc. 1*, 641-646 (2006)). *Arabidopsis* T1 seedlings were screened in 1 / 2 strength MS medium supplemented with 2% sucrose, 20 mg / L glufosinate-ammonia (PPT), and 250 mg / L cefotaxime. To evaluate resistance to atrazine or PPT, T2 or T3 seeds were sown in 1 / 2 strength MS medium containing 2% sucrose with or without atrazine (1 mg / L) and with or without PPT (20 mg), and allowed to grow for 2 weeks. All transgenic plants were cultivated under long-day conditions at 23°C (16 hours of light / 8 hours of darkness).

[0184] Example 3: Protoplast isolation and transfection Lettuce seeds (Lactuca sativa L., Cheongchima variety, Nongwoo Biotechnology Co., Ltd., South Korea) were first treated with 70% ethanol for 3 minutes, followed by sterilization with 0.4% sodium hypochlorite (commercially available, Clorox) for 10 minutes. After sterilization, the seeds were rinsed 5 times with sterile water and sown in 1 / 2 strength MS (Murashige and Skoog) basal medium supplemented with 0.4 mg / L thiamine, 100 mg / L inositol, 30 g / L sucrose, and 4 g / L gellan gum. The pH of the medium was adjusted to 5.7 using 1N NaOH. The seeds were allowed to grow for 6 days at 23°C under 16 hours of light / 8 hours of darkness. Protoplasts were isolated from the lettuce mother seed according to known methods. The protoplast isolation process included collecting the cotyledons of 6-day-old seedlings and decomposing them with a cell wall-degrading enzyme solution. The decomposition process is carried out in a dark environment at 25°C, with gentle shaking for 4-5 hours (50 rpm). The enzyme solution contained 1% Viscozyme, 0.5% Celluclast, 0.5% Pectinex, 9% mannitol, and 3 mM 2-(N-morpholino)ethanesulfonic acid (MES) as in the CPW solution. After enzyme treatment, the protoplast mixture was filtered through a nylon mesh (40 μm pore size) and then centrifuged at 100 g for 5 minutes in a round-bottom tube to collect the protoplasts. The purified protoplasts were washed with W5 solution (2 mM MES, 154 mM NaCl, 125 mM CaCl2, and 5 mM KCl, pH 5.7) and centrifuged at 100 g for 5 minutes to granulate. The protoplasts were then resuspended in W5 solution and counted under a microscope using a hemocytometer. The protoplasts were diluted in MMG solution (containing 0.4 M mannitol, 15 mM MgCl2, and 4 mM MES, pH 5.7). In 5.7), its density is made to reach 2.5 × 10⁻⁶. 6 Protoplasts / mL. For transfection, 5 × 10⁵ protoplasts were resuspended in 200 μl of MMG solution. 5One protoplast was transfected into a plasmid (30 μg / constructor) using a PEG solution (containing 40% w / v PEG 4000, 0.2 M mannitol, and 0.1 M CaCl2) and cultured for 10 minutes. Next, 400 μl of W5 solution was added and mixed, and the mixture was cultured for another 10 minutes. Then, 800 μl of W5 solution was added, and the protoplasts were collected by centrifugation at 80 g for 5 minutes. The protoplast mixture was then gently washed with 1 ml of W5 solution by inversion and centrifugation at 80 g for 5 minutes to further granulate it. The transfected protoplasts were resuspended in 1 / 2 B5 medium in 2 ml of protoplast medium containing 375 mg / L CaCl2·2H2O, 18.35 mg / L NaFe-EDTA, and 0.1 g / L MES. The mixture was then cultured in the dark at 25°C for 7 days before targeted deep sequencing was performed.

[0185] Example 4: Targeted deep sequencing for analyzing base editing frequencies For targeted deep sequencing, total DNA was extracted from actual leaves of selected plants or transfected lettuce protoplasts using the DNeasy Plant Mini kit (Qiagen). For large-scale analysis, DNA was extracted using 100 μL of cell lysis buffer (20 mM Tris-HCl, pH 8.0, Sigma-Aldrich, 5 mM EDTA, 400 mM NaCl, 0.05% sodium dodecyl sulfate) containing 2 μL of proteinase K (Qiagen). The lysates were cultured at 55°C for 2 hours, followed by incubation at 95°C for 10 minutes. Deep sequencing libraries were prepared by amplifying the target site using three PCR steps (first, second, and third) with TruSeq HT dual-index primers and PrimeSTAR® GXL DNA polymerase (TAKARA). Sequencing of the libraries was performed using Illumina. MiSeq paired-end sequencing technology was used. Base editing efficiency (frequency) was calculated as the percentage of sequencing reads that presented the desired base editing result out of all sequencing reads. The PCR primer sequences used for sequencing are shown in Table 2.

[0186] Table 2

[0187] The base editor fusion protein used has the following composition: CTS-3Xflag-LspsaA site 1 Left TALE (LspsaA site 1 Left TALE)-2aa lin-TadA-CDd-4aa lin 2-UGI-4aa lin 2-UGI.

[0188] CTS-3Xflag-LspsaA site 1 Left TALE (LspsaA site 1 Left TALE)-2aa lin-rApobec1-4aa lin 2-UGI-4aa lin 2-UGI.

[0189] CTS-3Xflag-LspsaA site 1 Left TALE (LspsaA site 1 Left TALE)-2aa lin-TadA8eE27R N46L-4aa lin 2-UGI-4aa lin 2-UGI.

[0190] CTS-3Xflag-LspsaA site 1 Left TALE (LspsaA site 1 Left TALE)-18aa lin-TadA-CDd-4aa lin 2-UGI-4aa lin 2-UGI.

[0191] CTS-3Xflag-LspsaA site 1 Left TALE (LspsaA site 1 Left TALE)-18aa lin-rApobec 1-4aa lin 2-UGI-4aa lin 2-UGI.

[0192] CTS-3Xflag-LspsaA site 1 Left TALE (LspsaA site 1 Left TALE)-18aa lin-TadA8eE27R N46L-4aa lin 2-UGI-4aa lin 2-UGI.

[0193] CTS-3Xflag-LspsaA site 1 Left TALE (LspsaA site 1 Left TALE)-2aa lin-TadA-Dual-4aa lin 2-UGI-4aa lin 2-UGI.

[0194] CTS-3Xflag-LspsaA site 1 Left TALE (LspsaA site 1 Left TALE)-2aa lin-TadA8e.

[0195] CTS-3Xflag-LspsaA site 1 Left TALE - 18aa lin-TadA8e.

[0196] CTS-3Xflag-LspsaA site 1 Right 1 TALE - 4aa lin-MutH.

[0197] CTS-3Xflag-LspsaA site 1 Right 2 TALE - 4aa lin-MutH.

[0198] CTS-3Xflag-LspsaA site 1 Left TALE - 4aa lin-1397N-4aa lin-UGI.

[0199] CTS-3Xflag-LspsaA site 1 Left TALE - 4aa lin-1397C-4aa lin-UGI.

[0200] CTS-3Xflag-LspsaA site 1 Right 1 TALE - 4aa lin-1397N-4aa lin-UGI.

[0201] CTS-3Xflag-LspsaA site 1 Right 1 TALE - 4aa lin-1397C-4aa lin-UGI.

[0202] CTS-3Xflag-LspsbA site Left 1 TALE - 2aa lin-TadA8e.

[0203] CTS-3Xflag-LspsbA site Left 2 TALE - 2aa lin-TadA8e.

[0204] CTS-3Xflag-LspsbA site Left 1 TALE - 2aa lin-TadA8e.

[0205] CTS-3Xflag-LspsbA site Left 2 TALE - 2aa lin-TadA8e.

[0206] CTS-3Xflag-LspsbA site Left TALE-2aa lin- TadA-CDd-4aa lin 2-UGI-4aa lin 2-UGI。

[0207] CTS-3Xflag-LspsbA site Right 1 TALE-4aa lin-Nt.BspD6I(C)。

[0208] CTS-3Xflag-LspsbA site Right 2 TALE-4aa lin-Nt.BspD6I(C)。

[0209] CTS-3Xflag-LspsbA site Right 3 TALE-4aa lin-Nt.BspD6I(C)。

[0210] CTS-3Xflag-LspsbA site Left 2 TALE- TadA-CDd-4aa lin 2-UGI-4aa lin 2-UGI。

[0211] CTS-3Xflag-LspsbA site Right 3 TALE-2aa lin- TadA8e -16aa lin-Nt.BspD6I(C)。

[0212] CTS-3Xflag-LspsaA site 2 Left TALE-4aa lin-1397N。

[0213] CTS-3Xflag-LspsaA site 2 Left TALE-4aa lin-1397C-16aa lin-TadA8e。

[0214] CTS-3Xflag-LspsaA site 2 Right TALE-4aa lin-1397N。

[0215] CTS-3Xflag-LspsaA site 2 Right TALE-4aa lin-1397C-16aa lin-TadA8e。

[0216] CTS-3Xflag-LspsaA site 2 Left TALE (LspsaA site 2 Left TALE)-4aa lin-1397N.

[0217] CTS-3Xflag-LspsaA site 2 Left TALE (LspsaA site 2 Left TALE)-4aa lin-1397C-16aa lin-TadA8e.

[0218] CTS-3Xflag-AtpsaA site (site) Left (Left) TALE-2aa lin-TadA-CDd-4aa lin 2-UGI-4aa lin 2-UGI.

[0219] CTS-3Xflag-AtpsaA site (site) Left (Left) TALE-2aa lin-rApobec 1-4aa lin 2-UGI-4aa lin 2-UGI.

[0220] CTS-3Xflag-AtpsaA site (site) right (Right) TALE-4aa lin-MutH.

[0221] CTS-3Xflag-AtpsaA site (site) Right (Right) TALE-17aa lin-Fok I E490K I538K-40aa lin-FoK I D450A Q486A I499L.

[0222] CTS-3Xflag-AtpsbA site (site) left (Left) TALE-2aa lin-TadA8e.

[0223] CTS-3Xflag-AtpsbA site (site) right (Right) TALE-4aa lin-Nt.BspD6I (C).

[0224] The sequences of the protein components used in the fusion protein are as follows: CTS (SEQ ID NO: 2): MDSQLVLSLKLNPSFTPLSPLFPFTPCSSFSPSLRFSSCYSRRLYSPVTVYAAK.

[0225] 3xflag (SEQ ID NO: 4): DYKDHDGDYKDHDIDYKDDDDK.

[0226] MutH (SEQ ID NO: 6): MSQPRPLLSPPETEEQLLAQAQQLSGYTLGELAALAGLVTPENLKRDKGWIGVLLEIWLGASAGSKPEQDFAALGVELKTIPVDSLGRPLETTFVCVAPLTGNSGVTWETSHVRHKLKRVLWIPVEGERSIPLAKRRVGSPLLWSPNEEEDRQLREDWEELMDMIVLGQIERITARHGEYLQIRPKAANAKALTEAIGARGERILTLPRGFYLKKNFTSALLARHFLIQ. <0oo00599>Nt.BspD6I (C) (SEQ ID NO: 8): MRQLEEVIDLLEVYHEKKNVIEEKIKARFIANKNTVFEWLTWNGFIILGNALEYKNNFVIDEELQPVTHAAGNQPDMEIIYEDFIVLGEVTTSKGATQFKMESEPVTRHYLNKKKELEKQGVEKELYCLFIAPEINKNTFEEFMKYNIVQNTRIIPLSLKQFNMLLMVQKKLIEKGRRLSSYDIKNLMVSLYRTTIECERKYTQIKAGLEETLNNWVVDKEVRF. <ooo00602>2aa linker: GS.

[0229] 4aa linker 1 (SEQ ID NO: 10): GSGS.

[0230] 4aa linker 2 (SEQ ID NO: 12): SGGS.

[0231] 4aa linker 3 (SEQ ID NO: 14): LVGS.

[0232] 16aa linker (SEQ ID NO: 16): SGSETPGTSESATPES.

[0233] 17aa linker (SEQ ID NO: 18): It should be noted that there are some "ooo00" in the original text which seem to be incorrect. I have translated them as "ooo00" as they are not clear what they should be. If they are actual errors, you may need to correct them in the original text for a more accurate translation.SGGGGSGGGGSGGGGSS。

[0234] 18aa linker (Linker) (SEQ ID NO: 20): GSSGSETPGTSESATPES。

[0235] TadA-CDd (SEQ ID NO: 22): MSEVEFSHEYWMRHALTLAKRARDERKAPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIIALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMINSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN。

[0236] rApobec 1 (SEQ ID NO: 24): SSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSIWRHTSQNTNKHVEVNFIEKFTTERYFCPNTRCSITWFLSWSPCGECSRAITEFLSRYPHVTLFIYIARLYHHADPRNRQGLRDLISSGVTIQIMTEQESGYCWRNFVNYSPSNEAHWPRYPHLWVRLYVLELYCIILGLPPCLNILRRKQPQLTFFTIALQSCHYQRLPPHILWATGLKSGSETPGTSESATPES。

[0237] TadA E27R N46L (SEQ ID NO: 26): SEVEFSHEYWMRHALTLAKRARDERRVPVGAVLVLNNRVIGEGWLRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN。

[0238] TadA8e (SEQ ID NO: 28): SEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN。

[0239] TadA - dual (SEQ ID NO: 30): MSEVEFSHEYWMRHALTLAKRARDEGEAPVGAVLVLNNRVIGEGWNRRIGLHDPTAHAEIIALRQGGLVMQNSRLIDATLYVTFEPCVMCAGAMINSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN。

[0240] FokI E490K I538K - 40aa - FokI D450A Q486E I499L (SEQ ID NO: 32): LVKSELEEKKSELRHKLKYVPHEYIELIEIARNSTQDRILEMKVMEFFMKVYGYRGKHLGGSRKPDGAIYTVGSPIDYGVIVDTRAYSGGYNLPIGQADEMQRYVKENQTRNKHINPNEWWKVYPSSVTEFKFLFVSGHFKGNYKAQLTRLNHKTNCNGAVLSVEELLIGGEMIKAGTLTLEEVRRKFNNGEINFSSGGGGSGGGGSGGGGSGGSGGGSSGGGGSGGGGSGGGGSLVKSELEEKKSELRHKLKYVPHEYIELIEIARNSTQDRILEMKVMEFFMKVYGYRGKHLGGSRKPAGAIYTVGSPIDYGVIVDTRAYSGGYNLPIGQADEMERYVEENQTRNKHLNPNEWWKVYPSSVTEFKFLFVSGHFKGNYKAQLTRLNHITNCNGAVLSVEELLIGGEMIKAGTLTLEEVRRKFNNGEINF。

[0241] DddA tox [[ID= GSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPEG。

[0242] DddA tox 1397C(SEQ ID NO:36): GSAIPVKRGATGETKVFTGNSNSPKSPTKGGC。

[0243] GSVG(SEQ ID NO:38): GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLEGKVFSSGGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPEGVIPVKRGATGETKVFTGNSNGPKSPTKGGC。

[0244] 2x UGI(SEQ ID NO:40): TNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKMLSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIK

[0245] UGI (SEQ ID NO:42): TNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML。

[0246] LspsaA (LspsaA site 1 Left TALE) (SEQ ID NO:44): 。

[0247] LspsaA site 1 Right 1 NEW(SEQ ID NO:46): 。

[0248] LspsaA site 1 Right 2 STEP(SEQ ID NO:48): 。

[0249] LspsaA site 2 Left TALE(SEQ ID NO:50): 。

[0250] LspsaA site 2 Right TALE (SEQ ID NO:52): DLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLG。

[0251] LspsbA Left 1 TALE (SEQ ID NO: 54): DLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLG。

[0252] LspsbA Left 2 TALE (SEQ ID NO: 56): DLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLG。

[0253] LspsbA Right 1 TALE (SEQ ID NO: 58): DLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLG。

[0254] LspsbA Right 2 TALE (SEQ ID NO: 60): DLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLG。

[0255] LspsbA Right 3 TALE (SEQ ID NO: 62): DLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLG。

[0256] AtpsaA Left TALE (SEQ ID NO: 64): DLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLG。

[0257] AtpsaA Right TALE (SEQ ID NO: 66): DLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLG。

[0258] AtpsbA Left (Left) TALE (AtpsbA Left TALE) (SEQ ID NO: 68): DLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLG。

[0259] AtpsbA Right TALE (SEQ ID NO: 70): .

[0260] The base editor of this invention, targeting the lettuce psaA gene, has been shown to effectively achieve C-to-T and A-to-G editing when MutH is used as a nicking enzyme. Figures 2 to 5 The target DNA sequence of the base editor used is as follows: Figure 1As shown. Specifically, when TadA-CDd, rApobec, and TadA E27R N46L were used as cytosine deaminases, C-to-T base editing at the target site was significantly more pronounced than in the control group, regardless of whether two amino acids or 18 amino acids were used as linkers. Figure 2 and Figure 3 Even when TadA8e is used as a deaminase, significant A-to-G base editing was confirmed at the target site in both cases using a two-amino acid linker and an 18-amino acid linker. Figure 4 ).

[0261] On the other hand, it was confirmed when using the base editor of the present invention that, compared with using DddA tox Unlike existing base editors, this one not only addresses the 5'-TC-3' motif but also edits cytosine (C) bases in the 5'-GC-3' motif into guanine bases. Figure 5 Specifically, such as Figure 5 As shown, among all the base editors used, the cytosine (C) bases of T7-C8 were edited to guanine, but only when using the base editor of the present invention (using MutH) was a significant editing of the cytosine (C) bases of G3-C4 (cytosine on the opposite strand of G3 and C4) confirmed compared to the control group.

[0262] Figure 6 This involves identifying other target sites of the psaA gene in lettuce (see...). Figure 6 The result of editing an A-to-G base for an object, when using an existing base editor (using DddA) tox Although the editing efficiency at T5 and T13 sites is low, the nicking enzyme base editor of this invention (using BspD6I) exhibits significantly superior base editing efficiency. Figure 6 ).

[0263] Furthermore, even when Nt.BspD6I(C) was used as a nicking enzyme, it was confirmed that targeted A-to-G and C-to-T editing of the lettuce psbA gene could be effectively achieved. Figures 8 to 12 The target DNA sequence of the base editor used is as follows: Figure 7 As shown. Specifically, when TadA8e was used as a deaminase, the A-to-G editing effect was confirmed in all experimental groups using amino acid linkers. Figure 8 and Figure 9 When TadA-CDd was used as a deaminase, C-to-T editing effects were also confirmed at the target site. Figure 10 ).

[0264] On the other hand, according to the present invention, when nicking enzyme and deaminase are used together, dual editing of A-to-G and C-to-T editing has been successfully confirmed. When cytosine deaminase (TadA-CDd) and adenine deaminase (TadA8e) are used together, A-to-G and C-to-T editing occur simultaneously at the target site of the psaA gene. Figure 11 When using the TadA8e variant with SEQ ID NO: 30, simultaneous A-to-G and C-to-T editing at the target site of the psbA gene was confirmed, even in the absence of other deaminases. Figure 12 ).

[0265] Even when the base editor of this invention is expressed in Arabidopsis thaliana via Agrobacterium, A-to-G and C-to-T editing are effectively performed. Using the target site of the psaA gene in Arabidopsis thaliana as the target, the base editor of this invention, employing cytosine deaminase (TadA-CDd or rApobec) and nickase (MutH), is applicable. Figure 13 The psaA gene encodes a protein that forms part of the photosystem I complex. When the C in the GATC sequence recognized by MutH is edited to T, it induces the production of a TAA base sequence corresponding to the stop codon, resulting in a light green leaf phenotype. That is, since the light green leaf phenotype indicates that C-to-T editing has occurred correctly, screening for T1 plants that have achieved base editing based on this phenotype showed an average editing efficiency of approximately 74.2% when TadA-CDd was used as the cytosine deaminase. Figure 13 and Figure 14 When rApobec is used as a cytosine deaminase, the editing efficiency is on average approximately 77.8%. Figure 13 and Figure 15 When used as a nickase, a heterodimer composed of FokI variants also achieved excellent C-to-T editing, with average base editing frequencies of approximately 56.1% (average editing efficiency at the G5 position) and 26.3% (average editing efficiency at the C8 position) confirmed in multiple individuals. Figure 16 and Figure 17 ).

[0266] Using the psaA gene of Arabidopsis thaliana as an example, when the base editor of this invention (using Nt-BspD6I(C) as a nicking enzyme and TadA8e as a deaminase) was used, A-to-G base editing at the target site was confirmed. Figure 18 and Figure 19 ).

[0267] The psbA gene encodes a D1 protein, an important subunit of the photosystem II complex. In various plant species, resistance to atrazine herbicides is often associated with the substitution of glycine for serine at position 264 in the D1 protein. A new TALE protein was constructed, which binds to the DNA of both the LspsaA left (Left) 2 TALE protein and the LspsaA right (Right) TALE protein targeted by the psbA gene in lettuce at the same sequence. This protein was introduced into Arabidopsis thaliana via a binary vector. T1 plants were screened in 1 / 2 strength MS medium supplemented with glufosinate (20 mg / L) under long-day conditions (16 hours light / 8 hours dark). Targeted deep sequencing confirmed the A-to-G editing frequency in the DNA base sequence corresponding to amino acid S264. Figure 18 Then, T2 seeds were collected from T1 plants and cultured in 1 / 2 strength MS medium containing atrazine (1 mg / L) to obtain 8 progeny with atrazine resistance. The base editing frequency was then confirmed by targeted deep sequencing. Figure 19 Next, under the same conditions as T2 seeds, T3 seeds of T2 plants were cultivated, and... Figure 20 The image shows plants exhibiting resistance to phosphonotricine (20 mg / L, also known as Basta) (plant #2-1) and resistance to atrazine (plants #2-3 and #2-4). Targeted deep sequencing of the atrazine-resistant T3 plants showed that editing the 6 bases of the a portion encoding amino acid S264 to G bases resulted in over 99% homogeneity. Figure 21 ).

[0268] Example 5: Genotyping analysis for confirming the presence of exogenous genes Whole DNA was extracted from actual leaves of T3 plants using the DNeasy Plant Mini kit (Qiagen). Regions of interest were then amplified throughout the DNA using PrimeSTAR® GXL DNA polymerase (TAKARA). The resulting PCR products were analyzed on a 1% agarose gel. The PCR primer sequences used are shown in Table 3.

[0269] To confirm that the atrazine-resistant lines #2-3-1 and #2-4-1 did not contain a foreign gene, genomic PCR was performed using primers specific to the asta resistance gene (part b) contained in the foreign gene. The results confirmed that, unlike #2-1-1, #2-3-1 and #2-4-1 plants did not contain the foreign gene. Figure 22This result demonstrates that base editing induced by the nicking enzyme-dependent base editor of this invention can be stably inherited and maintained.

[0270] Table 3

Claims

1. A method for base editing of plant DNA, comprising the step of expressing a DNA base editing composition in a target plant, plant cell, or protoplast, characterized in that, The base editing composition comprises one or more DNA-binding proteins and one or more enzyme proteins, or comprises a polynucleotide encoding said protein. The one or more enzyme proteins include nicking enzymes, and also include adenine deaminase and / or cytosine deaminase.

2. The base editing method according to claim 1, characterized in that, The DNA base editing composition comprises one or more fusion proteins or one or more polynucleotides encoding one or more fusion proteins. Each of the more than one fusion proteins independently contains a DNA-binding protein. At least one of the more than one fusion protein contains a nicking enzyme. At least one of the more than one fusion protein contains adenine deaminase and / or cytosine deaminase.

3. The base editing method according to claim 2, characterized in that, At least one of the more than one fusion protein contains a cytosine deaminase. The cytosine deaminase is an apolipoprotein B editing complex, an activation-induced deaminase, or a tRNA-specific adenosine deaminase, or a variant thereof.

4. The base editing method according to claim 3, characterized in that, Cytosine deaminase has deaminase activity for single-stranded DNA.

5. The base editing method according to claim 3, characterized in that, Cytosine deaminase is a tRNA-specific adenosine deaminase or a variant thereof.

6. The base editing method according to claim 3, characterized in that, Cytosine deaminase comprises an amino acid sequence selected from the group consisting of the amino acid sequences of SEQ ID NO: 22, SEQ ID NO: 24, SEQ ID NO: 26 and SEQ ID NO: 30, or a conserved amino acid substitute thereof.

7. The base editing method according to claim 2, characterized in that, At least one of the more than one fusion protein contains an adenine deaminase, which is a tRNA-specific adenosine deaminase or a variant thereof.

8. The base editing method according to claim 7, characterized in that, Adenine deaminase contains the amino acid sequence of SEQ ID NO: 28 or SEQ ID NO: 30 or its conserved amino acid substitutes.

9. The base editing method according to any one of claims 1 to 8, characterized in that, The nicking enzymes were selected from the group consisting of MutH, BspD6I, FokI, BsaI, BsmBI, BsmAI, BsrDI, CviPII, BspQI, AlwI, and I-TevI, their fragments, and their variants. The FokI comprises FokI-FokI homodimer or heterodimer.

10. The base editing method according to claim 9, characterized in that, The nicking enzyme comprises an amino acid sequence selected from the group consisting of SEQ ID NO: 6, SEQ ID NO: 8 and SEQ ID NO: 32, or a conserved amino acid substitute thereof.

11. The base editing method according to claim 1 or 2, characterized in that, Each of the aforementioned DNA-binding proteins is independently selected from the group consisting of zinc finger proteins, transcription activator-like effector proteins, and CRISPR-related nucleases.

12. The base editing method according to claim 1 or 2, characterized in that, One or more of adenine deaminase, cytosine deaminase, and nickase are expressed by a polynucleotide whose codons have been optimized for expression in plants.

13. The base editing method according to any one of claims 1 to 12, characterized in that, DNA is the organelle DNA of plants.

14. The base editing method according to any one of claims 3 to 6, characterized in that, Fusion proteins containing cytosine deaminase also contain uracil glycosylation inhibitors.

15. The base editing method according to claim 2, characterized in that, Each of the aforementioned fusion proteins also includes an extranuclear transport signal.

16. The base editing method according to claim 2, characterized in that, Each of the more than one fusion protein also contains a mitochondrial targeting sequence.

17. The base editing method according to claim 2, characterized in that, Each of the aforementioned fusion proteins further comprises a chloroplast transport signal.

18. The base editing method according to claim 2, characterized in that, Each of the more than one fusion protein contains a nuclear localization signal.

19. The base editing method according to claim 2, characterized in that, The DNA base editing composition comprises two fusion proteins or two polynucleotides encoding two fusion proteins, one of which contains adenine deaminase and / or cytosine deaminase, and the other contains a nicking enzyme.

20. The base editing method according to claim 2, characterized in that, DNA base editing compositions contain two fusion proteins or two polynucleotides encoding two fusion proteins. Of the two fusion proteins, (1) one contains adenine deaminase and the other contains cleavage enzyme and cytosine deaminase, or (2) one contains cytosine deaminase and the other contains cleavage enzyme and adenine deaminase.

21. The base editing method according to claim 2, characterized in that, The DNA base editing composition comprises a fusion protein containing a nicking enzyme and adenine deaminase and / or cytosine deaminase.

22. The base editing method according to any one of claims 1 to 21, characterized in that, Simultaneously, it enables targeted editing of cytosine (C) bases to thymine (T) bases and targeted editing of adenine (A) bases to guanine (G) bases.

23. The base editing method according to any one of claims 1 to 21, characterized in that, In the GC base sequence of the target DNA sequence, cytosine (C) bases are edited to thymine (T) bases.

24. A plant cell or protoplast, characterized in that, DNA base editing is achieved by the base editing method according to any one of claims 1 to 23.

25. A plant or a part of a plant, characterized in that, The plant cells or protoplasts described in claim 24 are used for growth or culture.

26. A plant or a part of a plant, characterized in that, It is the offspring or clone of the plant according to claim 25.

27. A seed, characterized in that, Obtained from the plant according to claim 25 or 26.

28. A plant, its offspring, or a part of a plant, characterized in that, It is grown from the seed according to claim 27.

29. A plant cell or protoplast, a plant or a part of a plant, or a seed, characterized in that, It is a plant cell or protoplast according to claim 24, a plant or part of a plant according to claim 25, 26 or 28, or a seed according to claim 27. In the wild-type target DNA sequence, cytosine (C) bases are edited to thymine (T) bases or adenine (A) bases are edited to guanine (G) bases.

30. The plant cell, protoplast, plant or part of a plant, or seed according to claim 29, characterized in that, In the GC sequence of the target DNA sequence, cytosine (C) bases are edited to thymine (T) bases.

31. A base editing composition, characterized in that, It contains one or more DNA-binding proteins and one or more enzyme proteins, or one or more polynucleotides encoding them. The one or more enzyme proteins include nicking enzymes, and also include adenine deaminase and / or cytosine deaminase.

32. The base editing composition according to claim 31, characterized in that, The one or more DNA-binding proteins and one or more enzyme proteins exist in the form of one or more fusion proteins. Each of the more than one fusion proteins independently contains a DNA-binding protein. At least one of the more than one fusion protein contains a nicking enzyme. At least one of the more than one fusion protein contains adenine deaminase and / or cytosine deaminase.

33. A delivery composition, characterized in that, It comprises one or more carriers, said carriers comprising the base editing composition according to claim 31 or 32.

34. A plant cell or protoplast, or a plant or part of a plant containing them, or a seed, characterized in that, It is derived from one or more carriers as described in claim 33.

Citation Information

Patent Citations

  • Targeted deaminase and base editing using same

    WO2022060185A1

  • Compositions and methods for the treatment of hereditary angioedema (HAE)

    WO2023086953A1