A base editor and its application

By developing a new base editor that includes sequence-specific DNA binding proteins, nicases, exonucleases and base-specific deaminases, the problem of difficulty in achieving single-strand specific and high-purity base editing in the prior art is solved, and efficient and safe base editing on cell nucleus, mitochondrial DNA or chloroplast DNA is achieved.

CN117384885BActive Publication Date: 2025-06-10INST OF GENETICS & DEVELOPMENTAL BIOLOGY CHINESE ACAD OF SCI +1
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202311622479.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2023-08-14
Filing Date
2023-11-30
Publication Date
2025-06-10
Estimated Expiration
2043-11-30

AI Technical Summary

Technical Problem

It is difficult to develop a high-purity base editor with single-strand specificity and can function on the nucleus and mitochondrial DNA or chloroplast DNA.

Method used

Provides a novel base editor that does not rely on CRISPR technology, including sequence-specific DNA binding proteins, nucleases, exonucleases and base-specific deaminases, which can perform high-purity base editing on nucleus, mitochondrial DNA or chloroplast DNA.

Benefits of technology

Base editing on selected single-stranded DNA is achieved, with good safety and accuracy, high editing product purity, low insertion deletion rate and off-target rate, suitable for base editing in cell nucleus and organelles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117384885B_ABST
    Figure CN117384885B_ABST
Patent Text Reader

Abstract

The present invention discloses a base editor and its applications. The present invention provides a nucleic acid base editor, specifically a base editor not based on CRISPR technology. The base editor includes a sequence-specific DNA-binding protein, a nickase, an exonuclease, and a base-specific deaminase. This base editor has single-strand specificity, and compared with traditional base editors, the base editor of the present invention has wide applicability in cells and can act on nuclear as well as mitochondrial DNA and / or chloroplast DNA. While achieving efficient base editing, this base editor also has the characteristics of high purity of base editing products and few indel by-products, which is conducive to being used as an efficient and safe gene editing tool.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Priority and Related Applications

[0002] This application claims priority to Chinese Patent Application No. 202211613160.4, entitled "A Base Editor and Its Applications", filed on December 15, 2022, and Chinese Patent Application No. 202311017698.3, entitled "A Base Editor and Its Applications", filed on August 14, 2023. The entire contents of the above applications, including the appendices, are incorporated herein by reference. Technical Field

[0003] The present invention relates to the field of gene editing, and particularly to a nucleic acid base editor, and more particularly to a base editor comprising a sequence-specific DNA binding protein, a nickase, an exonuclease, and a base-specific deaminase, and its applications. Background Art

[0004] It is known that mutations in genomic and mitochondrial DNA can cause various genetic diseases (Newby et al., 2021, Nature 595: 295 - 302), and correcting these mutations is expected to effectively treat or improve some serious diseases. In plants, some important agronomic traits are related to single nucleotide variations (SNVs) that occur in the plant genome, the plant mitochondrial genome, or the plant chloroplast genome; introducing these SNVs into plants can promote plant performance, molecular breeding, restore gene function to alleviate disease states, etc.

[0005] Genome editing has shown great potential for genome modification; among genome editing tools, base editing can achieve targeted base substitution without introducing DNA double-strand breaks (DSBs), thus enabling more precise and accurate editing (Gaudelli et al., 2017, Nature 551: 464 - 471; Komor et al., 2016, Nature 533: 420 - 424), and therefore has broad prospects in disease treatment and crop improvement.

[0006] Cytosine base editors (CBEs) (Komor et al., 2016, Nature 533:420-424) and adenine base editors (ABEs) (Gaudelli et al., 2017, Nature 551:464-471) are the most widely used base editors. In the CBE system, the CRISPR-Cas9 nickase (nCas9) with single-stranded DNA cleavage activity is guided by sgRNA to the target dsDNA. nCas9 cleaves the sgRNA-targeted strand, forming an R-loop. Subsequently, the single-strand-specific cytosine deaminase converts cytosine (C) to uracil (U) within a window of approximately five nucleotides in the single-stranded DNA bubble structure generated by nCas9. After DNA repair, U is replaced by T, resulting in the conversion of C:G base pairs to T:A base pairs. Additionally, the addition of uracil glycosylase inhibitor (UGI), which hinders uracil excision and its downstream processes, can improve base editing efficiency and product purity. Cytosine deaminases applicable to the Cas-mediated CBE system include, but are not limited to, APOBEC1, hAID, and hAPOBEC3A. Some new deaminase systems have recently been discovered that are also applicable to the deaminases described in the present invention (Huang, J. et al. Discovery of new deaminase functions by structure-based protein clustering. bioRxiv (2023).).

[0007] Fusing nCas9 with the artificially evolved single-stranded DNA adenine deaminase TadA produced the ABE system (Gaudelli et al., 2017, Nature 551:464-471). The working principle of ABE is similar to that of CBE. nCas9 will generate a nick on the DNA target strand under the guidance of sgRNA, and the adenine deaminase TadA converts adenine (A) to inosine (I), which is replaced by G after DNA repair, resulting in the conversion of A:T base pairs to G:C base pairs. However, the ABE system does not require UGI to improve its editing efficiency or product concentration because there is no uracil intermediate involved in the process.

[0008] The ABE and CBE mentioned above can work effectively in the nucleus, but they cannot work in chloroplasts or mitochondria because the sgRNA in the CRISPR system cannot be effectively transferred into these organelles.

[0009] In 2020, researchers developed a non-CRISPR base editor system consisting only of protein components. This new base editor system was named DdCBE (Mok et al., 2020, Nature 583: 631-637). The core components of DdCBE include the double-stranded DNA cytosine deaminase DddA, which can convert C on double-stranded DNA to U without the need for CRISPR-Cas9 to create single-stranded DNA. However, the intact DddA is cytotoxic, so it was split into two halves - DddA-N and DddA-C, which were respectively fused with a pair of TALE proteins. DddA-N and DddA-C are guided to the target DNA sequence by the TALE pair, and the two recombine to restore their cytosine deaminase activity; similar to the CRISPR-based CBE system, this system can also convert C:G base pairs to T:A base pairs; adding UGI can improve the base editing efficiency and product purity of DdCBE. Due to the all-protein component characteristics of the DdCBE system, it can function both in the nucleus and can be transported into chloroplasts and mitochondria to achieve targeted cytosine base editing of chloroplast DNA and mitochondrial DNA.

[0010] However, because the DddA toxin is a cytosine deaminase, it can only act on the cytosine base in the CBE system and cannot act on the adenine base required for the ABE system, severely limiting its scope of application. In 2022, researchers fused the artificially directed evolved adenine deaminase TadA-8e with DdCBE to generate the TALED system, which can achieve base editing of A to G (Cho et al., 2022, Cell 185: 1764-1776). In the TALED system, the adenine deaminase TadA-8e is fused with one of the split DddAs, and this combination successfully induces base conversions of C to T and A to G simultaneously in mitochondrial DNA. In addition, when the deaminase activity of DddA is inactivated, the A-to-G base editing mediated by TadA-8e is still effective.

[0011] Although the DdCBE and TALED systems have extended the application scope of base editing to mitochondrial DNA and / or chloroplast DNA, there are still some limitations. First, due to the inherent double-stranded DNA cytosine deaminase activity of DddA, cytosines in the deamination windows on both strands will undergo deamination reactions, which means that the deamination reaction cannot occur only on the selected single strand. Therefore, it is not safe and precise enough to be used safely. Second, compared with CBE- and ABE-mediated base editing in the nucleus, the base editing products of DddA contain a relatively high indel frequency, resulting in a lower product purity. Third, it has been reported that the DddA-based mitochondrial base editor causes extensive off-target mutations in the nucleus when performing mitochondrial base editing (Lei et al., 2022, Nature 606: 804-811). Notably, most of the off-target mutations are TALE-independent and caused by DddA. A large number of nuclear off-target mutations will have a significant adverse impact on the safety of using these base editors.

[0012] Therefore, there is an urgent need in the art to develop a new type of base editor with single-strand specificity that can function on nuclear, mitochondrial DNA, and / or chloroplast DNA and has a high product purity. Summary of the Invention

[0013] To solve the above technical problems, the present application provides a new type of CRISPR technology-independent base editor. This system has single-strand specificity, can function on nuclear, mitochondrial DNA, or chloroplast DNA, and can obtain highly pure editing products.

[0014] Specifically, the present invention provides a new nucleic acid base editor protein composition, a recombinant expression construct encoding the new synthetic nucleic acid base editor protein, a genetically engineered cell containing one or more recombinant expression constructs encoding the new synthetic nucleic acid base editor protein, and application methods for the above-mentioned new nucleic acid base editor protein, recombinant expression construct, and genetically engineered cell.

[0015] The nucleic acid base editor of the present invention comprises: a sequence-specific DNA binding protein; a nickase; an exonuclease; and a base-specific deaminase. In certain embodiments, the nucleic acid base editor further comprises a uracil glycosylase inhibitor. In specific embodiments, the sequence-specific DNA binding protein, the nickase, the exonuclease, and the base-specific deaminase form one or more fusion proteins. In a preferred embodiment of the nucleic acid base editor provided by the present invention, the sequence-specific DNA binding protein is selected from TALE proteins, ZFA proteins, Cas proteins, or meganucleases. And in certain specific embodiments, the sequence-specific DNA binding protein is preferably a TALE protein. In a specific embodiment of the nucleic acid base editor of the present invention, the nickase is a FokI nickase. The deaminase in the nucleic acid base editor of the present invention is selected from cytosine-specific deaminases or adenine-specific deaminases. In a preferred embodiment of the nucleic acid base editor of the present invention comprising a cytosine-specific deaminase, the cytosine deaminase is selected from hAPOBEC3A, rAPOBEC1, hAID, pmCDA1, or Sdd deaminase. In a preferred embodiment of the nucleic acid base editor of the present invention comprising an adenine-specific deaminase, the adenine deaminase is TadA-8e.

[0016] In another preferred embodiment, the composition provided by the present invention comprises one or more recombinant expression constructs encoding sequence-specific DNA-binding proteins, nickases, exonucleases, base-specific deaminases, wherein each of the sequence-specific DNA-binding protein, nickase, exonuclease, and base-specific deaminase can be expressed in cells. In certain embodiments, these nucleic acid compositions further comprise a recombinant expression construct encoding a uracil glycosylase inhibitor. In a specific embodiment, the composition comprises one or more recombinant expression constructs encoding a sequence-specific DNA-binding protein, nickase, exonuclease, base-specific deaminase as a fusion protein, wherein the fusion protein comprised can be expressed in cells. In a preferred embodiment of the nucleic acid base editor provided herein, the sequence-specific DNA-binding protein is selected from TALE proteins, ZFA proteins, Cas proteins, or meganucleases, and in certain specific embodiments, the sequence-specific DNA-binding protein is a TALE protein. In a specific embodiment of the nucleic acid base editor of the present invention, the nickase is a FokI nickase. The deaminases in the nucleic acid base editor of the present invention are selected from cytosine-specific deaminases or adenine-specific deaminases, preferably deaminases shown by the sequences of SEQ ID NO. 36-59, 80-86. In a beneficial embodiment of the above nucleic acid base editor comprising a cytosine-specific deaminase, the cytosine deaminase is selected from hAPOBEC3A, rAPOBEC1, hAID, pmCDA1, or Sdd deaminase. In an embodiment of the nucleic acid base editor of the present invention comprising an adenine-specific deaminase, the adenine deaminase is TadA-8e.

[0017] In another preferred embodiment, the present invention also provides a recombinant cell comprising one or more recombinant expression constructs containing coding sequences for a sequence-specific DNA binding protein, a nickase, an exonuclease, and a base-specific deaminase; wherein each of the sequence-specific DNA binding protein, the nickase, the exonuclease, and the base-specific deaminase can be expressed in the cell. In certain embodiments, these recombinant cells contain a nucleic acid composition further comprising a recombinant expression construct encoding a uracil glycosylase inhibitor. In a specific embodiment, the recombinant cell comprises one or more recombinant expression constructs encoding a sequence-specific DNA binding protein, a nickase, an exonuclease, and a base-specific deaminase as a fusion protein, wherein the fusion protein contained therein can be expressed in the cell. In a preferred embodiment of the recombinant cell provided herein, the sequence-specific DNA binding protein is selected from TALE proteins, ZFA proteins, Cas proteins, or meganucleases, and in certain specific embodiments, the sequence-specific DNA binding protein is a TALE protein. In a specific embodiment of the recombinant cell provided herein, the nickase is selected as FokI. Further provided is a recombinant cell of the present invention comprising one or more recombinant expression constructs encoding a deaminase, wherein the deaminase is a cytosine-specific deaminase or an adenine-specific deaminase, preferably selected from the deaminases shown in SEQ ID NOs. 36-59, 80-86. A preferred embodiment of the recombinant cell provided herein comprises one or more recombinant expression constructs encoding a cytosine-specific deaminase, wherein in a preferred embodiment, the cytosine deaminase is selected from hAPOBEC3A, rAPOBEC1, hAID, pmCDA1, or Sdd deaminase. In another preferred embodiment, the recombinant cell comprises one or more recombinant expression constructs encoding an adenine-specific deaminase, wherein in a non-limiting example, the adenine deaminase is selected as TadA-8e.

[0018] In another preferred embodiment, the present invention also provides a method for base editing in a cell, comprising the step of introducing a nucleic acid base editor, or a recombinant expression construct encoding the nucleic acid base editor of the present invention, or a fusion protein encoding the nucleic acid base editor of the present invention into the cell. In the practice of the method described herein, base editing is performed on a target nucleic acid recognized by a specific binding protein and results in the alteration of a cytosine residue or an adenine residue.

[0019] In another preferred embodiment, the present invention provides a nucleic acid base editor that is specific for base editing activity in the nucleus or organelles. Further, the nucleic acid base editor for the nucleus may include a nuclear localization signal (NLS). Further, the base editors for mitochondria or chloroplasts may include a mitochondrial targeting sequence (MTS) or a chloroplast transit peptide (CTP), respectively. In these embodiments, depending on the specific target organelle or base editor, NLS, MTS, or CTP may be substituted for each other, which will be further elaborated in detail herein.

[0020] The following are exemplary technical solutions of the present invention:

[0021] The first object of the present invention is to provide a nucleic acid base editor, which includes the following components: a) a sequence-specific DNA binding protein; b) a nickase; c) an exonuclease; d) a base-specific deaminase.

[0022] Preferably, the components of the nucleic acid base editor exist separately or form one or more fusion proteins.

[0023] Preferably, the sequence-specific DNA binding protein is one or more of a TALE protein, a ZFA protein, a Cas protein, and a meganuclease.

[0024] Preferably, the sequence-specific DNA binding protein is a TALE protein.

[0025] Preferably, the nickase is a dimer of the cleavage functional domain monomer FokICD of FokI or a mutant thereof, and the dimer of FokICD or a mutant thereof is composed of a pair of interacting cleavage functional domain monomers of FokI, and only one FokICD monomer in the dimer of FokICD or a mutant thereof has DNA endonuclease activity.

[0026] Preferably, the cleavage functional domain monomer of FokI is isolated from a mutant of the wild-type FokI protein, and the mutant of the wild-type FokI protein is mutated at position 450 and / or 467, or has an amino acid sequence that is at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the cleavage functional domain monomer of FokI.

[0027] More preferably, the mutation renders the FokICD monomer devoid of DNA endonuclease activity.

[0028] Preferably, the monomeric FokI cleavage functional domain FokICD is isolated from a mutant of the wild-type FokI protein, and the mutation renders the FokICD monomer unable to self-associate with FokICD monomers containing the same site mutation to form a dimer.

[0029] More preferably, the FoKICD monomer sequence is selected from SEQ ID No. 87-88.

[0030] Preferably, the amino acid sequence of the monomeric FokI cleavage functional domain FokICD is selected from SEQ No. 60-63.

[0031] Preferably, the base-specific deaminase is selected from cytosine-specific deaminase or adenine-specific deaminase.

[0032] More preferably, the base deaminase is selected from SEQ ID NO. 36-59, 80-86.

[0033] More preferably, the base-specific deaminase is cytosine-specific deaminase.

[0034] More preferably, the cytosine-specific deaminase is one or more of hAPOBEC3A, rAPOBEC1, hAID, pmCDA1 or Sdd deaminase.

[0035] Furthermore, the nucleic acid base editor further contains:

[0036] e) Uracil glycosylase inhibitor (UGI);

[0037] And the uracil glycosylase inhibitor exists alone or forms at least one fusion protein with other nucleic acid base editor components.

[0038] Preferably, the base-specific deaminase is adenine-specific deaminase.

[0039] Preferably, the adenine-specific deaminase is TadA-8e.

[0040] Furthermore, the nucleic acid base editor further contains:

[0041] f) γb;

[0042] The γb forms at least one fusion protein with other nucleic acid base editor components.

[0043] The second object of the present invention is to provide a fusion protein as a nucleic acid base editor, and the fusion protein contains the protein domain of the base editor as described in the first object.

[0044] Another object of the present invention is to provide a fusion protein as a nucleic acid base editor, which is arranged in a linear order starting from the amino terminus of the protein and comprises: an exonuclease, an XTEN spacer peptide, a base-specific deaminase, an XTEN spacer peptide, a uracil glycosylase inhibitor (UGI), and a nuclear localization signal.

[0045] Another object of the present invention is to provide a fusion protein as a nucleic acid base editor, which is arranged in a linear order starting from the amino terminus of the protein and comprises: an exonuclease, a 48-amino acid spacer peptide, a base-specific deaminase, an XTEN spacer peptide, a uracil glycosylase inhibitor (UGI), and a nuclear localization signal.

[0046] Another object of the present invention is to provide a composition of a fusion protein having nucleic acid base editor activity, the composition comprising:

[0047] A first fusion protein, which comprises: a nuclear localization signal (NLS), a sequence-specific DNA binding protein, and a base-specific deaminase;

[0048] A second fusion protein, which comprises: an exonuclease and a nuclear localization signal (NLS); and,

[0049] A third fusion protein, which comprises: a uracil glycosylase inhibitor (UGI) and a nuclear localization signal (NLS).

[0050] Another object of the present invention is to provide a composition of a fusion protein having nucleic acid base editor activity, the composition comprising:

[0051] A first fusion protein, which is arranged in a linear order starting from the amino terminus of the protein and comprises: a nuclear localization signal (NLS), a base-specific deaminase, a TALE-L protein, a FokI-L D450A protein, a T2A sequence, an NLS, a TALE-R protein, and a FokI-R protein;

[0052] A second fusion protein, which comprises: an exonuclease and a nuclear localization signal (NLS); and,

[0053] A third fusion protein, which comprises: a uracil glycosylase inhibitor (UGI) and a nuclear localization signal (NLS).

[0054] Another object of the present invention is to provide a composition of a fusion protein having nucleic acid base editor activity, the composition comprising:

[0055] A first fusion protein, which is arranged in a linear order starting from the amino terminus of the protein and comprises: a nuclear localization signal (NLS), a TALE-L protein, a FokI-L D450AProtein, T2A sequence, NLS, base-specific deaminase, 48-amino acid spacer peptide, TALE-R protein, and FokI-R protein;

[0056] A second fusion protein, comprising: an exonuclease and a nuclear localization signal (NLS); and,

[0057] A third fusion protein, comprising: uracil glycosylase inhibitor (UGI) and a nuclear localization signal (NLS).

[0058] Another object of the present invention is to provide a composition of a fusion protein having nucleic acid base editor activity, the composition comprising:

[0059] A first fusion protein, which, arranged in a linear order starting from the amino terminus of the protein, comprises: a nuclear localization signal (NLS), a sequence-specific DNA binding protein, a base-specific deaminase, and uracil glycosylase inhibitor (UGI); and,

[0060] A second fusion protein, comprising: an exonuclease and a nuclear localization signal (NLS).

[0061] Another object of the present invention is to provide a composition of a fusion protein having nucleic acid base editor activity, the composition comprising:

[0062] A first fusion protein, arranged in a linear order starting from the amino terminus of the protein, which comprises: a nuclear localization signal (NLS), a base-specific deaminase, a 48-amino acid spacer peptide, TALE-L protein, FokI-L D450A protein, T2A sequence, NLS, TALE-R protein, FokI-R protein, a 4-amino acid spacer peptide, and uracil glycosylase inhibitor (UGI); and,

[0063] A second fusion protein, comprising: an exonuclease and a nuclear localization signal (NLS).

[0064] Another object of the present invention is to provide a composition of a fusion protein having nucleic acid base editor activity, the composition comprising:

[0065] A first fusion protein, arranged in a linear order starting from the amino terminus of the protein, which comprises: a nuclear localization signal (NLS), uracil glycosylase inhibitor (UGI), a 4-amino acid spacer peptide,

[0066] a base-specific deaminase, a 48-amino acid spacer peptide, TALE-L protein, FokI-L D450A protein, T2A sequence, NLS, TALE-R protein, FokI-R protein; and,

[0067] A second fusion protein, comprising: an exonuclease and a nuclear localization signal (NLS).

[0068] Another object of the present invention is to provide a composition of a fusion protein having nucleic acid base editing activity and capable of performing base editing on mitochondria, the composition comprising:

[0069] A first fusion protein, arranged in a linear order starting from the amino terminus of the protein, comprising: a mitochondrial targeting sequence (MTS), a TALE-L protein, and a FokI-L D450A protein;

[0070] A second fusion protein, arranged in a linear order starting from the amino terminus of the protein, comprising: a mitochondrial targeting sequence (MTS), a TALE-R protein, and a FokI-R protein;

[0071] A third fusion protein, arranged in a linear order starting from the amino terminus of the protein, comprising: a mitochondrial targeting sequence (MTS) and an exonuclease;

[0072] A fourth fusion protein, arranged in a linear order starting from the amino terminus of the protein, comprising: a mitochondrial targeting sequence (MTS) and a base-specific deaminase; and,

[0073] A fifth fusion protein, arranged in a linear order starting from the amino terminus of the protein, comprising: a mitochondrial targeting sequence (MTS) and a uracil glycosylase inhibitor (UGI).

[0074] Another object of the present invention is to provide a composition of a fusion protein having nucleic acid base editing activity and capable of performing base editing on mitochondria, the composition comprising:

[0075] A first fusion protein, arranged in a linear order starting from the amino terminus of the protein, comprising: a mitochondrial targeting sequence (MTS), a TALE-L protein, and a FokI-L D450A protein;

[0076] A second fusion protein, arranged in a linear order starting from the amino terminus of the protein, comprising: a mitochondrial targeting sequence (MTS), a TALE-R protein, and a FokI-R protein;

[0077] A third fusion protein, arranged in a linear order starting from the amino terminus of the protein, comprising: a mitochondrial targeting sequence (MTS), γb, and an exonuclease;

[0078] A fourth fusion protein arranged in a linear order starting from the amino terminus of the protein, comprising: a mitochondrial targeting sequence (MTS) and a base-specific deaminase; and,

[0079] A fifth fusion protein, arranged in a linear order starting from the amino terminus of the protein, comprising: a mitochondrial targeting sequence (MTS), γb, and a uracil glycosylase inhibitor (UGI).

[0080] Another object of the present invention is to provide a composition of a fusion protein having nucleic acid base editor activity, the composition comprising:

[0081] A first fusion protein, comprising: a nuclear localization signal (NLS) / chloroplast transit peptide (CTP) / mitochondrial targeting sequence (MTS), a sequence-specific DNA binding protein, and a nickase;

[0082] A second fusion protein, comprising: an exonuclease and a nuclear localization signal (NLS) / chloroplast transit peptide (CTP) / mitochondrial targeting sequence (MTS); and,

[0083] A third fusion protein, comprising: a base-specific deaminase, a uracil glycosylase inhibitor (UGI), and a nuclear localization signal (NLS) / chloroplast transit peptide (CTP) / mitochondrial targeting sequence (MTS).

[0084] Another object of the present invention is to provide a composition of a fusion protein having nucleic acid base editing activity, the composition comprising:

[0085] A first fusion protein, arranged in a linear order starting from the amino terminus of the protein, comprising: a nuclear localization signal (NLS) / chloroplast transit peptide (CTP) / mitochondrial targeting sequence (MTS), a TALE-L protein, a FokI-L D450A protein, a T2A sequence, an NLS, a TALE-R protein, a FokI-R protein, or, comprising: a nuclear localization signal (NLS) / chloroplast transit peptide (CTP) / mitochondrial targeting sequence (MTS), a TALE-L protein, a FokI-L protein, a T2A sequence, an NLS, a TALE-R protein, a FokI-R D450A protein;

[0086] A second fusion protein, arranged in a linear order starting from the amino terminus of the protein, comprising: a nuclear localization signal (NLS) / chloroplast transit peptide (CTP) / mitochondrial targeting sequence (MTS) and an exonuclease; and,

[0087] A third fusion protein, arranged in a linear order starting from the amino terminus of the protein, comprising: a nuclear localization signal (NLS) / chloroplast transit peptide (CTP) / mitochondrial targeting sequence (MTS), a base-specific deaminase, an XTEN spacer peptide, and a uracil glycosylase inhibitor (UGI).

[0088] Another object of the present invention is to provide a composition of a fusion protein having nucleic acid base editing activity and capable of performing base editing on mitochondria, characterized in that the composition comprises:

[0089] A first fusion protein, arranged in a linear order starting from the amino terminus of the protein, which comprises: a mitochondrial targeting sequence (MTS), a base-specific deaminase, a spacer peptide of 48 amino acids, a TALE-L protein, a FokI-L D450A protein, a spacer peptide of 11 amino acids, and a uracil glycosylase inhibitor (UGI); and,

[0090] A second fusion protein, arranged in a linear order starting from the amino terminus of the protein, which comprises: a mitochondrial targeting sequence (MTS), a spacer peptide of 48 amino acids, a TALE-R protein, a uracil glycosylase inhibitor (UGI), a spacer peptide of 14 amino acids, and a FokI-R protein.

[0091] Another object of the present invention is to provide a recombinant expression construct for nucleic acid base editing, which is used to express the nucleic acid base editor described in the first object above or the fusion protein or composition described in the above other objects.

[0092] Another object of the present invention is to provide a genetically engineered cell, which is used to transform the recombinant expression construct described in the above object.

[0093] Another object of the present invention is to provide a method for nucleic acid base editing in a cell, introducing the nucleic acid base editor or recombinant expression construct described in the above object into the cell to edit the target gene.

[0094] Preferably, the target gene is selected from nuclear genomic DNA, mitochondrial genomic DNA, or chloroplast genomic DNA.

[0095] More preferably, the target gene is nuclear genomic DNA, and the nucleic acid base editor further comprises a nuclear localization signal (NLS).

[0096] More preferably, the target gene is mitochondrial genomic DNA, and the nucleic acid base editor further comprises a mitochondrial targeting sequence (MTS).

[0097] More preferably, the target gene is chloroplast genomic DNA, and the nucleic acid base editor further comprises a chloroplast transit peptide (CTP).

[0098] Another object of the present invention is to fuse γb to the end of each component.

[0099] Further preferably, γb is fused with UGI and Trex2 respectively.

[0100] Another object of the present invention is to provide an application of a base editing technology in base editing, using the base editor, fusion protein, composition, recombinant expression construct, genetically engineered cell or method described in the above object to perform base editing on DNA in a cell, and the cell is a mammalian cell, bacterium, protist, fungus, insect cell, yeast, unconventional yeast or plant cell.

[0101] Preferably, the plant cell is derived from a whole plant, seedling, meristem, ground tissue, vascular tissue, dermal tissue, seed, leaf, root, bud, stem, flower, fruit, stolon, bulb, tuber, corm, asexual terminal branch, bud, sprout, or tumor tissue of a monocotyledonous or dicotyledonous plant.

[0102] Preferably, the mammalian cell is selected from human neurons, muscle cells, endocrine / exocrine cells, epithelial cells, muscle cells, tumor cells, hematopoietic cells, bone cells, somatic cells, induced pluripotent stem cells.

[0103] Preferably, the editor is used to perform base editing on the nuclear genome or organelle genome.

[0104] Preferably, the organelle is a mitochondrion or a chloroplast.

[0105] Another object of the present invention is to provide the use of the base editor, fusion protein, composition, recombinant expression construct, genetically engineered cell described in the above object in the preparation of a pharmaceutical composition for treating a disease in a subject in need.

[0106] Another object of the present invention is to provide a pharmaceutical composition for treating a disease in a subject in need, the pharmaceutical composition comprising the base editor, fusion protein, composition, recombinant expression construct, genetically engineered cell described in the above object, and, optionally, a pharmaceutically acceptable carrier.

[0107] Another object of the present invention is to provide a method for producing a genetically modified plant, characterized in that the method comprises introducing the base editor, fusion protein, composition, recombinant expression construct, genetically engineered cell described in the above object into at least one of the plants.

[0108] The base editor and its application provided by the present invention have the following beneficial effects:

[0109] (1) The base editor of the present invention only performs base editing on the selected single strand, and has good safety and accuracy.

[0110] (2) The base editor of the present invention has high purity of editing products and a low rate of indel by-products, and has excellent editing efficiency.

[0111] (3) The base editor of the present invention has a low off-target rate, effectively improving its therapeutic effect and safety.

[0112] (4) The base editor of the present invention is not based on CRISPR technology, has a wider range of application scopes and scenarios, and its components can all play roles in cell organelles such as the cell nucleus, mitochondria, and chloroplasts. BRIEF DESCRIPTION OF THE DRAWINGS

[0113] For a better understanding of the technical solutions described in the present invention, the following drawings are now used for illustration.

[0114] Figure 1 is the functional schematic diagram of the nucleic acid base editor of the present invention. In the first step, the sequence-specific DNA binding protein (SSDBP) locates and binds to the target DNA sequence; in the second step, the nickase preferentially cleaves one DNA strand at the target site, and then the exonuclease digests the cleaved DNA strand from the nick to the SSDBP binding site. This will expose a ssDNA fragment in the complementary strand, which then becomes the substrate of the deaminase to achieve deamination, and then causes the corresponding base to be converted (C:G pairing to T:A pairing or A:T pairing to G:C pairing, and the conversion type depends on the deaminase used) after DNA repair.

[0115] Figure 2A and Figure 2B is the application effect of the nucleic acid base editor of the present invention in high-purity base editing in rice cell nuclei. Among them Figure 2A is the C to T base editing efficiency for the OsBADH2 target site in rice protoplasts under different treatment methods. Figure 2B is the C>T base editing efficiency and the frequency of generating indel by-products for the OsBADH2 target site in rice protoplasts under different treatment methods.

[0116] Figure 3A and Figure 3B is the analysis of the base editing window of the base editor of the present invention. After transforming the nucleic acid base editor of the present invention into rice protoplasts, DNA is extracted, and high-throughput sequencing is performed on the target site to obtain the editing efficiency of different bases on the target sequence. Figure 3A is the schematic diagram of the OsBADH2 target site sequence. The gray sequences on both sides are the TALE binding sites, and the black region in the middle is the spacer sequence. Figure 3B is the analysis of the base editing window of the base editor according to the high-throughput sequencing results. Among them, CK is the blank control without transforming any plasmid, and TALEN WTand TALEN WT +ExoI are the combinations of transforming only wild-type TALEN or TALEN + exonuclease ExoI respectively, and these two treatments serve as negative controls.

[0117] Figure 4A and Figure 4B show that after transforming rice protoplasts with the base editor of the present invention to target OsDEP1, the editing efficiency of cytosine nucleotides at the target site ([ Figure 4A ) and the occurrence frequency of insertion-deletion (indel) by-products ([ Figure 4B ) are analyzed by high-throughput sequencing. Among them, CK is the blank control without transforming any plasmid, TALEN WT and TALEN WT +ExoI are the combinations of transforming only wild-type TALEN or TALEN + exonuclease ExoI respectively, and these two treatments serve as negative controls.

[0118] Figure 5A and Figure 5B show the implementation effects of the base editor when using different combinations of FokI nickase, different exonucleases, and cytosine deaminase. Different editing windows are generated when using exonucleases with different digestion directions; when using different nickases, specific base editing is performed on different DNA single strands at the target site ([ Figure 5A ). And the analysis of the purity of the editing products and the occurrence frequency of by-products of the base editor of the present invention under different combinations ([ Figure 5B ).

[0119] Figure 6A and Figure 6B show the base editing efficiency and the insertion-deletion (indel) by-product frequency detected by high-throughput sequencing when the base editor of the present invention including a combination of cytosine deaminase and exonuclease introduces the target sequence (OsBADH2 in rice protoplasts). Among them, the exonuclease is a 5'-exonuclease or a 3'-exonuclease.

[0120] Figure 7A and Figure 7B show the base editing efficiency and the editing window detected by high-throughput sequencing when the combination base editor of the present invention including different cytosine deaminases and exonucleases introduces the target sequence (OsBADH2 in rice protoplasts).

[0121] Figure 8 show the base editing efficiency detected by high-throughput sequencing when the base editor of the present invention containing adenine deaminase introduces the target sequence (OsCKX2 in rice protoplasts).

[0122] Figure 9Schematic diagram of the base editor of the present invention, which comprises a fusion protein of an exonuclease, a deaminase, a uracil DNA glycosylase inhibitor, and a nuclear localization signal (NLS) separated by an XTEN spacer peptide or a 48-amino acid spacer peptide.

[0123] Figure 10A and Figure 10B Base editing efficiency ([ Figure 10A ) and editing window ([ Figure 10B ) of different base editors of the present invention introduced into the target sequence (OsDEP1 in rice protoplasts), detected by high-throughput sequencing.

[0124] Figure 11A and Figure 11B Schematic diagram of the base editor of the present invention containing a deaminase-TALE fusion protein vector. In each embodiment, the fusion protein of NLS-exonuclease and NLS-uracil glycosylase inhibitor (UGI) are provided in separate vectors.

[0125] Figure 12A and Figure 12B Bar graphs showing the base editing rate and insertion-deletion (indel) rate of the target sequence (OsDEP1 in rice protoplasts, Figure 12A ; OsCKX2 in rice protoplasts, Figure 12B ) introduced by the base editor of the present invention. The results of the fusion protein of the deaminase-TALE-FokI-R nickase protein are shown in Figure 12A , and the results of the fusion protein of the deaminase-TALE-FokI-L nickase protein are shown in Figure 12B .

[0126] Figure 13A and Figure 13B Schematic diagram of the base editor of the present invention containing a deaminase-TALE fusion protein. In each embodiment, the fusion protein of NLS and exonuclease is provided in an additional vector.

[0127] Figure 14 Base editing efficiency of the target sequence (OsDEP1 in rice protoplasts) using the fusion protein as shown in Figure 13A and Figure 13B by the base editor of the present invention and separately expressing each component.

[0128] Figure 15ASchematic diagram of the vector used by the base editor of the present invention in mitochondrial editing, which comprises a vector expressing MTS-deaminase, MTS-UGI, MTS-TALE-R-FokI-R (or MTS-TALE-R-FokI-R D450A )、MTS-TALE-L-FokI-L D450A (or MTS-TALE-L-FokI-L) constructs of nickase and MTS-exonuclease.

[0129] Figure 15B The base editor of the present invention is used Figure 15A Schematic diagram of the target sequence of the construct shown in, showing the TALE-R and TALE-L binding sites and cytosine residues targeted by certain nucleic acid base editors of the present invention, i.e., a schematic diagram of the mitochondrial ND6 target sequence and TALE binding site.

[0130] Figure 15C The use of the present invention Figure 15A Base mutation efficiency of the base editors of the constructs shown in the figure on the target sequence.

[0131] Figure 16A - Figure 16E is a representative representation of a recombinant expression construct encoding a base editor used in the Examples described herein in rice. Figure 16A - Figure 16E In the assay, FokK-L-nickase is equivalent to FOKI-L; FokI-R is equivalent to FOKI-R (D450A / D467A).

[0132] Figure 16A Represents the recombinant expression construct encoding the wild-type TALEN used in Example 2 and other examples (NLS-TALEN WT Schematic diagram of the vector, taking the TALE targeting OsBADH2 as an example), this vector can cause double-strand breaks in the target DNA and randomly induce insertion and deletion (indel) mutations, which is used as a control group in each embodiment. In this construct, a stable expression T-DNA vector with a UBI promoter and Nos terminator from corn is used to drive the expression of wild-type TALEN (including TALE-L-FokI-L and TALE-R-FokI-R fusion proteins, in which FokI does not contain D450A or D467A mutations), wherein the N- and C-terminal regions of TALE contain corresponding truncations (ΔN152 / C63) and are located on the flanks of the DNA binding domain of TALE. The TALE-L-FokI-L and TALE-R-FokI-R fusion proteins are connected by a T2A self-cleaving peptide. Other components shown in the figure include the CaMV 35S promoter (a cauliflower mosaic virus-derived promoter), the hygromycin resistance selection gene Hyg, the nopaline synthase terminator Nos of Agrobacterium tumefaciens, etc.

[0133] Figure 16B It is a diagram of a recombinant expression construct containing the sequence-specific DNA-binding proteins TALE-L, TALE-R, and the nicking enzyme FokI nicking enzyme (i.e., a partial schematic of a vector containing a nicking enzyme, exonuclease, and deaminase, taking the TALE targeting OsBADH2 as an example; according to different target sequences, corresponding TALE coding sequences can be designed), as well as two additional constructs: NLS-deaminase-UGI and exonuclease-NLS. A UBI promoter from maize and a Nos terminator are included in these constructs to drive the expression of the deaminase-UGI fusion protein and the exonuclease, respectively. UGI, a uracil-DNA glycosylase inhibitor from Bacillus subtilis phage, protects uracil in DNA by irreversibly inhibiting the key DNA repair enzyme uracil-DNA glycosylase. Other components shown in the figure include the CaMV 35S promoter (a cauliflower mosaic virus-derived promoter), the hygromycin resistance selection gene Hyg, the nopaline synthase terminator Nos of Agrobacterium tumefaciens, and the CaMV polyadenylation signal terminator.

[0134] Figure 16C It is a diagram of a recombinant expression construct containing the sequence-specific DNA-binding proteins TALE-L, TALE-R, the nicking enzyme (FokI nicking enzyme), and a fusion protein of deaminase (i.e., a partial schematic of a vector containing a nicking enzyme, exonuclease, deaminase, and uracil glycosylase inhibitor, taking the TALE targeting OsBADH2 as an example; according to different target sequences, corresponding TALE coding sequences can be designed), as well as two additional constructs: UGI-NLS and exonuclease-NLS. The recombinant expression constructs UGI-NLS and exonuclease-NLS each have a UBI promoter and a CaMV terminator to drive the expression of UGI and the exonuclease, respectively. UGI, a uracil-DNA glycosylase inhibitor from Bacillus subtilis phage, protects uracil in DNA by irreversibly inhibiting the key DNA repair enzyme uracil-DNA glycosylase. Other components shown in the figure include the CaMV 35S promoter (a cauliflower mosaic virus-derived promoter), the hygromycin resistance selection gene Hyg, the nopaline synthase terminator Nos of Agrobacterium tumefaciens, and the CaMV polyadenylation signal terminator.

[0135] Figure 16D It is a diagram of a recombinant expression construct containing the sequence-specific DNA-binding proteins TALE-L, TALE-R, the nicking enzyme FokI nicking enzyme, deaminase, and a fusion protein of UGI (i.e., a partial schematic of a vector containing NLS-deaminase-TALE-L-FokI- nickase-TALEN-R-UGI, Schematic diagram of the vector of exonuclease-NLS, taking the TALE targeting OsBADH2 as an example; according to different target sequences, corresponding TALE coding sequences can be designed), and an additional construct: exonuclease-NLS. The recombinant expression construct exonuclease-NLS has a UBI promoter and a CaMV terminator to drive the expression of the exonuclease. UGI, uracil-DNA glycosylase inhibitor from Bacillus subtilis phage, protects uracil in DNA by irreversibly inhibiting the key DNA repair enzyme uracil-DNA glycosylase. Other components shown in the figure include the CaMV 35S promoter (a cauliflower mosaic virus-derived promoter), the hygromycin resistance screening gene Hyg, the nopaline synthase terminator Nos of Agrobacterium tumefaciens, and the CaMV polyadenylation signal terminator.

[0136] Figure 16E is a schematic diagram of a recombinant expression construct of a fusion protein containing the sequence-specific DNA-binding proteins TALE-L, TALE-R, the nickase FokI nickase, the deaminase, the exonuclease, and UGI (NLS-deaminase-TALE-L-FokI- nickase -TALEN-R-UGI-exonuclease vector schematic diagram, taking the TALE targeting OsBADH2 as an example, according to different target sequences, corresponding TALE coding sequences can be designed), with the additional feature that UGI and the exonuclease are encoded in the construct rather than introduced into the cell in a separate construct.

[0137] Figure 17A to Figure 17H is a representative schematic diagram of a recombinant expression construct encoding a base editor applied to mitochondrial editing in human cells in the examples described herein.

[0138] Figure 17AIt is a representative of the mitochondrial recombinant expression construct MTS-TALE-L-FokI-L (schematic diagram of the MTS-TALE-L-FokI-L vector targeting mitochondrial ND6). The TALE sequence therein can be replaced accordingly according to different targets. The expression vector MTS-TALE-L-FokI-L has a CMV promoter and a bGH poly(A) signal terminator to drive the expression of the MTS-TALE-L-FokI-L fusion protein. The N and C terminal regions of TALE contain corresponding truncations (ΔN152 / C63) and are flanked by the DNA binding domain of TALE (see Mok et al., 2020, Nature 583: 631-637). MTS is the mitochondrial targeting sequence from human superoxide dismutase 2, which helps transport proteins into mitochondria. The CMV promoter is a promoter derived from human herpesvirus 5 and has been shown to be highly active in animal cells. The CMV enhancer is a fragment containing the cytomegalovirus promoter region, which can enhance the transcriptional efficiency of the CMV promoter. The bGH poly(A) signal is a terminator derived from a growth hormone polyadenylation signal.

[0139] Figure 17B It is a representative of the mitochondrial recombinant expression construct MTS-TALE-R-FokI-R (schematic diagram of the MTS-TALE-R-FokI-R vector targeting mitochondrial ND6). The TALE sequence therein can be replaced accordingly according to different targets. The expression vector MTS-TALE-R-FokI-R has a CMV promoter and a bGH poly(A) signal terminator to drive the expression of the MTS-TALE-R-FokI-R fusion protein. The N and C terminal regions of TALE contain corresponding truncations (ΔN152 / C63) and are flanked by the DNA binding domain of TALE (see Mok et al., 2020, Nature 583: 631-637). The MTS in this vector is the mitochondrial targeting sequence of cytochrome c oxidase subunit 8, which helps transport proteins into mitochondria. The CMV promoter is a promoter derived from human herpesvirus 5 and has been shown to be highly active in animal cells. The CMV enhancer is a fragment containing the cytomegalovirus promoter region, which can enhance the transcriptional efficiency of the CMV promoter. The bGH poly(A) signal is a terminator derived from a growth hormone polyadenylation signal.

[0140] Figure 17CSchematic diagram of the mitochondrial recombinant expression construct MTS-deaminase (schematic diagram of the MTS-deaminase vector). This recombinant expression construct has a CMV promoter and a bGH poly(A) signal terminator, which can drive the expression of MTS-deaminase in human mitochondria. MTS, CMV promoter, CMV enhancer, and bGH poly(A) signal terminator are as described in Figure 17A as described below.

[0141] Figure 17D Schematic diagram of the mitochondrial recombinant expression construct MTS-exonuclease (schematic diagram of the MTS-exonuclease vector). This recombinant expression construct has a CMV promoter and a bGH poly(A) signal terminator, which can drive the expression of MTS-exonuclease in human mitochondria. MTS, CMV promoter, CMV enhancer, and bGH poly(A) signal terminator are as described in Figure 17A as described below.

[0142] Figure 17E Schematic diagram of the mitochondrial recombinant expression construct MTS-UGI (schematic diagram of the MTS-UGI vector). This recombinant expression construct has a CMV promoter and a bGH poly(A) signal terminator, which can drive the expression of MTS-UGI (uracil glycosylase inhibitor from Bacillus subtilis phage) in human mitochondria. MTS, CMV promoter, CMV enhancer, and bGH poly(A) signal terminator are as described in Figure 17A as described below.

[0143] Figure 17F Schematic diagram of the mitochondrial recombinant expression construct MTS-deaminase-TALE-L-FokI-L (schematic diagram of the MTS-deaminase-TALE-L-FokI-L vector). The recombinant expression construct MTS-deaminase-TALE-L-FokI-L has a CMV promoter and a bGH poly(A) signal terminator to drive the expression of the MTS-deaminase-TALE-L fusion protein. Components such as MTS, CMV promoter, CMV enhancer, and bGH poly(A) signal terminator are as described in Figure 17A as described below.

[0144] Figure 17GSchematic diagram of mitochondrial recombinant expression construct MTS-exonuclease-TALE-R-FokI-R (schematic diagram of MTS-exonuclease-TALE-R-FokI-R vector). The recombinant expression construct MTS-exonuclease-TALE-R-FokI-R has a CMV promoter and a bGH poly(A) signal terminator to drive the expression of the MTS-exonuclease-TALE-R fusion protein. Components such as MTS, CMV promoter, CMV enhancer, and bGH poly(A) signal terminator are as described in Figure 17B as described below.

[0145] Figure 17H Schematic diagram of mitochondrial recombinant expression construct MTS-UGI-exonuclease-TALE-R-FokI-R (schematic diagram of MTS-UGI-exonuclease-TALE-R-FokI-R vector). The recombinant expression construct MTS-UGI-exonuclease-TALE-R-FokI-R has a CMV promoter and a bGH poly(A) signal terminator to drive the expression of the MTS-exonuclease-TALE-R fusion protein. Components such as MTS, CMV promoter, CMV enhancer, and bGH poly(A) signal terminator are as described in Figure 17B as described below.

[0146] Figure 18 Schematic diagram of CyDENT for nuclear genome editing.

[0147] Figure 19A C-to-T conversion frequency and indel frequency of nuCyDENT-R and TALEN at the OsDEP1, OsSD1, OsCKX2, and OsBADH2 loci in rice protoplasts.

[0148] Figure 19B Base editing window of CyDENT at the OsDEP1, OsSD1, OsCKX2, and OsBADH2 loci in rice protoplasts. The gray area in the figure represents the TALE binding site, and the middle area is the spacer region.

[0149] Figure 20 Base editing of CyDENT at the OsCKX2 and OsSD1 loci in rice protoplasts. The gray area is the TALE binding site.

[0150] Figure 21 Base editing of CyDENT at the human SIRT6 locus. The gray area is the TALE binding site.

[0151] Figure 22ASchematic overview of the modular CyDENT constructs used in chloroplast genome editing, taking cpCyDENT-R as an example.

[0152] Figure 22B Base editing window of CyDENT at the OsrbcL locus in rice protoplasts. The gray area is the TALE binding site.

[0153] Figure 23A Schematic diagram of the modular CyDENT structure used in mitochondria. Taking mtCyDENT-R as an example.

[0154] Figure 23B Base editing of mtCyDENT-L or mtCyDENT-R in various γb fusion states at the mitochondrial ND6 locus in HEK293T cells.

[0155] Figure 24 Editing frequencies of DdCBE, mtCyDENT-R, mtCyDENT1b-R, mtCyDENT-L, and mtCyDENT1b-L at the mitochondrial ND1.2, ND1.3, ND3, and ND6.2 loci in HEK293T cells.

[0156] Figure 25 Indel frequencies of DdCBE, mtCyDENT1b-R, and mtCyDENT1b-L at different loci in the mitochondria of HEK293T cells.

[0157] Figure 26 Base editing sites of mtCyDENT at different loci in the mitochondria of HEK293T cells. The gray area is the TALE binding site.

[0158] Figure 27 Editing frequencies of mtCyDENT1b-L and mtCyDENT1b-R using the Sdd7 deaminase at the ND5.1, ND6, and ND1.3 loci in HEK293T cells.

[0159] Figure 28A Schematic diagram of the mtCyDENT2 construct in the mitochondrial genome.

[0160] Figure 28B Editing efficiency of DdCBE and mtCyDENT2-L, mtCyDENT2-R containing different deaminases at the ND6 locus in HEK293T cells and the proportion of various editing events.

[0161] Figure 29Editing frequencies and editing strand preferences of DdCBE and mtCyDENT2-L containing Sdd3 deaminase at the ND1.2 and ND6.2 loci in HEK293T cells. Gray indicates the TALE binding sites.

[0162] Figure 30 Editing strand preference of mtCyDENT2-L (Sdd3 deaminase + TALE-L1 + TALE-R1) designed for the pathogenic mutation at the ND6.2 locus of Leigh's syndrome at the ND6.2 locus in HEK293T cells.

[0163] Figure 31A WGS and NGS analyses of the targeted editing frequencies of ND3 and ND6.2

[0164] Figure 31B Logo plots of off-target C:G to T:A and G:C to A:T base conversions for each editor.

[0165] Figure 31C Frequency distributions of SNVs and indels in potential TALE-dependent off-target sites. Detailed implementation manners

[0166] Terms:

[0167] Unless otherwise defined, all technical terms used herein have the same meaning as commonly understood by those skilled in the art.

[0168] Numeric ranges include the numbers defining the range and explicitly include each integer and non-integer fraction within the defined range. Unless otherwise indicated, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art.

[0169] The terms "structure", "recombinant expression structure" or "recombinant expression construct" used in the present invention refer to an artificially designed DNA fragment that can be used to introduce genetic material into a target cell (for example, using the recombinant expression structure to produce a base editor or its components). The term "expression" refers to the transcription and translation of a nucleic acid coding sequence to produce a polypeptide.

[0170] As used herein, the term "genetic engineering" refers to the alteration of the genetic makeup of a cell by biotechnology, including the transfer of genes within and across species, to produce modified or non-naturally occurring cells. In the specific use of this term, a construct encodes a base editor or a component thereof, and the genetically engineered cell produces the base editor. A cell containing an exogenous, recombinant, synthetic, and / or otherwise modified polynucleotide is considered a genetically engineered cell and thus is non-naturally occurring relative to any naturally occurring counterpart. In some cases, a genetically engineered cell contains one or more recombinant nucleic acids. In other cases, a genetically engineered cell contains one or more synthetic or genetically engineered nucleic acids (e.g., a nucleic acid containing at least one artificial insertion, deletion, inversion, or substitution of a sequence relative to its naturally occurring counterpart). Methods for producing genetically engineered cells are known in the art, e.g., Sambrook et al., Molecular Cloning, A Laboratory Manual (Fourth Edition), Cold Spring Harbor Press, Cold Spring Harbor, N.Y. (2012).

[0171] As used herein, the term "genetically engineered cell" or "genetically engineered host cell" or "recombinant expression host cell" can be a cell modified with gene editing technology. Gene editing refers to a form of genetic engineering in which DNA is inserted, deleted, modified, or replaced in the genome of a living cell. Compared with other genetic engineering techniques that can randomly insert genetic material into the host genome, gene editing can target the insertion to a specific location (e.g., the AAVS1 allele). Examples of gene editing techniques include, but are not limited to, restriction endonucleases, zinc finger nucleases, TALENs, and CRISPR-Cas9. The base editors disclosed herein are a particular example of gene editing that allows for one or more single nucleotide changes, especially those that result in a change in the cell phenotype.

[0172] As used herein, the terms "deaminase", "base-specific deaminase", or "deaminase domain" refer to a protein or enzyme that catalyzes a deamination reaction. The terms "deaminase" and "base-specific deaminase" are used interchangeably herein. In some embodiments, the deaminase or deaminase domain is a cytosine deaminase that catalyzes the hydrolytic deamination of cytosine or deoxycytosine to uracil, which is ultimately converted to thymine (T) during cellular modification and DNA replication. In some embodiments, the deaminase or deaminase domain is an adenine deaminase domain that catalyzes the hydrolytic deamination of adenine or deoxyadenine to hypoxanthine or deoxyhypoxanthine (I), which is ultimately converted to guanine or deoxyguanosine nucleotide (G) during cellular modification and DNA replication. In some embodiments, the deaminase or deaminase domain is a naturally occurring deaminase from an organism such as a microorganism, plant, or animal, e.g., human, chimpanzee, gorilla, monkey, cow, dog, rat, or mouse. In some embodiments, the deaminase or deaminase domain is a variant of a naturally occurring deaminase from an organism that does not exist in nature. For example, in some embodiments, the deaminase or deaminase domain is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a naturally occurring deaminase from an organism.

[0173] As used herein, the term "spacer peptide" or "Linker" refers to an element that links two molecules or moieties, e.g., two domains of a fusion protein. In some embodiments, the spacer peptide is an organic molecule, group, polymer, or chemical moiety. In some embodiments, the spacer peptide is 5-100 amino acids in length, e.g., 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acids. Longer or shorter spacer peptides are also contemplated.

[0174] As used herein, the term "mutation" refers to the replacement of a residue within a sequence (such as a nucleic acid or amino acid sequence) or the deletion or insertion of one or more residues within the sequence. The present invention generally describes mutations by identifying the initial residue, followed by the position of the residue within the sequence and the identity of the newly substituted residue. Various methods for generating the amino acid substitutions (mutations) provided herein are known in the art and are provided, for example, by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th Edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. (2012)).

[0175] As used herein, the term "uracil glycosylase inhibitor" or "UGI" refers to a protein capable of inhibiting uracil-DNA glycosylase as a base excision repair enzyme.

[0176] As used herein, the terms "top strand" or "strand A" and "bottom strand" or "strand B" are only for distinguishing the relative positions of the two strands of a DNA target site in a certain embodiment, so as to facilitate the exemplary description of the editing effect of the base editor of the present invention on single-stranded DNA, and have no special restrictive effect on the specific DNA double-stranded structure. Among them, "top strand" can be interchanged with "strand A", and "bottom strand" can be interchanged with "strand B". Unless otherwise specified, the "top strand" or "strand A" in accordance with the schematic diagram of the present application ( Figure 1 ) is the single-stranded DNA interacting with TALE-L, and correspondingly, the "bottom strand" or "strand B" is the single-stranded DNA interacting with TALE-R.

[0177] Various embodiments of the compositions and methods according to the present invention are now described in the following non-limiting examples. These examples are for illustrative purposes only and do not limit the scope of the present invention in any way.

[0178] Nucleic acid base editor

[0179] The base editing function of the nucleic acid base editor of the present invention is as Figure 1 shown. Its components include a sequence-specific DNA binding protein (SSDBP), a nickase, an exonuclease (having 5' or 3' exonuclease activity), a cytosine or adenine deaminase, and optionally a uracil glycosylase inhibitor (UGI), and optionally a targeting sequence. These components can be expressed by separate constructs or fused using appropriate spacer peptides in a single or multiple constructs.

[0180] Sequence-specific DNA binding protein

[0181] In the base editors disclosed herein, the SSDBP can be a TALE protein, a zinc finger protein (ZFA protein), a CRISPR-Cas endonuclease (Cas protein), or a meganuclease, and in some specific embodiments, a TALE protein is selected. The TALE protein (Transcription activator-like effector) is a transcription activator effector derived from Xanthomonas spp., which has been artificially engineered into a sequence-specific DNA-binding protein. The TALE protein contains 1 to 33 repeat units (or repeat modules) each with a length of 33 to 35 amino acid residues, and each repeat unit and the terminal half-repeat unit can specifically recognize and bind to a specific nucleotide target site. In each repeat sequence, the type of DNA base that TALE can recognize and bind is determined by two highly variable residues (referred to as repeat variable diresidues (RVDs)) at positions 12 and 13 targeting specific base pairs. The types or codes of RVDs for recognizing DNA have been deciphered: RVDs His / Asp (HD), Asn / Gly (NG), Asn / Asn (NN), and Asn / Ile (NI) recognize cytosine (C), thymine (T), guanine (G), and adenine (A), respectively (see, Boch & Bonas, 2010, Annu. Rev. Phytopathol. 48:419-436; Deng et al., 2012, Cell Res. 22:1502-1504). The TALE repeat units are modular, and RVDs can be artificially designed to target and bind to DNA. As disclosed in the present invention, a pair of TALE proteins (referred to as TALE-L or TALE-L protein, and TALE-R or TALE-R protein, respectively) are used to bind to DNA at two adjacent sites on the DNA, and the DNA sequence between the adjacent sites is the spacer sequence, also referred to as the target sequence; the binding sites of TALE-L and TALE-R are defined as the Left Binding Site and the Right Binding Site. The sequence specificity of the TALE protein is used to determine the target site in the base editors disclosed in the present invention. In addition, in some cases, only one TALE (instead of a pair) is required to bind and target dsDNA to achieve the function of base editing of the present invention.

[0182] Exemplary structures of TALE proteins that can be used as components of the base editors disclosed in the present invention include, but are not limited to, an N-terminus as shown in SEQ ID NO.1, a C-terminus as shown in SEQ ID NO.2, and repeat units as shown in SEQ ID NOs. 3-35.

[0183] TALE-NTD(Δ152):

[0184] MVDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVTYQHIITALPEATHEDIVGVGKQWSGARALEALLTDAGELRGPPLQLDTGQLVKIAKRGGVTAMEAVHASRNALTGAPLN(SEQ IDNO.1)

[0185] TALE-CTD(C63) :

[0186] SIVAQLSRPDPALAALTNDHLVALACLGGRPAMDAVKKGLPHAPELIRRVNRRIGERTSHRVA(SEQID NO.2)

[0187] OsBADH2-TALE-Left repeat :

[0188] LTPDQVVAIASNNGGKQALETVQRLLPVLCQDHG(SEQ ID NO.3)

[0189] LTPDQVVAIASHDGGKQALETVQRLLPVLCQDHG(SEQ ID NO.4)

[0190] LTPDQVVAIASNGGGKQALETVQRLLPVLCQDHG(SEQ ID NO.5)

[0191] LTPDQVVAIASNNGGKQALETVQRLLPVLCQDHG(SEQ ID NO.6)

[0192] LTPDQVVAIASNNGGKQALETVQRLLPVLCQDHG(SEQ ID NO.7)

[0193] LTPDQVVAIASNIGGKQALETVQRLLPVLCQDHG(SEQ ID NO.8)

[0194] LTPDQVVAIASNGGGKQALETVQRLLPVLCQDHG(SEQ ID NO.9)

[0195] LTPDQVVAIASNNGGKQALETVQRLLPVLCQDHG(SEQ ID NO.10)

[0196] LTPDQVVAIASHDGGKQALETVQRLLPVLCQDHG(SEQ ID NO.11)

[0197] LTPDQVVAIASNGGGKQALETVQRLLPVLCQDHG(SEQ ID NO.12)

[0198] LTPDQVVAIASNGGGKQALETVQRLLPVLCQDHG(SEQ ID NO.13)

[0199] LTPDQVVAIASNGGGKQALETVQRLLPVLCQDHG(SEQ ID NO.14)

[0200] LTPDQVVAIASNNGGKQALETVQRLLPVLCQDHG(SEQ ID NO.15)

[0201] LTPDQVVAIASNIGGKQALETVQRLLPVLCQDHG(SEQ ID NO.16)

[0202] LTPDQVVAIASNNGGKQALETVQRLLPVLCQDHG(SEQ ID NO.17)

[0203] LTPDQVVAIASNGGGKQALETVQRLLPVLCQDHG(SEQ ID NO.18)

[0204] LTPDQVVAIASNIGGKQALE(SEQ ID NO.19)

[0205] OsBADH2-TALE-Right repeat :

[0206] LTPDQVVAIASNNGGKQALETVQRLLPVLCQDHG(SEQ ID NO.20)

[0207] LTPDQVVAIASHDGGKQALETVQRLLPVLCQDHG(SEQ ID NO.21)

[0208] LTPDQVVAIASHDGGKQALETVQRLLPVLCQDHG(SEQ ID NO.22)

[0209] LTPDQVVAIASNGGGKQALETVQRLLPVLCQDHG(SEQ ID NO.23)

[0210] LTPDQVVAIASNGGGKQALETVQRLLPVLCQDHG(SEQ ID NO.24)

[0211] LTPDQVVAIASNGGGKQALETVQRLLPVLCQDHG(SEQ ID NO.25)

[0212] LTPDQVVAIASNGGGKQALETVQRLLPVLCQDHG(SEQ ID NO.26)

[0213] LTPDQVVAIASNNGGKQALETVQRLLPVLCQDHG(SEQ ID NO.27)

[0214] LTPDQVVAIASNGGGKQALETVQRLLPVLCQDHG(SEQ ID NO.28)

[0215] LTPDQVVAIASHDGGKQALETVQRLLPVLCQDHG(SEQ ID NO.29)

[0216] LTPDQVVAIASHDGGKQALETVQRLLPVLCQDHG(SEQ ID NO.30)

[0217] LTPDQVVAIASNIGGKQALETVQRLLPVLCQDHG(SEQ ID NO.31)

[0218] LTPDQVVAIASNIGGKQALETVQRLLPVLCQDHG(SEQ ID NO.32)

[0219] LTPDQVVAIASNNGGKQALETVQRLLPVLCQDHG(SEQ ID NO.33)

[0220] LTPDQVVAIASNNGGKQALETVQRLLPVLCQDHG(SEQ ID NO.34)

[0221] LTPDQVVAIASNIGGKQALETVQRLLPVLCQDHG(SEQ ID NO.35)

[0222] Nickase

[0223] The nickase used as a component of the base editor disclosed herein is capable of cleaving one of the double strands of the target DNA. In the base editors disclosed herein, an exemplary nickase is FokI from Flavobacterium okeanokoites, or the FokI protein, particularly an amino acid sequence variant in which its dsDNA cleavage activity is converted to a nick generated only on one strand of the target DNA, including but not limited to the D450A / D467A mutant. In addition, alternative nickases comprising bacterial type IIS restriction enzymes can also be used as components of the base editors disclosed herein.

[0224] Wild-type FokI consists of two functional domains, namely the recognition domain and the cleavage domain. The recognition domain is artificially removed to obtain FokICD that only retains the cleavage domain. When two molecules of the FokICD monomer interact to form a dimer, the cleavage activity of FokICD will be activated, and both strands of double-stranded DNA can be cleaved. The following are exemplary FokICD monomers that can be used in the present invention, including but not limited to those shown in SEQ ID NO.87-88: FokI-L:

[0225] QLVKSELEEKKSELRHKLKYVPHEYIELIEIARNSTQDRILEMKVMEFFMKVYGYRGKHLGGSRKPDGAIYTVGSPIDYGVIVDTKAYSGGYNLPIGQADEMQRYVKENQTRNKHINPNEWWKVYPSSVTEFKFLFVSGHFKGNYKAQLTRLNHKTNCNGAVLSVEELLIGGEMIKAGTLTLEEVRRKFNNGEINF (SEQ ID NO.87)

[0226] FokI-R:

[0227] QLVKSELEEKKSELRHKLKYVPHEYIELIEIARNSTQDRILEMKVMEFFMKVYGYRGKHLGGSRKPDGAIYTVGSPIDYGVIVDTKAYSGGYNLPIGQADEMERYVEENQTRNKHLNPNEWWKVYPSSVTEFKFLFVSGHFKGNYKAQLTRLNHITNCNGAVLSVEELLIGGEMIKAGTLTLEEVRRKFNNGEINF (SEQ ID NO.88)

[0228] When, in one FokICD monomer of a dimer, the aspartic acid at position 450 (counting the first amino acid of wild-type FokI when including the recognition domain as amino acid 1; if only counting the first amino acid of FokICD of the cleavage domain as amino acid 1, it is position 67) and / or position 467 (counting the first amino acid of wild-type FokI when including the recognition domain as amino acid 1; if only counting the first amino acid of FokICD of the cleavage domain as amino acid 1, it is position 84) is mutated to alanine (D450A or D467A), this FokICD monomer will lose its cleavage activity, while the other FokICD monomer in the dimer that has not undergone amino acid mutation still retains its cleavage activity. At this time, the obtained FokICD dimer can and can only cleave one strand of double-stranded DNA and cannot cleave the other strand. This dimer of FokICD is called FokI nickase , that is, FokI nickase. For ease of description, the inventors refer to the FokICD monomer fused with TALE-L as FokI-L (such as shown in SEQ ID NO.87), and the FokICD monomer fused with TALE-R as FokI-R (such as shown in SEQ ID NO.88). Further, the FokICD mutant monomers containing the FokI D450A and / or D467A mutations and thus losing cleavage activity are respectively called FokI-L D450A / D467A and FokI-R D450A / D467A . In the present invention, the FokICD dimer formed by the interaction of FokI-L D450A / D467A and FokI-R nickase only retains the cleavage activity of FokI-L, and this dimer is called FokI-L D450A / D467A (or FokI-L nickase); correspondingly, the FokICD dimer formed by the interaction of FokI-L D450A / D467A and FokI-R only retains the cleavage activity of FokI-R and is called FokI-R nickase (or FokI-R nickase).

[0229] It should be noted that FokI-L nickase and FokI-R nickase tend to cleave different single strands in double-stranded DNA, that is, FokI-L nickase and FokI-R nickase have single-strand specificity or preference when cleaving DNA. As Figure 1 shown, at this target site, the used FokI-R nickase tends to cleave strand B. Correspondingly, if FokI-L nickase is used, it tends to cleave strand A( Figure 1 shown). FokI-Lnickase and FokI-R nickase The strand specificity exhibited is conducive to selecting the required single-stranded DNA for the subsequent deamination step. Along with the sequence-specific binding of TALE-L and TALE-R to the left and right binding sites, FokI-L nickase or FokI-R nickase cuts the target sequence, leaving a nick on strand A or strand B respectively. The strand specificity of the nickase determines the further deamination of the single-stranded DNA under the action of the base editor of the present invention.

[0230] The nickase protein monomers that can be used as exemplary nucleic acid base editor components of the present invention are provided below, including but not limited to those shown in SEQ ID NO.60-63:

[0231] FokI-L D450A :

[0232] QLVKSELEEKKSELRHKLKYVPHEYIELIEIARNSTQDRILEMKVMEFFMKVYGYRGKHLGGSRKPAGAIYTVGSPIDYGVIVDTKAYSGGYNLPIGQADEMQRYVKENQTRNKHINPNEWWKVYPSSVTEFKFLFVSGHFKGNYKAQLTRLNHKTNCNGAVLSVEELLIGGEMIKAGTLTLEEVRRKFNNGEINF(SEQ ID NO.60)

[0233] FokI-L D467A :

[0234] QLVKSELEEKKSELRHKLKYVPHEYIELIEIARNSTQDRILEMKVMEFFMKVYGYRGKHLGGSRKPDGAIYTVGSPIDYGVIVATKAYSGGYNLPIGQADEMQRYVKENQTRNKHINPNEWWKVYPSSVTEFKFLFVSGHFKGNYKAQLTRLNHKTNCNGAVLSVEELLIGGEMIKAGTLTLEEVRRKFNNGEINF(SEQ ID NO.61)

[0235] FokI-R D450A :

[0236] QLVKSELEEKKSELRHKLKYVPHEYIELIEIARNSTQDRILEMKVMEFFMKVYGYRGKHLGGSRKPAGAIYTVGSPIDYGVIVDTKAYSGGYNLPIGQADEMERYVEENQTRNKHLNPNEWWKVYPSSVTEFKFLFVSGHFKGNYKAQLTRLNHITNCNGAVLSVEELLIGGEMIKAGTLTLEEVRRKFNNGEINF(SEQ ID NO.62)

[0237] FokI-R D467A :

[0238] QLVKSELEEKKSELRHKLKYVPHEYIELIEIARNSTQDRILEMKVMEFFMKVYGYRGKHLGGSRKPDGAIYTVGSPIDYGVIVATKAYSGGYNLPIGQADEMERYVEENQTRNKHLNPNEWWKVYPSSVTEFKFLFVSGHFKGNYKAQLTRLNHITNCNGAVLSVEELLIGGEMIKAGTLTLEEVRRKFNNGEINF(SEQ ID NO.63)

[0239] Exonuclease

[0240] Depending on the type of exonuclease used, the exonuclease component of the nucleic acid base editor of the present invention digests the cleaved DNA strand in the 5'→3' direction or in the 3'→5' direction from the nick position. After exonuclease digestion, short ssDNA fragments are exposed on the complementary DNA strand. The type of exonuclease determines the ssDNA region (or editing window) to be deaminated. Exonucleases that can be components of the base editors disclosed herein include, but are not limited to, DNA polymerases I and III (E. coli), mammalian p53 protein, exonucleases I-VII (E. coli) (such as exonucleases I and V (with 3'→5' exonuclease activity)), phage-derived polymerases (such as T4 DNA polymerase (with 3'→5' exonuclease activity)), Thermus aquaticus polymerase (with 5'→3' exonuclease activity), and the 3'→5' exonuclease reported by Shevelev and Hübscher (Shevelev & Hübscher, 2002, Nat. Rev. Molec. Cell Biol. 3:364-376).

[0241] Exonuclease proteins that can be used as exemplary base editor components of the present invention include, but are not limited to, the proteins shown in SEQ ID NO.64-67 and 153.

[0242] Exonuclease V (ExoV):

[0243] MAETGEEETASAEASGFSDLSDSELVEFLDLEEAKESAVSLSKPGPSAELPGKDDKPVSLQNWKGGLDVLSPMERFHLKYLYVTDLCTQNWCELQMVYGKELPGSLTPEKAAVLDTGASIHLAKELELHDLVTVPIATKEDAWAVKFLNILAMIPALQSEGRVREFPVFGEVEGIFLVGVIDELHYTSKGELELAELKTRRRPVLPLPAQKKKDYFQVSLYKYIFDAMVQGKVTPASLIHHTKLCLDKPLGPSVLRHARQGGVSVKSLGDLMELVFLSLTLSDLPAIDTLKLEYIHQETATILGTEIVAFEEKEVKSKVQHYVAYWMGHRDPQGVDVEEAWKCRTCDYVDICEWRRGSGVLSSSWEPKAKKFK(SEQ ID NO.153)

[0244] mExoI:

[0245] MGIQGLLQFIQEASEPVNVKKYKGQAVAVDTYCWLHKGAIACAEKLAKGEPTDRYVGFCMKFVNMLLSYGVKPILIFDGCTLPSKKEVERSRRERRQSNLLKGKQLLREGKVSEARDCFARSINITHAMAHKVIKAARALGVDCLVAPYEADAQLAYLNKAGIVQAVITEDSDLLAFGCKKVILKMDQFGNGLEVDQARLGMCKQLGDVFTEEKFRYMCILSGCDYLASLRGIGLAKACKVLRLANNPDIVKVIKKIGHYLRMNITVPEDYITGFIRANNTFLYQLVFDPIQRKLVPLNAYGDDVNPETLTYAGQYVGDSVALQIALGNRDVNTFEQIDDYSPDTMPAHSRSHSWNEKAGQKPPGTNSIWHKNYCPRLEVNSVSHAPQLKEKPSTLGLKQVISTKGLNLPRKSCVLKRPRNEALAEDDLLSQYSSVSKKIKENGCGDGTSPNSSKMSKSCPDSGTAHKTDAHTPSKMRNKFATFLQRRNEESGAVVVPGTRSRFFCSSQDFDNFIPKKESGQPLNETVATGKATTSLLGALDCPDTEGHKPVDANGTHNLSSQIPGNAAVSPEDEAQSSETSKLLGAMSPPSLGTLRSCFSWSGTLREFSRTPSPSASTTLQQFRRKSDPPACLPEASAVVTDRCDSKSEMLGETSQPLHELGCSSRSQESMDSSCGLNTSSLSQPSSRDSGSEESDCNNKSLDNQGEQNSKQHLPHFSKKDGLRRNKVPGLCRSSSMDSFSTTKIKPLVPARVSGLSKKSGSMQTRKHHDVENKPGLQTKISELWKNFGFKKDSEKLPSCKKPLSPVKDNIQLTPETEDEIFNKPECVRAQRAIFH(SEQ ID NO.64)

[0246] mTrex2:

[0247] MSEPPRAETFVFLDLEATGLPNMDPEIAEISLFAVHRSSLENPERDDSGSLVLPRVLDKLTLCMCPERPFTAKASEITGLSSESLMHCGKAGFNGAVVRTLQGFLSRQEGPICLVAHNGFDYDFPLLCTELQRLGAHLPQDTVCLDTLPALRGLDRAHSHGTRAQGRKSYSLASLFHRYFQAEPSAAHSAEGDVHTLLLIFLHRAPELLAWADEQARSWAHIEPMYVPPDGPSLEA(SEQ ID NO.65)

[0248] mArtimes:

[0249] MSSGMAYTSDRDRNKARAYSHCHKDHMKGRASKRRCSKVYCSVTKTSKYRWNRTTTSVDASGKVVVTAGHCGSVMGSNGTVYTGDRAKGASRMHSGGRVKDSVYDTTCDRYSRCRGVRSWVTRSHHVVWNCKAAYGYYTNSGVVHVDKDMKNMDHHTTDRNTHACRHKACWNKCGTSNKTAHTSKSTMWGRTRKTNVVRTGSSYRACSHSSSKDSYCVNVYNVVGTVDKVMDVKCRSSVKYKGKKRARTHDSDDDDDTRHKVYTSMKADRSGGCKASVWSSANDCSNSDSGTSGGGSTVNADDVDWVKRRDTGCHSSTGGSSKCSDSKCSDSKCSDSDGDSTHSSNSSSTHTDGSGWDSCDTVSSKSGGDSTSNKGAYKKKSSASDACDTHCDKSRAVNGACVDTSGRKSKTSSTRADSSSSDSTATHCYRKATGSVVKRKCSDS(SEQID NO.66)

[0250] T5 exo:

[0251] MSKSWGKFIEEEEAEMASRRNLMIVDGTNLGFRFKHNNSKKPFASSYVSTIQSLAKSYSARTTIVLGDKGKSVFRLEHLPEYKGNRDEKYAQRTEEEKALDEQFFEYLKDAFELCKTTFPTFTIRGVEADDMAAYIVKLIGHLYDHVWLISTDGDWDTLLTDKVSRFSFTTRREYHLRDMYEHHNVDDVEQFISLKAIMGDLGDNIRGVEGIGAKRGYNIIREFGNVLDIIDQLPLPGKQKYIQNLNASEELLFRNLILVDLPTYCVDAIAAVGQDVLDKFTKDILEIAEQ(SEQ ID NO.67)

[0252] Deaminase

[0253] Deaminases that can be used as components of the base editors of the present invention include cytosine deaminases and adenine deaminases. Cytosine deaminases include, but are not limited to, hAPOBEC3A (Zong et al., 2018, Nat. Biotechnol. Oct 1. doi:10.1038 / nbt.4261), rAPOBEC1, C57, Sdd (Huang J et al., 2023, Cell, doi:10.1101 / 2023.05.21.541555), which result in the conversion of the base site from C to T. Optional adenine deaminases include TadA-8e (Richter et al., 2020, Nat. Biotechnol. 38:883-891), which produces the conversion of the base site from A to G.

[0254] The following provides deaminases that can be used as exemplary base editor components of the present invention, including but not limited to the deaminases listed in Table 1 (such as the proteins shown in SEQ ID NO. 36-59, 80-86):

[0255] Table 1 Types of Deaminases

[0256]

[0257] Deaminase rAPOBEC1:

[0258] MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSIWRHTSQNTNKHVEVNFIEKFTTERYFCPNTRCSITWFLSWSPCGECSRAITEFLSRYPHVTLFIYIARLYHHADPRNRQGLRDLISSGVTIQIMTEQESGYCWRNFVNYSPSNEAHWPRYPHLWVRLYVLELYCIILGLPPCLNILRRKQPQLTFFTIALQSCHYQRLPPHILWATGLK(SEQ ID NO.36)

[0259] hAPOBEC3A:

[0260] MEASPASGPRHLMDPHIFTSNFNNGIGRHKTYLCYEVERLDNGTSVKMDQHRGFLHNQAKNLLCGFYGRHAELRFLDLVPSLQLDPAQIYRVTWFISWSPCFSWGCAGEVRAFLQENTHVRLRIFAARIYDYDPLYKEALQMLRDAGAQVSIMTYDEFKHCWDTFVDHQGCPFQPWDGLDEHSQALSGRLRAILQNQGN(SEQ ID NO.37)

[0261] hAPOBEC3G-CTD:

[0262] MDPPTFTFNFNNEPWVRGRHETYLCYEVERMHNDTWVLLNQRRGFLCNQAPHKHGFLEGRHAELCFLDVIPFWKLDLDQDYRVTCFTSWSPCFSCAQEMAKFISKNKHVSLCIFTARIYDDQGRCQEGLRTLAEAGAKISIMTYSEFKHCWDTFVDHQGCPFQPWDGLDEHSQDLSGRLRAILQNQEN(SEQ ID NO.38)

[0263] PmCDA1:

[0264] MTDAEYVRIHEKLDIYTFKKQFFNNKKSVSHRCYVLFELKRRGERRACFWGYAVNKPQSGTERGIHAEIFSIRKVEEYLRDNPGQFTINWYSSWSPCADCAEKILEWYNQELRGNGHTLKIWACKLYYEKNARNQIGLWNLRDNGVGLNVMVSEHYQCCRKIFIQSSHNQLNENRWLEKTLKRAEKRRSELSIMIQVKILHTTKSPAVSRGSG(SEQ IDNO.39)

[0265] tCDA1EQ:

[0266] SHRCYVLFELKRRGERRACFWGYAVNKPQSGTERGIHAEIFSIRKVEEYLRDNPGQFTINWYSSWSPCADCAEKILEWYNQELRGNGHTLKIEACKLYYEKNARNQIGLQNLRDNGVGLNV(SEQ ID NO.40)

[0267] hAID:

[0268] MDSLLMNRRKFLYQFKNVRWAKGRRETYLCYVVKRRDSATSFSLDFGYLRNKNGCHVELLFLRYISDWDLDPGRCYRVTWFTSWSPCYDCARHVADFLRGNPNLSLRIFTARLYFCEDRKAEPEGLRRLHRAGVQIAIMTFKDYFYCWNTFVENHERTFKAWEGLHENSVRLSRQLRRILLPLYEVDDLRDAFRTLGL(SEQ ID NO.41)

[0269] PpAPOBEC1:

[0270] MTSEKGPSTGDPTLRRRIESWEFDVFYDPRELRKETCLLYEIKWGMSRKIWRSSGKNTTNHVEVNFIKKFTSERRFHSSISCSITWFLSWSPCWECSQAIREFLSQHPGVTLVIYVARLFWHMDQRNRQGLRDLVNSGVTIQIMRASEYYHCWRNFVNYPPGDEAHWPQYPPLWMMLYALELHCIILSLPPCLKISRRWQNHLAFFRLHLQNCHYQTIPPHILLATGLIHPSVTWR(SEQ ID NO.42)

[0271] RrA3F:

[0272] MKPQIRDHRPNPMEAMYPHIFYFHFENLEKAYGRNETWLCFTVEIIKQYLPVPWKKGVFRNQVDPETHCHAEKCFLSWFCNNTLSPKKNYQVTWYTSWSPCPECAGEVAEFLAEHSNVKLTIYTARLYYFWDTDYQEGLRSLSEEGASVEIMDYEDFQYCWENFVYDDGEPFKRWKGLKYNFQSLTRRLREILQ(SEQ ID NO.43)

[0273] AmAPOBEC1:

[0274] MADSSEKMRGQYISRDTFEKNYKPIDGTKEAHLLCEIKWGKYGKPWLHWCQNQRMNIHAEDYFMNNIFKAKKHPVHCYVTWYLSWSPCADCASKIVKFLEERPYLKLTIYVAQLYYHTEEENRKGLRLLRSKKVIIRVMDISDYNYCWKVFVSNQNGNEDYWPLQFDPWVKENYSRLLDIFWESKCRSPNPW(SEQ ID NO.44)

[0275] SsAPOBEC3B:

[0276] MDPQRLRQWPGPGPASRGGYGQRPRIRNPEEWFHELSPRTFSFHFRNLRFASGRNRSYICCQVEGKNCFFQGIFQNQVPPDPPCHAELCFLSWFQSWGLSPDEHYYVTWFISWSPCCECAAKVAQFLEENRNVSLSLSAARLYYFWKSESREGLRRLSDLGAQVGIMSFQDFQHCWNNFVHNLGMPFQPWKKLHKNYQRLVTELKQILREEPATYGSPQAQGKVRIGSTAAGLRHSHSHTRSEAHLRPNHSSRQHRILNPPREARARTCVLVDASWICYR(SEQ ID NO.45)

[0277] hA3B:

[0278] MNPQIRNPMERMYRDTFYDNFENEPILYGRSYTWLCYEVKIKRGRSNLLWDTGVFRGQVYFKPQYHAEMCFLSWFCGNQLPAYKCFQITWFVSWTPCPDCVAKLAEFLSEHPNVTLTISAARLYYYWERDYRRALCRLSQAGARVKIMDYEEFAYCWENFVYNEGQQFMPWYKFDENYAFLHRTLKEILRYLMDPDTFTFNFNNDPLVLRRRQTYLCYEVERLDNGTWVLMDQHMGFLCNEAKNLLCGFYGRHAELRFLDLVPSLQLDPAQIYRVTWFISWSPCFSWGCAGEVRAFLQENTHVRLRIFAARIYDYDPLYKEALQMLRDAGAQVSIMTYDEFEYCWDTFVYRQGCPFQPWDGLEEHSQALSGRLRAILQNQGN(SEQ ID NO.46)

[0279] hA3C:

[0280] MNPQIRNPMKAMYPGTFYFQFKNLWEANDRNETWLCFTVEGIKRRSVVSWKTGVFRNQVDSETHCHAERCFLSWFCDDILSPNTKYQVTWYTSWSPCPDCAGEVAEFLARHSNVNLTIFTARLYYFQYPCYQEGLRSLSQEGVAVEIMDYEDFKYCWENFVYNDNEPFKPWKGLKTNFRLLKRRLRESLQ(SEQ ID NO.47)

[0281] hA3D:

[0282] MNPQIRNPMERMYRDTFYDNFENEPILYGRSYTWLCYEVKIKRGRSNLLWDTGVFRGPVLPKRQSNHRQEVYFRFENHAEMCFLSWFCGNRLPANRRFQITWFVSWNPCLPCVVKVTKFLAEHPNVTLTISAARLYYYRDRDWRWVLLRLHKAGARVKIMDYEDFAYCWENFVCNEGQPFMPWYKFDDNYASLHRTLKEILRNPMEAMYPHIFYFHFKNLLKACGRNESWLCFTMEVTKHHSAVFRKRGVFRNQVDPETHCHAERCFLSWFCDDILSPNTNYEVTWYTSWSPCPECAGEVAEFLARHSNVNLTIFTARLCYFWDTDYQEGLCSLSQEGASVKIMGYKDFVSCWKNFVYSDDEPFKPWKGLQTNFRLLKRRLREILQ(SEQ ID NO.48)

[0283] hA3F:

[0284] MKPHFRNTVERMYRDTFSYNFYNRPILSRRNTVWLCYEVKTKGPSRPRLDAKIFRGQVYSQPEHHAEMCFLSWFCGNQLPAYKCFQITWFVSWTPCPDCVAKLAEFLAEHPNVTLTISAARLYYYWERDYRRALCRLSQAGARVKIMDDEEFAYCWENFVYSEGQPFMPWYKFDDNYAFLHRTLKEILRNPMEAMYPHIFYFHFKNLRKAYGRNESWLCFTMEVVKHHSPVSWKRGVFRNQVDPETHCHAERCFLSWFCDDILSPNTNYEVTWYTSWSPCPECAGEVAEFLARHSNVNLTIFTARLYYFWDTDYQEGLRSLSQEGASVEIMGYKDFKYCWENFVYNDDEPFKPWKGLKYNFLFLDSKLQEILE(SEQ ID NO.49)

[0285] hA3G:

[0286] MKPHFRNTVERMYRDTFSYNFYNRPILSRRNTVWLCYEVKTKGPSRPPLDAKIFRGQVYSELKYHPEMRFFHWFSKWRKLHRDQEYEVTWYISWSPCTKCTRDMATFLAEDPKVTLTIFVARLYYFWDPDYQEALRSLCQKRDGPRATMKIMNYDEFQHCWSKFVYSQRELFEPWNNLPKYYILLHIMLGEILRHSMDPPTFTFNFNNEPWVRGRHETYLCYEVERMHNDTWVLLNQRRGFLCNQAPHKHGFLEGRHAELCFLDVIPFWKLDLDQDYRVTCFTSWSPCFSCAQEMAKFISKNKHVSLCIFTARIYDDQGRCQEGLRTLAEAGAKISIMTYSEFKHCWDTFVDHQGCPFQPWDGLDEHSQDLSGRLRAILQNQEN(SEQ ID NO.50)

[0287] hA3H:

[0288] MALLTAETFRLQFNNKRRLRRPYYPRKALLCYQLTPQNGSTPTRGYFENKKKCHAEICFINEIKSMGLDETQCYQVTCYLTWSPCSSCAWELVDFIKAHDHLNLRIFASRLYYHWCKPQQDGLRLLCGSQVPVEVMGFPEFADCWENFVDHEKPLSFNPYKMLEELDKNSRAIKRRLDRIKS(SEQ ID NO.51)

[0289] hA3Bctd:

[0290] MEILRYLMDPDTFTFNFNNDPLVLRRRQTYLCYEVERLDNGTWVLMDQHMGFLCNEAKNLLCGFYGRHAELRFLDLVPSLQLDPAQIYRVTWFISWSPCFSWGCAGEVRAFLQENTHVRLRIFAARIYDYDPLYKEALQMLRDAGAQVSIMTYDEFEYCWDTFVYRQGCPFQPWDGLEEHSQALSGRLRAILQNQGN(SEQ ID NO.52)

[0291] FERNY:

[0292] FERNYDPRELRKETYLLYEIKWGKSGKLWRHWCQNNRTQHAEVYFLENIFNARRFNPSTHCSITWYLSWSPCAECSQKIVDFLKEHPNVNLEIYVARLYYHEDERNRQGLRDLVNSGVTIRIMDLPDYNYCWKTFVSDQGGDEDYWPGHFAPWIKQYSLKL(SEQ ID NO.53)

[0293] ecTadA:

[0294] MSEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRVVFGARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKAQKKAQSSTD(SEQ ID NO.54)

[0295] mADA:

[0296] MAQTPAFNKPKVELHVHLDGAIKPETILYFGKKRGIALPADTVEELRNIIGMDKPLSLPGFLAKFDYYMPVIAGCREAIKRIAYEFVEMKAKEGVVYVEVRYSPHLLANSKVDPMPWNQTEGDVTPDDVVDLVNQGLQEGEQAFGIKVRSILCCMRHQPSWSLEVLELCKKYNQKTVVAMDLAGDETIEGSSLFPGHVEAYEGAVKNGIHRTVHAGEVGSPEVVREAVDILKTERVGHGYHTIEDEALYNRLLKENMHFEVCPWSSYLTGAWDPKTTHAVVRFKNDKANYSLNTDDPLIFKSTLDTDYQMTKKDMGFTEEEFKRLNINAAKSSFLPEEEKKELLERLYREYQ(SEQ ID NO.55)

[0297] hADAR2:

[0298] MHLDQTPSRQPIPSEGLQLHLPQVLADAVSRLVLGKFGDLTDNFSSPHARRKVLAGVVMTTGTDVKDAKVISVSTGTKCINGEYMSDRGLALNDCHAEIISRRSLLRFLYTQLELYLNNKDDQKRSIFQKSERGGFRLKENVQFHLYISTSPCGDARIFSPHEPILEEPADRHPNRKARGQLRTKIESGEGTIPVRSNASIQTWDGVLQGERLLTMSCSDKIARWNVVGIQGSLLSIFVEPIYFSSIILGSLYHGDHLSRAMYQRISNIEDLPPLYTLNKPLLSGISNAEARQPGKAPNFSVNWTVGDSAIEVINATTGKDELGRASRLCKHALYCRWMRVHGKVPSHLLRSKITKPNVYHESKLAAKEYQAAKARLFTAFIKAGLGAWVEKPTEQDQFSLTP(SEQ ID NO.56)

[0299] hADAT2:

[0300] MEAKAAPKPAASGACSVSAEETEKWMEEAMHMAKEALENTEVPVGCLMVYNNEVVGKGRNEVNQTKNATRHAEMVAIDQVLDWCRQSGKSPSEVFEHTVLYVTVEPCIMCAAALRLMKIPLVVYGCQNERFGGCGSVLNIASADLPNTGRPFQCIPGYRAEEAVEMLKTFYKQENPNAPKSKVRKKECQKS(SEQ ID NO.57)

[0301] ecTadA*(7.10):

[0302] MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD(SEQ ID NO.58)

[0303] TadA*ABE8e (TadA-8e):

[0304] MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN(SEQ ID NO.59)

[0305] Sdd2

[0306] MAPDSLVWFDPLGLIVLQQVPYNDHPLFGAVSEFIQGKSRSDLRGRNVAAVLLDDGTVIVRASEGGGNHAERVLMGLSEVDPAKVVAVYTERSPCTGRINCHDLLDSSLGADVPVYYTHEMIRGQEGKTAQQIEADRNQFCRGG(SEQ ID NO.80)

[0307] Sdd3

[0308] MSASAQLNTYLAAIGNSTTTVEAQPEAAPPPAAAESLDSTPRLPDGGIDFHALAKRLGLLEARPTEQPPFDPRRFNPACWQGLKPYDQAGTAEGNLFIAPGKRWNTRPMQASKLEVGPQSDLHPQWRSRKAPWHIEGKIAAYMRQKGFTDGCVYLNARPCSGPDGCARNLPDLLPVGSTLHVHARYIDRTGETRFYYREYRGTGKALT(SEQ IDNO.81)

[0309] Sdd4

[0310] MLDAMDAYLSEIAGGNAPARAGPKAPEPKQPGGSSSPRARDGRIDFRALLERLKAQGVVGLEGRSDDPIPDFDPKKQNPACYQGLAPRQKGKPVRGNLFFPDGRRWNDVALESSRGEPAFDLNIIKPEYRSLSPARGHLEGNVAAWMRSTFHQEMVLYINESPCRKHGKGCLYTLEHFLPRGYVLHVWSRNDRGEWRGNTFRGSGEAFTEGA(SEQ IDNO.82)

[0311] Sdd6

[0312] MVETRDKIIAAKSRSDAGLLAFQQATNGSIDSRPAEAIANLQRAKTHLDEAQRLVANSDAAVDNYINAILGGASAATAQPSAVIPASKPSRFKPMRTDPAKADEIRPHVGKDRAVATLWDADGNRVLGLHSADDDGPAATAAWKPPWRDYVRLRRHVEAHAAARMHQDGHKTMVMYINLPPCKYFDGCKLNLEDILPKGSTLWMHRVFQNGGTKIYQFNGTGRAYV(SEQ ID NO.83)

[0313] Sdd7(also designated as C57 in this specification)

[0314] MLEAVRARLIGEGGGPGAVPEGGDGPPAVPAEEVERLRGELPPPVVPGTGQKTHGRWIGPDGRVRAIVSGRDEDAALVHAQLAAKGIPDEPTRNSDVEQKLAAHMVANGIRHVTLVINHRPCRGFDDSCDTLVPIILPEGCTLTVHGQTDKGMRVRVRYTGGARPWWS(SEQ ID NO.84)

[0315] Sdd10

[0316] MLDAALGAVRRIIAALGTSGAERASPGANGSERVDELAERLPPTVVPNTSAKTHGWWFTGQGAAQELISGEGPDARAAYEALREEGYPRPGMPFVAMHVEIKLAAHMRRNDIEHATVVINNIPCPLVWGCENLIGVVLPEGSSLTVHGSNGYERTFTGGRKPPWPR(SEQ ID NO.85)

[0317] Sdd59

[0318] MLLTPPPRPAAPPTTRPKPLVARTGDAYPPGTEWALPLIVQPHPPVGGTVPVEGHVRALRPESQISHVFHPGGGHWTEQARARLRVLPGFGWAVNLGHHVELQIAAWMTACGIHHAELVLNRPPCGERYGLGCHQALPVLLPRGYRLTVSSTRGGPQPYQHHYEGKA(SEQ ID NO.86)

[0319] Uracil glycosylase inhibitor (UGI)

[0320] In some embodiments, when using a cytosine deaminase, a uracil glycosylase inhibitor (UGI) is fused to the N-terminus of the deaminase, while UGI is not required when using an adenine deaminase.

[0321] Exemplary UGI proteins that can be used in the base editor in the present invention are disclosed below, including but not limited to the protein shown in SEQ ID NO.68:

[0322] TNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML(SEQ ID NO.68)

[0323] Nuclear localization sequence (NLS)

[0324] In some embodiments of the present invention, the NLS of the fusion protein of the present invention can be located at the N-terminus and / or C-terminus. In some embodiments of the present invention, the NLS of the fusion protein of the present invention can be located between the adenine deaminase domain, cytosine deaminase domain, nucleic acid targeting domain, and / or UGI. In some embodiments, the fusion protein comprises about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLSs. In some embodiments, the fusion protein comprises about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLSs at or near the N-terminus. In some embodiments, the fusion protein comprises about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLSs at or near the C-terminus. In some embodiments, the polypeptide comprises a combination of these, such as one or more NLSs at the N-terminus and one or more NLSs at the C-terminus. When there are more than one NLSs, each can be selected independently of the other NLSs.

[0325] Generally, an NLS consists of one or more short sequences of positively charged lysine or arginine exposed on the protein surface, but other types of NLSs are also known. Non-limiting examples of NLSs include: KKRKV (SEQ ID NO. 150), PKKKRKV (SEQ ID NO. 151), or KRPAATKKAGQAKKKK (SEQ ID NO. 152).

[0326] Recombinant expression construct

[0327] Each component in the base editor of the present invention can be expressed separately, or can be expressed as one or more fusion proteins, or the above-mentioned elements or components can be expressed separately or jointly using a recombinant expression construct used in recombinant genetic engineering technology. Exemplary recombinant expression constructs of the present invention, for example Figure 16A to Figure 16E and Figure 17A to Figure 17H are described.

[0328] The types, functions, and references of the genes and regulatory elements in the above-mentioned exemplary recombinant expression constructs ( Figure 16A to Figure 16E and Figure 17A to 17H ) are explained and illustrated by examples as shown in Table 2 below:

[0329] Table 2: Examples of Construct Genes and Regulatory Elements

[0330]

[0331]

[0332]

[0333]

[0334]

[0335] Specifically, the genes and regulatory elements in the exemplary recombinant constructs used in the present invention include, but are not limited to, the following sequences: promoter sequences such as those shown in SEQ ID NO. 69-72; terminator sequences such as those shown in SEQ ID NO. 73-76; mitochondrial targeting sequences (MTS) such as those shown in SEQ ID NO. 77-78; chloroplast transit peptide (CTP) sequences such as that shown in SEQ ID NO. 79.

[0336] UBI promoter:

[0337]

[0338] CaMV 35S promoter(enhanced):

[0339] TGAGACTTTTCAACAAAGGGTAATATCGGGAAACCTCCTCGGATTCCATTGCCCAGCTATCTGTCACTTCATCAAAAGGACAGTAGAAAAGGAAGGTGGCACCTACAAATGCCATCATTGCGATAAAGGAAAGGCTATCGTTCAAGATGCCTCTGCCGACAGTGGTCCCAAAGATGGACCCCCACCCACGAGGAGCATCGTGGAAAAAGAAGACGTTCCAACCACGTCTTCAAAGCAAGTGGATTGATGTGATAACATGGTGGAGCACGACACTCTCGTCTACTCCAAGAATATCAAAGATACAGTCTCAGAAGACCAAAGGGCTATTGAGACTTTTCAACAAAGGGTAATATCGGGAAACCTCCTCGGATTCCATTGCCCAGCTATCTGTCACTTCATCAAAAGGACAGTAGAAAAGGAAGGTGGCACCTACAAATGCCATCATTGCGATAAAGGAAAGGCTATCGTTCAAGATGCCTCTGCCGACAGTGGTCCCAAAGATGGACCCCCACCCACGAGGAGCATCGTGGAAAAAGAAGACGTTCCAACCACGTCTTCAAAGCAAGTGGATTGATGTGATATCTCCACTGACGTAAGGGATGACGCACAATCCCACTATCCTTCGCAAGACCTTCCTCTATATAAGGAAGTTCATTTCATTTGGAGAGGACACGCTGA(SEQ ID NO.70)

[0340] CaMV 2 x 35S promoter

[0341] CCTGCAGGTCAACATGGTGGAGCACGACACACTTGTCTACTCCAAAAATATCAAAGATACAGTCTCAGAAGACCAAAGGGCAATTGAGACTTTTCAACAAAGGGTAATATCCGGAAACCTCCTCGGATTCCATTGCCCAGCTATCTGTCACTTTATTGTGAAGATAGTGGAAAAGGAAGGTGGCTCCTACAAATGCCATCATTGCGATAAAGGAAAGGCCATCGTTGAAGATGCCTCTGCCGACAGTGGTCCCAAAGATGGACCCCCACCCACGAGGAGCATCGTGGAAAAAGAAGACGTTCCAACCACGTCTTCAAAGCAAGTGGATTGATGTGATAACATGGTGGAGCACGACACACTTGTCTACTCCAAAAATATCAAAGATACAGTCTCAGAAGACCAAAGGGCAATTGAGACTTTTCAACAAAGGGTAATATCCGGAAACCTCCTCGGATTCCATTGCCCAGCTATCTGTCACTTTATTGTGAAGATAGTGGAAAAGGAAGGTGGCTCCTACAAATGCCATCATTGCGATAAAGGAAAGGCCATCGTTGAAGATGCCTCTGCCGACAGTGGTCCCAAAGATGGACCCCCACCCACGAGGAGCATCGTGGAAAAAGAAGACGTTCCAACCACGTCTTCAAAGCAAGTGGATTGATGTGATATCTCCACTGACGTAAGGGATGACGCACAATCCCACTATCCTTCGCAAGACCCTTCCTCTATATAAGGAAGTTCATTTCATTTGGAGAGGACCTCGACCTCAACACAACATATACAAAACAAACGAATCTCAAGCAATCAAGCATTCTACTTCTATTGCAGCAATTTAAATCATTTCTTTTAAAGCAAAAGCAATTTTCTGAAAATTTTCACCATTTACGAACGATA(SEQ ID NO.71)

[0342] CMV promoter:

[0343] GTGATGCGGTTTTGGCAGTACATCAATGGGCGTGGATAGCGGTTTGACTCACGGGGATTTCCAAGTCTCCACCCCATTGACGTCAATGGGAGTTTGTTTTGGCACCAAAATCAACGGGACTTTCCAAAATGTCGTAACAACTCCGCCCCATTGACGCAAATGGGCGGTAGGCGTGTACGGTGGGAGGTCTATATAAGCAGAGCT(SEQ ID NO.72)

[0344] Nos terminator:

[0345] GAATTTCCCCGATCGTTCAAACATTTGGCAATAAAGTTTCTTAAGATTGAATCCTGTTGCCGGTCTTGCGATGATTATCATATAATTTCTGTTGAATTACGTTAAGCATGTAATAATTAACATGTAATGCATGACGTTATTTATGAGATGGGTTTTTATGATTAGAGTCCCGCAATTATACATTTAATACGCGATAGAAAACAAAATATAGCGCGCAAACTAGGATAAATTATCGCGCGCGGTGTCATCTATGTTACT(SEQ ID NO.73)

[0346] E9 terminator:

[0347] AGAGCTTTCGTTCGTATCATCGGTTTCGACAACGTTCGTCAAGTTCAATGCATCAGTTTCATTGCGCACACACCAGAATCCTACTGAGTTTGAGTATTATGGCATTGGGAAAACTGTTTTTCTTGTACCATTTGTTGTGCTTGTAATTTACTGTGTTTTTTATTCGGTTTTCGCTATCGAACTGTGAAATGGAAATGGATGGAGAAGAGTTAATGAATGATATGGTCCTTTTGTTCATTCTCAAATTAATATTATTTGTTTTTTCTCTTATTTGTTGTGTGTTGAATTTGAAATTATAAGAGATATGCAAACATTTTGTTTTGAGTAAAAATGTGTCAAATCGTGGCCTCTAATGACCGAAGTTAATATGAGGAGTAAAACACTTGTAGTTGTACCATTATGCTTATTCACTAGGCAACAAATATATTTTCAGACCTAGAAAAGCTGCAAATGTTACTGAATACAAGTATGTCCTCTTGTGTTTTAGACATTTATGAACTTTCCTTTATGTAATTTTCCAGAATCCTTGTCAGATTCTAATCATTGCTTTATAATTATAGTTATACTCATGGATTTGTAGTTGAGTATGAAAATATTTTTTAATGCATTTTATGACTTGCCAATTGATTGACAAC(SEQ ID NO.74)

[0348] CaMV poly(A) signal:

[0349] TTTCTCCATAATAATGTGTGAGTAGTTCCCAGATAAGGGAATTAGGGTTCCTATAGGGTTTCGCTCATGTGTTGAGCATATAAGAAACCCTTAGTATGTATTTGTATTTGTAAAATACTTCTATCAATAAAATTTCTAATTCCTAAAACCAAAATCCAGTACTAAAATCCAGATC(SEQ ID NO.75)

[0350] bGH poly(A) signal:

[0351] CTGTGCCTTCTAGTTGCCAGCCATCTGTTGTTTGCCCCTCCCCCGTGCCTTCCTTGACCCTGGAAGGTGCCACTCCCACTGTCCTTTCCTAATAAAATGAGGAAATTGCATCGCATTGTCTGAGTAGGTGTCATTCTATTCTGGGGGGTGGGGTGGGGCAGGACAGCAAGGGGGAGGATTGGGAAGACAATAGCAGGCATGCTGGGGATGCGGTGGGCTCTATGG(SEQ ID NO.76).

[0352] SOD2 MTS:

[0353] MLSRAVCGTSRQLAPVLGYLGSRQKHSLPD(SEQ ID NO.77)

[0354] COX8 MTS:

[0355] MSVLTPLLLRGLTGSARRLPVPRAK(SEQ ID NO.78)

[0356] CTP:

[0357] MAPTVMMASSATAVAPFQGLKSAASLPVARRSTRSLGNVSNGGRIRCMQ(SEQ IDNO.79)

[0358] Target cells of interest

[0359] The recombinant expression constructs provided by the present invention can be produced according to genetic engineering methods known in the art. In some embodiments, a base editor or its recombinant expression construct is introduced into a cell to edit and express a target gene, thereby forming an edited genetically engineered cell.

[0360] Any cell from any organism can be used together with the nucleic acids, polypeptides, compositions, and methods described in the present invention. Cells include, but are not limited to, human, non-human, animal, mammalian, bacterial, protist, fungal, insect, yeast, unconventional yeast, and plant cells, including monocotyledonous and dicotyledonous plants and plant components, as well as plants and seeds produced by the methods described in the present invention. In some aspects, the cells of the organism are somatic cells.

[0361] In some embodiments, animal cells can include, but are not limited to: organisms of the following phyla, which include Chordata, Arthropoda, Mollusca, Annelida, Cnidaria, or Echinodermata; organisms of the following classes, which include mammals, insects, birds, amphibians, reptiles, or fish. In some aspects, the animal is a human, mouse, Caenorhabditis elegans, rat, Drosophila melanogaster, zebrafish, chicken, dog, cow, sheep, pig, guinea pig, hamster, chicken, medaka fish, sea lamprey, pufferfish, tree frog, monkey, or chimpanzee.

[0362] Specific animal cell types include neurons, muscle cells, endocrine or exocrine cells, epithelial cells, muscle cells, tumor cells, hematopoietic cells, bone cells, somatic cells, induced pluripotent stem cells. In some aspects, multiple cells from an organism can be used.

[0363] In some embodiments, the plant cells are derived from monocotyledonous and dicotyledonous plants. Examples of monocotyledonous plants that can be used include, but are not limited to, maize (Zea mays), rice (Oryza sativa), rye (Secale cereale), sorghum (Sorghum bicolor, Sorghum vulgare), millet (e.g., pearl millet, Pennisetum glaucum), proso millet (Panicum miliaceum), foxtail millet (Setaria italica), finger millet (Eleusine coracana), wheat (Triticum species, e.g., Triticum aestivum, Triticum monococcum), sugarcane (Saccharum spp.), oats (Avena), barley (Hordeum), switchgrass (Panicum virgatum), pineapple (Ananas comosus), banana (Musa spp.), palm, ornamental plants, turfgrass, and other grasses. Examples of dicotyledonous plants that can be used include, but are not limited to, soybean (Glycine max), Brassica species (e.g., but not limited to: rapeseed or canola) (Brassica napus, B. campestris, Brassica rapa, Brassica juncea), alfalfa (Medicago sativa), tobacco (Nicotiana tabacum), Arabidopsis (Arabidopsis thaliana), sunflower (Helianthus annuus), cotton (Gossypium arboreum, Gossypium barbadense), peanut (Arachis hypogaea), tomato (Solanum lycopersicum), potato (Solanum tuberosum).Additional plants that can be used include safflower (Carthamus tinctorius), sweet potato (Ipomoea batatas), cassava (Manihot esculenta), coffee (Coffee spp.), coconut (Cocos nucifera), citrus trees (Citrus spp.), cocoa (Theobroma cacao), tea (Camellia sinensis), banana (Musa spp.), avocado (Persea americana), fig (Ficus carica), guava (Psidium guajava), mango (Mangifera indica), olive (Olea europaea), papaya (Carica papaya), cashew (Anacardium occidentale), macadamia (Macadamia integrifolia), almond (Prunus amygdalus), sugar beets (Beta vulgaris), vegetables, ornamental plants, and conifers. Vegetables that can be used include tomato (Lycopersicon esculentum), lettuce (e.g., Lactuca sativa), green beans (Phaseolus vulgaris), lima bean (Phaseolus limensis), peas (Lathyrus spp.), and members of the cucumber genus such as cucumber (C. sativus), cantaloupe (C. cantalupensis), and musk melon (C. melo).Ornamental plants include azaleas (Rhododendron spp.), hydrangeas (Macrophylla hydrangea), hibiscus (Hibiscus rosa-sinensis), roses (Rosa spp.), tulips (Tulipa spp.), daffodils (Narcissus spp.), petunias (Petunia hybrida), carnations (Dianthus caryophyllus), poinsettias (Euphorbia pulcherrima), and chrysanthemums. Conifers that can be used include pines such as loblolly pine (Pinus taeda), slash pine (Pinus elliotii), ponderosa pine (Pinus ponderosa), lodgepole pine (Pinus contorta), and Monterey pine (Pinus radiata); Douglas fir (Pseudotsuga menziesii); Western hemlock (Tsuga canadensis); Sitka spruce (Picea glauca); redwood (Sequoia sempervirens); true firs such as silver fir (Abies amabilis) and balsam fir (Abies balsamea); and cedars such as Western red cedar (Thuja plicata) and Alaska yellow cedar (Chamaecyparis nootkatensis).

[0364] Particular plant cell types include, but are not limited to, cells obtained from whole plants, seedlings, meristems, ground tissue, vascular tissue, dermal tissue, seeds, leaves, roots, buds, stems, flowers, fruits, stolons, bulbs, tubers, corms, apomictic terminal branches, buds, sprouts, tumor tissues, and various forms of cells and cultures (e.g., single cells, protoplasts, embryos, callus). These can be present in plants or in plant organs, tissue cultures, or cell cultures

[0365] Therapeutic applications

[0366] The present invention also encompasses the use of the base editors of the present invention in the treatment of diseases.

[0367] By modifying disease-related genes with the base editor of the present invention, upregulation, downregulation, inactivation, activation, mutation correction, or introduction of disease-related loci of disease-related genes can be achieved, thereby realizing the prevention and / or treatment of diseases and / or the creation of disease-related models. For example, the target nucleic acid region described in the present invention can be located in the protein-coding region of a disease-related gene, or can be located, for example, in a gene expression regulatory region such as a promoter region or an enhancer region, so as to achieve functional modification of the disease-related gene or modification of the expression of the disease-related gene. Therefore, the modification of disease-related genes described herein includes the modification of the disease-related gene itself (such as the protein-coding region), and also includes the modification of its expression regulatory regions (such as promoters, enhancers, introns, etc.).

[0368] A "disease-related" gene refers to any gene that produces transcriptional or translational products at abnormal levels or in abnormal forms in cells derived from disease-affected tissues compared to tissues or cells of non-disease controls. In cases where the altered expression is associated with the occurrence and / or progression of a disease, it can be a gene that is expressed at an abnormally high level; it can be a gene that is expressed at an abnormally low level. A disease-related gene also refers to a gene having one or more mutations or genetic variations in linkage disequilibrium with one or more genes directly responsible for or associated with the etiology of a disease. The mutation or genetic variation is, for example, a single nucleotide variation (SNV). The transcriptional or translational product can be known or unknown and can be at normal or abnormal levels.

[0369] Therefore, the present invention also provides a method for treating a disease in a subject in need thereof, comprising delivering an effective amount of the base editor of the present invention to the subject to modify a gene related to the disease (for example, deaminating mitochondrial DNA by a fusion protein or multiple fusion proteins). The present invention also provides the use of a base editor in the preparation of a pharmaceutical composition for treating a disease in a subject in need thereof, wherein the base editor is used to modify a gene related to the disease. The present invention also provides a pharmaceutical composition for treating a disease in a subject in need thereof, which comprises the base editor of the present invention and optionally a pharmaceutically acceptable carrier, wherein the base editor is used to modify a gene related to the disease.

[0370] In some embodiments, the fusion proteins or base editors described in the present invention are used to introduce point mutations into nucleic acids by deaminating target nuclear bases (e.g., C residues). In some embodiments, deamination of the target nuclear base results in the correction of genetic defects, such as in correcting point mutations that result in loss of function in gene products. In some embodiments, the genetic defect is associated with a disease or disorder (e.g., lysosomal storage disease or metabolic disease, such as, for example, type I diabetes). In some embodiments, the methods provided herein can be used to introduce inactivating point mutations into genes or alleles encoding gene products associated with a disease or disorder.

[0371] In some embodiments, the aim of the protocols described in the present invention is to restore the function of a dysregulated gene via genome editing. The nuclear base editing proteins provided herein are for gene editing in human cells in vitro, such as correcting disease-related mutations in human cell cultures.

[0372] In some embodiments, the aim of the protocols described in the present invention is for the treatment of diseases associated with or caused by point mutations, which can be corrected by the DNA base editing fusion proteins provided herein. In some embodiments, the disease is a proliferative disease. In some embodiments, the disease is a genetic disease. In some embodiments, the disease is a neonatal disease. In some embodiments, the disease is a metabolic disease. In some embodiments, the disease is a lysosomal storage disease.

[0373] In some embodiments, the aim of the protocols described in the present invention is that they can be used to treat mitochondrial diseases or disorders. As used herein, "mitochondrial disease" refers to diseases caused by abnormal mitochondria, such as mitochondrial gene mutations, gene mutations in enzyme pathways, etc. Examples of diseases include, but are not limited to: neurological diseases, loss of motor control, muscle weakness and pain, gastrointestinal diseases and dysphagia, poor growth, heart disease, liver disease, diabetes, respiratory complications, epilepsy, visual / hearing problems, lactic acidosis, developmental delay, and susceptibility to infection.

[0374] Examples of the diseases described in the present invention include, but are not limited to, genetic diseases, circulatory system diseases, muscle diseases, brain, central nervous system and immune system diseases, Alzheimer's disease, secretase disorders, amyotrophic lateral sclerosis (ALS), autism, trinucleotide repeat expansion disorders, hearing diseases, gene targeting therapy for non-dividing cells (neurons, muscles), liver and kidney diseases, epithelial cell and lung diseases, cancer, Usher syndrome or retinitis pigmentosa-39, cystic fibrosis, HIV and AIDS, β-thalassemia, sickle cell disease, herpes simplex virus, autism, drug addiction, age-related macular degeneration, schizophrenia. Other diseases that are treated by correcting point mutations or introducing inactivating mutations into disease-related genes are known to those skilled in the art, and thus the present disclosure is not limited in this regard. In addition to the diseases described exemplarily in the present invention, other related diseases can also be treated with the strategies and fusion proteins provided by the present invention, and such applications are obvious to those skilled in the art. The diseases or targets to which the present invention can be applied refer to the related diseases applicable to the base editors listed in WO2015089465A1 (PCT / US2014 / 070135), WO2016205711A1 (PCT / US2016 / 038181), WO2018141835A1 (PCT / EP2018 / 052491), WO2020191234A1 (PCT / US2020 / 023713), WO2020191233A1 (PCT / US2020 / 023712), WO2019079347A1 (PCT / US2018 / 056146), WO2021155065A1 (PCT / US2021 / 015580).

[0375] Applications in plants

[0376] The base editing fusion proteins, base editors and methods of generating genetically modified cells of the present invention are particularly suitable for genetically modifying plants. Preferably, the plants are crop plants, including but not limited to wheat, rice, corn, soybean, sunflower, sorghum, rapeseed, alfalfa, cotton, barley, millet, sugarcane, tomato, tobacco, cassava and potato. More preferably, the plant is rice.

[0377] In another aspect, the present invention provides a method of generating a genetically modified plant, comprising introducing the base editor of the present invention into at least one of the plants, thereby resulting in one or more nucleotide substitutions in a target nucleic acid region in the genome of the at least one plant.

[0378] In some embodiments, the method further comprises screening the at least one plant for plants having the desired one or more nucleotide substitutions.

[0379] In the method of the present invention, the base editing composition can be introduced into plants by various methods well known to those skilled in the art. Methods that can be used to introduce the base editor of the present invention into plants include, but are not limited to: particle bombardment, PEG-mediated protoplast transformation, Agrobacterium tumefaciens-mediated transformation, plant virus-mediated transformation, pollen tube pathway method, and ovary injection method. Preferably, the base editing composition is introduced into plants by transient transformation.

[0380] In the method of the present invention, modification of the target sequence can be achieved by simply introducing or generating the base editing fusion protein in plant cells, and the modification can be stably inherited without stably transforming the plant with the exogenous polynucleotide encoding the components of the base editor. This avoids the potential off-target effects of the stably present (continuously produced) base editing composition and also avoids the integration of the exogenous nucleotide sequence into the plant genome, thus having higher biosafety.

[0381] In some preferred embodiments, the introduction is carried out in the absence of a selection pressure, thereby avoiding the integration of the exogenous nucleotide sequence into the plant genome.

[0382] In some embodiments, the introduction includes transforming the base editor of the present invention into isolated plant cells or tissues, and then regenerating the transformed plant cells or tissues into whole plants. Preferably, the regeneration is carried out in the absence of a selection pressure, that is, no selection agent for the selection gene carried on the expression vector is used during tissue culture. Not using a selection agent can improve the regeneration efficiency of plants and obtain modified plants without exogenous nucleotide sequences.

[0383] In some other embodiments, the base editor of the present invention can be transformed into specific parts of a whole plant, such as leaves, shoot tips, pollen tubes, young panicles, or hypocotyls. This is particularly suitable for the transformation of plants that are difficult to regenerate by tissue culture.

[0384] Therefore, in some embodiments, using the method of the present invention for genetic modification and breeding of plants can obtain plants whose genomes have no integration of exogenous polynucleotides, that is, modified plants that are transgene-free.

[0385] In some embodiments of the present invention, the modified target nucleic acid region is related to plant traits such as agronomic traits, whereby the one or more nucleotide substitutions result in the plant having altered (preferably improved) traits relative to the wild-type plant, such as agronomic traits.

[0386] In some embodiments, the method further comprises the step of screening plants having one or more desired nucleotide substitutions and / or desired traits such as agronomic traits.

[0387] In some embodiments of the present invention, the method further comprises obtaining progeny of the genetically modified plant. Preferably, the genetically modified plant or its progeny has one or more desired nucleotide substitutions and / or desired traits such as agronomic traits.

[0388] In another aspect, the present invention also provides a genetically modified plant or its progeny or a part thereof, wherein the plant is obtained by the method described above in the present invention. In some embodiments, the genetically modified plant or its progeny or a part thereof is non-transgenic. Preferably, the genetically modified plant or its progeny has a desired genetic modification and / or desired traits such as agronomic traits.

[0389] In another aspect, the present invention also provides a plant breeding method, comprising crossing a genetically modified first plant having one or more nucleotide substitutions in a target nucleic acid region obtained by the method described above in the present invention with a second plant that does not contain the one or more nucleotide substitutions, thereby introducing the one or more nucleotide substitutions into the second plant. Preferably, the genetically modified first plant has desired traits such as agronomic traits.

[0390] Examples

[0391] A further understanding of the present invention can be obtained by reference to some specific examples given herein. These examples are only for illustrative purposes and are not intended to limit the scope of the present invention in any way. Obviously, various modifications and changes can be made to the present invention without departing from the essence thereof. Therefore, these modifications and changes are also within the scope of the present application.

[0392] The sequences of some of the elements used in the subsequent examples are shown below:

[0393] OsBADH2 Left TALErepeat

[0394] LTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALE

[0395] (SEQ ID NO.89)

[0396] OsBADH2 Right TALErepeat

[0397] LTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALE

[0398] (SEQ ID NO.90)

[0399] OsDEP1 Left TALErepeat

[0400] LTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALE

[0401] (SEQ ID NO.91)

[0402] OsDEP1 Right TALErepeat

[0403] LTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALE

[0404] (SEQ ID NO.92)

[0405] OsCKX2 Left TALE repeat

[0406] LTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALE

[0407] (SEQ ID NO.93)

[0408] OsCKX2 Right TALE repeat

[0409] LTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALE

[0410] (SEQ ID NO.94)

[0411] OsSD1 Left TALE repeat

[0412] LTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALE

[0413] (SEQ ID NO.95)

[0414] OsSD1 Right TALE repeat

[0415] LTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALE

[0416] (SEQ ID NO.96)

[0417] SIRT6 Left TALE repeat

[0418] LTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIANNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPEQVVAIANNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIANNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIANNNGGKQALETVQRLLPVLCQDHGLTPEQVVAIANNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIANNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIANNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIANNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPEQVVAIASHDGGRPALE

[0419] (SEQ ID NO.97)

[0420] SIRT6 Right TALE repeat

[0421] LTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIANNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIANNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIANNNGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIANNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIANNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIANNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNNGGRPALE

[0422] (SEQ ID NO.98)

[0423] OsRbcL Left TALE repeat

[0424] LTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQAVETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQAVETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQAVETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETLQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQAVETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQAVETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQAVETVQRLLPVLCQDHGLTPEQVVAIASNGGGKQALE

[0425] (SEQ ID NO.99)

[0426] OsRbcL Right TALE repeat

[0427] LTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQAVETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQAVETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQAVETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETLQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQAVETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQAVETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQAVETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNGGGKQALE

[0428] (SEQ ID NO.100)

[0429] ND6 Left TALE repeat

[0430] LTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGRPALE

[0431] (SEQ ID NO.101)

[0432] ND6 Right TALE repeat

[0433] LTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGRPALE

[0434] (SEQ ID NO.102)

[0435] ND5.1 Left TALE repeat

[0436] LTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGRPALE

[0437] (SEQ ID NO.103)

[0438] ND5.1 Right TALE repeat

[0439] LTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGRPALE

[0440] (SEQ ID NO.104)

[0441] ND3 Left TALE repeat

[0442] LTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPQQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPQQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPQQVVAIASNIGGRPALE

[0443] (SEQ ID NO.105)

[0444] ND3 Right TALE repeat

[0445] LTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPQQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQLETVQRLLPVLCQAHGLTPQQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPQQVVAIASHDGGRPALE

[0446] (SEQ ID NO.106)

[0447] ND1.3 Left TALE repeat

[0448] LTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPQQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPQQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPQQVVAIASNNGGRPALE

[0449] (SEQ ID NO.107)

[0450] ND1.3 Right TALE repeat

[0451] LTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPQQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPQQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPQQVVAIASHDGGRPALE

[0452] (SEQ ID NO.108)

[0453] ND1.2 Left TALE repeat

[0454] LTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPQQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPQQVVAIASNNGGRPALE

[0455] (SEQ ID NO.109)

[0456] ND1.2 Right TALE repeat

[0457] LTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPQQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPQQVVAIASHDGGRPALE

[0458] (SEQ ID NO.110)

[0459] ND6.2 Left TALE repeat(TALE-L2)

[0460] LTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPQQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPQQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPQQVVAIASNIGGRPALE

[0461] (SEQ ID NO.111)

[0462] ND6.2 Right TALE repeat(TALE-R2)

[0463] LTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPQQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPQQVVAIASHDGGRPALE

[0464] (SEQ ID NO.112)

[0465] ND6.2 Left TALE repeat(TALE-L1)

[0466] LTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHG

[0467] (SEQ ID NO.185)

[0468] ND6.2 Left TALErepeat (TALE-L3)

[0469] LTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHG

[0470] (SEQ ID NO.186)

[0471] ND6.2 Right TALErepeat (TALE-R1)

[0472] LTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHG

[0473] (SEQ ID NO.187)

[0474] XTEN spacer peptide

[0475] NSGSETPGTSESATPES

[0476] (SEQ ID NO.113)

[0477] 48 - amino - acid spacer peptide

[0478] SGSETPGTSESATPESSGGSSGGSSGSETPGTSESATPESSGGSSGGS

[0479] (SEQ ID NO.114)

[0480] 16 - amino - acid spacer peptide

[0481] SGSETPGTSESATPES

[0482] (SEQ ID NO.115)

[0483] 14 - amino - acid spacer peptide

[0484] SGGGSGGSGGSGGS

[0485] (SEQ ID NO.116)

[0486] 11 - amino - acid spacer peptide

[0487] SGGSGGSGGSS

[0488] (SEQ ID NO.117)

[0489] Spacer peptide of 4 amino acids

[0490] SGGS

[0491] (SEQ ID NO.118)

[0492] γb

[0493] MMATFSCVCCGTLTTSTYCGKRCERKHVYSETRNKRLELYKKYLLEPQKCALNGIVGHSCGMPCSIAEEACDQLPIVSRFCGQKHADLYDSLLKRSEQELLLEFLQKKMQELKLSHIVKMAKLESEVNAIRKSVASSFEDSVGCDDSSSVSK

[0494] (SEQ ID NO.119)

[0495] Figure 16A to Figure 16E and Figure 17A to 17H The amino acid sequences of the vectors or elements involved are shown below. In subsequent examples, unless otherwise specified, Figure 16A to Figure 16E and Figure 17A to 17H According to the schematic diagrams of the constructs shown, corresponding fusion proteins can be constructed by combining the sequences disclosed in this specification.

[0496] OsBADH2-NLS-TALEN WT ( Figure 16A )

[0497]

[0498] (SEQ ID NO.120)

[0499] OsBADH2-NLS-TALE-L-FokI-L-T2A-TALE-R-FokI-RD 450A ( Figure 16B )

[0500]

[0501] (SEQ ID NO.121)

[0502] OsBADH2-NLS-TALE-L-FokI-L D450A -T2A-TALE-R-FokI-R( Figure 16B )

[0503]

[0504] (SEQ ID NO.122)

[0505] NLS-A3A-XTEN-UGI( Figure 16B )

[0506] MKRTADGSEFESPKKKRKVMEASPASGPRHLMDPHIFTSNFNNGIGRHKTYLCYEVERLDNGTSVKMDQHRGFLHNQAKNLLCGFYGRHAELRFLDLVPSLQLDPAQIYRVTWFISWSPCFSWGCAGEVRAFLQENTHVRLRIFAARIYDYDPLYKEALQMLRDAGAQVSIMTYDEFKHCWDTFVDHQGCPFQPWDGLDEHSQALSGRLRAILQNQGNSGSETPGTSESATPESTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML

[0507] (SEQ ID NO.123)

[0508] NLS-UGI( Figure 16B )

[0509] MTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKMLMKRTADGSEFESPKKKRKV

[0510] (SEQ ID NO.163)

[0511] NLS-C57-XTEN-UGI( Figure 16B )

[0512] MKRTADGSEFESPKKKRKVLEAVRARLIGEGGGPGAVPEGGDGPPAVPAEEVERLRGELPPPVVPGTGQKTHGRWIGPDGRVRAIVSGRDEDAALVHAQLAAKGIPDEPTRNSDVEQKLAAHMVANGIRHVTLVINHRPCRGFDDSCDTLVPIILPEGCTLTVHGQTDKGMRVRVRYTGGARPWWSNSGSETPGTSESATPESTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML

[0513] (SEQ ID NO.124)

[0514] NLS-rAPOBEC1-XTEN-UGI( Figure 16B )

[0515] MKRTADGSEFESPKKKRKVSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSIWRHTSQNTNKHVEVNFIEKFTTERYFCPNTRCSITWFLSWSPCGECSRAITEFLSRYPHVTLFIYIARLYHHADPRNRQGLRDLISSGVTIQIMTEQESGYCWRNFVNYSPSNEAHWPRYPHLWVRLYVLELYCIILGLPPCLNILRRKQPQLTFFTIALQSCHYQRLPPHILWATGLKNSGSETPGTSESATPESTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML

[0516] (SEQ ID NO.164)

[0517] TadA8e-NLS( Figure 16B )

[0518] MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSINSGGSMKRTADGSEFESPKKKRKV

[0519] (SEQ ID NO.166)

[0520] mExoI-NLS( Figure 16B )

[0521] MGIQGLLQFIQEASEPVNVKKYKGQAVAVDTYCWLHKGAIACAEKLAKGEPTDRYVGFCMKFVNMLLSYGVKPILIFDGCTLPSKKEVERSRRERRQSNLLKGKQLLREGKVSEARDCFARSINITHAMAHKVIKAARALGVDCLVAPYEADAQLAYLNKAGIVQAVITEDSDLLAFGCKKVILKMDQFGNGLEVDQARLGMCKQLGDVFTEEKFRYMCILSGCDYLASLRGIGLAKACKVLRLANNPDIVKVIKKIGHYLRMNITVPEDYITGFIRANNTFLYQLVFDPIQRKLVPLNAYGDDVNPETLTYAGQYVGDSVALQIALGNRDVNTFEQIDDYSPDTMPAHSRSHSWNEKAGQKPPGTNSIWHKNYCPRLEVNSVSHAPQLKEKPSTLGLKQVISTKGLNLPRKSCVLKRPRNEALAEDDLLSQYSSVSKKIKENGCGDGTSPNSSKMSKSCPDSGTAHKTDAHTPSKMRNKFATFLQRRNEESGAVVVPGTRSRFFCSSQDFDNFIPKKESGQPLNETVATGKATTSLLGALDCPDTEGHKPVDANGTHNLSSQIPGNAAVSPEDEAQSSETSKLLGAMSPPSLGTLRSCFSWSGTLREFSRTPSPSASTTLQQFRRKSDPPACLPEASAVVTDRCDSKSEMLGETSQPLHELGCSSRSQESMDSSCGLNTSSLSQPSSRDSGSEESDCNNKSLDNQGEQNSKQHLPHFSKKDGLRRNKVPGLCRSSSMDSFSTTKIKPLVPARVSGLSKKSGSMQTRKHHDVENKPGLQTKISELWKNFGFKKDSEKLPSCKKPLSPVKDNIQLTPETEDEIFNKPECVRAQRAIFHMKRTADGSEFESPKKKRKV

[0522] (SEQ ID NO.125)

[0523] Trex2-NLS( Figure 16B )

[0524] MSEPPRAETFVFLDLEATGLPNMDPEIAEISLFAVHRSSLENPERDDSGSLVLPRVLDKLTLCMCPERPFTAKASEITGLSSESLMHCGKAGFNGAVVRTLQGFLSRQEGPICLVAHNGFDYDFPLLCTELQRLGAHLPQDTVCLDTLPALRGLDRAHSHGTRAQGRKSYSLASLFHRYFQAEPSAAHSAEGDVHTLLLIFLHRAPELLAWADEQARSWAHIEPMYVPPDGPSLEAMKRTADGSEFESPKKKRKV

[0525] (SEQ ID NO.126)

[0526] OsBADH2-NLS-A3A-TALE-L-FokI-L-T2A-TALE-R-FokI-R D450A ( Figure 16C )

[0527]

[0528] (SEQ ID NO.127)

[0529] OsBADH2-NLS-A3A-TALE-L-FokI-L D450A -T2A-TALE-R-FokI-R( Figure 16C )

[0530]

[0531] (SEQ ID NO.128)

[0532] mExoI-NLS( Figure 16C )

[0533] MGIQGLLQFIQEASEPVNVKKYKGQAVAVDTYCWLHKGAIACAEKLAKGEPTDRYVGFCMKFVNMLLSYGVKPILIFDGCTLPSKKEVERSRRERRQSNLLKGKQLLREGKVSEARDCFARSINITHAMAHKVIKAARALGVDCLVAPYEADAQLAYLNKAGIVQAVITEDSDLLAFGCKKVILKMDQFGNGLEVDQARLGMCKQLGDVFTEEKFRYMCILSGCDYLASLRGIGLAKACKVLRLANNPDIVKVIKKIGHYLRMNITVPEDYITGFIRANNTFLYQLVFDPIQRKLVPLNAYGDDVNPETLTYAGQYVGDSVALQIALGNRDVNTFEQIDDYSPDTMPAHSRSHSWNEKAGQKPPGTNSIWHKNYCPRLEVNSVSHAPQLKEKPSTLGLKQVISTKGLNLPRKSCVLKRPRNEALAEDDLLSQYSSVSKKIKENGCGDGTSPNSSKMSKSCPDSGTAHKTDAHTPSKMRNKFATFLQRRNEESGAVVVPGTRSRFFCSSQDFDNFIPKKESGQPLNETVATGKATTSLLGALDCPDTEGHKPVDANGTHNLSSQIPGNAAVSPEDEAQSSETSKLLGAMSPPSLGTLRSCFSWSGTLREFSRTPSPSASTTLQQFRRKSDPPACLPEASAVVTDRCDSKSEMLGETSQPLHELGCSSRSQESMDSSCGLNTSSLSQPSSRDSGSEESDCNNKSLDNQGEQNSKQHLPHFSKKDGLRRNKVPGLCRSSSMDSFSTTKIKPLVPARVSGLSKKSGSMQTRKHHDVENKPGLQTKISELWKNFGFKKDSEKLPSCKKPLSPVKDNIQLTPETEDEIFNKPECVRAQRAIFHMKRTADGSEFESPKKKRKV

[0534] (SEQ ID NO.129)

[0535] Trex2-NLS( Figure 16C )

[0536] MSEPPRAETFVFLDLEATGLPNMDPEIAEISLFAVHRSSLENPERDDSGSLVLPRVLDKLTLCMCPERPFTAKASEITGLSSESLMHCGKAGFNGAVVRTLQGFLSRQEGPICLVAHNGFDYDFPLLCTELQRLGAHLPQDTVCLDTLPALRGLDRAHSHGTRAQGRKSYSLASLFHRYFQAEPSAAHSAEGDVHTLLLIFLHRAPELLAWADEQARSWAHIEPMYVPPDGPSLEAMKRTADGSEFESPKKKRKV

[0537] (SEQ ID NO.130)

[0538] UGI-NLS( Figure 16C )

[0539] MKRTADGSEFESPKKKRKVTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML

[0540] (SEQ ID NO.131)

[0541] OsBADH2-NLS-A3A-TALE-L-FokI-L-T2A-TALE-R-FokI-R D450A -UGI( Figure 16D )

[0542]

[0543] (SEQ ID NO.132)

[0544] OsBADH2-NLS-A3A-TALE-L-FokI-L D450A -T2A-TALE-R-FokI-R-UGI( Figure 16D )

[0545]

[0546] (SEQ ID NO.133)

[0547] mExoI-NLS( Figure 16D )

[0548] MGIQGLLQFIQEASEPVNVKKYKGQAVAVDTYCWLHKGAIACAEKLAKGEPTDRYVGFCMKFVNMLLSYGVKPILIFDGCTLPSKKEVERSRRERRQSNLLKGKQLLREGKVSEARDCFARSINITHAMAHKVIKAARALGVDCLVAPYEADAQLAYLNKAGIVQAVITEDSDLLAFGCKKVILKMDQFGNGLEVDQARLGMCKQLGDVFTEEKFRYMCILSGCDYLASLRGIGLAKACKVLRLANNPDIVKVIKKIGHYLRMNITVPEDYITGFIRANNTFLYQLVFDPIQRKLVPLNAYGDDVNPETLTYAGQYVGDSVALQIALGNRDVNTFEQIDDYSPDTMPAHSRSHSWNEKAGQKPPGTNSIWHKNYCPRLEVNSVSHAPQLKEKPSTLGLKQVISTKGLNLPRKSCVLKRPRNEALAEDDLLSQYSSVSKKIKENGCGDGTSPNSSKMSKSCPDSGTAHKTDAHTPSKMRNKFATFLQRRNEESGAVVVPGTRSRFFCSSQDFDNFIPKKESGQPLNETVATGKATTSLLGALDCPDTEGHKPVDANGTHNLSSQIPGNAAVSPEDEAQSSETSKLLGAMSPPSLGTLRSCFSWSGTLREFSRTPSPSASTTLQQFRRKSDPPACLPEASAVVTDRCDSKSEMLGETSQPLHELGCSSRSQESMDSSCGLNTSSLSQPSSRDSGSEESDCNNKSLDNQGEQNSKQHLPHFSKKDGLRRNKVPGLCRSSSMDSFSTTKIKPLVPARVSGLSKKSGSMQTRKHHDVENKPGLQTKISELWKNFGFKKDSEKLPSCKKPLSPVKDNIQLTPETEDEIFNKPECVRAQRAIFHMKRTADGSEFESPKKKRKV

[0549] (SEQ ID NO.134)

[0550] Trex2-NLS( Figure 16D )

[0551] MSEPPRAETFVFLDLEATGLPNMDPEIAEISLFAVHRSSLENPERDDSGSLVLPRVLDKLTLCMCPERPFTAKASEITGLSSESLMHCGKAGFNGAVVRTLQGFLSRQEGPICLVAHNGFDYDFPLLCTELQRLGAHLPQDTVCLDTLPALRGLDRAHSHGTRAQGRKSYSLASLFHRYFQAEPSAAHSAEGDVHTLLLIFLHRAPELLAWADEQARSWAHIEPMYVPPDGPSLEAMKRTADGSEFESPKKKRKV

[0552] (SEQ ID NO.135)

[0553] OsBADH2-NLS-A3A-TALE-L-FokI-L-T2A-TALE-R-FokI-R D450A -UGI--mExoI-NLS( Figure 16E )

[0554]

[0555] (SEQ ID NO.136)

[0556] OsBADH2-NLS-A3A-TALE-L-FokI-L D450A -T2A-TALE-R-FokI-R-UGI--mExoI-NLS( Figure 16E )

[0557]

[0558] (SEQ ID NO.137)

[0559] ND6-MTS-TALE-L-FokI-L( Figure 17A )

[0560] MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAMDIADLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGRPALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSQLVKSELEEKKSELRHKLKYVPHEYIELIEIARNSTQDRILEMKVMEFFMKVYGYRGKHLGGSRKPDGAIYTVGSPIDYGVIVDTKAYSGGYNLPIGQADEMQRYVKENQTRNKHINPNEWWKVYPSSVTEFKFLFVSGHFKGNYKAQLTRLNHKTNCNGAVLSVEELLIGGEMIKAGTLTLEEVRRKFNNGEINF

[0561] (SEQ ID NO.138)

[0562] ND6-MTS-TALE-R-FokI-RD 450A (Figure 17B )

[0563] MASVLTPLLLRGLTGSARRLPVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDKMDIADLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKRGAGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGRPALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSQLVKSELEEKKSELRHKLKYVPHEYIELIEIARNSTQDRILEMKVMEFFMKVYGYRGKHLGGSRKPAGAIYTVGSPIDYGVIVDTKAYSGGYNLPIGQADEMERYVEENQTRNKHLNPNEWWKVYPSSVTEFKFLFVSGHFKGNYKAQLTRLNHITNCNGAVLSVEELLIGGEMIKAGTLTLEEVRRKFNNGEINF

[0564] (SEQ ID NO.139)

[0565] ND6-MTS-TALE-L-FokI-L D450A ( Figure 17A )

[0566] MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAMDIADLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGRPALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSQLVKSELEEKKSELRHKLKYVPHEYIELIEIARNSTQDRILEMKVMEFFMKVYGYRGKHLGGSRKPAGAIYTVGSPIDYGVIVDTKAYSGGYNLPIGQADEMQRYVKENQTRNKHINPNEWWKVYPSSVTEFKFLFVSGHFKGNYKAQLTRLNHKTNCNGAVLSVEELLIGGEMIKAGTLTLEEVRRKFNNGEINF

[0567] (SEQ ID NO.140)

[0568] ND6-MTS-TALE-R-FokI-R( Figure 17B )

[0569] MASVLTPLLLRGLTGSARRLPVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDKMDIADLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKRGAGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGRPALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSQLVKSELEEKKSELRHKLKYVPHEYIELIEIARNSTQDRILEMKVMEFFMKVYGYRGKHLGGSRKPDGAIYTVGSPIDYGVIVDTKAYSGGYNLPIGQADEMERYVEENQTRNKHLNPNEWWKVYPSSVTEFKFLFVSGHFKGNYKAQLTRLNHITNCNGAVLSVEELLIGGEMIKAGTLTLEEVRRKFNNGEINF

[0570] (SEQ ID NO.141)

[0571] MTS-mExoI( Figure 17D )

[0572] MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDMGIQGLLQFIQEASEPVNVKKYKGQAVAVDTYCWLHKGAIACAEKLAKGEPTDRYVGFCMKFVNMLLSYGVKPILIFDGCTLPSKKEVERSRRERRQSNLLKGKQLLREGKVSEARDCFARSINITHAMAHKVIKAARALGVDCLVAPYEADAQLAYLNKAGIVQAVITEDSDLLAFGCKKVILKMDQFGNGLEVDQARLGMCKQLGDVFTEEKFRYMCILSGCDYLASLRGIGLAKACKVLRLANNPDIVKVIKKIGHYLRMNITVPEDYITGFIRANNTFLYQLVFDPIQRKLVPLNAYGDDVNPETLTYAGQYVGDSVALQIALGNRDVNTFEQIDDYSPDTMPAHSRSHSWNEKAGQKPPGTNSIWHKNYCPRLEVNSVSHAPQLKEKPSTLGLKQVISTKGLNLPRKSCVLKRPRNEALAEDDLLSQYSSVSKKIKENGCGDGTSPNSSKMSKSCPDSGTAHKTDAHTPSKMRNKFATFLQRRNEESGAVVVPGTRSRFFCSSQDFDNFIPKKESGQPLNETVATGKATTSLLGALDCPDTEGHKPVDANGTHNLSSQIPGNAAVSPEDEAQSSETSKLLGAMSPPSLGTLRSCFSWSGTLREFSRTPSPSASTTLQQFRRKSDPPACLPEASAVVTDRCDSKSEMLGETSQPLHELGCSSRSQESMDSSCGLNTSSLSQPSSRDSGSEESDCNNKSLDNQGEQNSKQHLPHFSKKDGLRRNKVPGLCRSSSMDSFSTTKIKPLVPARVSGLSKKSGSMQTRKHHDVENKPGLQTKISELWKNFGFKKDSEKLPSCKKPLSPVKDNIQLTPETEDEIFNKPECVRAQRAIFH

[0573] (SEQ ID NO.142)

[0574] MTS-Trex2( Figure 17D )

[0575] MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDMSEPPRAETFVFLDLEATGLPNMDPEIAEISLFAVHRSSLENPERDDSGSLVLPRVLDKLTLCMCPERPFTAKASEITGLSSESLMHCGKAGFNGAVVRTLQGFLSRQEGPICLVAHNGFDYDFPLLCTELQRLGAHLPQDTVCLDTLPALRGLDRAHSHGTRAQGRKSYSLASLFHRYFQAEPSAAHSAEGDVHTLLLIFLHRAPELLAWADEQARSWAHIEPMYVPPDGPSLEA

[0576] (SEQ ID NO.143)

[0577] MTS-A3A( Figure 17C )

[0578] MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDEASPASGPRHLMDPHIFTSNFNNGIGRHKTYLCYEVERLDNGTSVKMDQHRGFLHNQAKNLLCGFYGRHAELRFLDLVPSLQLDPAQIYRVTWFISWSPCFSWGCAGEVRAFLQENTHVRLRIFAARIYDYDPLYKEALQMLRDAGAQVSIMTYDEFKHCWDTFVDHQGCPFQPWDGLDEHSQALSGRLRAILQNQGN

[0579] (SEQ ID NO.144)

[0580] MTS-C57 / Sdd7( Figure 17C )

[0581] MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDLEAVRARLIGEGGGPGAVPEGGDGPPAVPAEEVERLRGELPPPVVPGTGQKTHGRWIGPDGRVRAIVSGRDEDAALVHAQLAAKGIPDEPTRNSDVEQKLAAHMVANGIRHVTLVINHRPCRGFDDSCDTLVPIILPEGCTLTVHGQTDKGMRVRVRYTGGARPWWS

[0582] (SEQ ID NO.145)

[0583] MTS-UGI( Figure 17E )

[0584] MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDGSSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML

[0585] (SEQ ID NO.146)

[0586] ND6-MTS-A3A-TALE-L-FokI-L( Figure 17F )

[0587]

[0588] (SEQ ID NO.147)

[0589] ND6-MTS-Trex2-TALE-R-FokI-R D450A ( Figure 17G )

[0590]

[0591] (SEQ ID NO.148)

[0592] ND6-MTS-UGI-Trex2-TALE-R-FokI-R D450A ( Figure 17H )

[0593]

[0594] (SEQ ID NO.149)

[0595] In the examples, the exemplary amino acid sequences of the components or fusion proteins are shown below. Unless otherwise specified in the subsequent examples, Figure 16A to Figure 16E and Figure 17A to 17H the schematic diagrams of the constructs shown, and the exemplary sequences shown below, can be used to construct the corresponding fusion proteins in combination with the sequences disclosed in this specification.

[0596] In the subsequent examples, the nickase used in the experiment of editing OsBADH2:

[0597] TALEN WT (SEQ ID NO.154)

[0598]

[0599] TALE-FokI-R nickase(D450A) or TALE-FokI-R nickase (SEQ ID NO.155)

[0600]

[0601] TALE-FokI-R nickase(D467A) (SEQ ID NO.156)

[0602]

[0603] Edit the nickase used in the OsDEP1 experiment:

[0604] TALEN WT (SEQ ID NO.157)

[0605]

[0606] TALE-FokI-R nickase(D450A) or TALE-FokI-R nickase (SEQ ID NO.158)

[0607]

[0608] TALE-FokI-R nickase(D467A) (SEQ ID NO.159)

[0609]

[0610] Editing the nickase used in the OsCKX2 experiment:

[0611] TALEN WT (SEQ ID NO.160)

[0612]

[0613] TALE-FokI-R nickase (SEQ ID NO.161)

[0614]

[0615] TALE-FokI-L nickase (SEQ ID NO.162)

[0616]

[0617] In Examples 1-6, mExoI: namely, the aforementioned mExoI-NLS( Figure 16B ), SEQ ID NO.125; A3A-UGI: namely, the aforementioned: NLS-A3A-XTEN-UGI( Figure 16B ), SEQ ID NO.123; Trex2: namely, the aforementioned Trex2-NLS( Figure 16B ), SEQ IDNO.126.

[0618] In Examples 1-6, the amino acid sequence of UGI is the aforementioned NLS-UGI( Figure 16B )(SEQ ID NO.163).

[0619] In Example 4, the amino acid sequence of APOBEC1-UGI is the aforementioned NLS-rAPOBEC1-XTEN-UGI( Figure 16B )(SEQ ID NO.164).

[0620] In Example 1, the amino acid sequence of ExoV (ExoV-NLS) (SEQ ID NO.165):

[0621] MAETGEEETASAEASGFSDLSDSELVEFLDLEEAKESAVSLSKPGPSAELPGKDDKPVSLQNWKGGLDVLSPMERFHLKYLYVTDLCTQNWCELQMVYGKELPGSLTPEKAAVLDTGASIHLAKELELHDLVTVPIATKEDAWAVKFLNILAMIPALQSEGRVREFPVFGEVEGIFLVGVIDELHYTSKGELELAELKTRRRPVLPLPAQKKKDYFQVSLYKYIFDAMVQGKVTPASLIHHTKLCLDKPLGPSVLRHARQGGVSVKSLGDLMELVFLSLTLSDLPAIDTLKLEYIHQETATILGTEIVAFEEKEVKSKVQHYVAYWMGHRDPQGVDVEEAWKCRTCDYVDICEWRRGSGVLSSSWEPKAKKFKMKRTADGSEFESPKKKRKV

[0622] In Example 5, the amino acid sequence of TadA-8e is the aforementioned TadA8e-NLS( Figure 16B )(SEQ ID NO.166).

[0623] In Example 6:

[0624] The amino acid sequence of mExoI-16aa-A3A-UGI (SEQ ID NO.167):

[0625]

[0626] The amino acid sequence of mExoI-48aa-A3A-UGI (SEQ ID NO.168):

[0627]

[0628] A3A-TALE-FokI-R nickase (SEQ ID NO.169)

[0629]

[0630] APOBEC1-TALE-FokI-R nickase (SEQ ID NO.170)

[0631]

[0632] A3A-TALE-FokI-L nickase (SEQ ID NO.171)

[0633]

[0634] APOBEC1-TALE-FokI-L nickase (SEQ ID NO.172)

[0635]

[0636] SIRT6-NLS-TALE-L-DddA N -UGI(SEQ ID NO.173)

[0637]

[0638] SIRT6-NLS-TALE-R-DddA C -UGI(SEQ ID NO.174)

[0639] MASVLTPLLLRGLTGSARRLPVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDKMDIADLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKRGAGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIANNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIANNNGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIANNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIANNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIANNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNNGGRPALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML

[0640] In Examples 11, 14, and 15:

[0641] ND6-MTS-TALE-L-DddAN -UGI (SEQ ID NO.175)

[0642] MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAMDIADLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGRPALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML

[0643] ND6 - MTS - TALE - R - DddA C -UGI (SEQ ID NO.176)

[0644] MASVLTPLLLRGLTGSARRLPVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDKMDIADLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKRGAGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGRPALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML

[0645] ND1.2-MTS-TALE-L-DddA N -UGI(SEQ ID NO.177)

[0646] MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAMDIADLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPQQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPQQVVAIASNNGGRPALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML

[0647] ND1.2-MTS-TALE-R-DddA C -UGI(SEQ ID NO.178)

[0648] MASVLTPLLLRGLTGSARRLPVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDKMDIADLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKRGAGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPQQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGSIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML

[0649] ND1.3-MTS-TALE-L-DddA N -UGI(SEQ ID NO.179)

[0650] MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAMDIADLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPQQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPQQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPQQVVAIASNNGGRPALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML

[0651] ND1.3-MTS-TALE-R-DddA C -UGI (SEQ ID NO.180)

[0652] MASVLTPLLLRGLTGSARRLPVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDKMDIADLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKRGAGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPQQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPQQVVAIASHDGGRPALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML

[0653] ND6.2-MTS-TALE-L-DddA N -UGI(SEQ ID NO.181)

[0654] MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAMDIADLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPQQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPQQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPQQVVAIASNIGGRPALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML

[0655] ND6.2-MTS-TALE-R-DddA C -UGI(SEQ ID NO.182)

[0656] MASVLTPLLLRGLTGSARRLPVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDKMDIADLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKRGAGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPQQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPQQVVAIASHDGGRPALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML

[0657] ND3-MTS-TALE-L-DddA N -UGI (SEQ ID NO.183)

[0658] MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAMDIADLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPQQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPQQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPQQVVAIASNIGGRPALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML

[0659] ND3-MTS-TALE-R-DddA C -UGI(SEQ ID NO.184)

[0660] MASVLTPLLLRGLTGSARRLPVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDKMDIADLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKRGAGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPQQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQLETVQRLLPVLCQAHGLTPQQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPQQVVAIASHDGGRPALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML

[0661] The target sequences in the following examples and the accompanying drawings are as follows:

[0662] Target A chain of OsBADH2 in the figure

[0663] SEQ ID NO.188

[0664] GCTGGATGCTTTGAGTACTTTGCAGATCTTGCAGAATCCTTGGACAAAAGGC

[0665] Target B chain of OsBADH2 in the figure

[0666] SEQ ID NO.189

[0667] CGACCTACGAAACTCATGAAACGTCTAGAACGTCTTAGGAACCTGTTTTCCG

[0668] Target A chain of OsDEP1 in the figure

[0669] SEQ ID NO.190

[0670] GCAAAAGACCAAGGTGCCTCAATTGTTCTTGCAGCTCATGCTGCGACGAGCC

[0671] Target B chain of OsDEP1 in the figure

[0672] SEQ ID NO.191

[0673] CGTTTTCTGGTTCCACGGAGTTAACAAGAACGTCGAGTACGACGCTGCTCGG

[0674] Target A chain of OsCKX2 in the figure

[0675] SEQ ID NO.192

[0676] CCTGGACCGCGTCCACGACGGCGAGCTCAAGCTCCGCGCCGCGGGGCTCTGGG

[0677] Target B chain of OsCKX2 in the figure

[0678] SEQ ID NO.193

[0679] GGACCTGGCGCAGGTGCTGCCGCTCGAGTTCGAGGCGCGGCGCCCCGAGACCC

[0680] Target A chain of Human ND6 in the figure

[0681] SEQ ID NO.194

[0682] CCCCTGACCCCCATGCCTCAGGATACTCCTCAATAGCCATCGCTGTA

[0683] Human ND6 target B chain in the figure

[0684] SEQ ID NO.195

[0685] GGGGACTGGGGGTACGGAGTCCTATGAGGAGTTATCGGTAGCGACAT

[0686] OsSD1 target A chain in the figure

[0687] SEQ ID NO.196

[0688] CCAGGACGACGTCGGCGGCCTCGAGGTCCTCGTCGACGGCGAATGGCGCCCCGTC

[0689] OsSD1 target B chain in the figure

[0690] SEQ ID NO.197

[0691] GGTCCTGCTGCAGCCGCCGGAGCTCCAGGAGCAGCTGCCGCTTACCGCGGGGCAG

[0692] SIRT6 target A chain in the figure

[0693] SEQ ID NO.198

[0694] TACGCGGCGGGGCTGTCGCCGTACGCGGACAAGGGCAAGTGCGGCCTCCCGG

[0695] SIRT6 target B chain in the figure

[0696] SEQ ID NO.199

[0697] ATGCGCCGCCCCGACAGCGGCATGCGCCTGTTCCCGTTCACGCCGGAGGGCC

[0698] OsRbcL target A chain in the figure

[0699] SEQ ID NO.200

[0700] TTACCAAAGATGATGAAAACGTAAACTCACAACCATTTATGCGTTGG

[0701] The B chain of the OsRbcL target in the figure

[0702] SEQ ID NO.201

[0703] AATGGTTTCTACTACTTTTGCATTTGAGTGTTGGTAAATACGCAACC

[0704] The A chain of the ND6.2 target in the figure

[0705] SEQ ID NO.202

[0706] GACCCCCATGCCTCAGGATACTCCTCAATAGCCATCGCTGTAGTATATCCAA

[0707] The B chain of the ND6.2 target in the figure

[0708] SEQ ID NO.203

[0709] CTGGGGGTACGGAGTCCTATGAGGAGTTATCGGTAGCGACATCATATAGGTT

[0710] The A chain of the ND1.2 target in the figure

[0711] SEQ ID NO.204

[0712] CCTATTTATTCTAGCCACCTCTAGCCTAGCCGTTTACTCA

[0713] The B chain of the ND1.2 target in the figure

[0714] SEQ ID NO.205

[0715] GGATAAATAAGATCGGTGGAGATCGGATCGGCAAATGAGT

[0716] The A chain of the ND1.3 target in the figure

[0717] SEQ ID NO.206

[0718] TCTCCACACTAGCAGAGACCAACCGAACCCCCTTCGACCTTGCCGAAGGGG

[0719] The B chain of the ND1.3 target in the figure

[0720] SEQ ID NO.207

[0721] AGAGGTGTGATCGTCTCTGGTTGGCTTGGGGGAAGCTGGAACGGCTTCCCC

[0722] ND3 target A chain in the figure

[0723] SEQ ID NO.208

[0724] ACGAGTGCGGCTTCGACCCTATATCCCCCGCCCGCGTCCCTTTCTCCAT

[0725] ND3 target B chain in the figure

[0726] SEQ ID NO.209

[0727] TGCTCACGCCGAAGCTGGGATATAGGGGGCGGGCGCAGGGAAAGAGGTA

[0728] ND1 target A chain in the figure

[0729] SEQ ID NO.210

[0730] CTAGCCTAGCCGTTTACTCAATCCTCTCATCAGGGTGAGCATCAAACTC

[0731] ND1 target B chain in the figure

[0732] SEQ ID NO.211

[0733] GATCGGATCGGCAAATGAGTTAGGAGACTAGTCCCACTCGTAGTTTGAG

[0734] ND4 target A chain in the figure

[0735] SEQ ID NO.212

[0736] GCTAGTAACCACGTTCTCCTGATCAAATATCACTCTCCTACTTACAGG

[0737] ND4 target B chain in the figure

[0738] SEQ ID NO.213

[0739] CGATCATTGGTGCAAGAGGACTAGTTTATAGTGAGAGGATGAATGTCC

[0740] The A chain of the ND5.1 target in the figure

[0741] SEQ ID NO.214

[0742] AGCATTAGCAGGAATACCTTTCCTCACAGGTTTCTACTCCAAAG

[0743] The B chain of the ND5.1 target in the figure

[0744] SEQ ID NO.215

[0745] TCGTAATCGTCCTTATGGAAAGGAGTGTCCAAAGATGAGGTTTC

[0746] Example 1: Synthesis and detection of base editors

[0747] The synthesis strategy of the base editor of the present invention is as Figure 1 shown.

[0748] To verify the above strategy, target sites in the rice OsBADH2 gene were selected, and two groups of TALE-encoding vectors modified to target these sites were constructed, and the above components are listed in Table 3.

[0749] Table 3 Specific examples of combinations of base editors in the examples

[0750]

[0751] The FokICD (or mutant) monomer was separately fused to the C-terminus of TALE-L and TALE-R, and wild-type FokI (without D450A or D467A mutations) was used as a control ([[]] Figure 16A ). The application of two exonucleases (Exonuclease I (rat exonuclease I, abbreviated as mExoI) and Exonuclease V (abbreviated as ExoV)) and one deaminase (hAPOBEC3A, abbreviated as hA3A or A3A) in the novel base editor was evaluated, where in each group UGI was fused to the carboxyl terminus of the deaminase with an XTEN spacer peptide ([[]] Figure 16B ). The nuclear localization signal (NLS, i.e., the SV40 NLS in Table 2) was fused to the end of the protein.

[0752] Through PEG-mediated transformation, the recombinant expression constructs encoding these components were transformed into rice protoplasts. The constructs are as Figure 16A - Figure 16BAs shown. Different construct combinations were transformed into rice protoplasts to target the OsBADH2 locus, and next-generation sequencing technology (NGS) was used to determine the C>T base editing frequency. The sequencing results ( Figure 2A ) showed that for the combination containing FokI nickase, deaminase, exonuclease, and UGI, the targeted cytosine base editing frequency could reach up to approximately 10%. Importantly, the detection results also showed that the novel nucleic acid base editor only induced very low indel by-products (as Figure 2B shown). The above results indicate that the novel base editor has the characteristic of high product purity, which is important for precise genome editing.

[0753] In Figure 2A and Figure 2B , the experimental treatments or construct combinations involved in the figures and the schematic diagrams of the associated vectors are shown as follows:

[0754]

[0755] Example 2: Characterization of the single-strand cleavage performance of the base editor

[0756] The base editing window of the base editor tested in Example 1 was analyzed. Among the four C sites (C1, C6, C11, and C15, with the first base adjacent to TALE-L in the spacer sequence of the two TALEs counted as 1) present in the A strand of the target gene (as Figure 3A shown), C6 and C11 cytosines were effectively edited ( Figure 3B ).

[0757] In Figure 3B , the experimental treatments or construct combinations involved in the figures and the schematic diagrams of the associated vectors are shown as follows:

[0758]

[0759] These results showed that the base editor containing FokI-R nickase (in which FokI-L in the dimer nickase composed of FokI-L and FokI-R undergoes D450A or D467A mutation) was more inclined to cut the B strand through the nickase, and then the exonuclease digested the cut single strand, leaving a short ssDNA fragment on the A strand. The digestion direction depended on the enzymatic catalytic direction of the exonuclease (5' to 3' or 3 to 5').

[0760] To verify the above results, the inventors evaluated the nucleic acid base editor at another locus, OsDEP1, in this example. This locus contains 5 C bases (C1, C9, C13, C16, and C18) in the A strand. The results of NGS analysis of different construct combinations transformed into rice protoplasts to target the OsDEP1 locus showed that the base editing window was mainly located near the 5' region (C9 and C1) of the A strand, and slight editing also occurred at C13 and C16 (as Figure 4A shown), which was caused by the instantaneous generation of a 3' flap structure after nicking. Importantly, similar to the OsBADH2 locus, only extremely low indel by-products appeared in the labeled products at the OsDEP1 locus (as Figure 4B shown). The above results indicate that the novel base editor has the advantage of high product purity.

[0761] In Figure 4A and Figure 4B , the experimental treatments or construct combinations involved in the figures and the schematic diagrams of the associated vectors are as follows:

[0762]

[0763] Example 3: Effects of Exonuclease Digestion Direction and Nicking Enzyme Single-Strand Preference on Editing Results

[0764] For exonucleases with 5'→3' digestion directionality (e.g., rat exonuclease I (mExoI)), cytosine residues near the 5' region of the target site in the complementary strand can be exposed and deaminated by the deaminase; while for 3' exonucleases, cytosine residues near the 3' region of the target site in the complementary strand are exposed and deaminated by the deaminase. To verify that the base editor disclosed in the present invention can produce the expected effects for different exonuclease digestion directions, the inventors simultaneously tested the 5' exonuclease mExoI and a 3' exonuclease, human Trex2 exonuclease, at the OsCKX2 target site, and analyzed the editing window of the resulting base editor by NGS. The experimental results showed that for base editing mediated by FokI-R nickase , when the 5' exonuclease mExoI was used, the editing window was mainly located in the 5' region (C9 and C11) of the A strand of the target site; on the contrary, when the 3' exonuclease Trex2 was used, the editing window switched to the 3' adjacent region (C11 and C15) of the A strand of the OsCKX2 target site, and the cytosine residues on the B strand were not edited (as Figure 5A , Figure 5B shown). Further, the inventors also evaluated the effect of the single-strand preference of the nicking enzyme used on the single strand where base editing can occur. FokI-R nickaseReplace with FokI-L that prefers to cleave the A strand nickase , as expected, the single strand where base editing occurred switched from the A strand to the B strand ( Figure 5A ); meanwhile, for the editing window, when the 5'-exonuclease mExoI was used, the 5'-adjacent region (C6 and C8) of the B strand at the OsCKX2 target site in the editing window, correspondingly, when the 3'-exonuclease Trex2 was used, the editing window could be switched to the 3'-adjacent region (C3 and C6) of the B strand at the OsCKX2 target site, and the cytosine residues on the A strand were not edited ( Figure 5A ). Thus, it can be seen that the base editor of the present invention can be used with exonucleases in different digestion directions and play the digestion role of the corresponding exonucleases, so as to selectively edit the target site.

[0765] Different construct combinations were transformed into rice protoplasts to target the OsCKX2 locus, and the C>T base editing efficiency and indel by-product frequencies detected by NGS. In Figure 5A and Figure 5B , the experimental treatments or construct combinations involved in the figures and the schematic diagrams of the associated vectors are shown as follows:

[0766]

[0767] Example 4: Influence of the type of cytosine deaminase

[0768] The novel base editor of the present invention is independent of the type of deaminase and can be compatible with different types of deaminases. To exclude the dependence of the editing ability of the novel base editor on the deaminase hAPOBEC3A (A3A), the inventors tested another cytosine deaminase rAPOBEC1 (APOBEC1) in this example. The NGS analysis results showed that in the presence of exonucleases, such as mExoI (as shown in Figure 6A ) and Trex2 (as shown in Figure 6B ), after replacing hAPOBEC3A with rAPOBEC1 at the OsBADH2 locus, high-purity targeted base editing was also achieved, indicating that different types of deaminases are applicable to the base editor of the present invention.

[0769] In Figure 6A , different construct combinations were transformed into rice protoplasts to target the OsBADH2 locus, and the C>T base editing efficiency and indel by-product frequencies detected by NGS. The experimental treatments or construct combinations involved in the figures and the schematic diagrams of the associated vectors are shown as follows:

[0770]

[0771] InFigure 6B In this case, different construct combinations were transformed into rice protoplasts to target the OsDEP1 locus, and the C>T base editing efficiency and the frequency of insertion / deletion (indel) by-products detected by NGS are as follows. The experimental treatments or construct combinations involved in the figure and the schematic diagrams of the associated vectors are shown below:

[0772]

[0773] When analyzing the editing windows of these base editors, cytosine residues near the 5' region of the target site on the complementary strand of the single strand where cleavage occurs were effectively edited in the group containing mExoI (as Figure 7A shown); while cytosine residues near the 3' region of the target site in the complementary strand were effectively edited in the group containing TREX2 (as Figure 7B shown), which is consistent with the results in the above embodiments. These results indicate that the base editing methods and base editors disclosed in the present invention can be compatible with different cytosine deaminases.

[0774] In Figure 7A the base editing window of the base editor was analyzed according to the NGS results. The experimental treatments or construct combinations involved in the figure and the schematic diagrams of the associated vectors are shown below:

[0775]

[0776] In Figure 7B the base editing window of the base editor was analyzed according to the NGS results. The experimental treatments or construct combinations involved in the figure and the schematic diagrams of the associated vectors are shown below:

[0777]

[0778] Example 5: Base editors containing adenine deaminase

[0779] To expand the range of target sequences editable by the base editors of the present invention, in this example, an adenine deaminase TadA-8e that uses deoxyadenine (A) in single-stranded DNA as a substrate was used as the deaminase to target A1, A7, A12, and A13 of the OsCKX2 locus (as Figure 8 shown). In this example, UGI is not an essential component of the base editor being tested because it is not necessary for adenine base editing. The adenine base editing window of the base editor was analyzed according to the NGS results. NGS analysis showed that effective targeted A to G conversion occurred at the target( Figure 8) This indicates that the base editor of the present invention can be compatible with adenine deaminase and is used for adenine base editing. From Examples 4 and 5, it can be seen that the base editing method and base editor disclosed in the present invention can be compatible with different deaminases and play their corresponding editing roles.

[0780] In Figure 8 the experimental treatments or construct combinations involved in the figure and the schematic diagrams of the associated vectors are shown as follows:

[0781]

[0782] Example 6: Base Editor of a Fusion Protein Containing Base Editing Components

[0783] After the functions and effects of the base editor of the present invention were demonstrated in the above examples, this example verified whether the transformation efficiency (and correspondingly the editing efficiency) could be improved by fusing modular components into a single vector. The structures of two examples of such base editors containing fusion components are as Figure 9 shown, where the exonuclease is fused to the amino terminus of the deaminase-UGI fusion protein through an XTEN spacer peptide or a 48-amino acid spacer peptide (48aa) to target the OsDEP1 gene, that is, the deaminase is fused with the exonuclease.

[0784] Different construct combinations were transformed into rice protoplasts to target the OsDEP1 locus, and the C>T base editing efficiency and insertion-deletion (indel) byproduct frequencies detected by NGS. NGS analysis showed that targeted base editing could be achieved by fusing the exonuclease and deaminase, but the efficiency under this vector structure was close to that when the exonuclease and deaminase were expressed separately (as Figure 10A shown). When using this base editor, the editing window tended to be more towards C1 and C9 (as Figure 10B shown), which was consistent with the catalytic direction of the mExoI exonuclease.

[0785] Figure 10A and Figure 10B the experimental treatments or construct combinations involved in the figure and the schematic diagrams of the associated vectors are shown as follows:

[0786]

[0787] In addition, the inventors also tested other fusion protein structures. The structures of the above base editors are shown in Figure 11A and Figure 11B where the deaminase (hAPOBEC3A or rAPOBEC1) is linked to TALE-L ( Figure 11A ) or TALE-R ( Figure 11B) is fused to the amino terminus, and UGI and exonuclease are expressed by separate vectors, that is, the deaminase, TALE protein, and nickase are fused.

[0788] For deaminase-TALE-FokI-R nickase , OsDEP1 was selected as the target gene to be characterized (as Figure 12A shown), and for deaminase-TALE-FokI-L nickase , OsCKX2 was selected as the target gene to be characterized (as Figure 12B shown). NGS analysis showed that both deaminase-TALE-FokI-L / R nickase achieved the conversion of C to T at the target site. It indicates that the deaminase can form a fusion with the TALE protein and nickase without interfering with their respective functions. In addition, the experimental results further showed that base editing could occur using both deaminase hAPOBEC3A and rAPOBEC1 (as Figure 12A and Figure 12B shown).

[0789] In Figure 12A , the experimental treatments or construct combinations involved in the figure and the schematic diagrams of the associated vectors are shown as follows:

[0790]

[0791] In Figure 12B , the experimental treatments or construct combinations involved in the figure and the schematic diagrams of the associated vectors are shown as follows:

[0792]

[0793] To investigate the effects brought about by the fusion of UGI or exonuclease, in the deaminase-TALE-FokI-R nickase construct with the same target specificity as the present invention, its base editor is connected to UGI at the carboxyl terminus of FokI-L D450A (as Figure 13A shown) or at the amino terminus of the deaminase (as Figure 13B shown) through a spacer peptide of 48 amino acids or 4 amino acids. NGS analysis showed that the effect of connecting UGI to the fusion protein was similar to that of the embodiment expressing UGI alone ( Figure 14 ). In addition, in the deaminase-TALE-FokI-R nickase construct, the embodiments in which the exonuclease is fused to the carboxyl terminus of FokI-R through a spacer peptide of 4 amino acids, 16 amino acids, or 48 amino acids also had similar editing efficiencies ( Figure 14 ). Therefore, both the separate expression of UGI / exonuclease and their co-expression by fusing them to the vector are technical solutions that can be adopted in the present invention.

[0794] In Figure 14 this study, different construct combinations were transformed into rice protoplasts to target the OsDEP1 locus. By analyzing the DNA strands and editing windows with base editing events through high-throughput sequencing results, the experimental treatments or construct combinations involved in the figure and the schematic diagrams of the associated vectors are shown as follows:

[0795]

[0796] Based on the above results, the modular components of the base editor of the present invention can be expressed individually or form one or more fusion proteins with each other.

[0797] Example 7: Base editing of plant nuclear genomes

[0798] In the above examples, the functions and characteristics of the base editor of the present invention were verified, that is, a composition containing modular components of deaminase, exonuclease, nickase, and DNA-binding protein TALE can achieve efficient and precise editing of DNA. For ease of description, the above base editor was named DENT (Deaminase-Exonuclease-Nickase-TALE), and was named CyDENT (Cytidine Deaminase-Exonuclease-Nickase-TALE) and AdDENT (Adenine Deaminase-Exonuclease-Nickase-TALE) respectively according to the types of deaminases. In this example, the applicable environments and scenarios of the base editor of the present invention were analyzed.

[0799] The inventors selected the rice protoplast nuclear genome to evaluate the editing effect of the base editor of the present invention. In this example, 4 pairs of TALE proteins were designed for the rice endogenous gene loci OsDEP1, OsCKX2, OsBADH2, and OsSD1 respectively. Exonucleases with 5'→3' (mExol) or 3'→5' (Trex2) enzymatic cleavage preferences were used to evaluate the effect of combining exonuclease with nickase to generate ssDNA intermediates. In this example, the highly efficient cytosine deaminase hAPOBEC3A (hA3A) was selected to perform cytosine deamination on the ssDNA intermediate, and the uracil glycosylase inhibitor (UGI) peptide was fused to its C-terminus to further improve the editing efficiency by minimizing the influence of DNA base excision repair. The nuclear localization signal (NLS) was fused to the N-terminus of each component to directly edit the nuclear genome. This combination of base editors targeting the nuclear genome is hereby referred to as nuCyDENT, and an exemplary construct schematic diagram is as Figure 18As shown, the nuCyDENT was targeted to the rice OsDEP1, OsCKX2, OsBADH2, and OsSD1 loci and introduced into rice protoplasts, and the editing efficiency was evaluated after 2 days. Targeted cytosine base editing was evaluated within the 18-bp spacer region between the TALE binding sites at all four nuclear genomic loci using NGS analysis. We observed an editing efficiency of 3 - 18% and a lower indel frequency compared to the corresponding wild-type TALEN system ( Figure 19A and Figure 19B ). These results indicate that the base editor of the present invention can achieve efficient base editing in the nuclear genome while only inducing low levels of indel by-products.

[0800] In terms of single-strand editing performance, the inventors used nuCyDENT-L (nuCyDENT containing the FokI-L nickase structure) and nuCyDENT-R (nuCyDENT containing the FokI-R nickase structure) for base editing at the rice genomic loci OsCKX2 and OsSD1, respectively. The results showed that when edited with nuCyDENT-R, the DNA top strand was edited; when edited with nuCyDENT-L, the DNA bottom strand was edited ( Figure 20 ). This conclusion is the same as that in Example 2, also showing the single-strand editing performance of CyDENT in the nuclear genome.

[0801] In Figure 19A , Figure 19B , Figure 20 , the experimental treatments or construct combinations involved in the figures are shown as follows:

[0802]

[0803] Example 8: Base Editing of Animal Nuclear Genomes

[0804] This example compared the base editing effects of CyDENT and DdCBE at the target site of the human SIRT6 gene. The inventors designed TALE proteins for the SIRT6 target, designed nuCyDENT-L according to the method of Example 7, and designed the DddA-dependent DdCBE according to the method in the prior art (Nakazato, I. et al. Targeted base editing in the mitochondrial genome of Arabidopsis thaliana. Proc. Natl. Acad. Sci. USA. 119, e2121177119 (2022)). The experimental results showed that nuCyDENT-L had a higher base editing efficiency at the target site than DdCBE ( Figure 21)。It shows that the base editing system of the present invention has good base editing performance in the nuclear genome of animals.

[0805] In Figure 21 , the experimental treatments or construct combinations involved in the figure are shown as follows:

[0806]

[0807] Example 9: Base editing of DNA in organelles - Chloroplast

[0808] The base editor of the present invention can be used for base editing of mitochondrial DNA and chloroplast DNA, and has advantages compared with CRISPR base editors that require nucleic acid components. The protein components in the base editor of the present invention can be transferred into mitochondria and chloroplasts through mitochondrial targeting sequence (MTS) and chloroplast transit peptide (CTP), respectively. In these examples, MTS or CTP can be selected to replace NLS according to the type of target organelle.

[0809] First, we attempted to perform base editing on plant chloroplast DNA using the CyDENT base editing strategy. Plant chloroplast DNA is an important organelle unique to plants, with its own genomic DNA (cpDNA), and cannot be edited using CRISPR-derived base editors either. The inventors replaced NLS with chloroplast transit peptide (CTP) in nuCyDENT designed according to the method of Example 7 (Kang, B.C. et al. Chloroplast and mitochondrial DNA editing in plants. Nat Plants 7, 899 - 905 (2021).)( Figure 22A ), and named it cpCyDENT. The inventors used cpCyDENT-L (containing FokI-L nickase ) and cpCyDENT-R (containing FokI-R nickase ) containing TALE proteins targeting the large subunit gene (rbcL) of endogenous ribulose-1,5-bisphosphate carboxylase / oxygenase (RuBisCO) to transform rice protoplasts. Base editing of the rbcL target was detected in the cpCyDENT-L treatment ( Figure 22B ). It is worth noting that precise editing of specific bases can be achieved by regulating the types and directions of the nickase and exonuclease of cpCyDENT. For example, for the G 1 base (counting the nucleotide at the 5'-most position of the spacer region as position 1, see Figure 22B) It can efficiently edit this base only when using the cpCyDENT-L(mExol) tool, that is, when containing FokI-Lnickase and the 5'→3' mExol exonuclease, and the editing efficiency is about 1.67%. This result is consistent with the conclusion of the foregoing embodiment. These results indicate that cpCyDENT can selectively and precisely perform base editing on the DNA strand in the chloroplast genome.

[0810] In Figure 22B the experimental treatments or construct combinations involved in the figure are shown as follows:

[0811]

[0812] Example 10: Base Editing of DNA in Organelles - Mitochondria

[0813] In this example, the inventors evaluated the effect of CyDENT base editing on the base editing of human cell mitochondrial DNA (mtDNA), replaced the NLS with a mitochondrial targeting sequence (MTS), and selected a promoter and terminator suitable for expression in HEK293T cells to obtain a base editor for mtDNA called mtCyDENT. The mtCyDENT construct was generated as Figure 15A shown (TALE-FokI-R nickase and TALE-FokI-L nickase ).

[0814] First, a target site in the ND6 gene of human mitochondrial DNA was selected, and TALE-FokI-R nickase and TALE-FokI-L nickase expression vectors were constructed, in which the TALE protein was modified to target this site, and were co-transfected into HEK293T cells together with vectors expressing deaminase (hAPOBEC3A or C57), exonuclease (mExoI or Trex2), and UGI, where the mitochondrial targeting sequence (MTS) was fused to the end of the protein. NGS was used to determine the base editing frequency after transfection of the base editor. The results showed that targeted cytosine base editing was achieved with an efficiency of approximately 6.0% in the mitochondrial DNA target of human cells ( Figure 15C ). The results indicate that the base editor of the present invention can be used for base editing of organelle genomes.

[0815] In Figure 15C different construct combinations were transfected into HEK293T cells to target the mitochondrial ND6 locus, and the DNA strand and editing window where base editing occurred were analyzed through high-throughput sequencing results. The experimental treatments or construct combinations involved in the figure and the associated vector schematic diagrams are shown as follows:

[0816]

[0817] Example 11: Influence of Base Editor Fusion Status on Mitochondrial DNA Editing

[0818] Next, the inventors verified the effects of separately expressed deaminase, exonuclease, UGI, and TALE-FokI nickase on the base editing efficiency of mtDNA.

[0819] For this purpose, the inventors used a small peptide called γb fused to the N-terminus of the domain of one or more modular components in mtCyDENT to drive the recruitment of each protein component ( Figure 23A ). γb is an RNA silencing suppressor derived from Barley stripe mosaic virus (BSMV) and has self-interaction (Jiang, Z., Yang, M., Zhang, Y., Jackson, A. O. & Li, D. in Encyclopedia of Virology 420 - 429 (2021)). The exonuclease selected by the inventors in this experiment was Trex2. The inventors designed multiple fusion schemes of γb with each component to screen out the base editor composition with the best editing effect ( Figure 23B ). Considering the size of the protein components entering the mitochondria, in this example, the constructs of 5 proteins / fusion proteins as shown in Figure 23A were used for expression. The proteins / fusion proteins were respectively: the fusion protein of TALE-L and FokI-L (abbreviated as TALE-L-FokI-L, TALEL-FL or TALEL-FokI-L), the fusion protein of TALE-R and FokI-R (abbreviated as TALE-R-FokI-R, TALEL-FR or TALER-FokI-R), hA3A deaminase protein, Trex2 exonuclease protein, and UGI protein. Among them, the tail label D450A was used to represent its mutant, and WT represented the wild type.

[0820] The experimental results showed that when γb was only fused with UGI and Trex2, it had a high editing effect. The base editor composition with γb fused to UGI and Trex2 structures respectively was named mtCyDENT1b.

[0821] Next, we evaluated mtCyDENT and mtCyDENT1b at another 7 endogenous mtDNA genomic sites. We observed that the average editing frequency of mtCyDENT was 1.16 - 11.7%, and mtCyDENT1b further increased the average editing efficiency by 2.42 - 6.18 times, reaching 4.55 - 39.3% ( Figure 24)。Moreover, the editing efficiency of mtCyDENT1b at the ND1.2, ND1.3, ND3, and ND6.2 target sites with the same TALE sequence is higher than that of DdCBE. In addition, we also noticed that when using CyDENT base editing at the mtDNA target site, the indel frequency is lower than that of DdCBE ( Figure 25 ). In summary, both mtCyDENT and mtCyDENT1b can achieve efficient base editing in human mitochondrial DNA.

[0822] In Figure 23B , the experimental treatments or construct combinations involved in the figure are shown as follows (from top to bottom):

[0823]

[0824]

[0825] In Figure 24 - 27 , the experimental treatments or construct combinations involved in the figure are shown as follows:

[0826]

[0827]

[0828] Example 12: Improving the Editing Efficiency and Precision of CyDENT

[0829] As described in the aforementioned Example 4, the base editor of the present invention can be self-assembled from multiple functional modules and is compatible with different types of deaminases. Therefore, the deaminase domain in the base editor can be replaced with known deaminases in the prior art to utilize the unique characteristics of each deaminase, thereby enhancing the activity or further improving the precision of the editing strand. A newly discovered single-stranded DNA (ssDNA)-specific cytidine deaminase, Sdd7, was found to have higher editing activity than other deaminases (Huang, J. et al. Discovery of new deaminase functions by structure-based protein clustering. bioRxiv (2023)). In this example, the inventors took the mtCyDENT1b composition as an example and used Sdd7 as the deaminase of this editor to evaluate the editing efficiency at the mtDNA target sites ND5.1, ND6, and ND1.3. We observed that 87.5% of the base editing induced by Sdd7-mtCyDENT1b-L occurred on only one DNA strand, and 93.0% of the base editing induced by Sdd7-mtCyDENT1b-R occurred on only one DNA strand. This result further demonstrates the superior base editing strand specificity of CyDENTFigure 26 )。The average editing efficiency of these two editors for the target strand at the bottom of DNA is between 4.88 - 9.13% ( Figure 27 ). These results further verify the ability of the base editor of the present invention to replace the deaminase domain during modular assembly.

[0830] Example 13: Improvement of the base editor

[0831] In the foregoing examples, the inventors verified through experiments that the base editor composition of the present invention has technical advantages such as single-strand editing specificity, modular assembly, efficient, precise and controllable base editing, and low indel frequency. In the subsequent examples, the inventors further optimized the base editor in order to obtain a base editor combination with more excellent functions.

[0832] In this example, the inventors fused the deaminase and exonuclease domains with a 48-amino acid spacer peptide (flexible linker) to the N-terminus of TALE-L and TALE-R, and UGI was fused to the C-terminus and N-terminus of FokI-L and FokI-R respectively. This construct structure is hereinafter referred to as mtCyDENT2 ( Figure 28A ). The base editing effect was detected using mtCyDENT2-L (containing FokI-L nickase ) on ND6 ( Figure 28B ), and 94.5% of the base editing occurred only on the top strand, demonstrating the good single-strand specific editing ability of the CyDENT system.

[0833] In Figures 28A - 28B , the experimental treatments or construct combinations involved in the figure are shown as follows:

[0834]

[0835] Example 14 mtCyDNET base editing for the G C motif

[0836] Since the DddA-dependent DdCBE system has strict context restrictions on the T C motif for cytidine deamination, and it was found that in G CEditing in sequence context occurs less frequently (Nakazato, I. et al. Targeted base editing in the mitochondrial genome of Arabidopsis thaliana. Proc. Natl. Acad. Sci. USA. 119, e2121177119 (2022)). A phage-assisted non-continuous and continuous evolution was used to evolve "wild-type" DddA (Mok, B. Y. et al. CRISPR-free base editors with enhanced activity and expanded targeting scope in mitochondrial and nuclear DNA. Nat. Biotechnol. 40, 1378-1387 (2022)). The evolved DddA11 variant has better A C and C C sequence motif compatibility, but editing of DddA11 on GC sequence motifs remains challenging. In this example, modular replacement of the deaminase domain of CyDENT was utilized to achieve efficient and strand-selective G C sequence motif editing.

[0837] The inventors introduced a single-stranded DNA-specific cytidine deaminase with G C sequence motif editing activity, thereby developing a G C compatible mtCyDENT base editor. Recently, a newly discovered single-stranded DNA-specific, G C and A C compatible cytidine deaminase Sdd3, which exhibits higher editing activity on G C sequence motifs than other deaminases (Huang, J. et al. Discovery of new deaminase functions by structure-based protein clustering. bioRxiv (2023)).

[0838] Therefore, the present invention designed TALE arrays ( Figure 29 ) targeting the ND1.2 and ND6.2 sites in HEK293T cells to evaluate the editing preferences of sequence motifs that are difficult to edit by the prior art. Notably, at the ND1.2 and ND6.2 sites, G CThe editing efficiencies of chain-specific cytosine base editing observed under the sequence motifs reached 21.0% and 20.0% respectively, which are editing efficiencies that the prior art DdCBE cannot achieve at the same target sites; at the ND1.2 site, 96.9% of the editing selectively occurred on the top DNA strand, while at the ND6.2 site, 92.0% of the editing selectively occurred on the bottom DNA strand( Figure 29 ).

[0839] Subsequently, the inventors adjusted the TALE binding sites and observed that Sdd3-mtCyDENT had an editing efficiency of 2.06% at the ND6.2 site( Figure 30 ). It is reported that this specific mutation (m.14453G>A) is directly related to the development of Leigh syndrome, and the prior art DdCBE cannot achieve editing in the background of this same target sequence. Therefore, mtCyDENT and future optimized products can be used as a superior base editing method for precise editing of pathogenic mutations on mtDNA.

[0840] In Figure 29 , 30 , the experimental treatments or construct combinations involved in the figures are shown as follows:

[0841]

[0842] Example 15: Off-target analysis of mtCyDENT

[0843] Editing mitochondria with the prior art DdCBE can induce a large number of nuclear off-target edits. To evaluate the off-target rates of CyDENT in the whole nuclear genome and the whole mitochondrial genome, in this example, 2.25 Tb of pure bases were obtained, with an average of 281.13 Gb per sample, and the average depth of mitochondrial genome sequencing was approximately 6362-fold. The human reference genome used was hg19.

[0844] In this example, DdCBE plasmids targeting ND3 and mtCyDENT1b-R(hA3A) plasmids, as well as mtCyDENT2-L(Sdd3) plasmids targeting ND6.2 were designed and transfected into HEK293T cells. Whole genome sequencing (WGS) and NGS analysis showed that these plasmids were able to perform editing on the sequence motifs on G C . Figure 31A) Subsequently, the off-target rates of the entire mitochondrial genome and the entire nuclear genome were analyzed. The results showed that the average C·G-to-T·A and G·C-to-A·T base conversion frequencies in the mitochondrial genomes of the untreated negative control group, DdCBE, mtCyDENT1b-R(hA3A), and mtCyDENT2-L(Sdd3) treatment groups were 4.8%, 6.9%, 16.5%, and 5.9%, respectively. Compared with the control group, we found an average of 32, 678, and 16 single nucleotide variations (SNVs) in the mitochondrial genomes of the DdCBE, mtCyDENT1b-R(hA3A), and mtCyDENT2-L(Sdd3) treatment groups, respectively. By analyzing the 5 bp regions upstream and downstream of each potential off-target SNV, a conserved TC motif was found in the DdCBE and mtCyDENT1b-R(hA3A) groups, while a conserved GC / AC motif was found in the mtCyDENT2-L(Sdd3) group( Figure 31B )。

[0845] In the nuclear genome, the inventors analyzed the TALE-dependent off-target effects. A total of 74,963 potential off-target regions (containing 0 - 3 regions that do not match the ND3 and ND6.2 TALE binding sites) were identified. We observed that there were no differences in the allelic SNV frequencies and indel frequencies at the ND3 or ND6.2 loci in the control group, DdCBE, mtCyDENT1b-R(hA3A), and mtCyDENT2-L(Sdd3) treatment groups( Figure 31C )。These results indicate that the modular assembly and optimization of CyDENT can minimize off-target effects in mitochondrial and nuclear genomes. mtCyDENT is a valuable tool for mitochondrial genome editing.

[0846] In Figures 31A - 31C , the experimental treatments or construct combinations involved in the figures are shown as follows:

[0847]

[0848] The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.

Claims

1. A nucleic acid base editor, characterized in that, the nucleic acid base editor comprises the following components: a) A sequence-specific DNA-binding protein; b) A nickase; c) An exonuclease; d) A base-specific deaminase; the sequence-specific DNA-binding protein is a TALE protein; the nickase is a dimer of a cleavage functional domain monomer of FokI or a mutant thereof, the dimer or mutant is composed of a pair of interacting cleavage functional domain monomers of FokI, and only one of the dimer or mutant has DNA endonuclease activity; The cleavage functional domain monomer of FokI with DNA endonuclease activity is selected from the FokI-L protein with the amino acid sequence shown in SEQ ID No. 87 and the FokI-R protein with the amino acid sequence shown in SEQ ID No. 88; The cleavage functional domain monomer of FokI lacking DNA endonuclease activity is selected from the FokI-L protein with an amino acid sequence as shown in SEQ ID No. 60, the FokI-L protein with an amino acid sequence as shown in SEQ ID No. 61, the FokI-R protein with an amino acid sequence as shown in SEQ ID No. 62, and the FokI-R protein with an amino acid sequence as shown in SEQ ID No. 63; D450A the FokI-L protein with an amino acid sequence as shown in SEQ ID No. 61, D467A the FokI-R protein with an amino acid sequence as shown in SEQID No. 62, D450A the FokI-R protein with an amino acid sequence as shown in SEQ ID No. 63; D467A protein; the base-specific deaminase is selected from a cytosine-specific deaminase or an adenine-specific deaminase; the exonuclease is a 5'-exonuclease or a 3'-exonuclease.

2. The nucleic acid base editor according to claim 1, characterized in that, each component of the nucleic acid base editor exists alone or forms one or more fusion proteins.

3. The nucleic acid base editor according to claim 1, characterized in that, the cleavage functional domain monomer of FokI is isolated from a mutant of the wild-type FokI protein, and the mutant of the wild-type FokI protein is mutated at position 450 and / or position 467.

4. The nucleic acid base editor according to claim 3, characterized in that, the mutation causes the cleavage functional domain monomer of FokI to lose DNA endonuclease activity.

5. The nucleic acid base editor according to claim 1, characterized in that, the cleavage functional domain monomer of FokI is isolated from a mutant of the wild-type FokI protein, and the mutation causes the cleavage functional domain monomer of FokI to be unable to self-assemble with the cleavage functional domain monomer of FokI containing the same site mutation to form a dimer.

6. The nucleic acid base editor according to claim 1 or 2, characterized in that, the amino acid sequence of the base-specific deaminase is selected from SEQ ID NO. 36-59, 80-86.

7. The nucleic acid base editor according to claim 1 or 2, characterized in that, the base-specific deaminase is a cytosine-specific deaminase.

8. The nucleic acid base editor according to claim 7, characterized in that, the cytosine-specific deaminase is one or more of hAPOBEC3A, rAPOBEC1, hAID, pmCDA1 or Sdd deaminase.

9. The nucleic acid base editor according to claim 7, characterized in that, the nucleic acid base editor further contains: e) Uracil glycosylase inhibitor (UGI); and the uracil glycosylase inhibitor exists alone or forms at least one fusion protein with other nucleic acid base editor components.

10. The nucleic acid base editor according to claim 1 or 2, characterized in that, The base-specific deaminase is an adenine-specific deaminase.

11. The nucleic acid base editor according to claim 10, wherein, the adenine-specific deaminase is TadA-8e.

12. The nucleic acid base editor according to any one of claims 1-5, 8, 9, and 11, wherein, the nucleic acid base editor further contains: f) γb; the γb forms at least one fusion protein with other nucleic acid base editor components.

13. The nucleic acid base editor according to claim 1 or 2, wherein, the amino acid sequence of the exonuclease is selected from SEQ ID NO.64-67, 153.

14. A fusion protein as a nucleic acid base editor, wherein, the fusion protein contains the protein domain of the base editor described in any one of claims 1-13.

15. The fusion protein as a nucleic acid base editor according to claim 14, wherein, the fusion protein is arranged in a linear order starting from the amino terminus of the protein, and it contains: exonuclease, XTEN spacer peptide, base-specific deaminase, XTEN spacer peptide, uracil glycosylase inhibitor (UGI), and nuclear localization signal.

16. The fusion protein as a nucleic acid base editor according to claim 14, wherein, the fusion protein is arranged in a linear order starting from the amino terminus of the protein, and it contains: exonuclease, a 48-amino acid spacer peptide, base-specific deaminase, XTEN spacer peptide, uracil glycosylase inhibitor (UGI), and nuclear localization signal.

17. The nucleic acid base editor according to claim 2, wherein, the fusion protein includes: a first fusion protein, which contains: nuclear localization signal (NLS), sequence-specific DNA binding protein, nickase, and base-specific deaminase; a second fusion protein, which contains: exonuclease and nuclear localization signal (NLS); and, a third fusion protein, which contains: uracil glycosylase inhibitor (UGI) and nuclear localization signal (NLS).

18. The nucleic acid base editor according to claim 2, wherein, the fusion protein includes: A first fusion protein, arranged in a linear order starting from the amino terminus of the protein, comprising: a nuclear localization signal (NLS), a base-specific deaminase, a TALE-L protein, FokI-L D450A Protein / FokI-L D467A Protein / FokI-L protein, T2A sequence, NLS, TALE-R protein, and FokI-R protein / FokI-R D450A Protein / FokI-R D467A Protein; a second fusion protein, which contains: exonuclease and nuclear localization signal (NLS); and, a third fusion protein, which contains: uracil glycosylase inhibitor (UGI) and nuclear localization signal (NLS).

19. The nucleic acid base editor according to claim 2, wherein, the fusion protein includes: A first fusion protein, arranged in a linear order starting from the amino terminus of the protein, comprising: a nuclear localization signal (NLS), a TALE-L protein, a FokI-L D450A protein / FokI-L D467A protein / FokI-L protein, a T2A sequence, an NLS, a base-specific deaminase, a 48-amino acid spacer peptide, a TALE-R protein, and a FokI-R protein / FokI-R D450A protein / FokI-R D467A protein; a second fusion protein, which contains: exonuclease and nuclear localization signal (NLS); and, a third fusion protein, which contains: uracil glycosylase inhibitor (UGI) and nuclear localization signal (NLS).

20. The nucleic acid base editor according to claim 2, wherein, the fusion protein includes: a first fusion protein, which contains: nuclear localization signal (NLS), sequence-specific DNA binding protein, nickase, base-specific deaminase, and uracil glycosylase inhibitor (UGI); and, A second fusion protein, comprising: an exonuclease and a nuclear localization signal (NLS).

21. The nucleic acid base editor according to claim 2, wherein, the fusion protein comprises: A first fusion protein, arranged in a linear order starting from the amino terminus of the protein, comprising: a nuclear localization signal (NLS), a base-specific deaminase, a spacer peptide of 48 amino acids, a TALE-L protein, a FokI-L D450A protein, a T2A sequence, an NLS, a TALE-R protein, a FokI-R protein, a spacer peptide of 4 amino acids, and a uracil glycosylase inhibitor (UGI); and, a second fusion protein, comprising: an exonuclease and a nuclear localization signal (NLS), alternatively, the fusion protein comprises: A first fusion protein, arranged in a linear order starting from the amino terminus of the protein, comprising: a nuclear localization signal (NLS), a uracil glycosylase inhibitor (UGI), a spacer peptide of 4 amino acids, a base-specific deaminase, a spacer peptide of 48 amino acids, a TALE-L protein, a FokI-L D450A protein, a T2A sequence, an NLS, a TALE-R protein, a FokI-R protein; and, a second fusion protein, comprising: an exonuclease and a nuclear localization signal (NLS).

22. The nucleic acid base editor according to claim 2, wherein, the fusion protein comprises: A first fusion protein, arranged in a linear order starting from the amino terminus of the protein, comprising: a mitochondrial targeting sequence (MTS), a TALE-L protein, and a FokI-L D450A protein / FokI-L D467A protein / FokI-L protein; A second fusion protein, arranged in a linear order starting from the amino terminus of the protein, comprising: a mitochondrial targeting sequence (MTS), a TALE-R protein, and a FokI-R protein / FokI-R D450A protein / FokI-R D467A protein; a third fusion protein, arranged linearly starting from the amino terminus of the protein, comprising: a mitochondrial targeting sequence (MTS) and an exonuclease; a fourth fusion protein, arranged linearly starting from the amino terminus of the protein, comprising: a mitochondrial targeting sequence (MTS) and a base-specific deaminase; and, a fifth fusion protein, arranged linearly starting from the amino terminus of the protein, comprising: a mitochondrial targeting sequence (MTS) and a uracil glycosylase inhibitor (UGI).

23. The nucleic acid base editor according to claim 2, wherein, the fusion protein comprises: A first fusion protein, arranged in a linear order starting from the amino terminus of the protein, comprising: a mitochondrial targeting sequence (MTS), a TALE-L protein, and a FokI-L D450A protein / FokI-L D467A protein / FokI-L protein; A second fusion protein, arranged in a linear sequence starting from the amino terminus of the protein, comprising: a mitochondrial targeting sequence (MTS), a TALE-R protein, and a FokI-R protein / FokI-R D450A protein / FokI-R D467A protein; a third fusion protein, arranged linearly starting from the amino terminus of the protein, comprising: a mitochondrial targeting sequence (MTS), γb and an exonuclease; a fourth fusion protein arranged linearly starting from the amino terminus of the protein, comprising: a mitochondrial targeting sequence (MTS) and a base-specific deaminase; and, a fifth fusion protein, arranged linearly starting from the amino terminus of the protein, comprising: a mitochondrial targeting sequence (MTS), γb and a uracil glycosylase inhibitor (UGI).

24. The nucleic acid base editor according to claim 2, wherein, the fusion protein comprises: a first fusion protein, comprising: a nuclear localization signal (NLS) / chloroplast transit peptide (CTP) / mitochondrial targeting sequence (MTS), a sequence-specific DNA binding protein and a nickase; a second fusion protein, comprising: an exonuclease and a nuclear localization signal (NLS) / chloroplast transit peptide (CTP) / mitochondrial targeting sequence (MTS); and, a third fusion protein, comprising: a base-specific deaminase, a uracil glycosylase inhibitor (UGI) and a nuclear localization signal (NLS) / chloroplast transit peptide (CTP) / mitochondrial targeting sequence (MTS).

25. The nucleic acid base editor according to claim 2, wherein, the fusion protein comprises: A first fusion protein, arranged in a linear order starting from the amino terminus of the protein, comprising: a nuclear localization signal (NLS) / chloroplast transit peptide (CTP) / mitochondrial targeting sequence (MTS), a TALE-L protein, a FokI-L D450A protein / FokI-L D467A protein, a T2A sequence, an NLS, a TALE-R protein, a FokI-R protein, or, comprising: a nuclear localization signal (NLS) / chloroplast transit peptide (CTP) / mitochondrial targeting sequence (MTS), a TALE-L protein, a FokI-L protein, a T2A sequence, an NLS, a TALE-R protein, a FokI-R D450A protein / FokI-R D467A protein; a second fusion protein, arranged linearly starting from the amino terminus of the protein, comprising: a nuclear localization signal (NLS) / chloroplast transit peptide (CTP) / mitochondrial targeting sequence (MTS) and an exonuclease; and, a third fusion protein, arranged linearly starting from the amino terminus of the protein, comprising: a nuclear localization signal (NLS) / chloroplast transit peptide (CTP) / mitochondrial targeting sequence (MTS), a base-specific deaminase, an XTEN spacer peptide and a uracil glycosylase inhibitor (UGI).

26. The nucleic acid base editor according to claim 2, wherein, The fusion protein comprises: A first fusion protein, arranged in a linear order starting from the amino terminus of the protein, comprising: a mitochondrial targeting sequence (MTS), a base-specific deaminase, a 48-amino acid spacer peptide, a TALE-L protein, a FokI-L D450A protein / FokI-L D467A protein / FokI-L protein, an 11-amino acid spacer peptide, and a uracil glycosylase inhibitor (UGI); and, A second fusion protein, arranged in a linear order starting from the amino terminus of the protein, comprising: a mitochondrial targeting sequence (MTS), a 48-amino acid spacer peptide, a TALE-R protein, a uracil glycosylase inhibitor (UGI), a 14-amino acid spacer peptide, and a FokI-R protein / FokI-R D450A protein / FokI-R D467A protein.

27. A recombinant expression construct for nucleic acid base editing, characterized in that the recombinant expression construct is used for expressing the nucleic acid base editor according to any one of claims 1-13 and 17-26 or the fusion protein according to any one of claims 14-16.

28. A method for nucleic acid base editing in cells for non-therapeutic purposes, characterized in that introducing the nucleic acid base editor according to any one of claims 1-13 or the recombinant expression construct according to claim 27 into cells to edit the target gene.

29. The method for nucleic acid base editing according to claim 28, characterized in that the target gene is selected from nuclear genomic DNA, mitochondrial genomic DNA, or chloroplast genomic DNA.

30. The method for nucleic acid base editing according to claim 28 or 29, characterized in that the target gene is nuclear genomic DNA, and the nucleic acid base editor further comprises a nuclear localization signal (NLS).

31. The method for nucleic acid base editing according to claim 28 or 29, characterized in that the target gene is mitochondrial genomic DNA, and the nucleic acid base editor further comprises a mitochondrial targeting sequence (MTS).

32. The method for nucleic acid base editing according to claim 28 or 29, characterized in that the target gene is chloroplast genomic DNA, and the nucleic acid base editor further comprises a chloroplast transit peptide (CTP).

33. Use of the base editor according to any one of claims 1-13 and 17-26, the fusion protein according to any one of claims 14-16, or the recombinant expression construct according to claim 27 in the preparation of a reagent for base editing of DNA in cells, wherein the cells are mammalian cells, bacteria, protists, fungi, insect cells or plant cells.

34. The use according to claim 33, characterized in that the plant cells are derived from the whole plant of monocotyledonous plants or dicotyledonous plants.

35. The use according to claim 33, characterized in that the plant cells are derived from the seedlings of monocotyledonous plants or dicotyledonous plants.

36. The use according to claim 33, characterized in that the plant cells are derived from the meristematic tissue, fundamental tissue, vascular tissue or dermal tissue of monocotyledonous plants or dicotyledonous plants.

37. The use according to claim 33, characterized in that the plant cells are derived from the seeds, leaves, roots, buds, stems, flowers or fruits of monocotyledonous plants or dicotyledonous plants.

38. The use according to claim 33, characterized in that the plant cells are derived from the stolons, bulbs, tubers, corms, sprouts or vegetative terminal branches of monocotyledonous plants or dicotyledonous plants.

39. The use according to claim 33, characterized in that the plant cells are derived from the tumor tissues of monocotyledonous plants or dicotyledonous plants.

40. The use according to claim 33, characterized in that the mammalian cells are selected from human somatic cells.

41. The use according to claim 33, It is characterized in that the mammalian cells are selected from human neurons, muscle cells, endocrine / exocrine cells, epithelial cells, hematopoietic cells, and bone cells.

42. The application according to claim 33, It is characterized in that the mammalian cells are selected from human induced pluripotent stem cells.

43. The application according to claim 33, It is characterized in that the mammalian cells are selected from human tumor cells.

44. The application according to claim 33, It is characterized in that the fungi are selected from yeast or non-conventional yeast.

45. A pharmaceutical composition for treating a disease in a subject in need thereof, It is characterized in that the pharmaceutical composition comprises the base editor according to any one of claims 1-13 and 17-26, the fusion protein according to any one of claims 14-16, or the recombinant expression construct according to claim 27.

46. The pharmaceutical composition according to claim 45, It is characterized in that the pharmaceutical composition further comprises a pharmaceutically acceptable carrier.

47. A method for generating a genetically modified plant, It is characterized in that the method comprises introducing the base editor according to any one of claims 1-13 and 17-26, the fusion protein according to any one of claims 14-16, or the recombinant expression construct according to claim 27 into at least one of the plants.

Citation Information

Patent Citations

  • Delivery, use and therapeutic applications of the crispr-CAS systems and compositions for HBV and viral diseases and disorders

    WO2015089465A1

  • Novel crispr enzymes and systems

    WO2016205711A1

  • Compounds, compositions and methods for cancer treatment

    WO2018141835A1

  • Uses of adenosine base editors

    WO2019079347A1

  • Methods and compositions for editing nucleotide sequences

    WO2020191233A1