A base editor and applications thereof
By developing a CRISPR-independent nucleic acid base editor, and utilizing components such as sequence-specific DNA-binding proteins, nicking enzymes, and deaminases, the safety and accuracy issues of existing base editors on nuclear, mitochondrial, and chloroplast DNA have been resolved. This has achieved high-purity editing with low off-target rates, expanding the application scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- INST OF GENETICS & DEVELOPMENTAL BIOLOGY CHINESE ACAD OF SCI
- Filing Date
- 2023-11-30
- Publication Date
- 2026-05-05
AI Technical Summary
Existing base editors suffer from insufficient safety and precision in nuclear, mitochondrial, and chloroplast DNA, resulting in high off-target mutation frequencies, low product purity, and limited application scope. In particular, the CRISPR system cannot effectively transfer these organelles.
Develop a nucleic acid base editor that does not rely on CRISPR technology, which includes sequence-specific DNA-binding proteins, nicking enzymes, exonucleases, and base-specific deaminases, combined with uracil glycosylation inhibitors, to achieve high-precision and high-purity base editing through specific DNA binding and cleavage.
It enables highly safe and precise base editing of nucleus, mitochondrial DNA, and chloroplast DNA, reduces off-target rates, improves the purity and efficiency of edited products, and expands the scope of applications.
Smart Images

Figure CN121065143B_ABST
Abstract
Description
[0001] This application is a divisional application of Chinese patent application No. 202311622479.8, filed on November 30, 2023, entitled "A Base Editor and Its Application".
[0002] Priority and related applications
[0003] This application claims priority to Chinese Patent Application No. 202211613160.4, filed on December 15, 2022, entitled “A Base Editor and Its Application”, and Chinese Patent Application No. 202311017698.3, filed on August 14, 2023, entitled “A Base Editor and Its Application”, the entire contents of which, including appendices, are incorporated herein by reference. Technical Field
[0004] This invention relates to the field of gene editing, specifically to a nucleic acid base editor, and more particularly to a base editor comprising a sequence-specific DNA-binding protein, a nicking enzyme, an exonuclease, and a base-specific deaminase, and its applications. Background Technology
[0005] Mutations in the genome and mitochondrial DNA are known to cause various genetic diseases (Newby et al., 2021, Nature 595: 295-302), and correcting these mutations holds promise for effectively treating or improving some serious diseases. In plants, some important agronomic traits are associated with single nucleotide variants (SNVs) occurring in the plant genome, mitochondrial genome, or chloroplast genome; introducing these SNVs into plants can promote plant performance, molecular breeding, and restore gene function to alleviate disease states.
[0006] Genome editing has demonstrated enormous potential for genome modification; among genome editing tools, base editing enables targeted base substitutions without introducing DNA double-strand breaks (DSBs), thus achieving more precise and accurate editing (Gaudelli et al., 2017, Nature 551: 464-471; Komor et al., 2016, Nature 533: 420-424), and therefore holds great promise for disease treatment and crop improvement.
[0007] Cytosine base editors (CBE) (Komor et al., 2016, Nature 533: 420-424) and adenine base editors (ABE) (Gaudelli et al., 2017, Nature 551: 464-471) are the most widely used base editors. In the CBE system, a CRISPR-Cas9 cleavage enzyme (nCas9) with single-strand DNA cleaving activity is guided by sgRNA to the target dsDNA. nCas9 cleaves the sgRNA target strand, forming an R loop. Subsequently, a single-strand-specific cytosine deaminase converts cytosine (C) to uracil (U) within a window of approximately five nucleotides in the vesicular structure of the single-stranded DNA generated by nCas9. After DNA repair, U is replaced by T, resulting in a C:G base pair to T:A base pair conversion. Furthermore, the addition of uracil glycosylase inhibitors (UGIs) that inhibit uracil excision and its downstream processes can improve base editing efficiency and product purity. Cytosine deaminases suitable for the Cas-mediated CBE system include, but are not limited to, APOBEC1, hAID, and hAPOBEC3A. Several new deaminase systems have recently been discovered that are also suitable for the deaminases described in this invention (Huang, J. et al. Discovery of new deaminase functions by structure-based protein clustering. bioRxiv (2023).).
[0008] The ABE system was created by fusing nCas9 with an artificially evolved single-stranded DNA adenine deaminase, TadA (Gaudelli et al., 2017, Nature 551: 464-471). The working principle of ABE is similar to CBE: nCas9, guided by sgRNA, cuts a gap in the target DNA strand, while the adenine deaminase TadA converts adenine (A) to hypoxanthine (I), which is then replaced by G after DNA repair, resulting in an A:T base pair conversion to a G:C base pair. However, the ABE system does not require UGI to improve its editing efficiency or product concentration because the process does not involve uracil intermediates.
[0009] The ABE and CBE mentioned above can work effectively in the cell nucleus, but they cannot work in chloroplasts or mitochondria because the sgRNA in the CRISPR system cannot be efficiently transferred to these organelles.
[0010] In 2020, researchers developed a non-CRISPR base editor system composed solely of protein components. This novel base editor system was named DdCBE (Mok et al., 2020, Nature 583: 631-637). The core component of DdCBE includes the double-stranded DNA cytosine deaminase DddA, which can convert C to U on double-stranded DNA without requiring CRISPR-Cas9 to create single-stranded DNA. However, intact DddA is cytotoxic, so it is split into two halves—DddA-N and DddA-C—each fused to a TALE protein pair. DddA-N and DddA-C are guided by the TALE pairs to the target DNA sequence, where they recombine to restore their cytosine deaminase activity. Similar to CRISPR-based CBE systems, this system can also convert C:G base pairs to T:A base pairs. Adding UGI can improve the base editing efficiency and product purity of DdCBE. Due to the full-protein component characteristics of the DdCBE system, it can function both in the cell nucleus and be transported to chloroplasts and mitochondria to achieve targeted cytosine base editing of chloroplast DNA and mitochondrial DNA.
[0011] However, because DddA toxin is a cytosine deaminase, it can only act on cytosine bases in the CBE system and not on the adenine bases required for the ABE system, severely limiting its application. In 2022, researchers fused the artificially directed-evolutionary adenine deaminase TadA-8e with DdCBE to generate the TALED system, which enables A-to-G base editing (Cho et al., 2022, Cell 185: 1764-1776). In the TALED system, the adenine deaminase TadA-8e is fused with one of the split DddA bases, and this combination successfully induces simultaneous C-to-T and A-to-G base conversions in mitochondrial DNA. Furthermore, TadA-8e-mediated A-to-G base editing remains effective even when the deaminase activity of DddA is inactivated.
[0012] While the DdCBE and TALED systems extend the application of base editing to mitochondrial DNA and / or chloroplast DNA, several limitations remain. First, due to the inherent double-stranded DNA cytosine deaminase activity of DddA, cytosine in the deamination window on both strands undergoes deamination, meaning deamination cannot occur only on the selected single strand, thus lacking safety and precision, and thus unsafe for use. Second, compared to CBE and ABE-mediated base editing in the nucleus, DddA-based base editing products contain a relatively high indel frequency, resulting in lower product purity. Third, it has been reported that DddA-based mitochondrial base editors induce widespread off-target mutations in the nucleus during mitochondrial base editing (Lei et al., 2022, Nature 606: 804-811). Notably, most off-target mutations are independent of TALED and are caused by DddA itself. A large number of nuclear off-target mutations can significantly adversely affect the safety of using these base editors.
[0013] Therefore, there is an urgent need in this field to develop novel base editors that are single-strand specific, can function on cell nucleus and mitochondrial DNA and / or chloroplast DNA, and produce products with high purity. Summary of the Invention
[0014] To address the aforementioned technical problems, this application provides a novel base editor that does not rely on CRISPR technology. This system is single-strand specific, can function on nuclear, mitochondrial, or chloroplast DNA, and can produce high-purity edited products.
[0015] Specifically, the present invention provides a novel nucleic acid base editor protein composition, a recombinant expression construct encoding a novel synthetic nucleic acid base editor protein, a genetically engineered cell comprising one or more recombinant expression constructs encoding the novel synthetic nucleic acid base editor protein, and a method for using the above-mentioned novel nucleic acid base editor protein, recombinant expression construct and genetically engineered cell.
[0016] The nucleic acid base editor of the present invention comprises: a sequence-specific DNA-binding protein; a nicking enzyme; an exonuclease; and a base-specific deaminase. In some embodiments, the nucleic acid base editor further comprises a uracil glycosylation inhibitor. In a particular embodiment, the sequence-specific DNA-binding protein, the nicking enzyme, the exonuclease, and the base-specific deaminase form one or more fusion proteins. In an advantageous embodiment of the nucleic acid base editor provided by the present invention, the sequence-specific DNA-binding protein is selected from TALE protein, ZFA protein, Cas protein, or a wide range of nucleases. And in some specific embodiments, the sequence-specific DNA-binding protein is preferably a TALE protein. In a specific embodiment of the nucleic acid base editor of the present invention, the nicking enzyme is a FokI nicking enzyme. The deaminase in the nucleic acid base editor of the present invention is selected from cytosine-specific deaminases or adenine-specific deaminases. In an advantageous embodiment of the nucleic acid base editor of the present invention that includes a cytosine-specific deaminase, the cytosine deaminase is selected from hAPOBEC3A, rAPOBEC1, hAID, pmCDA1, or Sdd deaminase. In an advantageous embodiment of the nucleic acid base editor of the present invention, which includes an adenine-specific deaminase, the adenine deaminase is TadA-8e.
[0017] In another preferred embodiment, the compositions provided by the present invention comprise recombinant expression constructs encoding one or more sequence-specific DNA-binding proteins, nickases, exonucleases, and base-specific deaminases, wherein each of the sequence-specific DNA-binding protein, nickase, exonuclease, and base-specific deaminase is expressible in cells. In some embodiments, these nucleic acid compositions further comprise recombinant expression constructs encoding uracil glycosylase inhibitors. In particular embodiments, the composition comprises recombinant expression constructs encoding one or more sequence-specific DNA-binding proteins, nickases, exonucleases, and base-specific deaminases as fusion proteins, wherein the fusion proteins contained herein are expressible in cells. In advantageous embodiments of the nucleic acid base editor provided herein, the sequence-specific DNA-binding protein is selected from TALE proteins, ZFA proteins, Cas proteins, or a wide range of nucleases, and in certain specific embodiments, the sequence-specific DNA-binding protein is a TALE protein. In a specific embodiment of the nucleic acid base editor of the present invention, the nickase is a FokI nickase. The deaminase in the nucleic acid base editor of the present invention is selected from cytosine-specific deaminases or adenine-specific deaminases, preferably from the deaminases shown in SEQ ID NO. 36-59 and 80-86. In the above-described advantageous embodiments of the nucleic acid base editor containing a cytosine-specific deaminase, the cytosine deaminase is selected from hAPOBEC3A, rAPOBEC1, hAID, pmCDA1, or Sdd deaminase. In embodiments of the nucleic acid base editor of the present invention containing an adenine-specific deaminase, the adenine deaminase is TadA-8e.
[0018] In another preferred embodiment, the present invention also provides a recombinant cell comprising one or more recombinant expression constructs encoding a sequence-specific DNA-binding protein, a nickase, an exonuclease, and a base-specific deaminase; wherein each of the sequence-specific DNA-binding protein, nickase, exonuclease, and base-specific deaminase is expressible in the cell. In some embodiments, these recombinant cells contain a nucleic acid composition further comprising a recombinant expression construct encoding a uracil glycosylation inhibitor. In a specific embodiment, the recombinant cell comprises one or more recombinant expression constructs encoding a sequence-specific DNA-binding protein, a nickase, an exonuclease, and a base-specific deaminase as a fusion protein, wherein the fusion protein is expressible in the cell. In an advantageous embodiment of the recombinant cell provided herein, the sequence-specific DNA-binding protein is selected from TALE protein, ZFA protein, Cas protein, or a wide range of nucleases, and in certain specific embodiments, the sequence-specific DNA-binding protein is TALE protein. In a specific embodiment of the recombinant cell provided herein, the nickase is selected as FokI. Further, the present invention provides recombinant cells comprising one or more recombinant expression constructs encoding deaminases, wherein the deaminases are cytosine-specific deaminases or adenine-specific deaminases, preferably from the deaminases shown in SEQ ID NO. 36-59, 80-86. Advantageous embodiments of the recombinant cells provided herein comprise one or more recombinant expression constructs encoding cytosine-specific deaminases, wherein in an advantageous embodiment the cytosine deaminase is selected from hAPOBEC3A, rAPOBEC1, hAID, pmCDA1, or Sdd deaminase. In another advantageous embodiment, the recombinant cells comprise one or more recombinant expression constructs encoding adenine-specific deaminases, wherein, in a non-limiting example, the adenine deaminase is selected as TadA-8e.
[0019] In another preferred embodiment, the present invention also provides a method for base editing in cells, comprising the step of introducing a nucleic acid base editor, or a recombinant expression construct encoding the nucleic acid base editor of the present invention, or a fusion protein encoding the nucleic acid base editor of the present invention, into the cell. In practice of the methods described herein, base editing is performed on target nucleic acids recognized by specifically binding proteins and results in alterations to cytosine or adenine residues.
[0020] In another preferred embodiment, the present invention provides a nucleic acid base editor that is specific for base editing activity in the cell nucleus or organelles. Further, the nucleic acid base editor targeting the cell nucleus may include a nuclear localization signal (NLS). Further, the base editor targeting mitochondria or chloroplasts may respectively include a mitochondrial targeting sequence (MTS) or a chloroplast transport peptide (CTP). In these embodiments, NLS, MTS, or CTP may be substituted for each other depending on the specific target organelle or base editor, as will be described in further detail herein.
[0021] The following are exemplary technical solutions of the present invention:
[0022] The first objective of this invention is to provide a nucleic acid base editor, which includes the following components: a) a sequence-specific DNA-binding protein; b) a nicking enzyme; c) an exonuclease; and d) a base-specific deaminase.
[0023] Preferably, the components of the nucleic acid base editor exist individually or form one or more fusion proteins.
[0024] Preferably, the sequence-specific DNA-binding protein is one or more of TALE protein, ZFA protein, Cas protein, and a wide range of nucleases.
[0025] Preferably, the sequence-specific DNA-binding protein is the TALE protein.
[0026] Preferably, the cleavage enzyme is a dimer of FokI's cleavage domain monomer FokICD or a mutant thereof, wherein the FokICD dimer or mutant thereof consists of a pair of interacting FokI cleavage domain monomers, and only one FokICD monomer in the FokICD dimer or mutant thereof has DNA endonuclease activity.
[0027] Preferably, the cleavage domain monomer of FokI is isolated from a mutant of wild-type FokI protein, wherein the mutant of wild-type FokI protein has a mutation at position 450 and / or position 467, or has an amino acid sequence that is at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the cleavage domain monomer of FokI.
[0028] More preferably, the mutation causes the FokICD monomer to lose its DNA endonuclease activity.
[0029] Preferably, the FokI cleavage domain monomer FokICD is isolated from a mutant of the wild-type FokI protein, the mutation preventing the FokICD monomer from self-polymerizing with FokICD monomers containing the same site mutation to form a dimer.
[0030] More preferably, the FoKICD monomer sequence is selected from SEQ ID No. 87-88.
[0031] Preferably, the amino acid sequence of the FokI cleavage functional domain monomer FokICD is selected from SEQ No. 60-63.
[0032] Preferably, the base-specific deaminase is selected from cytosine-specific deaminase or adenine-specific deaminase.
[0033] More preferably, the base deaminase is selected from SEQ ID NO. 36-59, 80-86.
[0034] More preferably, the base-specific deaminase is a cytosine-specific deaminase.
[0035] More preferably, the cytosine-specific deaminase is one or more of hAPOBEC3A, rAPOBEC1, hAID, pmCDA1, or Sdd deaminase.
[0036] Furthermore, the nucleic acid base editor also contains:
[0037] e) Uracil glycosylation inhibitor (UGI);
[0038] Furthermore, the uracil glycosylation enzyme inhibitor exists alone or forms at least one fusion protein with other nucleic acid base editor components.
[0039] Preferably, the base-specific deaminase is an adenine-specific deaminase.
[0040] Preferably, the adenine-specific deaminase is TadA-8e.
[0041] Furthermore, the nucleic acid base editor also contains:
[0042] f) γb;
[0043] The γb, together with other nucleic acid base editor components, forms at least one fusion protein.
[0044] A second objective of the present invention is to provide a fusion protein as a nucleic acid base editor, the fusion protein containing the protein domain of the base editor as described in the first objective.
[0045] Another object of the present invention is to provide a fusion protein as a nucleic acid base editor, the fusion protein being arranged in a linear order starting from the amino terminus of a protein, comprising: an exonuclease, an XTEN spacer peptide, a base-specific deaminase, an XTEN spacer peptide, a uracil glycosylation inhibitor (UGI), and a nuclear localization signal.
[0046] Another object of the present invention is to provide a fusion protein as a nucleic acid base editor, the fusion protein being arranged in a linear order starting from the amino terminus of a protein, comprising: an exonuclease, a 48-amino acid spacer peptide, a base-specific deaminase, an XTEN spacer peptide, a uracil glycosylation inhibitor (UGI), and a nuclear localization signal.
[0047] Another object of the present invention is to provide a composition of a fusion protein having nucleic acid base editor activity, said composition comprising:
[0048] The first fusion protein comprises: a nuclear localization signal (NLS), a sequence-specific DNA-binding protein, and a base-specific deaminase;
[0049] The second fusion protein comprises: an exonuclease and a nuclear localization signal (NLS); and,
[0050] The third fusion protein comprises: a uracil glycosylase inhibitor (UGI) and a nuclear localization signal (NLS).
[0051] Another object of the present invention is to provide a composition of a fusion protein having nucleic acid base editor activity, said composition comprising:
[0052] The first fusion protein, arranged in a linear sequence starting from the amino terminus of the protein, includes: nuclear localization signal (NLS), base-specific deaminase, TALE-L protein, and FokI-L. D450A Protein, T2A sequence, NLS, TALE-R protein, and FokI-R protein;
[0053] The second fusion protein comprises: an exonuclease and a nuclear localization signal (NLS); and,
[0054] The third fusion protein comprises: a uracil glycosylase inhibitor (UGI) and a nuclear localization signal (NLS).
[0055] Another object of the present invention is to provide a fusion protein composition having nucleic acid base editor activity, said composition comprising:
[0056] The first fusion protein, arranged in a linear sequence starting from the amino terminus of the protein, includes: nuclear localization signal (NLS), TALE-L protein, and FokI-L.D450A Protein, T2A sequence, NLS, base-specific deaminase, 48-amino acid spacer peptide, TALE-R protein, and FokI-R protein;
[0057] The second fusion protein comprises: an exonuclease and a nuclear localization signal (NLS); and,
[0058] The third fusion protein comprises: a uracil glycosylase inhibitor (UGI) and a nuclear localization signal (NLS).
[0059] Another object of the present invention is to provide a composition of a fusion protein having nucleic acid base editor activity, said composition comprising:
[0060] The first fusion protein comprises: a nuclear localization signal (NLS), a sequence-specific DNA-binding protein, a base-specific deaminase, and a uracil glycosylation inhibitor (UGI); and,
[0061] The second fusion protein contains an exonuclease and a nuclear localization signal (NLS).
[0062] Another object of the present invention is to provide a fusion protein composition having nucleic acid base editor activity, said composition comprising:
[0063] The first fusion protein, arranged linearly from the amino terminus, comprises: a nuclear localization signal (NLS), a base-specific deaminase, a 48-amino acid spacer peptide, TALE-L protein, and FokI-L. D450A Protein, T2A sequence, NLS, TALE-R protein, FokI-R protein, 4-amino acid spacer peptide, and uracil glycosylation inhibitor (UGI); and,
[0064] The second fusion protein contains an exonuclease and a nuclear localization signal (NLS).
[0065] Another object of the present invention is to provide a fusion protein composition having nucleic acid base editor activity, said composition comprising:
[0066] The first fusion protein, arranged linearly from the amino terminus, includes: a nuclear localization signal (NLS), a uracil glycosylation inhibitor (UGI), and a 4-amino acid spacer peptide.
[0067] Base-specific deaminase, 48-amino acid spacer peptide, TALE-L protein, FokI-L D450A Protein, T2A sequence, NLS, TALE-R protein, FokI-R protein; and,
[0068] The second fusion protein contains an exonuclease and a nuclear localization signal (NLS).
[0069] Another object of the present invention is to provide a composition of a fusion protein having nucleic acid base editing activity and capable of base editing on mitochondria, said composition comprising:
[0070] The first fusion protein, arranged in a linear sequence starting from the amino terminus of the protein, comprises: the mitochondrial targeting sequence (MTS), the TALE-L protein, and the FokI-L protein. D450A protein;
[0071] The second fusion protein, arranged in a linear sequence starting from the amino terminus of the protein, comprises: the mitochondrial targeting sequence (MTS), the TALE-R protein, and the FokI-R protein;
[0072] The third fusion protein, arranged in a linear sequence starting from the amino terminus of the protein, contains: a mitochondrial targeting sequence (MTS) and an exonuclease;
[0073] The fourth fusion protein, arranged linearly from the amino terminus of the protein, comprises: a mitochondrial targeting sequence (MTS) and a base-specific deaminase; and,
[0074] The fifth fusion protein, arranged in a linear sequence starting from the amino terminus of the protein, contains: a mitochondrial targeting sequence (MTS) and a uracil glycosylation inhibitor (UGI).
[0075] Another object of the present invention is to provide a composition of a fusion protein having nucleic acid base editing activity and capable of base editing on mitochondria, said composition comprising:
[0076] The first fusion protein, arranged in a linear sequence starting from the amino terminus of the protein, comprises: the mitochondrial targeting sequence (MTS), the TALE-L protein, and the FokI-L protein. D450A protein;
[0077] The second fusion protein, arranged in a linear sequence starting from the amino terminus of the protein, comprises: the mitochondrial targeting sequence (MTS), the TALE-R protein, and the FokI-R protein;
[0078] The third fusion protein, arranged in a linear sequence starting from the amino terminus of the protein, contains: mitochondrial targeting sequence (MTS), γb, and exonuclease;
[0079] The fourth fusion protein is arranged in a linear sequence starting from the amino terminus of the protein, and it includes: a mitochondrial targeting sequence (MTS) and a base-specific deaminase; and,
[0080] The fifth fusion protein, arranged in a linear sequence starting from the amino terminus of the protein, contains: mitochondrial targeting sequence (MTS), γb, and uracil glycosylation inhibitor (UGI).
[0081] Another object of the present invention is to provide a composition of a fusion protein having nucleic acid base editor activity, said composition comprising:
[0082] The first fusion protein comprises: nuclear localization signal (NLS) / chloroplast transport peptide (CTP) / mitochondrial targeting sequence (MTS), sequence-specific DNA-binding protein, and cleavage enzyme;
[0083] The second fusion protein comprises: an exonuclease and a nuclear localization signal (NLS) / chloroplast transport peptide (CTP) / mitochondrial targeting sequence (MTS); and,
[0084] The third fusion protein comprises: a base-specific deaminase, a uracil glycosylation inhibitor (UGI), and a nuclear localization signal (NLS) / chloroplast transport peptide (CTP) / mitochondrial targeting sequence (MTS).
[0085] Another object of the present invention is to provide a composition of a fusion protein having nucleic acid base editing activity, said composition comprising:
[0086] The first fusion protein, arranged linearly from the amino terminus, contains: nuclear localization signal (NLS), chloroplast transport peptide (CTP), mitochondrial targeting sequence (MTS), TALE-L protein, and FokI-L. D450A Protein, T2A sequence, NLS, TALE-R protein, FokI-R protein, or containing: nuclear localization signal (NLS) / chloroplast transport peptide (CTP) / mitochondrial targeting sequence (MTS), TALE-L protein, FokI-L protein, T2A sequence, NLS, TALE-R protein, FokI-R D450A protein;
[0087] The second fusion protein, arranged linearly from the amino terminus of the protein, comprises: a nuclear localization signal (NLS) / chloroplast transport peptide (CTP) / mitochondrial targeting sequence (MTS) and an exonuclease; and,
[0088] The third fusion protein, arranged in a linear sequence starting from the amino terminus of the protein, contains: nuclear localization signal (NLS) / chloroplast transport peptide (CTP) / mitochondrial targeting sequence (MTS), base-specific deaminase, XTEN spacer peptide, and uracil glycosylation inhibitor (UGI).
[0089] Another object of the present invention is to provide a composition of a fusion protein having nucleic acid base editing activity and capable of base editing on mitochondria, characterized in that the composition comprises:
[0090] The first fusion protein, arranged linearly from the amino terminus, comprises: a mitochondrial targeting sequence (MTS), a base-specific deaminase, a 48-amino acid spacer peptide, TALE-L protein, and FokI-L. D450A Protein, 11-amino acid spacer peptide, and uracil glycosylation inhibitor (UGI); and,
[0091] The second fusion protein, arranged in a linear sequence starting from the amino terminus of the protein, comprises: a mitochondrial targeting sequence (MTS), a 48-amino acid spacer peptide, TALE-R protein, a uracil glycosylation inhibitor (UGI), a 14-amino acid spacer peptide, and FokI-R protein.
[0092] Another object of the present invention is to provide a recombinant expression construct for nucleic acid base editing, the recombinant expression construct being used to express the nucleic acid base editor described in the first object above or the fusion protein or composition described in the other objects above.
[0093] Another object of the present invention is to provide a genetically engineered cell for transforming the recombinant expression construct as described above.
[0094] Another objective of this invention is to provide a method for nucleic acid base editing in cells, by introducing the nucleic acid base editor or recombinant expression construct described above into cells to edit the target gene.
[0095] Preferably, the target gene is selected from nuclear genomic DNA, mitochondrial genomic DNA, or chloroplast genomic DNA.
[0096] More preferably, the target gene is nuclear genomic DNA, and the nucleic acid base editor further includes a nuclear localization signal (NLS).
[0097] More preferably, the target gene is mitochondrial genomic DNA, and the nucleic acid base editor further includes a mitochondrial targeting sequence (MTS).
[0098] More preferably, the target gene is chloroplast genomic DNA, and the nucleic acid base editor further includes chloroplast transport peptides (CTPs).
[0099] Another objective of this invention is to fuse γb to the ends of each component.
[0100] Further preferred, γb is fused with UGI and Trex2 respectively.
[0101] Another objective of this invention is to provide an application of base editing technology in base editing, using the base editor, fusion protein, composition, recombinant expression construct, genetically engineered cell or method described above to perform base editing on DNA in cells, wherein the cells are mammalian cells, bacteria, protists, fungi, insect cells, yeast, unconventional yeast or plant cells.
[0102] Preferably, the plant cells are derived from the whole plant, seedling, meristem, ground tissue, vascular tissue, cortical tissue, seed, leaf, root, bud, stem, flower, fruit, stolon, bulb, tuber, corm, asexual terminal branch, bud, young bud, or tumor tissue of monocotyledonous or dicotyledonous plants.
[0103] Preferably, the mammalian cells are selected from human neurons, muscle cells, endocrine / exocrine cells, epithelial cells, muscle cells, tumor cells, hematopoietic cells, bone cells, somatic cells, and induced pluripotent stem cells.
[0104] Preferably, the editor is used to perform base editing on the nuclear genome or organelle genome.
[0105] Preferably, the organelle is a mitochondrial or a chloroplast.
[0106] Another object of the present invention is to provide the use of the base editor, fusion protein, composition, recombinant expression construct, and genetically engineered cell described above in the preparation of a pharmaceutical composition for treating diseases in subjects in need.
[0107] Another object of the present invention is to provide a pharmaceutical composition for treating diseases in subjects of need, said pharmaceutical composition comprising the base editor, fusion protein, composition, recombinant expression construct, genetically engineered cell, and, optionally, pharmaceutically acceptable vector described in the above objects.
[0108] Another object of the present invention is to provide a method for producing genetically modified plants, characterized in that the method includes introducing the base editor, fusion protein, composition, recombinant expression construct, or genetically engineered cell described in the above objects into at least one of the plants.
[0109] The base editor and its applications provided by this invention have the following beneficial effects:
[0110] (1) The base editor of the present invention only performs base editing on selected single strands, and has good safety and accuracy.
[0111] (2) The base editor of the present invention has high purity of the edited product and low insertion / deletion byproduct rate, and has excellent editing efficiency.
[0112] (3) The base editor of the present invention has a low off-target rate, which effectively improves its therapeutic effect and safety.
[0113] (4) The base editor of the present invention is not based on CRISPR technology and has a wider range of applications and scenarios. Its components can all play a role in the cell nucleus or organelles such as mitochondria and chloroplasts. Attached Figure Description
[0114] To better understand the technical solutions described in this invention, the following description is provided in conjunction with the accompanying drawings.
[0115] Figure 1 This is a functional schematic diagram of the nucleic acid base editor of the present invention. In the first step, a sequence-specific DNA-binding protein (SSDBP) locates and binds to the target DNA sequence. In the second step, a cleavage enzyme preferentially cleaves one DNA strand at the target site, and then an exonuclease digests the cleaved DNA strand from the cleavage site into the SSDBP binding site. This exposes an ssDNA fragment in the complementary strand, which then becomes a substrate for deaminase to achieve deamination. Following DNA repair, this leads to a base conversion (C:G to T:A or A:T to G:C, depending on the type of deaminase used).
[0116] Figure 2A and Figure 2B This invention demonstrates the application effect of the nucleic acid base editor in high-purity base editing of rice cell nuclei. Figure 2A To investigate the C-to-T base editing efficiency of the OsBADH2 target in rice protoplasts under different treatment methods. Figure 2B To investigate the C>T base editing efficiency and the frequency of indel generation byproducts at the OsBADH2 target site in rice protoplasts under different treatment methods.
[0117] Figure 3A and Figure 3B This paper analyzes the base editing window of the base editor of the present invention. After transforming rice protoplasts with the nucleic acid base editor described in this invention, DNA was extracted, and high-throughput sequencing was performed on the target sites to obtain the editing efficiency of different bases on the target sequence. Figure 3A This is a schematic diagram of the OsBADH2 target sequence. The gray sequences on both sides are TALE binding sites, and the black area in the middle is the spacer sequence. Figure 3B To analyze the base editing window of the base editor based on high-throughput sequencing results. CK represents a blank control without any plasmid transformation, and TALEN... WTand TALEN WT +ExoI represents either transformation of wild-type TALEN alone or a combination of TALEN + exonuclease ExoI, with these two treatments serving as negative controls.
[0118] Figure 4A and Figure 4B The efficiency of cytosine nucleotide editing at the target site was analyzed by high-throughput sequencing after the base editor of the present invention was transformed into rice protoplasts targeting OsDEP1. Figure 4A and the frequency of insertion / deletion byproducts ( ) and the occurrence frequency of insertion / deletion byproducts ( Figure 4B ), where CK is the blank control without any plasmid transformation, TALEN WT and TALEN WT +ExoI represents either transformation of wild-type TALEN alone or a combination of TALEN + exonuclease ExoI, with these two treatments serving as negative controls.
[0119] Figure 5A and Figure 5B The implementation effects of a base editor using different combinations of FokI nickases, different exonucleases, and cytosine deaminases are shown. Using exonucleases with different digestion directions produces different editing windows; using different nickases allows for specific base editing of different DNA single strands at the target site. Figure 5A And the analysis of the purity of the editing products and the frequency of byproducts of the base editor described in this invention under different combinations. Figure 5B ).
[0120] Figure 6A and Figure 6B The display shows the base editing efficiency and indel frequency detected by high-throughput sequencing of the base editor of this invention, which includes a combination of cytosine deaminase and exonuclease, introducing a target sequence (OsBADH2 in rice protoplasts). The exonuclease is either a 5' or 3' exonuclease.
[0121] Figure 7A and Figure 7B This invention demonstrates the efficiency and editing window of the base editing system, which incorporates a combination of different pyrimidine deaminases and exonucleases, to introduce a target sequence (OsBADH2 in rice protoplasts), as detected by high-throughput sequencing.
[0122] Figure 8 The base editing efficiency, as detected by high-throughput sequencing, is shown by introducing the target sequence (OsCKX2 in rice protoplasts) using the base editor containing adenine deaminase of the present invention.
[0123] Figure 9This is a schematic diagram of the base editor of the present invention, which comprises a fusion protein of an exonuclease, a deaminase, a uracil DNA glycosylase inhibitor, and a nuclear localization signal (NLS) separated by an XTEN spacer peptide or a 48-amino acid spacer peptide.
[0124] Figure 10A and Figure 10B The base editors of different constructs of this invention introduce the target sequence (OsDEP1 in rice protoplasts), and the base editing efficiency obtained by high-throughput sequencing is shown in Figure 10A, as well as the editing windows of different base editors (Figure 10B).
[0125] Figure 11A and Figure 11B This is a schematic diagram of a base editor for a vector containing the deaminase-TALE fusion protein of the present invention. In each embodiment, the NLS-exonuclease fusion protein and the NLS-uracil glycosylation inhibitor (UGI) are provided in separate vectors.
[0126] Figure 12A and Figure 12B This is a bar chart showing the base editing rate and indel rate of the target sequences (OsDEP1 in rice protoplasts, Figure 12A; OsCKX2 in rice protoplasts, Figure 12B) introduced by the base editor of the fusion protein of this invention. Deaminase-TALE-FokI-R nickase The results of the protein fusion protein are shown in Figure 12A Deaminase-TALE-FokI-L nickase The results of the protein fusion protein are shown in Figure 12B As shown.
[0127] Figure 13A and Figure 13B This is a schematic diagram of the base editor containing the deaminase-TALE fusion protein of the present invention. In each embodiment, the fusion protein of NLS and exonuclease is provided in a separate vector.
[0128] Figure 14 Using the base editor of this invention, such as Figure 13A and Figure 13B The fusion protein and the individual expression of each component show the base editing efficiency of the target sequence (OsDEP1 in rice protoplasts).
[0129] Figure 15A This is a schematic diagram of the vector used in mitochondrial editing by the base editor of the present invention, which includes expression of MTS-deaminase, MTS-UGI, MTS-TALE-R-FokI-R (or MTS-TALE-R-FokI-R). D450A MTS-TALE-L-FokI-LD450A (or MTS-TALE-L-FokI-L) nickase and MTS-exonuclease constructs.
[0130] Figure 15B This invention uses the base editor. Figure 15A The schematic diagram of the target sequence of the construct shown illustrates the TALE-R and TALE-L binding sites and cytosine residues targeted by certain nucleic acid base editors of the present invention, namely the mitochondrial ND6 target sequence and TALE binding site.
[0131] Figure 15C This is the use of the present invention Figure 15A The base editor of the constructed shown demonstrates the base mutation efficiency of the target sequence.
[0132] Figure 16A- Figure 16E This is a representative illustration of the recombinant expression construct of the base editor used in the embodiments described herein, in rice. Figures 16A-16E In this context, FokK-L-nickase is equivalent to FOKI-L; FokI-R is equivalent to FOKI-R (D450A / D467A).
[0133] Figure 16A This refers to the recombinant expression construct (NLS-TALEN) encoding wild-type TALEN used in Example 2 and other examples. WT The vector diagram (using TALE targeting OsBADH2 as an example) illustrates how this vector can induce double-strand breaks in the target DNA and randomly trigger insertion and deletion (indel) mutations, serving as a control group in various embodiments. In this construct, a stable expression T-DNA vector with a UBI promoter and Nos terminator derived from maize is used to drive the expression of wild-type TALENs (including TALE-L-FokI-L and TALE-R-FokI-R fusion proteins, where FokI does not contain D450A or D467A mutations). The N and C-terminal regions of TALE contain corresponding truncations (ΔN152 / C63) located flanking the DNA-binding domain of TALE. The TALE-L-FokI-L and TALE-R-FokI-R fusion proteins are linked via a T2A self-cleaving peptide. Other components shown in the figure include the CaMV 35S promoter (a promoter derived from cauliflower mosaic virus), the hygroscopic enzyme resistance selection gene Hyg, and the Nos terminator of Agrobacterium tumefaciens carmine synthase.
[0134] Figure 16BThis diagram illustrates recombinant expression constructs containing the sequence-specific DNA-binding proteins TALE-L and TALE-R, and the nicking enzyme FokI (i.e., a schematic diagram of a vector partially containing the nicking enzyme, exonuclease, and deaminase, using TALE targeting OsBADH2 as an example; different TALE coding sequences can be designed depending on the target sequence), as well as two additional constructs: NLS-deaminase-UGI and exonuclease-NLS. Each of these constructs contains a UBI promoter from maize and a Nos terminator, driving the expression of the deaminase-UGI fusion protein and the exonuclease, respectively. UGI, a uracil-DNA glycosylation inhibitor derived from Bacillus subtilis phage, protects uracil in DNA by irreversibly inhibiting the key DNA repair enzyme uracil-DNA glycosylation. Other components shown in the figure include the CaMV 35S promoter (a promoter derived from cauliflower mosaic virus), the hygroscopic enzyme resistance selection gene Hyg, the Nos terminator of Agrobacterium tumefaciens, and the CaMV polyadenylate signal terminator.
[0135] Figure 16C This diagram illustrates a recombinant expression construct containing the sequence-specific DNA-binding proteins TALE-L and TALE-R, a cleavage enzyme (FokI cleavage enzyme), and a deaminase fusion protein (i.e., a schematic diagram of a vector partially containing the cleavage enzyme, exonuclease, deaminase, and uracil glycosylation inhibitor, using TALE targeting OsBADH2 as an example; different TALE coding sequences can be designed according to different target sequences), as well as two additional constructs: UGI-NLS and exonuclease-NLS. The recombinant expression constructs UGI-NLS and exonuclease-NLS each possess a UBI promoter and a CaMV terminator to drive the expression of UGI and the exonuclease, respectively. UGI, a uracil-DNA glycosylation inhibitor derived from Bacillus subtilis phage, protects uracil in DNA by irreversibly inhibiting the key DNA repair enzyme uracil-DNA glycosylation. Other components shown in the figure include the CaMV 35S promoter (a promoter derived from cauliflower mosaic virus), the hygroscopic enzyme resistance selection gene Hyg, the Nos terminator of Agrobacterium tumefaciens, and the CaMV polyadenylate signal terminator.
[0136] Figure 16D is a diagram of a recombinant expression construct containing the sequence-specific DNA-binding proteins TALE-L and TALE-R, the nicking enzyme FokI, a deaminase, and UGI (i.e., partially containing NLS-deaminase-TALE-L-FokI- nickaseA schematic diagram of the vector for TALEN-R-UGI, exonuclease-NLS (using TALE targeting OsBADH2 as an example; different TALE coding sequences can be designed depending on the target sequence), and an additional construct: exonuclease-NLS. The recombinant expression construct exonuclease-NLS has a UBI promoter and a CaMV terminator to drive the expression of the exonuclease. UGI, a uracil-DNA glycosylase inhibitor derived from Bacillus subtilis phage, protects uracil in DNA by irreversibly inhibiting the key DNA repair enzyme uracil-DNA glycosylase. Other components shown in the diagram include the CaMV 35S promoter (a promoter derived from cauliflower mosaic virus), the hygroscopic resistance selection gene Hyg, the nos terminator of Agrobacterium tumefaciens carmine synthase, and the CaMV polyadenylate signal terminator.
[0137] Figure 16E This is a diagram of a recombinant expression construct containing sequence-specific DNA-binding proteins TALE-L and TALE-R, the nicking enzyme FokI, a deaminase, an exonuclease, and a UGI fusion protein (NLS-deaminase-TALE-L-FokI-). nickase - Schematic diagram of the TALEN-R-UGI-exonuclease vector (taking TALE targeting OsBADH2 as an example; depending on the target sequence, a corresponding TALE coding sequence can be designed), which has the additional feature of encoding UGI and exonuclease in the construct rather than introducing cells into a separate construct.
[0138] Figures 17A to 17H This is a representative illustration of a recombinant expression construct encoding a base editor applied to human cell mitochondrial editing in the embodiments described herein.
[0139] Figure 17AThis is a representative mitochondrial recombinant expression construct, MTS-TALE-L-FokI-L (illustrated diagram of the MTS-TALE-L-FokI-L vector targeting mitochondrial ND6). The TALE sequence can be replaced according to different targets. The expression vector MTS-TALE-L-FokI-L has a CMV promoter and a bGH poly(A) signal terminator to drive the expression of the MTS-TALE-L-FokI-L fusion protein. The N and C terminal regions of TALE contain corresponding truncations (ΔN152 / C63) and are located flanking the DNA-binding domain of TALE (see Mok et al., 2020, Nature 583:631-637). MTS is a mitochondrial targeting sequence derived from human superoxide dismutase 2, which facilitates protein transport into mitochondria. The CMV promoter is a human herpesvirus 5-derived promoter and has been shown to be highly active in animal cells. CMV enhancers are fragments containing the cytomegalovirus promoter region that enhance the transcriptional efficiency of the CMV promoter. The bGH poly(A) signal is a terminator derived from the growth hormone polyadenylation signal.
[0140] Figure 17B This is a representative mitochondrial recombinant expression construct, MTS-TALE-R-FokI-R (illustrated diagram of the MTS-TALE-R-FokI-R vector targeting mitochondrial ND6). The TALE sequence can be replaced according to different targets. The expression vector MTS-TALE-R-FokI-R has a CMV promoter and a bGH poly(A) signal terminator to drive the expression of the MTS-TALE-R-FokI-R fusion protein. The N and C-terminal regions of TALE contain corresponding truncations (ΔN152 / C63) and are located flanking the DNA-binding domain of TALE (see Mok et al., 2020, Nature 583:631-637). In this vector, MTS is the mitochondrial targeting sequence of cytochrome c oxidase subunit 8, which facilitates protein transport into mitochondria. The CMV promoter is a human herpesvirus 5-derived promoter and has been shown to be highly active in animal cells. CMV enhancers are fragments containing the cytomegalovirus promoter region that enhance the transcriptional efficiency of the CMV promoter. The bGH poly(A) signal is a terminator derived from the growth hormone polyadenylation signal.
[0141] Figure 17C is a schematic diagram of the mitochondrial recombinant expression construct MTS-deaminase (MTS-deaminase vector schematic diagram). This recombinant expression construct has a CMV promoter and a bGH poly(A) signal terminator, which can drive the expression of MTS-deaminase in human mitochondria. The MTS, CMV promoter, CMV enhancer, and bGH poly(A) signal terminator are shown in Figure 17A.
[0142] Figure 17D shows a representative mitochondrial recombinant expression construct of MTS-exonuclease (schematic diagram of the MTS-exonuclease vector). This recombinant expression construct possesses a CMV promoter and a bGH poly(A) signal terminator, which can drive the expression of MTS-exonuclease in human mitochondria. The MTS, CMV promoter, CMV enhancer, and bGH poly(A) signal terminator are shown below. Figure 17A As described in [the text].
[0143] Figure 17E This is a representative example of the mitochondrial recombinant expression construct MTS-UGI (Schematic diagram of the MTS-UGI vector). This recombinant expression construct possesses a CMV promoter and a bGH poly(A) signal terminator, which can drive the expression of MTS-UGI (a uracil glycosylase inhibitor derived from Bacillus subtilis phage) in human mitochondria. The MTS, CMV promoter, CMV enhancer, and bGH poly(A) signal terminator are shown in Figure 17A.
[0144] Figure 17F is a schematic diagram of the mitochondrial recombinant expression construct MTS-deaminase-TALE-L-FokI-L (schematic diagram of the MTS-deaminase-TALE-L-FokI-L vector). The recombinant expression construct MTS-deaminase-TALE-L-FokI-L possesses a CMV promoter and a bGH poly(A) signal terminator to drive the expression of the MTS-deaminase-TALE-L fusion protein. Components such as MTS, CMV promoter, CMV enhancer, and bGH poly(A) signal terminator are shown in Figure 17A.
[0145] Figure 17G is a schematic diagram of the mitochondrial recombinant expression construct MTS-exonuclease-TALE-R-FokI-R (schematic diagram of the MTS-exonuclease-TALE-R-FokI-R vector). The recombinant expression construct MTS-exonuclease-TALE-R-FokI-R possesses a CMV promoter and a bGH poly(A) signal terminator to drive the expression of the MTS-exonuclease-TALE-R fusion protein. Components such as MTS, CMV promoter, CMV enhancer, and bGH poly(A) signal terminator are shown in Figure 17B.
[0146] Figure 17H is a schematic diagram of the mitochondrial recombinant expression construct MTS-UGI-exonuclease-TALE-R-FokI-R (schematic diagram of the MTS-UGI-exonuclease-TALE-R-FokI-R vector). The recombinant expression construct MTS-UGI-exonuclease-TALE-R-FokI-R possesses a CMV promoter and a bGH poly(A) signal terminator to drive the expression of the MTS-exonuclease-TALE-R fusion protein. Components such as MTS, CMV promoter, CMV enhancer, and bGH poly(A) signal terminator are shown below. Figure 17B As described in [the text].
[0147] Figure 18 This is a schematic diagram of the CyDENT structure used for nuclear genome editing.
[0148] Figure 19A The C-to-T conversion frequencies and indel frequencies of nuCyDENT-R and TALEN at the OsDEP1, OsSD1, OsCKX2, and OsBADH2 sites in rice protoplasts are measured.
[0149] Figure 19B This is the base editing window of CyDENT at the OsDEP1, OsSD1, OsCKX2, and OsBADH2 sites in rice protoplasts. The gray area in the figure represents the TALE binding site, and the middle area is the spacer region.
[0150] Figure 20 This represents CyDENT base editing at the OsCKX2 and OsSD1 sites in rice protoplasts. The gray area represents the TALE binding site.
[0151] Figure 21 This represents CyDENT base editing at the human SIRT6 site. The gray area represents the TALE binding site.
[0152] Figure 22AThis is a schematic overview of the modular CyDENT construct used in chloroplast genome editing, using cpCyDENT-R as an example.
[0153] Figure 22B CyDENT base editing window at the OsrbcL site in rice protoplasts. The gray area represents the TALE binding site.
[0154] Figure 23A This is a schematic diagram of the modular CyDENT structure used in mitochondria. The example is mtCyDENT-R.
[0155] Figure 23B It refers to the base editing of the mitochondrial ND6 site in HEK293T cells by mtCyDENT-L or mtCyDENT-R in multiple γb fusion states.
[0156] Figure 24 The editing frequencies of DdCBE, mtCyDENT-R, mtCyDENT1b-R, mtCyDENT-L, and mtCyDENT1b-L at the mitochondrial ND1.2, ND1.3, ND3, and ND6.2 sites in HEK293T cells.
[0157] Figure 25 The indel frequencies of DdCBE, mtCyDENT1b-R, and mtCyDENT1b-L at different sites in the mitochondria of HEK293T cells.
[0158] Figure 26 These are the base editing sites of mtCyDENT at different sites in the mitochondria of HEK293T cells. The gray areas are TALE binding sites.
[0159] Figure 27 The editing frequencies at ND5.1, ND6, and ND1.3 sites in HEK293T cells using Sdd7 deaminases mtCyDENT1b-L and mtCyDENT1b-R are shown.
[0160] Figure 28A This is a schematic diagram of the mtCyDENT2 construct in the mitochondrial genome.
[0161] Figure 28B The efficiency of DdCBE and mtCyDENT2-L and mtCyDENT2-R containing different deaminases in base editing at the ND6 site in HEK293T cells, as well as the proportion of various editing events.
[0162] Figure 29This represents the editing frequency and editing chain preference of DdCBE and mtCyDENT2-L containing Sdd3 deaminase at ND1.2 and ND6.2 sites in HEK293T cells. Gray indicates TALE binding sites.
[0163] Figure 30 The mtCyDENT2-L (Sdd3 deaminase + TALE-L1 + TALE-R1) edited strand preference at the ND6.2 site in HEK293T cells is designed to target the pathogenic mutation at the ND6.2 site in Leigh's syndrome.
[0164] Figure 31A This is a WGS and NGS analysis of the target editing frequencies of ND3 and ND6.2.
[0165] Figure 31B These are the logo images for off-target C:G to T:A and G:C to A:T base conversions in various editors.
[0166] Figure 31C This represents the frequency distribution of SNVs and indels in potential TALE-dependent off-target sites. Detailed Implementation
[0167] the term:
[0168] Unless otherwise defined, all technical terms used herein have the same meaning as commonly understood by those skilled in the art.
[0169] A numerical range includes the numbers that define the range, and explicitly includes every integer and non-integer fraction within the defined range. Unless otherwise stated, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art.
[0170] As used in this invention, the terms "structure," "recombinant expression structure," or "recombinant expression construct" refer to an artificially designed DNA fragment that can be used to introduce genetic material into target cells (e.g., using a recombinant expression structure to produce a base editor or a component thereof). The term "expression" refers to the transcription and translation of a nucleic acid coding sequence to produce a coding polypeptide.
[0171] As used in this invention, the term "genetic engineering" refers to altering the genetic composition of cells through biotechnology, including intra-species and inter-species gene transfer, to produce modified or non-naturally occurring cells. In its specific application, the construct encodes a base editor or a component thereof, and the genetically engineered cell produces the base editor. Cells containing exogenous, recombinant, synthetic, and / or other modified polynucleotides are considered genetically engineered cells, and therefore are non-naturally occurring relative to any naturally occurring counterpart. In some cases, genetically engineered cells contain one or more recombinant nucleic acids. In other cases, genetically engineered cells contain one or more synthetic or genetically engineered nucleic acids (e.g., a nucleic acid containing at least one artificially inserted, deleted, reversed, or substituted sequence relative to its naturally occurring counterpart). Methods for producing genetically engineered cells are known in the art, for example, Sambrook et al., Molecular Cloning, A Laboratory Manual (Fourth Edition), Cold Spring Harbor Press, Cold Spring Harbor, NY (2012).
[0172] As used in this invention, the terms "genetically engineered cell," "genetically engineered host cell," or "recombinant expression host cell" can refer to cells modified using gene editing techniques. Gene editing refers to a form of genetic engineering that inserts, deletes, modifies, or replaces DNA in the genome of a living cell. Compared to other genetic engineering techniques that can randomly insert genetic material into the host genome, gene editing can target the insertion to a specific location (e.g., the AAVS1 allele). Examples of gene editing techniques include, but are not limited to, restriction endonucleases, zinc finger nucleases, TALENs, and CRISPR-Cas9. The base editor disclosed herein is a specific example of gene editing that allows for changes with one or more single nucleotides, particularly resulting in alterations to the cell phenotype.
[0173] As used herein, the terms "deaminase," "base-specific deaminase," or "deaminase domain" refer to proteins or enzymes that catalyze deamination reactions. In this invention, "deaminase" and "base-specific deaminase" are used interchangeably. In some embodiments, the deaminase or deaminase domain is a cytosine deaminase that catalyzes the hydrolysis and deamination of cytosine or deoxycytosine to uracil, which is ultimately converted to thymine (T) during cell modification and DNA replication. In some embodiments, the deaminase or deaminase domain is an adenine deaminase domain that catalyzes the hydrolysis and deamination of adenine or deoxyadenine to hypoxanthine or deoxyhypoxanthine (I), which is ultimately converted to guanine or deoxyguanine nucleotides (G) during cell modification and DNA replication. In some embodiments, the deaminase or deaminase domain is a naturally occurring deaminase derived from an organism, such as a microorganism, plant, or animal, such as a human, chimpanzee, gorilla, monkey, cattle, dog, rat, or mouse. In some embodiments, the deaminase or deaminase domain is a variant of a naturally occurring deaminase from an organism that does not exist in nature. For example, in some embodiments, the deaminase or deaminase domain is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a naturally occurring deaminase from an organism.
[0174] As used herein, the term "spacer peptide" or "linker" refers to an element that connects two molecules or parts, such as two domains of a fusion protein. In some embodiments, the spacer peptide is an organic molecule, group, polymer, or chemical moiety. In some embodiments, it is a spacer peptide of 5-100 amino acids in length, for example, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acids. Longer or shorter spacer peptides are also considered.
[0175] As used herein, the term "mutation" refers to the substitution of a residue within a sequence (e.g., a nucleic acid or amino acid sequence) by another residue, or the deletion or insertion of one or more residues within the sequence. Mutations are generally described in this invention by identifying the initial residue, followed by the position of the residue within the sequence and the identity of the newly substituted residue. Various methods for generating the amino acid substitutions (mutations) provided herein are well known in the art and are provided, for example, by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012)).
[0176] As used in this invention, the term "uracil glycosylation inhibitor" or "UGI" refers to a protein that can inhibit uracil-DNA glycosylation enzyme as a base cleavage repair enzyme.
[0177] The terms "top strand" or "A strand" and "bottom strand" or "B strand" used in this invention are merely to distinguish the relative positions of the two strands of the DNA target site in a particular embodiment, so as to exemplarily describe the editing effect of the base editor of this invention on single-stranded DNA, and have no special limitation on specific DNA double-stranded structures. "Top strand" can be interchanged with "A strand," and "bottom strand" can be interchanged with "B strand." Unless otherwise specified, the schematic diagrams of this application (…) Figure 1 The "top strand" or "A strand" of TALE-L is a single strand of DNA that interacts with TALE-L, while the "bottom strand" or "B strand" is a single strand of DNA that interacts with TALE-R.
[0178] Various embodiments of the compositions and methods according to the present invention are now described in the following non-limiting examples. These examples are for illustrative purposes only and do not limit the scope of the invention in any way.
[0179] Nucleic acid base editor
[0180] The base editing function of the nucleic acid base editor of this invention is as follows: Figure 1 As shown. Its components include a sequence-specific DNA-binding protein (SSDBP), a nicking enzyme, an exonuclease (with 5' or 3' exonuclease activity), a cytosine or adenine deaminase, and optionally a uracil glycosylation inhibitor (UGI), as well as an optional positioning sequence. These components can be expressed in individual constructs or fused into one or more constructs using appropriate spacer peptides.
[0181] Sequence-specific DNA-binding proteins
[0182] In the base editors disclosed herein, SSDBP can be a TALE protein, a zinc finger protein (ZFA protein), a CRISPR-Cas endonuclease (Cas protein), or a wide range of nucleases, with TALE protein selected in some specific embodiments. The TALE protein (Transcription activator-like effector) is a transcription activator effector derived from Xanthomonas spp., artificially engineered as a sequence-specific DNA-binding protein. The TALE protein contains 1 to 33 repeating units (or repeat cells) of 33–35 amino acid residues in length, each repeating unit and the terminal half-repeat unit specifically recognizing and binding to a specific nucleotide target site. The type of DNA bases that TALE can recognize and bind in each repeat sequence is determined by the targeting of specific base pairs by two highly variable residues at positions 12 and 13 (called repeat variable double residues (RVD)). The DNA types or codes that RVDs recognize have been deciphered: RVDs His / Asp (HD), Asn / Gly (NG), Asn / Asn (NN), and Asn / Ile (NI) recognize cytosine (C), thymine (T), guanine (G), and adenine (A), respectively (see, Boch & Bonas, 2010, Annu. Rev. Phytopathol. 48: 419-436; Deng et al., 2012, Cell Res. 22: 1502-1504). TALE repeat units are modular, allowing for the artificial design of RVDs to target and bind to DNA. As disclosed in this invention, a pair of TALE proteins (referred to as TALE-L or TALE-L protein, and TALE-R or TALE-R protein, respectively) bind to DNA at two adjacent sites, where the DNA sequence between adjacent sites is the spacer sequence, also known as the target sequence; the binding sites of TALE-L and TALE-R are defined as the left binding site and the right binding site. The sequence specificity of the TALE proteins is used to determine the target sites in the base editor disclosed in this invention. Furthermore, in some cases, only one TALE (instead of a pair) is needed to bind to and target dsDNA to achieve the base editing function of this invention.
[0183] The following provides exemplary structures of TALE proteins that can be used as components of the base editor disclosed in this invention, including but not limited to an N-terminus as shown in SEQ ID NO. 1, a C-terminus as shown in SEQ ID NO. 2, and repeating units as shown in SEQ ID NO. 3-35.
[0184] TALE-NTD (Δ152) :
[0185] MVDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVTYQHIITALPEATHEDIVGVGKQWSGARALEALLTDAGELRGPPLQLDTGQLVKIAKRGGVTAMEAVHASRNALTGAPLN (SEQ ID NO. 1)
[0186] TALE-CTD (C63) :
[0187] SIVAQLSRPDPALAALTNDHLVALACLGGRPAMDAVKKGLPHAPELIRRVNRRIGERTSHRVA (SEQID NO. 2)
[0188] OsBADH2-TALE-Left repeat :
[0189] LTPDQVVAIASNNGGKQALETVQRLLPVLCQDHG (SEQ ID NO. 3)
[0190] LTPDQVVAIASHDGGKQALETVQRLLPVLCQDHG (SEQ ID NO. 4)
[0191] LTPDQVVAIASNGGGKQALETVQRLLPVLCQDHG (SEQ ID NO. 5)
[0192] LTPDQVVAIASNNGGKQALETVQRLLPVLCQDHG (SEQ ID NO. 6)
[0193] LTPDQVVAIASNNGGKQALETVQRLLPVLCQDHG (SEQ ID NO. 7)
[0194] LTPDQVVAIASNIGGKQALETVQRLLPVLCQDHG (SEQ ID NO. 8)
[0195] LTPDQVVAIASNGGGKQALETVQRLLPVLCQDHG (SEQ ID NO. 9)
[0196] LTPDQVVAIASNNGGKQALETVQRLLPVLCQDHG (SEQ ID NO. 10)
[0197] LTPDQVVAIASHDGGKQALETVQRLLPVLCQDHG (SEQ ID NO. 11)
[0198] LTPDQVVAIASNGGGKQALETVQRLLPVLCQDHG (SEQ ID NO. 12)
[0199] LTPDQVVAIASNGGGKQALETVQRLLPVLCQDHG (SEQ ID NO. 13)
[0200] LTPDQVVAIASNGGGKQALETVQRLLPVLCQDHG (SEQ ID NO. 14)
[0201] LTPDQVVAIASNNGGKQALETVQRLLPVLCQDHG (SEQ ID NO. 15)
[0202] LTPDQVVAIASNIGGKQALETVQRLLPVLCQDHG (SEQ ID NO. 16)
[0203] LTPDQVVAIASNNGGKQALETVQRLLPVLCQDHG (SEQ ID NO. 17)
[0204] LTPDQVVAIASNGGGKQALETVQRLLPVLCQDHG (SEQ ID NO. 18)
[0205] LTPDQVVAIASNIGGKQALE (SEQ ID NO.19)
[0206] OsBADH2-TALE-Right repeat :
[0207] LTPDQVVAIASNNGGKQALETVQRLLPVLCQDHG (SEQ ID NO.20)
[0208] LTPDQVVAIASHDGGKQALETVQRLLPVLCQDHG (SEQ ID NO.21)
[0209] LTPDQVVAIASHDGGKQALETVQRLLPVLCQDHG (SEQ ID NO.22)
[0210] LTPDQVVAIASNGGGKQALETVQRLLPVLCQDHG (SEQ ID NO.23)
[0211] LTPDQVVAIASNGGGKQALETVQRLLPVLCQDHG (SEQ ID NO.24)
[0212] LTPDQVVAIASNGGGKQALETVQRLLPVLCQDHG (SEQ ID NO.25)
[0213] LTPDQVVAIASNGGGKQALETVQRLLPVLCQDHG (SEQ ID NO.26)
[0214] LTPDQVVAIASNNGGKQALETVQRLLPVLCQDHG (SEQ ID NO.27)
[0215] LTPDQVVAIASNGGGKQALETVQRLLPVLCQDHG (SEQ ID NO.28)
[0216] LTPDQVVAIASHDGGKQALETVQRLLPVLCQDHG (SEQ ID NO.29)
[0217] LTPDQVVAIASHDGGKQALETVQRLLPVLCQDHG (SEQ ID NO.30)
[0218] LTPDQVVAIASNIGGKQALETVQRLLPVLCQDHG (SEQ ID NO.31)
[0219] LTPDQVVAIASNIGGKQALETVQRLLPVLCQDHG (SEQ ID NO.32)
[0220] LTPDQVVAIASNNGGKQALETVQRLLPVLCQDHG (SEQ ID NO.33)
[0221] LTPDQVVAIASNNGGKQALETVQRLLPVLCQDHG (SEQ ID NO.34)
[0222] LTPDQVVAIASNIGGKQALETVQRLLPVLCQDHG (SEQ ID NO.35)
[0223] Cutting enzyme
[0224] The nicking enzymes used as components of the base editor disclosed herein are capable of cleaving one of the double strands of the target DNA. In the base editor disclosed herein, exemplary nicking enzymes are FokI from Flavobacterium okeanokoites, or the FokI protein, specifically amino acid sequence variants whose dsDNA cleaving activity is converted to create a gap only on one strand of the target DNA, including but not limited to the D450A / D467A mutant. Furthermore, alternative nicking enzymes comprising bacterial IIS-type restriction enzymes may also be used as components of the base editor disclosed herein.
[0225] Wild-type FokI consists of two functional domains: a recognition domain and a cleavage domain. The recognition domain is artificially removed, resulting in FokICD that retains only the cleavage domain. When two FokICD monomers interact to form a dimer, the cleavage activity of FokICD is activated, enabling the cleavage of both strands of double-stranded DNA. The following provide FokICD monomers that can be used as examples of this invention, including but not limited to those shown in SEQ ID NO. 87-88:
[0226] FokI-L:
[0227] QLVKSELEEKKSELRHKLKYVPHEYIELIEIARNSTQDRILEMKVMEFFMKVYGYRGKHLGGSRKPDGAIYTVGSPIDYGVIVDTKAYSGGYNLPIGQADEMQRYVKENQTRNKHINPNEWWKVYPSSVTEFKFLFVSGHFKGNYKAQLTRLNHKTNCNGAVLSVEELLIGGEMIKAGTLTLEEVRRKFNNGEINF (SEQ ID NO.87)
[0228] FokI-R:
[0229] QLVKSELEEKKSELRHKLKYVPHEYIELIEIARNSTQDRILEMKVMEFFMKVYGYRGKHLGGSRKPDGAIYTVGSPIDYGVIVDTKAYSGGYNLPIGQADEMERYVEENQTRNKHLNPNEWWKVYPSSVTEFKFLFVSGHFKGNYKAQLTRLNHITNCNGAVLSVEELLIGGEMIKAGTLTLEEVRRKFNNGEINF (SEQ ID NO.88)
[0230] When an aspartic acid at position 450 (or position 67 if the first amino acid of the wild-type FokI containing the recognition domain is used as the first amino acid, or position 84 if the first amino acid of the FokI containing the cleavage domain is used as the first amino acid) and / or position 467 (or position 84 if the first amino acid of the FokI containing the recognition domain is used as the first amino acid) is mutated to alanine (D450A or D467A) in one of the FokICD monomers of the dimer, that FokICD monomer loses its cleavage activity, while the other FokICD monomer in the dimer, which has not undergone the amino acid mutation, retains its cleavage activity. The resulting FokICD dimer can cleave only one strand of double-stranded DNA and cannot cleave the other strand. This type of FokICD dimer is called FokI. nickase This refers to the FokI cleavage enzyme. For ease of description, the inventors refer to the FokICD monomer fused with TALE-L as FokI-L (e.g., shown in SEQ ID NO. 87), and the FokICD monomer fused with TALE-R as FokI-R (e.g., shown in SEQ ID NO. 88). Further, FokICD mutant monomers containing FokI D450A and / or D467A mutations that result in loss of cleavage activity are respectively referred to as FokI-L. D450A / D467A and FokI-R D450A / D467A In this invention, FokI-L and FokI-R D450A / D467A The FokICD dimer formed by the interaction retains only the cleavage activity of FokI-L; this dimer is called FokI-L. nickase (or FokI-L cleavage enzyme); correspondingly, FokI-L D450A / D467A The FokICD dimer formed by the interaction with FokI-R retains only the cleavage activity of FokI-R and is called FokI-R. nickase (or FokI-R cleavage enzyme).
[0231] It should be noted that FokI-L nickase and FokI-R nickase It tends to cleave different single strands in double-stranded DNA, i.e., FokI-L nickase and FokI-R nickase It exhibits single-strand specificity or preference when cutting DNA. For example... Figure 1 As shown, at this target point, the FokI-R used nickase It tends to cut the B chain; correspondingly, if FokI-L is used... nickase Then they tend to cut chain A ( Figure 1(As shown). FokI-L nickase and FokI-R nickase The exhibited strand specificity facilitates the selection of the desired DNA single strand for subsequent deamination steps. Along with the sequence-specific binding of TALE-L and TALE-R to the left and right binding sites, FokI-L... nickase Or FokI-R nickase The target sequence is cut, leaving a gap in either strand A or strand B. The strand specificity of the cleavage enzyme determines that the DNA single strand is further deaminated under the action of the base editor of this invention.
[0232] The following are nicking enzyme protein monomers that can be used as exemplary components of the nucleic acid base editor of this invention, including but not limited to those shown in SEQ ID NO. 60-63:
[0233] FokI-L D450A :
[0234] QLVKSELEEKKSELRHKLKYVPHEYIELIEIARNSTQDRILEMKVMEFFMKVYGYRGKHLGGSRKPAGAIYTVGSPIDYGVIVDTKAYSGGYNLPIGQADEMQRYVKENQTRNKHINPNEWWKVYPSSVTEFKFLFVSGHFKGNYKAQLTRLNHKTNCNGAVLSVEELLIGGEMIKAGTLTLEEVRRKFNNGEINF (SEQ ID NO.60)
[0235] FokI-L D467A :
[0236] QLVKSELEEKKSELRHKLKYVPHEYIELIEIARNSTQDRILEMKVMEFFMKVYGYRGKHLGGSRKPDGAIYTVGSPIDYGVIVATKAYSGGYNLPIGQADEMQRYVKENQTRNKHINPNEWWKVYPSSVTEFKFLFVSGHFKGNYKAQLTRLNHKTNCNGAVLSVEELLIGGEMIKAGTLTLEEVRRKFNNGEINF (SEQ ID NO.61)
[0237] FokI-R D450A :
[0238] QLVKSELEEKKSELRHKLKYVPHEYIELIEIARNSTQDRILEMKVMEFFMKVYGYRGKHLGGSRKPAGAIYTVGSPIDYGVIVDTKAYSGGYNLPIGQADEMERYVEENQTRNKHLNPNEWWKVYPSSVTEFKFLFVSGHFKGNYKAQLTRLNHITNCNGAVLSVEELLIGGEMIKAGTLTLEEVRRKFNNGEINF (SEQ ID NO.62)
[0239] FokI-R D467A :
[0240] QLVKSELEEKKSELRHKLKYVPHEYIELIEIARNSTQDRILEMKVMEFFMKVYGYRGKHLGGSRKPDGAIYTVGSPIDYGVIVATKAYSGGYNLPIGQADEMERYVEENQTRNKHLNPNEWWKVYPSSVTEFKFLFVSGHFKGNYKAQLTRLNHITNCNGAVLSVEELLIGGEMIKAGTLTLEEVRRKFNNGEINF (SEQ ID NO.63)
[0241] Nucleotide exonuclease
[0242] Depending on the type of exonuclease used, the exonuclease component of the nucleic acid base editor of this invention digests the cleaved DNA strand from the nick position along the 5'→3' direction or along the 3'→5' direction. After exonuclease digestion, the short ssDNA fragment is exposed on the complementary DNA strand. The type of exonuclease determines the deaminated ssDNA region (or editing window). Exonucleases that may be used as components of the basic editor disclosed herein include, but are not limited to, DNA polymerases I and III (E. coli), mammalian p53 protein, exonucleases I-VII (E. coli) (such as exonucleases I and V (with 3'→5' exonuclease activity)), phage-derived polymerases (such as T4 DNA polymerase (with 3'→5' exonuclease activity)), thermophilic aquatic polymerases (with 5'→3' exonuclease activity), and 3'→5' exonucleases reported by Shevelev and Hübscher (Shevelev & Hübscher, 2002, Nat. Rev. Molec. Cell Biol. 3: 364-376).
[0243] The following provides exonuclease proteins that can be used as exemplary base editor components of the present invention, including but not limited to proteins as shown in sequences SEQ ID NO. 64-67, 153.
[0244] Exonuclease V (ExoV):
[0245] MAETGEEETASAEASGFSDLSDSELVEFLDLEEAKESAVSLSKPGPSAELPGKDDKPVSLQNWKGGLDVLSPMERFHLKYLYVTDLCTQNWCELQMVYGKELPGSLTPEKAAVLDTGASIHLAKELELHDLVTVPIATKEDAWAVKFLNILAMIPALQSEGRVREFPVFGEVEGIFLVGVIDELHYTS KGELELAELKTRRRPVLPLPAQKKKDYFQVSLYKYIFDAMVQGKVTPASLIHHTKLCLDKPLGPSVLRHARQGGVSVKSLGDLMELVFLSLTLSDLPAIDTLKLEYIHQETATILGTEIVAFEEKEVKSKVQHYVAYWMGHRDPQGVDVEEAWKCRTCDYVDICEWRRGSGVLSSSWEPKAKKFK(SEQ ID NO. 153)
[0246] mExoI:
[0247] MGIQGLLQFIQEASEPVNVKKYKGQAVAVDTYCWLHKGAIACAEKLAKGEPTDRYVGFCMKFVNMLLSYGVKPILIFDGCTLPSKKEVERSRRERRQSNLLKGKQLLREGKVSEARDCFARSINITHAMAHKVIKAARALGVDCLVAPYEADAQLAYLNKAGIVQAVITEDSDLLAFGCKKVILKMDQFGNGLEVDQARLGMCKQLGDVFTEEKFRYMCILSGCDYLASLRGIGLAKACKVLRLANNPDIVKVIKKIGHYLRMNITVPEDYITGFIRANNTFLYQLVFDPIQRKLVPLNAYGDDVNPETLTYAGQYVGDSVALQIALGNRDVNTFEQIDDYSPDTMPAHSRSHSWNEKAGQKPPGTNSIWHKNYCPRLEVNSVSHAPQLKEKPSTLGLKQVISTKGLNLPRKSCVLKRPRNEALAEDDLLSQYSSVSKKIKENGCGDGTSPNSSKMSKSCPDSGTAHKTDAHTPSKMRNKFATFLQRRNEESGAVVVPGTRSRFFCSSQDFDNFIPKKESGQPLNETVATGKATTSLLGALDCPDTEGHKPVDANGTHNLSSQIPGNAAVSPEDEAQSSETSKLLGAMSPPSLGTLRSCFSWSGTLREFSRTPSPSASTTLQQFRRKSDPPACLPEASAVVTDRCDSKSEMLGETSQPLHELGCSSRSQESMDSSCGLNTSSLSQPSSRDSGSEESDCNNKSLDNQGEQNSKQHLPHFSKKDGLRRNKVPGLCRSSSMDSFSTTKIKPLVPARVSGLSKKSGSMQTRKHHDVENKPGLQTKISELWKNFGFKKDSEKLPSCKKPLSPVKDNIQLTPETEDEIFNKPECVRAQRAIFH (SEQ ID NO.64)
[0248] mTrex2:
[0249] MSEPPRAETFVFLDLEATGLPNMDPEIAEISLFAVHRSSLENPERDDSGSLVLPRVLDKLTLCMCPERPFTAKASEITGLSSESLMHCGKAGFNGAVVRTLQGFLSRQEGPICLVAHNGFDYDFPLLCTELQRLGAHLPQDTVCLDTLPALRGLDRAHSHGTRAQGRKSYSLASLFHRYFQAEPSAAHSAEGDVHTLLLIFLHRAPELLAWADEQARSWAHIEPMYVPPDGPSLEA (SEQ ID NO.65)
[0250] mArtimes:
[0251] MSSGMAYTSDRDRNKARAYSHCHKDHMKGRASKRRCSKVYCSVTKTSKYRWNRTTTSVDASGKVVVTAGHCGSVMGSNGTVYTGDRAKGASRMHSGGRVKDSVYDTTCDRYSRCRGVRSWVTRSHHVVWNCKAAYGYYTNSGVVHVDKDMKNMDHHTTDRNTHACRHKACWNKCGTSNKTAHTSKSTMWGRTRKTNVVRTGSSYRACSHSSSKDSYCVNVYNVVGTVDKVMDVKCRSSVKYKGKKRARTHDSDDDDDTRHKVYTSMKADRSGGCKASVWSSANDCSNSDSGTSGGGSTVNADDVDWVKRRDTGCHSSTGGSSKCSDSKCSDSKCSDSDGDSTHSSNSSSTHTDGSGWDSCDTVSSKSGGDSTSNKGAYKKKSSASDACDTHCDKSRAVNGACVDTSGRKSKTSSTRADSSSSDSTATHCYRKATGSVVKRKCSDS(SEQ ID NO.66)
[0252] T5 exo:
[0253] MSKSWGKFIEEEEAEMASRRNLMIVDGTNLGFRFKHNNSKKPFASSYVSTIQSLAKSYSARTTIVLGDKGKSVFRLEHLPEYKGNRDEKYAQRTEEEKALDEQFFEYLKDAFELCKTTFPTFTIRGVEADDMAAYIVKLIGHTYD HVWLISTDGDWDTLLTDKVSRFSFTTRREYHLRDMYEHHNVDDVEQFISLKAIMGDLGDNIRGVEGIGAKRGYNIIREFGNVLDIIDQLPLPGKQKYIQNLNASEELLFRNLILVDLPTYCVDAIAAVGQDVLDKFTKDILEIAEQ (SEQ ID NO.67)
[0254] Deaminase
[0255] Deaminases that can be used as components of the base editor of this invention include cytosine deaminases and adenine deaminases. Cytosine deaminases include, but are not limited to, hAPOBEC3A (Zong et al., 2018, Nat. Biotechnol. Oct 1. doi: 10.1038 / nbt.4261), rAPOBEC1, C57, and Sdd (Huang J et al., 2023, Cell, doi:10.1101 / 2023.05.21.541555), which result in a C-to-T base transition. Optional adenine deaminases include TadA-8e (Richter et al., 2020, Nat. Biotechnol. 38: 883-891), which produces an A-to-G base transition.
[0256] The following are deaminases that can be used as exemplary base editor components of the present invention, including but not limited to the deaminases listed in Table 1 (such as proteins shown in SEQ ID NO. 36-59, 80-86):
[0257] Table 1 Types of deaminases
[0258]
[0259] deaminase rAPOBEC1:
[0260] MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSIWRHTSQNTNKHVEVNFIEKFTTERYFCPNTRCSITWFLSWSPCGECSRAITEFLSRYPHVTLFIYIARLYHHADPRNRQGLRDLISSGVTIQIMTEQESGYCWRNFVNYSPSNEAHWPRYPHLWVRLYVLELYCIILGLPPCLNILRRKQPQLTFFTIALQSCHYQRLPPHILWATGLK (SEQ ID NO. 36)
[0261] hAPOBEC3A:
[0262] MEASPASGPRHLMDPHIFTSNFNNGIGRHKTYLCYEVERLDNGTSVKMDQHRGFLHNQAKNLLCGFYGRHAELRFLDLVPSLQLDPAQIYRVTWFISWSPCFSWGCAGEVRAFLQENTHVRLRIFAARIYDYDPLYKEALQMLRDAGAQVSIMTYDEFKHCWDTFVDHQGCPFQPWDGLDEHSQALSGRLRAILQNQGN (SEQ ID NO. 37)
[0263] hAPOBEC3G-CTD:
[0264] MDPPTFTFNFNNEPWVRGRHETYLCYEVERMHNDTWVLLNQRRGFLCNQAPHKHGFLEGRHAELCFLDVIPFWKLDLDQDYRVTCFTSWSPCFSCAQEMAKFISKNKHVSLCIFTARIYDDQGRCQEGLRTLAEAGAKISIMTYSEFKHCWDTFVDHQGCPFQPWDGLDEHSQDLSGRLRAILQNQEN (SEQ ID NO. 38)
[0265] PmCDA1:
[0266] MTDAEYVRIHEKLDIYTFKKQFFNNKKSVSHRCYVLFELKRRGERRACFWGYAVNKPQSGTERGIHAEIFSIRKVEEYLRDNPGQFTINWYSSWSPCADCAEKILEWYNQELRGNGHTLKIWACKLYYEKNARNQIGLWNLRDNGVGLNVMVSEHYQCCRKIFIQSSHNQLNENRWLEKTLKRAEKRRSELSIMIQVKILHTTKSPAVSRGSG (SEQ IDNO. 39)
[0267] tCDA1EQ:
[0268] SHRCYVLFELKRRGERRACFWGYAVNKPQSGTERGIHAEIFSIRKVEEYLRDNPGQFTINWYSSWSPCADCAEKILEWYNQELRGNGHTLKIEACKLYYEKNARNQIGLQNLRDNGVGLNV (SEQ ID NO. 40)
[0269] hAID:
[0270] MDSLLMNRRKFLYQFKNVRWAKGRRETYLCYVVKRRDSATSFSLDFGYLRNKNGCHVELLFLRYISDWDLDPGRCYRVTWFTSWSPCYDCARHVADFLRGNPNLSLRIFTARLYFCEDRKAEPEGLRRLHRAGVQIAIMTFKDYFYCWNTFVENHERTFKAWEGLHENSVRLSRQLRRILLPLYEVDDLRDAFRTLGL (SEQ ID NO. 41)
[0271] PpAPOBEC1:
[0272] MTSEKGPSTGDPTLRRRIESWEFDVFYDPRELRKETCLLYEIKWGMSRKIWRSSGKNTTNHVEVNFIKKFTSERRFHSSISCSITWFLSWSPCWECSQAIREFLSQHPGVTLVIYVARLFWHMDQRNRQGLRDLVNSGVTIQIMRASEYYHCWRNFVNYPPGDEAHWPQYPPLWMMLYALELHCIILSLPPCLKISRRWQNHLAFFRLHLQNCHYQTIPPHILLATGLIHPSVTWR (SEQ ID NO. 42)
[0273] RrA3F:
[0274] MKPQIRDHRPNPMEAMYPHIFYFHFENLEKAYGRNETWLCFTVEIIKQYLPVPWKKGVFRNQVDPETHCHAEKCFLSWFCNNTLSPKKNYQVTWYTSWSPCPECAGEVAEFLAEHSNVKLTIYTARLYYFWDTDYQEGLRSLSEEGASVEIMDYEDFQYCWENFVYDDGEPFKRWKGLKYNFQSLTRRLREILQ (SEQ ID NO. 43)
[0275] AmAPOBEC1:
[0276] MADSSEKMRGQYISRDTFEKNYKPIDGTKEAHLLCEIKWGKYGKPWLHWCQNQRMNIHAEDYFMNNIFKAKKHPVHCYVTWYLSWSPCADCASKIVKFLEERPYLKLTIYVAQLYYHTEEENRKGLRLLRSKKVIIRVMDISDYNYCWKVFVSNQNGNEDYWPLQFDPWVKENYSRLLDIFWESKCRSPNPW (SEQ ID NO. 44)
[0277] SsAPOBEC3B:
[0278] MDPQRLRQWPGPGPASRGGYGQRPRIRNPEEWFHELSPRTFSFHFRNLRFASGRNRSYICCQVEGKNCFFQGIFQNQVPPDPPCHAELCFLSWFQSWGLSPDEHYYVTWFISWSPCCECAAKVAQFLEENRNVSLSLSAARLYYFWKSESREGLRRLSDLGAQVGIMSFQDFQHCWNNFVHNLGMPFQPWKKLHKNYQRLVTELKQILREEPATYGSPQAQGKVRIGSTAAGLRHSHSHTRSEAHLRPNHSSRQHRILNPPREARARTCVLVDASWICYR (SEQ ID NO. 45)
[0279] hA3B:
[0280] MNPQIRNPMERMYRDTFYDNFENEPILYGRSYTWLCYEVKIKRGRSNLLWDTGVFRGQVYFKPQYHAEMCFLSWFCGNQLPAYKCFQITWFVSWTPCPDCVAKLAEFLSEHPNVTLTISAARLYYYWERDYRRALCRLSQAGARVKIMDYEEFAYCWENFVYNEGQQFMPWYKFDENYAFLHRTLKEILRYLMDPDTFTFNFNNDPLVLRRRQTYLCYEVERLDNGTWVLMDQHMGFLCNEAKNLLCGFYGRHAELRFLDLVPSLQLDPAQIYRVTWFISWSPCFSWGCAGEVRAFLQENTHVRLRIFAARIYDYDPLYKEALQMLRDAGAQVSIMTYDEFEYCWDTFVYRQGCPFQPWDGLEEHSQALSGRLRAILQNQGN (SEQ ID NO. 46)
[0281] hA3C:
[0282] MNPQIRNPMKAMYPGTFYFQFKNLWEANDRNETWLCFTVEGIKRRSVVSWKTGVFRNQVDSETHCHAERCFLSWFCDDILSPNTKYQVTWYTSWSPCPDCAGEVAEFLARHSNVNLTIFTARLYYFQYPCYQEGLRSLSQEGVAVEIMDYEDFKYCWENFVYNDNEPFKPWKGLKTNFRLLKRRLRESLQ (SEQ ID NO. 47)
[0283] hA3D:
[0284] MNPQIRNPMERMYRDTFYDNFENEPILYGRSYTWLCYEVKIKRGRSNLLWDTGVFRGPVLPKRQSNHRQEVYFRFENHAEMCFLSWFCGNRLPANRRFQITWFVSWNPCLPCVVKVTKFLAEHPNVTLTISAARLYYYRDRDWRWVLLRLHKAGARVKIMDYEDFAYCWENFVCNEGQPFMPWYKFDDNYASLHRTLKEILRNPMEAMYPHIFYFHFKNLLKACGRNESWLCFTMEVTKHHSAVFRKRGVFRNQVDPETHCHAERCFLSWFCDDILSPNTNYEVTWYTSWSPCPECAGEVAEFLARHSNVNLTIFTARLCYFWDTDYQEGLCSLSQEGASVKIMGYKDFVSCWKNFVYSDDEPFKPWKGLQTNFRLLKRRLREILQ (SEQ ID NO. 48)
[0285] hA3F:
[0286] MKPHFRNTVERMYRDTFSYNFYNRPILSRRNTVWLCYEVKTKGPSRPRLDAKIFRGQVYSQPEHHAEMCFLSWFCGNQLPAYKCFQITWFVSWTPCPDCVAKLAEFLAEHPNVTLTISAARLYYYWERDYRRALCRLSQAGARVKIMDDEEFAYCWENFVYSEGQPFMPWYKFDDNYAFLHRTLKEILRNPMEAMYPHIFYFHFKNLRKAYGRNESWLCFTMEVVKHHSPVSWKRGVFRNQVDPETHCHAERCFLSWFCDDILSPNTNYEVTWYTSWSPCPECAGEVAEFLARHSNVNLTIFTARLYYFWDTDYQEGLRSLSQEGASVEIMGYKDFKYCWENFVYNDDEPFKPWKGLKYNFLFLDSKLQEILE(SEQ ID NO. 49)
[0287] hA3G:
[0288] MKPHFRNTVERMYRDTFSYNFYNRPILSRRNTVWLCYEVKTKGPSRPPLDAKIFRGQVYSELKYHPEMRFFHWFSKWRKLHRDQEYEVTWYISWSPCTKCTRDMATFLAEDPKVTLTIFVARLYYFWDPDYQEALRSLCQKRDGPRATMKIMNYDEFQHCWSKFVYSQRELFEPWNNLPKYYILLHIMLGEILRHSMDPPTFTFNFNNEPWVRGRHETYLCYEVERMHNDTWVLLNQRRGFLCNQAPHKHGFLEGRHAELCFLDVIPFWKLDLDQDYRVTCFTSWSPCFSCAQEMAKFISKNKHVSLCIFTARIYDDQGRCQEGLRTLAEAGAKISIMTYSEFKHCWDTFVDHQGCPFQPWDGLDEHSQDLSGRLRAILQNQEN (SEQ ID NO. 50)
[0289] hA3H:
[0290] MALLTAETFRLQFNNKRRLRRPYYPRKALLCYQLTPQNGSTPTRGYFENKKKCHAEICFINEIKSMGLDETQCYQVTCYLTWSPCSSCAWELVDFIKAHDHLNLRIFASRLYYHWCKPQQDGLRLLCGSQVPVEVMGFPEFADCWENFVDHEKPLSFNPYKMLEELDKNSRAIKRRLDRIKS (SEQ ID NO. 51)
[0291] hA3Bctd:
[0292] MEILRYLMDPDTFTFNFNNDPLVLRRRQTYLCYEVERLDNGTWVLMDQHMGFLCNEAKNLLCGFYGRHAELRFLDLVPSLQLDPAQIYRVTWFISWSPCFSWGCAGEVRAFLQENTHVRLRIFAARIYDYDPLYKEALQMLRDAGAQVSIMTYDEFEYCWDTFVYRQGCPFQPWDGLEEHSQALSGRLRAILQNQGN (SEQ ID NO. 52)
[0293] FERNY:
[0294] FERNYDPRELRKETYLLYEIKWGKSGKLWRHWCQNNRTQHAEVYFLENIFNARRFNPSTHCSITWYLSWSPCAECSQKIVDFLKEHPNVNLEIYVARLYYHEDERNRQGLRDLVNSGVTIRIMDLPDYNYCWKTFVSDQGGDEDYWPGHFAPWIKQYSLKL (SEQ ID NO. 53)
[0295] ecTadA:
[0296] MSEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRVVFGARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKAQKKAQSSTD (SEQ ID NO. 54)
[0297] mADA:
[0298] MAQTPAFNKPKVELHVHLDGAIKPETILYFGKKRGIALPADTVEELRNIIGMDKPLSLPGFLAKFDYYMPVIAGCREAIKRIAYEFVEMKAKEGVVYVEVRYSPHLLANSKVDPMPWNQTEGDVTPDDVVDLVNQGLQEGEQAFGIKVRSILCCMRHQPSWSLEVLELCKKYNQKTVVAMDLAGDETIEGSSLFPGHVEAYEGAVKNGIHRTVHAGEVGSPEVVREAVDILKTERVGHGYHTIEDEALYNRLLKENMHFEVCPWSSYLTGAWDPKTTHAVVRFKNDKANYSLNTDDPLIFKSTLDTDYQMTKKDMGFTEEEFKRLNINAAKSSFLPEEEKKELLERLYREYQ (SEQ ID NO. 55)
[0299] hADAR2:
[0300] MHLDQTPSRQPIPSEGLQLHLPQVLADAVSRLVLGKFGDLTDNFSSPHARRKVLAGVVMTTGTDVKDAKVISVSTGTKCINGEYMSDRGLALNDCHAEIISRRSLLRFLYTQLELYLNNKDDQKRSIFQKSERGGFRLKENVQFHLYISTSPCGDARIFSPHEPILEEPADRHPNRKARGQLRTKIESGEGTIPVRSNASIQTWDGVLQGERLLTMSCSDKIARWNVVGIQGSLLSIFVEPIYFSSIILGSLYHGDHLSRAMYQRISNIEDLPPLYTLNKPLLSGISNAEARQPGKAPNFSVNWTVGDSAIEVINATTGKDELGRASRLCKHALYCRWMRVHGKVPSHLLRSKITKPNVYHESKLAAKEYQAAKARLFTAFIKAGLGAWVEKPTEQDQFSLTP (SEQ ID NO. 56)
[0301] hADAT2:
[0302] MEAKAAPKPAASGACSVSAEETEKWMEEAMHMAKEALENTEVPVGCLMVYNNEVVGKGRNEVNQTKNATRHAEMVAIDQVLDWCRQSGKSPSEVFEHTVLYVTVEPCIMCAAALRLMKIPLVVYGCQNERFGGCGSVLNIASADLPNTGRPFQCIPGYRAEEAVEMLKTFYKQENPNAPKSKVRKKECQKS (SEQ ID NO. 57)
[0303] ecTadA*(7.10):
[0304] MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD (SEQ ID NO. 58)
[0305] TadA*ABE8e (TadA-8e):
[0306] MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN (SEQ ID NO. 59)
[0307] Sdd2
[0308] MAPDSLVWFDPLGLIVLQQVPYNDHPLFGAVSEFIQGKSRSDLRGRNVAAVLLDDGTVIVRASEGGGNHAERVLMGLSEVDPAKVVAVYTERSPCTGRINCHDLLDSSLGADVPVYYTHEMIRGQEGKTAQQIEADRNQFCRGG(SEQ ID NO. 80)
[0309] Sdd3
[0310] MSASAQLNTYLAAIGNSTTTVEAQPEAAPPPAAAESLDSTPRLPDGGIDFHALAKRLGLLEARPTEQPPFDPRRFNPACWQGLKPYDQAGTAEGNLFIAPGKRWNTRPMQASKLEVGPQSDLHPQWRSRKAPWHIEGKIAAYMRQKGFTDGCVYLNARPCSGPDGCARNLPDLLPVGSTLHVHARYIDRTGETRFYYREYRGTGKALT (SEQ ID NO.81)
[0311] Sdd4
[0312] MLDAMDAYLSEIAGGNAPARAGPKAPEPKQPGGSSSPRARDGRIDFRALLERLKAQGVVGLEGRSDDPIPDFDPKKQNPACYQGLAPRQKGKPVRGNLFFPDGRRWNDVALESSRGEPAFDLNIIKPEYRSLSPARGHLEGNVAAWMRSTFHQEMVLYINESPCRKHGKGCLYTLEHFLPRGYVLHVWSRNDRGEWRGNTFRGSGEAFTEGA (SEQ IDNO. 82)
[0313] Sdd6
[0314] MVETRDKIIAAKSRSDAGLLAFQQATNGSIDSRPAEAIANLQRAKTHLDEAQRLVANSDAAVDNYINAILGGASAATAQPSAVIPASKPSRFKPMRTDPAKADEIRPHVGKDRAVATLWDADGNRVLGLHSADDDGPAATAAWKPPWRDYVRLRRHVEAHAAARMHQDGHKTMVMYINLPPCKYFDGCKLNLEDILPKGSTLWMHRVFQNGGTKIYQFNGTGRAYV (SEQ ID NO. 83)
[0315] Sdd seventh (also denoted as C57 in this specification)
[0316] MLEAVRARLIGEGGGPGAVPEGGDGPPAVPAEEVERLRGELPPPVVPGTGQKTHGRWIGPDGRVRAIVSGRDEDAALVHAQLAAKGIPDEPTRNSDVEQKLAAHMVANGIRHVTLVINHRPCRGFDDSCDTLVPIILPEGCTLTVHGQTDKGMRVRVRYTGGARPWWS (SEQ ID NO. 84)
[0317] Sdd10
[0318] MLDAALGAVRRIIAALGTSGAERASPGANGSERVDELAERLPPTVVPNTSAKTHGWWFTGQGAAQELISGEGPDARAAYEALREEGYPRPGMPFVAMHVEIKLAAHMRRNDIEHATVVINNIPCPLVWGCENLIGVVLPEGSSLTVHGSNGYERTFTGGRKPPWPR (SEQ ID NO. 85)
[0319] Sdd59
[0320] MLLTPPPRPAAPPTTRPKPLVARTGDAYPPGTEWALPLIVQPHPPVGGTVPVEGHVRALRPESQISHVFHPGGGHWTEQARARLRVLPGFGWAVNLGHHVELQIAAWMTACGIHHAELVLNRPCGERYGLGCHQALPVLLPRGYRLTVSSTRGGPQPYQHHYEGKA (SEQ ID NO. 86)
[0321] Uracil glycosylation enzyme inhibitor (UGI)
[0322] In some implementations, when cytosine deaminase is used, a uracil glycosylase inhibitor (UGI) is fused to the N-terminus of the deaminase, while when adenine deaminase is used, UGI is not required.
[0323] The following discloses exemplary UGI proteins that can be used in base editors in this invention, including but not limited to the protein shown in SEQ ID NO. 68:
[0324] TNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML (SEQ ID NO.68)
[0325] Nuclear localization sequence (NLS)
[0326] In some embodiments of the invention, the NLS of the fusion protein of the invention may be located at the N-terminus and / or the C-terminus. In some embodiments of the invention, the NLS of the fusion protein of the invention may be located between the adenine deamination domain, the cytosine deamination domain, the nucleic acid targeting domain, and / or the UGI. In some embodiments, the fusion protein comprises about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLS. In some embodiments, the fusion protein comprises about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLS at or near the N-terminus. In some embodiments, the fusion protein comprises about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLS at or near the C-terminus. In some embodiments, the polypeptide comprises a combination of these, such as one or more NLS at the N-terminus and one or more NLS at the C-terminus. When more than one NLS is present, each may be selected independently of the other NLS.
[0327] Generally, NLS consists of one or more short sequences of positively charged lysine or arginine exposed on the surface of a protein, but other types of NLS are also known. Non-limiting examples of NLS include: KKRKV (SEQ ID NO. 150), PKKKRKV (SEQ ID NO. 151), or KRPAATKKAGQAKKKK (SEQ ID NO. 152).
[0328] Recombinant expression constructs
[0329] The components in the base editor of this invention can be expressed individually, or they can be combined to form one or more fusion proteins for expression, or the above elements or components can be expressed individually or jointly using recombinant expression constructs used in recombinant genetic engineering technology. Exemplary recombinant expression constructs of this invention include, for example... Figures 16A to 16E and Figures 17A to 17H This is explained in the text.
[0330] The following describes the exemplary recombination expression constructs described above ( Figures 16A to 16E and Figures 17A to 17H The types, functions, and references of genes and regulatory elements in the gene sequence are explained and illustrated in Table 2 below:
[0331] Table 2: Examples of construct genes and regulatory elements
[0332]
[0333]
[0334]
[0335]
[0336]
[0337] Specifically, the genes and regulatory elements in the exemplary recombinant constructs used in this invention include, but are not limited to, the following sequences: promoter sequences as shown in SEQ ID NO. 69-72; terminator sequences as shown in SEQ ID NO. 73-76; mitochondrial targeting sequences (MTS) as shown in SEQ ID NO. 77-78; and chloroplast transport peptide (CTP) sequences as shown in SEQ ID NO. 79.
[0338] UBI promoter:
[0339]
[0340] CaMV 35S promoter (enhanced):
[0341] TGAGACTTTTCAACAAAGGGTAATATCGGGAAACCTCCTCGGATTCCATTGCCCAGCTATCTGTCACTTCATCAAAAGGACAGTAGAAAAGGAAGGTGGCACCTACAAATGCCATCATTGCGATAAAGGAAAGGCTATCGTTCAAGATGCCTCTGCCGACAGTGGTCCCAAAGATGGACCCCCACCCACGAGGAGCATCGTGGAAAAAGAAGACGTTCCAACCACGTCTTCAAAGCAAGTGGATTGATGTGATAACATGGTGGAGCACGACACTCTCGTCTACTCCAAGAATATCAAAGATACAGTCTCAGAAGACCAAAGGGCTATTGAGACTTTTCAACAAAGGGTAATATCGGGAAACCTCCTCGGATTCCATTGCCCAGCTATCTGTCACTTCATCAAAAGGACAGTAGAAAAGGAAGGTGGCACCTACAAATGCCATCATTGCGATAAAGGAAAGGCTATCGTTCAAGATGCCTCTGCCGACAGTGGTCCCAAAGATGGACCCCCACCCACGAGGAGCATCGTGGAAAAAGAAGACGTTCCAACCACGTCTTCAAAGCAAGTGGATTGATGTGATATCTCCACTGACGTAAGGGATGACGCACAATCCCACTATCCTTCGCAAGACCTTCCTCTATATAAGGAAGTTCATTTCATTTGGAGAGGACACGCTGA (SEQ ID NO. 70)
[0342] CaMV 2 x 35S promoter
[0343] CCTGCAGGTCAACATGGTGGAGCACGACACACTTGTCTACTCCAAAAATATCAAAGATACAGTCTCAGAAGACCAAAGGGCAATTGAGACTTTTCAACAAAGGGTAATATCCGGAAACCTCCTCGGATTCCATTGCCCAGCTATCTGTCACTTTATTGTGAAGATAGTGGAAAAGGAAGGTGGCTCCTACAAATGCCATCATTGCGATAAAGGAAAGGCCATCGTTGAAGATGCCTCTGCCGACAGTGGTCCCAAAGATGGACCCCCACCCACGAGGAGCATCGTGGAAAAAGAAGACGTTCCAACCACGTCTTCAAAGCAAGTGGATTGATGTGATAACATGGTGGAGCACGACACACTTGTCTACTCCAAAAATATCAAAGATACAGTCTCAGAAGACCAAAGGGCAATTGAGACTTTTCAACAAAGGGTAATATCCGGAAACCTCCTCGGATTCCATTGCCCAGCTATCTGTCACTTTATTGTGAAGATAGTGGAAAAGGAAGGTGGCTCCTACAAATGCCATCATTGCGATAAAGGAAAGGCCATCGTTGAAGATGCCTCTGCCGACAGTGGTCCCAAAGATGGACCCCCACCCACGAGGAGCATCGTGGAAAAAGAAGACGTTCCAACCACGTCTTCAAAGCAAGTGGATTGATGTGATATCTCCACTGACGTAAGGGATGACGCACAATCCCACTATCCTTCGCAAGACCCTTCCTCTATATAAGGAAGTTCATTTCATTTGGAGAGGACCTCGACCTCAACACAACATATACAAAACAAACGAATCTCAAGCAATCAAGCATTCTACTTCTATTGCAGCAATTTAAATCATTTCTTTTAAAGCAAAAGCAATTTTCTGAAAATTTTCACCATTTACGAACGATA (SEQ ID NO.71)
[0344] CMV promoter:
[0345] GTGATGCGGTTTTGGCAGTACATCAATGGGCGTGGATAGCGGTTTGACTCACGGGGATTTCCAAGTCTCCACCCCATTGACGTCAATGGGAGTTTGTTTTGGCACCAAAATCAACGGGACTTTCCAAAATGTCGTAACAACTCCGCCCCATTGACGCAAATGGGCGGTAGGCGTGTACGGTGGGAGGTCTATATAAGCAGAGCT (SEQ ID NO. 72)
[0346] Nos terminator:
[0347] GAATTTCCCCGATCGTTCAAACATTTGGCAATAAAGTTTCTTAAGATTGAATCCTGTTGCCGGTCTTGCGATGATTATCATATAATTTCTGTTGAATTACGTTAAGCATGTAATAATTAACATGTAATGCATGACGTTATTTATGAGATGGGTTTTTATGATTAGAGTCCCGCAATTATACATTTAATACGCGATAGAAAACAAAATATAGCGCGCAAACTAGGATAAATTATCGCGCGCGGTGTCATCTATGTTACT (SEQ ID NO. 73)
[0348] E9 terminator:
[0349] AGAGCTTTCGTTCGTATCATCGGTTTCGACAACGTTCGTCAAGTTCAATGCATCAGTTTCATTGCGCACACACCAGAATCCTACTGAGTTTGAGTATTATGGCATTGGGAAAACTGTTTTTCTTGTACCATTTGTTGTGCTTGTAATTTACTGTGTTTTTTATTCGGTTTTCGCTATCGAACTGTGAAATGGAAATGGATGGAGAAGAGTTAATGAATGATATGGTCCTTTTGTTCATTCTCAAATTAATATTATTTGTTTTTTCTCTTATTTGTTGTGTTGAATTTGAAATTATAAGAGATATGCAAACATTTT GTTTTGAGTAAAAATGTGTCAAATCGTGGCCTCTAATGACCGAAGTTAATATGAGGAGTAAAACACTTGTAGTTGTACCATTATGCTTATTCACTAGGCAACAAATATATTTTCAGACCTAGAAAAGCTGCAAATGTTACTGAATACAAGTATGTCCTCTTGTGTTTTAGACATTTATGAACTTTCCTTTATGTAATTTTCCAGAATCCTTGTCAGATTCTAATCATTGCTTTATAATTATAGTTATACTCATGGATTTGTAGTTGAGTATGAAAATATTTTTTAATGCATTTTATGACTTGCCAATTGATTGACAAC (SEQ ID NO. 74)
[0350] CaMV poly(A) signal:
[0351] TTTCTCCATAATAATGGTGAGTAGTTCCCAGATAAGGGAATTAGGGTTCCTATAGGGTTTCGCTCATGTGTTGAGCATATAAGAAACCCTTAGTATGTATTTGTATTTGTAAAATACTTCTATCAATAAAATTTCTAATTCCTAAAACCAAAATCCAGTACTAAAATCCAGATC (SEQ ID NO. 75)
[0352] bGH poly(A) signal:
[0353] CTGTGCCTTCTAGTTGCCAGCCATCTGTTGTTTGCCCCTCCCCCGTGCCTTCCTTGACCCTGGAAGGTGCCACTCCCACTGTCCTTCCTAATAAAATGAGGAAATTGCATCGCATTGTCTGAGTAGGTGTCATTCTATTCTGGGGGGTGGGGTGGCAGGACAGCAAGGGGGAGGATTGGGAAGACAATAGCAGGCATGCTGGGGATGCGGTGGGCTCTATGG (SEQ ID NO. 76).
[0354] SOD2 MTS:
[0355] MLSRAVCGTSRQLAPVLGYLGSRQKHSLPD (SEQ ID NO. 77)
[0356] COX8 MTS:
[0357] MSVLTPLLLRGLTGSARLPVPRAK (SEQ ID NO. 78)
[0358] CTP:
[0359] MAPTVMMASSATAVAPFQGLKSAASLPVARRSTRSLGNVSNGGRIRCMQ (SEQ ID NO. 79)
[0360] Target cells of interest
[0361] The recombinant expression constructs provided by this invention can be produced according to gene engineering methods known in the art. In some embodiments, a base editor or its recombinant expression construct is introduced into cells to edit and express the target gene, forming edited genetically engineered cells.
[0362] Any cell from any organism can be used with the nucleic acids, peptides, compositions, and methods described in this invention. Cells include, but are not limited to, human, non-human, animal, mammalian, bacterial, protozoan, fungal, insect, yeast, unconventional yeast, and plant cells, including monocotyledonous and dicotyledonous plants and plant elements, as well as plants and seeds produced by the methods described in this invention. In some respects, the cells of the organism are somatic cells.
[0363] In some embodiments, animal cells may include, but are not limited to: organisms belonging to the following phyla, including Chordata, Arthropoda, Molluscs, Annelids, Cnidaria, or Echinodermata; and organisms belonging to the following classes, including mammals, insects, birds, amphibians, reptiles, or fish. In some aspects, the animals are humans, mice, *Caenorhabditis elegans*, rats, fruit flies, zebrafish, chickens, dogs, cattle, sheep, pigs, guinea pigs, hamsters, chickens, Japanese rice fish, lampreys, pufferfish, tree frogs, monkeys, or chimpanzees.
[0364] Specific animal cell types include neurons, muscle cells, endocrine or exocrine cells, epithelial cells, tumor cells, hematopoietic cells, bone cells, somatic cells, and induced pluripotent stem cells. In some cases, multiple cells from an organism can be used.
[0365] In some implementations, the plant cells are derived from monocotyledonous and dicotyledonous plants. Examples of monocotyledonous plants that can be used include, but are not limited to, maize (corn), rice (Oryza sativa), rye (Secalecereale), sorghum (Sorghum bicolor, Sorghum vulgare), millet (e.g., pearl millet, Pennisetum glaucum), foxtail millet (Panicum miliaceum), foxtail millet (Setaria italica), millet (Eleusine coracana), wheat (wheat species, such as Triticum estivum, Triticum monococcum), sugarcane (sugarcane species (Saccharum spp.)), oats (Avena), barley (Hordeum), switchgrass (Panicum virgatum), and pineapple (Ananas pineapple). comosus), bananas (Musaspp. species), palms, ornamental plants, turfgrass, and other grasses. Examples of dicotyledonous plants that may be used include, but are not limited to, soybean (Glycine max), Brassica species (e.g., but not limited to: rapeseed or canola rapeseed) (European rapeseed (Brassica napus), Chinese rapeseed (B. campestris), turnip (Brassica rapa), mustard (Brassica juncea)), alfalfa (Medicago sativa)), tobacco (Nicotiana tabacum)), Arabidopsis (Arabidopsis thaliana)), sunflower (Helianthus annuus)), cotton (Gossypium arboreum, Gossypium barbadense)), peanut (Arachis hypogaea)), tomato (Solanum lycopersicum)), and potato (Solanum tuberosum)).Other plants that can be used include safflower (Carthamustinctorius), sweet potato (Ipomoea batatas), cassava (Manihot esculenta), coffee (Coffeaspp.), coconut (Cocos nucifera), citrus (Citrus spp.), cocoa (Theobroma cacao), tea (Camellia sinensis), banana (Musaspp.), avocado (Persea americana), fig (Ficuscasica), guava (Psidium guajava), mango (Mangifera indica), olive (Olea europaea), papaya (Carica papaya), cashew (Anacardium occidentale), and macadamia (Macadamia). The plant species include *Integrifolia*, almonds (*Prunus amygdalus*), sugar beets (*Beta vulgaris*), vegetables, ornamental plants, and conifers. Vegetables that can be used include tomatoes (*Lycopersicon esculentum*), lettuce (e.g., *Lactuca sativa*), green beans (*Phaseolus vulgaris*), lima beans (*Phaseolus limensis*), peas (*Lathyrus* spp.)), and members of the genus *Cucumis* such as cucumber (*C. sattivus*), cantaloupe (*C. cantalupensis*), and honeydew melon (*C. melon*).Ornamental plants include azaleas (Rhododendron spp.), hydrangeas (Macrophylla hydrangea), hibiscus (Hibiscus rosasanensis), roses (Rosa spp.), tulips (Tulipa spp.), daffodils (Narcissus spp.), petunias (Petunia hybrida), carnations (Dianthus caryophyllus), poinsettias (Euphorbia pulcherrima), and chrysanthemums. Suitable coniferous trees include pine trees such as loblolly pine (Pinus taeda), slash pine (Pinus elliotii), ponderosa pine (Pinus ponderosa), lodgepole pine (Pinus contorta), and Monterey pine (Pinus radiata); Douglas fir (Pseudotsuga menziesii); Western hemlock (Tsugacanadensis); North American spruce (Picea glauca); redwood (Sequoia sempervirens); and true fir trees, such as silver fir (Abies amabilis) and balsam fir (Abies sempervirens). balsamea); and cedars, such as Western Red Cedar (Thujaplicata) and Alaskan Yellow Cedar (Chamaecyparis nootkatensis).
[0366] Specific plant cell types include, but are not limited to, cells derived from the following substances: whole plant, seedling, meristem, ground tissue, vascular tissue, cortical tissue, seeds, leaves, roots, buds, stems, flowers, fruits, stolons, bulbs, tubers, corms, asexual terminal shoots, buds, young shoots, tumor tissue, and various forms of cells and cultures (e.g., single cells, protoplasts, embryos, callus). They can exist in plants or in plant organs, tissue cultures, or cell cultures.
[0367] Therapeutic applications
[0368] This invention also covers the application of the base editor of this invention in disease treatment.
[0369] By modifying disease-related genes using the base editor of this invention, it is possible to achieve upregulation, downregulation, inactivation, activation, mutation correction, or introduction of disease-related sites, thereby enabling disease prevention and / or treatment and / or the creation of disease-related models. For example, the target nucleic acid region described in this invention can be located within the protein-coding region of the disease-related gene, or, for example, within gene expression regulatory regions such as promoter regions or enhancer regions, thereby enabling modification of the function or expression of the disease-related gene. Therefore, the modification of disease-related genes described herein includes modification of the disease-related gene itself (e.g., protein-coding region), as well as modification of its expression regulatory regions (e.g., promoters, enhancers, introns, etc.).
[0370] "Disease-associated" genes are any genes that produce transcriptional or translational products at abnormal levels or in abnormal forms in cells derived from tissues affected by a disease, compared to tissues or cells from non-disease control groups. In cases where altered expression is associated with the onset and / or progression of the disease, it can be a gene expressed at abnormally high levels; it can also be a gene expressed at abnormally low levels. Disease-associated genes also refer to genes with one or more mutations or genetic variations that are directly responsible for or linked to one or more genes responsible for the etiology of the disease in disequilibrium. Such mutations or genetic variations are, for example, single nucleotide variants (SNVs). The transcribed or translated products can be known or unknown and can be at normal or abnormal levels.
[0371] Therefore, the present invention also provides a method for treating a disease in a subject of need, comprising delivering an effective amount of the base editor of the present invention to the subject to modify a gene associated with the disease (e.g., deamination of mitochondrial DNA via a fusion protein or multiple fusion proteins). The present invention also provides the use of the base editor in the preparation of a pharmaceutical composition for treating a disease in a subject of need, wherein the base editor is used to modify a gene associated with the disease. The present invention also provides a pharmaceutical composition for treating a disease in a subject of need, comprising the base editor of the present invention, and optionally a pharmaceutically acceptable carrier, wherein the base editor is used to modify a gene associated with the disease.
[0372] In some embodiments, the fusion protein or base editor described in this invention is used to introduce point mutations into nucleic acids by deaminating a target nucleobase (e.g., a C residue). In some embodiments, the deamination of the target nucleobase results in the correction of a genetic defect, such as in the correction of a point mutation that results in loss of function in a gene product. In some embodiments, the genetic defect is associated with a disease or condition (e.g., lysosomal storage disease or metabolic disease, such as, for example, type 1 diabetes). In some embodiments, the methods provided herein can be used to introduce inactive point mutations into a gene or allele encoding a gene product associated with a disease or condition.
[0373] In some embodiments, the purpose of the schemes described in this invention is to restore the function of dysfunctional genes via genome editing. The nucleobase editing proteins provided herein are intended for use in vitro gene editing in human cells, such as correcting disease-related mutations in human cell cultures.
[0374] In some embodiments, the purpose of the schemes described in this invention is to treat diseases associated with or caused by point mutations, which can be corrected by the DNA base-editing fusion proteins provided herein. In some embodiments, the disease is a proliferative disease. In some embodiments, the disease is a genetic disease. In some embodiments, the disease is a neonatal disease. In some embodiments, the disease is a metabolic disease. In some embodiments, the disease is a lysosomal storage disease.
[0375] In some embodiments, the purposes of the solutions described in this invention are for the treatment of mitochondrial diseases or disorders. As used herein, "mitochondrial disease" refers to diseases caused by abnormal mitochondria, such as mitochondrial gene mutations, gene mutations in enzyme pathways, etc. Examples of diseases include, but are not limited to: neurological disorders, loss of motor control, muscle weakness and pain, gastrointestinal disorders and dysphagia, poor growth, heart disease, liver disease, diabetes, respiratory complications, epilepsy, visual / hearing problems, lactic acidosis, developmental delay, and susceptibility to infection.
[0376] Examples of diseases described in this invention include, but are not limited to, genetic diseases, circulatory system diseases, muscle diseases, brain, central nervous system and immune system diseases, Alzheimer's disease, secretase disorders, amyotrophic lateral sclerosis (ALS), autism, trinucleotide repeat amplification disorders, hearing disorders, gene-targeted therapy for non-dividing cells (neurons, muscles), liver and kidney diseases, epithelial cell and lung diseases, cancer, Usher syndrome or retinitis pigmentosa-39, cystic fibrosis, HIV and AIDS, β-thalassemia, sickle cell disease, herpes simplex virus, autism, drug addiction, age-related macular degeneration, and schizophrenia. Other diseases treated by correcting point mutations or introducing inactive mutations into disease-related genes are known to those skilled in the art, and therefore the scope of this invention is not limited thereto. In addition to the diseases exemplarily described in this invention, other related diseases can also be treated with the strategies and fusion proteins provided by this invention, and this application will be apparent to those skilled in the art. The diseases or targets to which this invention can be applied refer to the diseases for which the base editor is applicable as listed in WO2015089465A1 (PCT / US2014 / 070135), WO2016205711A1 (PCT / US2016 / 038181), WO2018141835A1 (PCT / EP2018 / 052491), WO2020191234A1 (PCT / US2020 / 023713), WO2020191233A1 (PCT / US2020 / 023712), WO2019079347A1 (PCT / US2018 / 056146), and WO2021155065A1 (PCT / US2021 / 015580).
[0377] Applications in plants
[0378] The base-editing fusion protein, base editor, and method for generating genetically modified cells of the present invention are particularly suitable for genetic modification of plants. Preferably, the plant is a crop plant, including but not limited to wheat, rice, corn, soybean, sunflower, sorghum, rapeseed, alfalfa, cotton, barley, millet, sugarcane, tomato, tobacco, cassava, and potato. More preferably, the plant is rice.
[0379] In another aspect, the present invention provides a method for producing genetically modified plants, comprising introducing a base editor of the present invention into at least one of said plants, thereby causing substitution of one or more nucleotides in a target nucleic acid region in the genome of said at least one plant.
[0380] In some embodiments, the method further includes screening from the at least one plant for plants having one or more desired nucleotide substitutions.
[0381] In the method of this invention, the base editing composition can be introduced into plants using various methods well known to those skilled in the art. Methods for introducing the base editor of this invention into plants include, but are not limited to: gene gun method, PEG-mediated protoplast transformation, Agrobacterium-mediated transformation, plant virus-mediated transformation, pollen tube pathway method, and ovary injection method. Preferably, the base editing composition is introduced into plants via transient transformation.
[0382] In the method of this invention, modification of the target sequence can be achieved simply by introducing or generating the base-editing fusion protein in plant cells, and the modification can be stably inherited without the need for stable transformation of the plant with exogenous polynucleotides encoding the base editor components. This avoids the potential off-target effects of a stably existing (continuously generated) base-editing composition and also avoids the integration of exogenous nucleotide sequences into the plant genome, thus providing higher biosafety.
[0383] In some preferred embodiments, the introduction is performed without selection pressure, thereby avoiding the integration of exogenous nucleotide sequences into the plant genome.
[0384] In some embodiments, the introduction includes converting the base editor of the present invention into isolated plant cells or tissues, and then regenerating the converted plant cells or tissues into complete plants. Preferably, the regeneration is performed without selection pressure, that is, without using any selecting agents targeting the selection genes carried on the expression vector during tissue culture. Not using selecting agents can improve the regeneration efficiency of the plants, resulting in modified plants free of exogenous nucleotide sequences.
[0385] In other embodiments, the base editor of the present invention can be transformed into specific parts of a whole plant, such as leaves, shoot tips, pollen tubes, young spikelets, or hypocotyls. This is particularly suitable for the transformation of plants that are difficult to regenerate through tissue culture.
[0386] Therefore, in some embodiments, using the methods of the present invention to genetically modify and breed plants can yield plants whose genomes are free of foreign polynucleotide integration, i.e., non-transgene-free modified plants.
[0387] In some embodiments of the invention, the modified target nucleic acid region is associated with plant traits such as agronomic traits, whereby the substitution of one or more nucleotides results in the plant having altered (preferably improved) traits, such as agronomic traits, relative to the wild-type plant.
[0388] In some embodiments, the method further includes the step of screening plants having one or more desired nucleotide substitutions and / or desired traits such as agronomic traits.
[0389] In some embodiments of the invention, the method further includes obtaining offspring of the genetically modified plant. Preferably, the genetically modified plant or its offspring has one or more desired nucleotide substitutions and / or desired traits such as agronomic traits.
[0390] In another aspect, the present invention also provides genetically modified plants or their offspring or portions thereof, wherein said plants are obtained by the methods described above. In some embodiments, the genetically modified plants or their offspring or portions thereof are non-GMO. Preferably, the genetically modified plants or their offspring have the desired genetic modification and / or desired traits such as agronomic traits.
[0391] In another aspect, the present invention also provides a plant breeding method, comprising crossing a genetically modified first plant, obtained by the method described above, containing one or more nucleotide substitutions in a target nucleic acid region, with a second plant not containing the one or more nucleotide substitutions, thereby introducing the one or more nucleotide substitutions into the second plant. Preferably, the genetically modified first plant has desired traits such as agronomic traits.
[0392] Example
[0393] A further understanding of the invention can be obtained by referring to some specific embodiments given herein, which are for illustrative purposes only and are not intended to limit the scope of the invention in any way. Obviously, various modifications and variations can be made to the invention without departing from its spirit; therefore, such modifications and variations are also within the scope of protection claimed in this application.
[0394] The partial component sequences used in subsequent embodiments are shown below:
[0395] OsBADH2 Left TALE repeat
[0396] LTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALE
[0397] (SEQ ID NO. 89)
[0398] OsBADH2 Right TALE repeat
[0399] LTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALE
[0400] (SEQ ID NO. 90)
[0401] OsDEP1 Left TALE repeat
[0402] LTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALE
[0403] (SEQ ID NO. 91)
[0404] OsDEP1 Right TALE repeat
[0405] LTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALE
[0406] (SEQ ID NO. 92)
[0407] OsCKX2 Left TALE repeat
[0408] LTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALE
[0409] (SEQ ID NO. 93)
[0410] OsCKX2 Right TALE repeat
[0411] LTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALE
[0412] (SEQ ID NO. 94)
[0413] OsSD1 Left TALE repeat
[0414] LTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALE
[0415] (SEQ ID NO. 95)
[0416] OsSD1 Right TALE repeat
[0417] LTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALE
[0418] (SEQ ID NO. 96)
[0419] SIRT6 Left TALE repeat
[0420] LTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIANNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPEQVVAIANNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIANNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIANNNGGKQALETVQRLLPVLCQDHGLTPEQVVAIANNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIANNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIANNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIANNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPEQVVAIASHDGGRPALE
[0421] (SEQ ID NO. 97)
[0422] SIRT6 Right TALE repeat
[0423] LTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIANNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIANNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIANNNGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIANNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIANNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIANNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNNGGRPALE
[0424] (SEQ ID NO. 98)
[0425] OsRbcL Left TALE repeat
[0426] LTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQAVETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQAVETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQAVETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETLQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQAVETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQAVETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQAVETVQRLLPVLCQDHGLTPEQVVAIASNGGGKQALE
[0427] (SEQ ID NO. 99)
[0428] OsRbcL Right TALE repeat
[0429] LTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQAVETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQAVETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQAVETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETLQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQAVETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQAVETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQAVETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNGGGKQALE
[0430] (SEQ ID NO. 100)
[0431] ND6 Left TALE repeat
[0432] LTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHD
[0433] (SEQ ID NO. 101)
[0434] ND6 Right TALE repeat
[0435] LTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNN
[0436] (SEQ ID NO. 102)
[0437] ND5.1 Left TALE repeat
[0438] LTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNI
[0439] (SEQ ID NO. 103)
[0440] ND5.1 Right TALE repeat
[0441] LTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNI
[0442] (SEQ ID NO. 104)
[0443] ND3 Left TALE repeat
[0444] LTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPQQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPQQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPQQVVAIASNIGGRPALE
[0445] (SEQ ID NO. 105)
[0446] ND3 Right TALE repeat
[0447] LTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPQQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQLETVQRLLPVLCQAHGLTPQQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPQQVVAIASHDGGRPALE
[0448] (SEQ ID NO. 106)
[0449] ND1.3 Left TALE repeat
[0450] LTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPQQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPQQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPQQVVAIASNNGGRPALE
[0451] (SEQ ID NO. 107)
[0452] ND1.3 Right TALE repeat
[0453] LTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPQQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPQQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPQQVVAIASHDGGRPALE
[0454] (SEQ ID NO. 108)
[0455] ND1.2 Left TALE repeat
[0456] LTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPQQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPQQVVAIASNNGGRPALE
[0457] (SEQ ID NO. 109)
[0458] ND1.2 Right TALE repeat
[0459] LTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPQQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPQQVVAIASHDGGRPALE
[0460] (SEQ ID NO. 110)
[0461] ND6.2 Left TALE repeat(TALE-L2)
[0462] LTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPQQVVAIASHDGGKQALETVQRLPVLCQAHGLTPQQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLC QAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLP VLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQ ALLPVLCQAHGLTPEQVVAIASNNGGKALETVQRLPVLCQAHGLTPQQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIAASngGGKQALETVQALLPVLCQAHGLTQQVVAIASNIGGRPALE
[0463] (SEQ ID NO. 111)
[0464] ND6.2 Right TALE repeat(TALE-R2)
[0465] LTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPQQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPQQVVAIASHDGGRPALE
[0466] (SEQ ID NO. 112)
[0467] ND6.2 Left TALE repeat (TALE-L1)
[0468] LTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHG
[0469] (SEQ ID NO. 185)
[0470] ND6.2 Left TALE repeat (TALE-L3)
[0471] LTPDQVVAIASNNGGQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQV IASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLPVLCQDHGLTQVVAIASNNGGQALETVQRLLPVLCQDHGLTQVVAIASNNGGQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQ ALETVQRLPVLCQDHGLTQVVAIASNGGGKQALETVQRLPVLCQDHGLTQVVAIASNIGGKQALETVQRLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRL LPVLCQDHGLTQVVAIASHDGGKALETVQRLLPVLCQDHGLTQVVAIASHDGGKQALETVQRLLPQDHGLTPDQVVAIASNGGGQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQGRLLPVLCQ
[0472] (SEQ ID NO. 186)
[0473] ND6.2 Right TALE repeat (TALE-R1)
[0474] LTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHG LTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHG LTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHG LTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHG
[0475] (SEQ ID NO. 187)
[0476] XTEN spacer peptide
[0477] NSGSETPGTSESATPES
[0478] (SEQ ID NO. 113)
[0479] 48 amino acid spacer peptide
[0480] SGSETPGTSESATPESSGGSSGGSSGSETPGTSESATPESSGGSSGGS
[0481] (SEQ ID NO. 114)
[0482] 16-amino acid spacer peptide
[0483] SGSETPGTSESATPES
[0484] (SEQ ID NO. 115)
[0485] 14 amino acid spacer peptide
[0486] SGGGSGGSGGSGGS
[0487] (SEQ ID NO. 116)
[0488] 11 amino acid spacer peptide
[0489] SGGSGGSGGSS
[0490] (SEQ ID NO. 117)
[0491] 4-amino acid spacer peptide
[0492] SGGS
[0493] (SEQ ID NO. 118)
[0494] γb
[0495] MMATFSCVCCGTLTTSTYCGKRCERKHVYSETRNKRLELYKKYLLEPQKCALNGIVGHSCGMPCSIAEEACDQLPIVSRFCGQKHADLYDSLLKRSEQELLLEFLQKKMQELKLSHIVKMAKLESEVNAIRKSVASSFEDSVGCDDSSSVSK
[0496] (SEQ ID NO. 119)
[0497] Figures 16A to 16E and Figures 17A to 17H The amino acid sequences of the carriers or elements involved are shown below. Unless otherwise specified in subsequent embodiments, they can be represented as follows. Figures 16A to 16E and Figures 17A to 17H The schematic diagram of the construct shown is used to construct the corresponding fusion protein in conjunction with the sequence disclosed in this specification.
[0498] OsBADH2-NLS-TALEN WT ( Figure 16A )
[0499]
[0500] (SEQ ID NO. 120)
[0501] OsBADH2-NLS-TALE-L-FokI-L-T2A-TALE-R-FokI-RD 450A ( Figure 16B )
[0502]
[0503] (SEQ ID NO. 121)
[0504] OsBADH2-NLS-TALE-L-FokI-L D450A -T2A-TALE-R-FokI-R( Figure 16B )
[0505]
[0506] (SEQ ID NO. 122)
[0507] NLS-A3A-XTEN-UGI( Figure 16B )
[0508] MKRTADGSEFESPKKKRKVMEASPASGPRHLMDPHIFTSNFNNGIGRHKTYLCYEVERLDNGTSVKMDQHRGFLHNQAKNLLCGFYGRHAELRFLDLVPSLQLDPAQIYRVTWFISWSPCFSWGCAGEVRAFLQENTHVRLRIFAARIYDYDPLYKEALQMLRDAGAQVSIMTYDEFKHCWDTFVDHQGCPFQPWDGLDEHSQALSGRLRAILQNQGNSGSETPGTSESATPESTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML
[0509] (SEQ ID NO. 123)
[0510] NLS-UGI( Figure 16B )
[0511] MTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKMLMKRTADGSEFESPKKKRKV
[0512] (SEQ ID NO. 163)
[0513] NLS-C57-XTEN-UGI( Figure 16B )
[0514] MKRTADGSEFESPKKKRKVLEAVRARLIGEGGGPGAVPEGGDGPPAVPAEEVERLRGELPPPVVPGTGQKTHGRWIGPDGRVRAIVSGRDEDAALVHAQLAAKGIPDEPTRNSDVEQKLAAHMVANGIRHVTLVINHRPCRGFDDSCDTLVPIILPEGCTLTVHGQTDKGMRVRVRYTGGARPWWSNSGSETPGTSESATPESTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML
[0515] (SEQ ID NO. 124)
[0516] NLS-rAPOBEC1-XTEN-UGI( Figure 16B )MKRTADGSEFESPKKKRKVSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSIWRHTSQNTNKHVEVNFIEKFTTERYFCPNTRCSITWFLSWSPCGECSRAITEFLSRYPHVTLFIYIARLYHHADPRNRQGLRDLISSGVTIQIMTEQESGYCWRNFVNYSPSNEAHWPRYPHLWVRLYVLELYCIILGLPPCLNILRRKQPQLTFFTIALQSCHYQRLPPHILWATGLKNSGSETPGTSESATPESTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML
[0517] (SEQ ID NO. 164)
[0518] TadA8e-NLS( Figure 16B )
[0519] MSEVEFSHEYWMRHALTLAKRARDEREVPVGVAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSINSGGSMKRTADGSEFESPKKKRKV
[0520] (SEQ ID NO. 166)
[0521] mExoI-NLS( Figure 16B )
[0522] MGIQGLLQFIQEASEPVNVKKYKGQAVAVDTYCWLHKGAIACAEKLAKGEPTDRYVGFCMKFVNMLLSYGVKPILIFDGCTLPSKKEVERSRRERRQSNLLKGKQLLREGKVSEARDCFARSINITHAMAHKVIKAARALGVDCLVAPYEADAQLAYLNKAGIVQAVITEDSDLLAFGCKKVILKMDQFGNGLEVDQARLGMCKQLGDVFTEEKFRYMCILSGCDYLASLRGIGLAKACKVLRLANNPDIVKVIKKIGHYLRMNITVPEDYITGFIRANNTFLYQLVFDPIQRKLVPLNAYGDDVNPETLTYAGQYVGDSVALQIALGNRDVNTFEQIDDYSPDTMPAHSRSHSWNEKAGQKPPGTNSIWHKNYCPRLEVNSVSHAPQLKEKPSTLGLKQVISTKGLNLPRKSCVLKRPRNEALAEDDLLSQYSSVSKKIKENGCGDGTSPNSSKMSKSCPDSGTAHKTDAHTPSKMRNKFATFLQRRNEESGAVVVPGTRSRFFCSSQDFDNFIPKKESGQPLNETVATGKATTSLLGALDCPDTEGHKPVDANGTHNLSSQIPGNAAVSPEDEAQSSETSKLLGAMSPPSLGTLRSCFSWSGTLREFSRTPSPSASTTLQQFRRKSDPPACLPEASAVVTDRCDSKSEMLGETSQPLHELGCSSRSQESMDSSCGLNTSSLSQPSSRDSGSEESDCNNKSLDNQGEQNSKQHLPHFSKKDGLRRNKVPGLCRSSSMDSFSTTKIKPLVPARVSGLSKKSGSMQTRKHHDVENKPGLQTKISELWKNFGFKKDSEKLPSCKKPLSPVKDNIQLTPETEDEIFNKPECVRAQRAIFHMKRTADGSEFESPKKKRKV
[0523] (SEQ ID NO. 125)
[0524] Trex2-NLS( Figure 16B )
[0525] MSEPPRAETFVFLDLEATGLPNMDPEIAEISLFAVHRSSLENPERDDSGSLVLPRVLDKLTLCCMCPERPFTAKASEITGLSSESLMHCGKAGFNGAVVRTLQGFLSRQEGPICLVAHNGFDYDFPLL CTELQRLGAHLPQDTVCLDTLPALRGLDRAHSHGTRAQGRKSYSLASLFHRYFQAEPSAAHSAEGDVHTLLLIFLHRAPELLAWADEQARSWAHIEPMYVPPDGPSLEAMKRTADGSEFESPKKKKV
[0526] (SEQ ID NO. 126)
[0527] OsBADH2-NLS-A3A-TALE-L-FokI-L-T2A-TALE-R-FokI-R D450A (( Figure 16C ) )
[0528]
[0529] (SEQ ID NO. 127)
[0530] OsBADH2-NLS-A3A-TALE-L-FokI-L D450A -T2A-TALE-R-FokI-R( Figure 16C )
[0531]
[0532] (SEQ ID NO. 128)
[0533] mExoI-NLS( Figure 16C )
[0534] MGIQGLLQFIQEASEPVNVKKYKGQAVAVDTYCWLHKGAIACAEKLAKGEPTDRYVGFCMKFVNMLLSYGVKPILIFDGCTLPSKKEVERSRRERRQSNLLKGKQLLREGKVSEARDCFARSINITHAMAHKVIKAARALGVDCLVAPYEADAQLAYLNKAGIVQAVITEDSDLLAFGCKKVILKMDQFGNGLEVDQARLGMCKQLGDVFTEEKFRYMCILSGCDYLASLRGIGLAKACKVLRLANNPDIVKVIKKIGHYLRMNITVPEDYITGFIRANNTFLYQLVFDPIQRKLVPLNAYGDDVNPETLTYAGQYVGDSVALQIALGNRDVNTFEQIDDYSPDTMPAHSRSHSWNEKAGQKPPGTNSIWHKNYCPRLEVNSVSHAPQLKEKPSTLGLKQVISTKGLNLPRKSCVLKRPRNEALAEDDLLSQYSSVSKKIKENGCGDGTSPNSSKMSKSCPDSGTAHKTDAHTPSKMRNKFATFLQRRNEESGAVVVPGTRSRFFCSSQDFDNFIPKKESGQPLNETVATGKATTSLLGALDCPDTEGHKPVDANGTHNLSSQIPGNAAVSPEDEAQSSETSKLLGAMSPPSLGTLRSCFSWSGTLREFSRTPSPSASTTLQQFRRKSDPPACLPEASAVVTDRCDSKSEMLGETSQPLHELGCSSRSQESMDSSCGLNTSSLSQPSSRDSGSEESDCNNKSLDNQGEQNSKQHLPHFSKKDGLRRNKVPGLCRSSSMDSFSTTKIKPLVPARVSGLSKKSGSMQTRKHHDVENKPGLQTKISELWKNFGFKKDSEKLPSCKKPLSPVKDNIQLTPETEDEIFNKPECVRAQRAIFHMKRTADGSEFESPKKKRKV
[0535] (SEQ ID NO. 129)
[0536] Trex2-NLS( Figure 16C ) )
[0537] MSEPPRAETFVFLDLEATGLPNMDPEIAEISLFAVHRSSLENPERDDSGSLVLPRVLDKLTLCCMCPERPFTAKASEITGLSSESLMHCGKAGFNGAVVRTLQGFLSRQEGPICLVAHNGFDYDFPLL CTELQRLGAHLPQDTVCLDTLPALRGLDRAHSHGTRAQGRKSYSLASLFHRYFQAEPSAAHSAEGDVHTLLLIFLHRAPELLAWADEQARSWAHIEPMYVPPDGPSLEAMKRTADGSEFESPKKKKV
[0538] (SEQ ID NO. 130)
[0539] UGI-NLS( Figure 16C ) )
[0540] MKRTADGSEFESPKKKRKVTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML
[0541] (SEQ ID NO. 131)
[0542] OsBADH2-NLS-A3A-TALE-L-FokI-L-T2A-TALE-R-FokI-R D450A -UGI( Figure 16D ) )
[0543]
[0544] (SEQ ID NO. 132)
[0545] OsBADH2-NLS-A3A-TALE-L-FokI-L D450A -T2A-TALE-R-FokI-R-UGI( Figure 16D )
[0546]
[0547] (SEQ ID NO. 133)
[0548] mExoI-NLS( Figure 16D )
[0549] MGIQGLLQFIQEASEPVNVKKYKGQAVAVDTYCWLHKGAIACAEKLAKGEPTDRYVGFCMKFVNMLLSYGVKPILIFDGCTLPSKKEVERSRRERRQSNLLKGKQLLREGKVSEARDCFARSINITHAMAHKVIKAARALGVDCLVAPYEADAQLAYLNKAGIVQAVITEDSDLLAFGCKKVILKMDQFGNGLEVDQARLGMCKQLGDVFTEEKFRYMCILSGCDYLASLRGIGLAKACKVLRLANNPDIVKVIKKIGHYLRMNITVPEDYITGFIRANNTFLYQLVFDPIQRKLVPLNAYGDDVNPETLTYAGQYVGDSVALQIALGNRDVNTFEQIDDYSPDTMPAHSRSHSWNEKAGQKPPGTNSIWHKNYCPRLEVNSVSHAPQLKEKPSTLGLKQVISTKGLNLPRKSCVLKRPRNEALAEDDLLSQYSSVSKKIKENGCGDGTSPNSSKMSKSCPDSGTAHKTDAHTPSKMRNKFATFLQRRNEESGAVVVPGTRSRFFCSSQDFDNFIPKKESGQPLNETVATGKATTSLLGALDCPDTEGHKPVDANGTHNLSSQIPGNAAVSPEDEAQSSETSKLLGAMSPPSLGTLRSCFSWSGTLREFSRTPSPSASTTLQQFRRKSDPPACLPEASAVVTDRCDSKSEMLGETSQPLHELGCSSRSQESMDSSCGLNTSSLSQPSSRDSGSEESDCNNKSLDNQGEQNSKQHLPHFSKKDGLRRNKVPGLCRSSSMDSFSTTKIKPLVPARVSGLSKKSGSMQTRKHHDVENKPGLQTKISELWKNFGFKKDSEKLPSCKKPLSPVKDNIQLTPETEDEIFNKPECVRAQRAIFHMKRTADGSEFESPKKKRKV
[0550] (SEQ ID NO. 134)
[0551] Trex2-NLS( Figure 16D ) )
[0552] MSEPPRAETFVFLDLEATGLPNMDPEIAEISLFAVHRSSLENPERDDSGSLVLPRVLDKLTLCCMCPERPFTAKASEITGLSSESLMHCGKAGFNGAVVRTLQGFLSRQEGPICLVAHNGFDYDFPLL CTELQRLGAHLPQDTVCLDTLPALRGLDRAHSHGTRAQGRKSYSLASLFHRYFQAEPSAAHSAEGDVHTLLLIFLHRAPELLAWADEQARSWAHIEPMYVPPDGPSLEAMKRTADGSEFESPKKKKV
[0553] (SEQ ID NO. 135)
[0554] OsBADH2-NLS-A3A-TALE-L-FokI-L-T2A-TALE-R-FokI-R D450A -UGI--mExoI-NLS( Figure 16E ) )
[0555]
[0556] (SEQ ID NO. 136)
[0557] OsBADH2-NLS-A3A-TALE-L-FokI-L D450A -T2A-TALE-R-FokI-R-UGI--mExoI-NLS( Figure 16E )
[0558]
[0559] (SEQ ID NO. 137)
[0560] ND6-MTS-TALE-L-FokI-L( Figure 17A ) )
[0561] MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAMDIADLRTLGYSQQQQEKPKVRSTVAQHHEALV GHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRN ALTGAPLNLTPEQVVAIASNNGGQQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGQAL ETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHD GGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQV VAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGRPALESIVAQLSRPDPALALTNDHLVALACLGGRPALDAVKKGLGGSQLVKS ELEEKKSELRHKLKYVPHEYIEIARNSTQDRILEMKVMEFFMKVYGYRGKHLGGSRKPDGAIYTVGSPIDYGVIVDTKAYSGGYNLPIGQAD EMQRYVKENQTRNKHINPNEWWKVYPSSVTEFKFLFVSGHFKGNYKAQLTRLNHKTNCNGAVLSVEELLIGGEMIKAGTLEEVRRKFNNGEINF
[0562] (SEQ ID NO. 138)
[0563] ND6-MTS-TALE-R-FokI-RD 450A (( Figure 17B ) )
[0564] MASVLTPPLLLRGLTGSARRLPVPRAKIHSLDYKDHDGDYKDHDYDYKDDDDKMDIADLRTLGYSQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVA LSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKRGAGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQV VAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTP EQVVAIASNNGGQQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGQQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHG LTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLC QAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGRPALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLG GSQLVXSELEEKKSELRHKLKYVPHEYIEIARNSTQDRILEMKVMEFFMKVYRGKHLGGSRKPAGAIYTVGSPIDYGVIVDTKAYSGGYNLPIG QADEMERYVEENQTRNKHLNPNEWWKVYPSSVTEFKFLFVSGHFKGNYKAQLTRLNHITNCNGAVLSVEELLIGGEMIKAGTLEEVRRKFNNGEINF
[0565] (SEQ ID NO. 139)
[0566] ND6-MTS-TALE-L-FokI-L D450A (( Figure 17A ) )
[0567] MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAMDIADLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGRPALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSQLVKSELEEKKSELRHKLKYVPHEYIELIEIARNSTQDRILEMKVMEFFMKVYGYRGKHLGGSRKPAGAIYTVGSPIDYGVIVDTKAYSGGYNLPIGQADEMQRYVKENQTRNKHINPNEWWKVYPSSVTEFKFLFVSGHFKGNYKAQLTRLNHKTNCNGAVLSVEELLIGGEMIKAGTLTLEEVRRKFNNGEINF
[0568] (SEQ ID NO. 140)
[0569] ND6-MTS-TALE-R-FokI-R( Figure 17B )
[0570] MASVLTPLLLRGLTGSARRLPVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDKMDIADLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKRGAGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGRPALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSQLVKSELEEKKSELRHKLKYVPHEYIELIEIARNSTQDRILEMKVMEFFMKVYGYRGKHLGGSRKPDGAIYTVGSPIDYGVIVDTKAYSGGYNLPIGQADEMERYVEENQTRNKHLNPNEWWKVYPSSVTEFKFLFVSGHFKGNYKAQLTRLNHITNCNGAVLSVEELLIGGEMIKAGTLTLEEVRRKFNNGEINF
[0571] (SEQ ID NO. 141)
[0572] MTS-mExoI( Figure 17D )
[0573] MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDMGIQGLLQFIQEASEPVNVKKYKGQAVAVDTYCWLHKGAIACAEKLAKGEPTDRYVGFCMKFVNMLLSYGVKPILIFDGCTLPSKKEVERSRRERRQSNLLKGKQLLREGKVSEARDCFARSINITHAMAHKVIKAARALGVDCLVAPYEADAQLAYLNKAGIVQAVITEDSDLLAFGCKKVILKMDQFGNGLEVDQARLGMCKQLGDVFTEEKFRYMCILSGCDYLASLRGIGLAKACKVLRLANNPDIVKVIKKIGHYLRMNITVPEDYITGFIRANNTFLYQLVFDPIQRKLVPLNAYGDDVNPETLTYAGQYVGDSVALQIALGNRDVNTFEQIDDYSPDTMPAHSRSHSWNEKAGQKPPGTNSIWHKNYCPRLEVNSVSHAPQLKEKPSTLGLKQVISTKGLNLPRKSCVLKRPRNEALAEDDLLSQYSSVSKKIKENGCGDGTSPNSSKMSKSCPDSGTAHKTDAHTPSKMRNKFATFLQRRNEESGAVVVPGTRSRFFCSSQDFDNFIPKKESGQPLNETVATGKATTSLLGALDCPDTEGHKPVDANGTHNLSSQIPGNAAVSPEDEAQSSETSKLLGAMSPPSLGTLRSCFSWSGTLREFSRTPSPSASTTLQQFRRKSDPPACLPEASAVVTDRCDSKSEMLGETSQPLHELGCSSRSQESMDSSCGLNTSSLSQPSSRDSGSEESDCNNKSLDNQGEQNSKQHLPHFSKKDGLRRNKVPGLCRSSSMDSFSTTKIKPLVPARVSGLSKKSGSMQTRKHHDVENKPGLQTKISELWKNFGFKKDSEKLPSCKKPLSPVKDNIQLTPETEDEIFNKPECVRAQRAIFH
[0574] (SEQ ID NO. 142)
[0575] MTS-Trex2( Figure 17D )
[0576] MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDMSEPPRAETFVFLDLEATGLPNMDPEIAEISLFAVHRSSLENPERDDSGSLVLPRVLDKLTLCMCPERPFTAKASEITGLSSESLMHCGKAGFNGAVVRTLQGFLSRQEGPICLVAHNGFDYDFPLLCTELQRLGAHLPQDTVCLDTLPALRGLDRAHSHGTRAQGRKSYSLASLFHRYFQAEPSAAHSAEGDVHTLLLIFLHRAPELLAWADEQARSWAHIEPMYVPPDGPSLEA
[0577] (SEQ ID NO. 143)
[0578] MTS-A3A( Figure 17C )
[0579] MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDEASPASGPRHLMDPHIFTSNFNNGIGRHKTYLCYEVERLDNGTSVKMDQHRGFLHNQAKNLLCGFYGRHAELRFLDLVPSLQLDPAQIYRVTWFISWSPCFSWGCAGEVRAFLQENTHVRLRIFAARIYDYDPLYKEALQMLRDAGAQVSIMTYDEFKHCWDTFVDHQGCPFQPWDGLDEHSQALSGRLRAILQNQGN
[0580] (SEQ ID NO. 144)
[0581] MTS-C57 / Sdd7( Figure 17C )
[0582] MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDLEAVRARLIGEGGGPGAVPEGGDGPPAVPAEEVERLRGELPPPVVPGTGQKTHGRWIGPDGRVRAIVSGRDEDAALVHAQLAAKGIPDEPTRNSDVEQKLAAHMVANGIRHVTLVINHRPCRGFDDSCDTLVPIILPEGCTLTVHGQTDKGMRVRVRYTGGARPWWS
[0583] (SEQ ID NO. 145)
[0584] MTS-UGI ( Figure 17E )
[0585] MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDGSSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML
[0586] (SEQ ID NO. 146)
[0587] ND6-MTS-A3A-TALE-L-FokI-L( Figure 17F )
[0588]
[0589] (SEQ ID NO. 147)
[0590] ND6-MTS-Trex2-TALE-R-FokI-R D450A ( Figure 17G )
[0591]
[0592] (SEQ ID NO. 148)
[0593] ND6-MTS-UGI-Trex2-SPEECH-R-FokI-R D450A ( 2010 ) Figure 17H )
[0594]
[0595] (SEQ ID NO. 149)
[0596] In this embodiment, the exemplary amino acid sequence of the element or fusion protein is shown below. Unless otherwise specified in subsequent embodiments, it can be represented as follows. Figures 16A to 16E and Figures 17A to 17H The schematic diagram of the construct shown below, and the exemplary sequences shown below, are used to construct the corresponding fusion proteins in conjunction with the sequences disclosed in this specification.
[0597] In subsequent embodiments, the nicking enzyme used in the OsBADH2 experiment was edited:
[0598] TALEN WT (SEQ ID NO. 154)
[0599]
[0600] TALE-FokI-R nickase (D450A) Or it can be called TALE-FokI-R nickase (SEQ ID NO. 155)
[0601]
[0602] TALE-FokI-R nickase (D467A) (SEQ ID NO. 156)
[0603]
[0604] Edit the nicking enzyme used in the OsDEP1 experiment:
[0605] TALEN WT (SEQ ID NO. 157)
[0606]
[0607] TALE-FokI-R nickase (D450A) Or it can be called TALE-FokI-R nickase (SEQ ID NO. 158)
[0608]
[0609] TALE-FokI-R nickase (D467A) (SEQ ID NO. 159)
[0610]
[0611] Edit the nicking enzyme used in the OsCKX2 experiment:
[0612] TALEN WT (SEQ ID NO. 160)
[0613]
[0614] TALE-FokI-R nickase (SEQ ID NO. 161)
[0615]
[0616] BACK-FORK-L nickase (SEQ ID NO. 162)
[0617]
[0618] In Examples 1-6, mExoI: i.e., mExoI-NLS mentioned above ( Figure 16B ), SEQ ID NO. 125; A3A-UGI: that is, the previous text: NLS-A3A-XTEN-UGI ( Figure 16B ), SEQ ID NO. 123; Trex2: i.e., the Trex2-NLS mentioned above ( Figure 16B ), SEQ ID NO. 126.
[0619] In Examples 1-6, the amino acid sequence of UGI is as described above for NLS-UGI ( Figure 16B (SEQ ID NO. 163).
[0620] In Example 4, the amino acid sequence of APOBEC1-UGI is as described above: NLS-rAPOBEC1-XTEN-UGI ( Figure 16B (SEQ ID NO. 164).
[0621] The amino acid sequence of ExoV (ExoV-NLS) in Example 1 (SEQ ID NO. 165):
[0622] MAETGEEETASAEASGFSDLSDSELVEFLDLEEAKESAVSLSKPGPSAELPGKDDKPVSLQNWKGGLDVLSPMERFHLKYLYVTDLCTQNWCELQMVYGKELPGSLTPEKAAVLDTGASIHLAKELELHDLVTVPIATKEDAWAVKFLNILAMIPALQSEGRVREFPVFGEVEGIFLVGVIDELHYTSKGELELAE LKTRRRPVLPLPAQKKKDYFQVSLYKYIFDAMVQGKVTPASLIHHTKLCLDKPLGPSVLRHARQGGVSVKSLGDLMELVFLSLTLSDLPAIDTLKLEY IHQETATILGTEIVAFEEKEVKSKVQHYVAYWMGHRDPQGVDVEEAWKCRTCDYVDICEWRRGSGVLSSSWEPKAKKFKMKRTADGSEFESPKKKRKV
[0623] In Example 5, the amino acid sequence of TadA-8e is the same as that described above for TadA8e-NLS ( Figure 16B (SEQ ID NO. 166).
[0624] In Example 6:
[0625] The amino acid sequence of mExoI-16aa-A3A-UGI (SEQ ID NO. 167):
[0626]
[0627] The amino acid sequence of mExoI-48aa-A3A-UGI (SEQ ID NO. 168):
[0628]
[0629] A3A-TALE-FokI-R nickase (SEQ ID NO. 169)
[0630]
[0631] APOBEC1-TALE-FoxI-R nickase (SEQ ID NO. 170)
[0632]
[0633] A3A-TALE-FokI-L nickase (SEQ ID NO. 171)
[0634]
[0635] APOBEC1-TALE-FoxI-L nickase (SEQ ID NO. 172)
[0636]
[0637] SIRT6-NLS-TALE-L-DddA N -UGI (SEQ ID NO. 173)
[0638]
[0639] SIRT6-NLS-TALE-R-DddA C -UGI(SEQ ID NO. 174)
[0640] MASVLTPLLLRGLTGSARRLPVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDKMDIADLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKRGAGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIANNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIANNNGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIANNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIANNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIANNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNNGGRPALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML
[0641] In Examples 11, 14 and 15:
[0642] ND6-MTS-TALE-L-DddAN -UGI(SEQ ID NO. 175)
[0643] MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAMDIADLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGRPALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML
[0644] ND6-MTS-TALE-R-DddA C -UGI(SEQ ID NO. 176)
[0645] MASVLTPLLLRGLTGSARRLPVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDKMDIADLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKRGAGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGRPALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML
[0646] ND1.2-MTS-TALE-L-DddA N -UGI(SEQ ID NO. 177)
[0647] MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAMDIADLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPQQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPQQVVAIASNNGGRPALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML
[0648] ND1.2-MTS-TALE-R-DddA C -UGI(SEQ ID NO. 178)
[0649] MASVLTPLLLRGLTGSARRLPVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDKMDIADLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKRGAGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPQQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGSIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML
[0650] ND1.3-MTS-TALE-L-DddA N -UGI(SEQ ID NO. 179)
[0651] MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAMDIADLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPQQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPQQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPQQVVAIASNNGGRPALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML
[0652] ND1.3-MTS-TALE-R-DddA C -UGI(SEQ ID NO. 180)
[0653] MASVLTPLLLRGLTGSARRLPVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDKMDIADLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKRGAGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPQQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPQQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPQQVVAIASHDGGRPALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML
[0654] ND6.2-MTS-TALE-L-DddA N -UGI(SEQ ID NO. 181)
[0655] MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAMDIADLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPQQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPQQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPQQVVAIASNIGGRPALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML
[0656] ND6.2-MTS-TALE-R-DddA C -UGI(SEQ ID NO. 182)
[0657] MASVLTPLLLRGLTGSARRLPVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDKMDIADLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKRGAGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPQQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPQQVVAIASHDGGRPALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML
[0658] ND3-MTS-TALE-L-DddA N -UGI(SEQ ID NO. 183)
[0659] MALSRAVCGTSRQLAPVLGYLGSRQKHSLPDYPYDVPDYAGYPYDVPDYAGYPYDVPDYAMDIADLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPQQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASHDGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQALLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPQQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPQQVVAIASNIGGRPALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML
[0660] ND3-MTS-TALE-R-DddA C -UGI(SEQ ID NO. 184)
[0661] MASVLTPLLLRGLTGSARRLPVPRAKIHSLDYKDHDGDYKDHDIDYKDDDDKMDIADLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKRGAGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPQQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPQQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNIGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPEQVVAIASNIGGKQLETVQRLLPVLCQAHGLTPQQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQALLPVLCQAHGLTPQQVVAIASHDGGRPALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLGGSAIPVKRGATGETKVFTGNSNSPKSPTKGGCSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML
[0662] The target sequence in the following embodiments and their accompanying drawings is as follows:
[0663] OsBADH2 target A chain in the figure
[0664] SEQ ID NO. 188
[0665] GCTGGATGCTTTGAGTACTTTGCAGATCTTGCAGAATCCTTGGACAAAAGGC
[0666] The OsBADH2 target B chain in the figure
[0667] SEQ ID NO. 189
[0668] CGACCTACGAAACTCATGAAACGTCTAGAACGTCTTAGGAACCTGTTTTCCG
[0669] The OsDEP1 target A chain in the figure
[0670] SEQ ID NO. 190
[0671] GCAAAAGACCAAGGTGCCTCAATTGTTCTTGCAGCTCATGCTGCGACGAGCC
[0672] The OsDEP1 target B chain in the figure
[0673] SEQ ID NO. 191
[0674] CGTTTTCTGGTTCCACGGAGTTAACAAGAACGTCGAGTACGACGCTGCTCGG
[0675] OsCKX2 target A chain in the figure
[0676] SEQ ID NO. 192
[0677] CCTGGACCGCGTCCACGACGGCGAGCTCAAGCTCCGCGCCGCGGGGCTCTGG
[0678] The OsCKX2 target B chain in the figure
[0679] SEQ ID NO. 193
[0680] GGACCTGGCGCAGGTGCTGCCGCTCGAGTTCGAGGCGCGGCGCCCCGAGACCC
[0681] Human ND6 target A chain in the figure
[0682] SEQ ID NO. 194
[0683] CCCCTGACCCCCATGCCTCAGGATACTCCTCAATAGCCATCGCTGTA
[0684] Human ND6 target B chain in the image
[0685] SEQ ID NO. 195
[0686] GGGGACTGGGGGGTACGGAGTCCTATGAGGAGTTATCGGTAGCGACAT
[0687] OsSD1 target A chain in the figure
[0688] SEQ ID NO. 196
[0689] CCAGGACGACGTCGGCGGCCTCGAGGTCCTCGTCGACGGCGAATGGCGCCCGTC
[0690] The OsSD1 target B chain in the figure
[0691] SEQ ID NO. 197
[0692] GGTCCTGCTGCAGCCGCCGGAGCTCCAGGAGCAGCTGCCGCTTACCGCGGGGCAG
[0693] SIRT6 target A chain in the figure
[0694] SEQ ID NO. 198
[0695] TACGCGGCGGGGCTGTCGCCGTACGCGGACAAGGGCAAGTGCGGCCTCCCGG
[0696] SIRT6 target B chain in the figure
[0697] SEQ ID NO. 199
[0698] ATGCGCCGCCCCGACAGCGGCATGCGCCTGTTCCCGTTCACGCCGGAGGGCC
[0699] In the figure, OsRbcL target A chain
[0700] SEQ ID NO. 200
[0701] TTACCAAAGATGATGAAAACGTAAACTCACAACCATTTATGCGTTGG
[0702] The OsRbcL target B chain in the figure
[0703] SEQ ID NO. 201
[0704] AATGGTTTCTACTACTTTTGCATTTGAGTGTTGGTAAATACGCAACC
[0705] The A chain of the ND6.2 target in the figure.
[0706] SEQ ID NO. 202
[0707] GACCCCCATGCCTCAGGATACTCCTCAATAGCCATCGCTGTAGTATATCCAA
[0708] The B chain of the ND6.2 target in the figure.
[0709] SEQ ID NO. 203
[0710] CTGGGGGGTACGGAGTCCTATGAGGAGTTATCGGTAGCGACATCATATAGGTT
[0711] ND1.2 target A chain in the figure
[0712] SEQ ID NO. 204
[0713] CCTATTTATTCTAGCCACCTCTAGCCTAGCCGTTTACTCA
[0714] The B chain of the ND1.2 target in the figure
[0715] SEQ ID NO. 205
[0716] GGATAAATAAGATCGGTGGAGATCGGATCGGCAAATGAGT
[0717] The A chain of the ND1.3 target in the figure
[0718] SEQ ID NO. 206
[0719] TCTCCACACTAGCAGAGACCAACCGAACCCCCTTCGACCTTGCCGAAGGGG
[0720] The B chain of the ND1.3 target in the figure
[0721] SEQ ID NO. 207
[0722] AGAGGTGTGATCGTCTCTGGTTGGCTTGGGGGAAGCTGGAACGGCTTCCCC
[0723] The ND3 target A chain in the figure
[0724] SEQ ID NO. 208
[0725] ACGAGTGCGGCTTCGACCCTATATCCCCCGCCCGCGTCCCTTTCTCCAT
[0726] The B chain of the ND3 target in the figure
[0727] SEQ ID NO. 209
[0728] TGCTCACGCCGAAGCTGGGATATAGGGGGCGGGCAGGGAAAGAGGTA
[0729] The A chain of ND1 target in the figure
[0730] SEQ ID NO. 210
[0731] CTAGCCTAGCCGTTTACTCAATCCTCTCATCAGGGTGAGCATCAAACTC
[0732] The B chain of the ND1 target in the figure
[0733] SEQ ID NO. 211
[0734] GATCGGATCGGCAAATGAGTTAGGAGACTAGTCCCACTCGTAGTTTGAG
[0735] The ND4 target A chain in the figure
[0736] SEQ ID NO. 212
[0737] GCTAGTAACCACGTTCTCCTGATCAAATATCACTCTCCTACTTACAGG
[0738] The B chain of the ND4 target in the figure
[0739] SEQ ID NO. 213
[0740] CGATCATGGTGCAAGAGGACTAGTTTATAGTGAGAGGATGAATGTCC
[0741] The A chain of the ND5.1 target in the figure
[0742] SEQ ID NO. 214
[0743] AGCATTAGCAGGAATACCTTTCCTCACAGGTTTCTACTCCAAAG
[0744] The B chain of the ND5.1 target in the figure
[0745] SEQ ID NO. 215
[0746] TCGTAATCGTCCTTATGGAAAGGAGTGTCCAAAGATGAGGTTTC
[0747] Example 1: Synthesis and Detection of Base Editor
[0748] The synthesis strategy of the base editor of the present invention is shown in Figure 1.
[0749] To verify the above strategy, target sites in the rice OsBADH2 gene were selected, and two sets of modified TALE coding vectors that can target these sites were constructed. The components are listed in Table 3.
[0750] Table 3. Specific examples of combinations of base editors in the embodiments.
[0751]
[0752] FokICD (or mutant) monomers were fused to the C-terminus of TALE-L and TALE-R, respectively, and wild-type FokI (without D450A or D467A mutations) was used as a control group. Figure 16A The application of two exonucleases (exonuclease I (mExoI) and exonuclease V (ExoV)) and one deaminase (hAPOBEC3A (hA3A or A3A)) in the novel base editor was evaluated, wherein in each group, UGI was fused to the carboxyl terminus of the deaminase having an XTEN spacer peptide. Figure 16B Nuclear localization signal (NLS, i.e., SV40 NLS in Table 2) is fused to the end of the protein.
[0753] Recombinant expression constructs encoding these components were transformed into rice protoplasts via PEG-mediated transformation. The corresponding constructs are shown in Figure 16A-. Figure 16B As shown in Figure 2A, different construct combinations were transformed into rice protoplasts targeting the OsBADH2 site, and next-generation sequencing (NGS) was used to determine the C>T base editing frequency. Sequencing results (Figure 2A) showed that for combinations containing FokI nickase, deaminase, exonuclease, and UGI, the targeted cytosine base editing frequency reached a maximum of approximately 10%. Importantly, the results also indicated that the novel nucleic acid base editor elicited only very low levels of insertion / deletion byproducts (Figure 2B), demonstrating the high product purity of the novel base editor, which is crucial for precise genome editing.
[0754] exist Figure 2A and Figure 2B The diagram below illustrates the experimental treatments or construct combinations and their associated carriers.
[0755]
[0756] Example 2: Characterization of the single-strand cleavage performance of the base editor
[0757] The base editing window of the base editor tested in Example 1 was analyzed. At the four C sites (C1, C6, C11, and C15) present in the A strand of the target gene, the first base adjacent to TALE-L in the spacer sequence of the two TALEs was counted as 1 (e.g., ...). Figure 3A As shown), C6 and C11 cytosine were effectively edited ( Figure 3B ).
[0758] exist Figure 3B The diagram below illustrates the experimental treatments or construct combinations and their associated carriers.
[0759]
[0760] These results show that FokI-R contains nickase The base editor (in which the FokI-L dimer nicking enzyme composed of FokI-L and FokI-R has a D450A or D467A mutation) tends to cleave the B strand with the nicking enzyme, and then the exonuclease digests the cleaved single strand, leaving a short ssDNA fragment in the A strand. The direction of digestion depends on the direction of enzyme catalysis of the exonuclease (5' to 3' or 3 to 5').
[0761] To verify the above results, the inventors evaluated the nucleic acid base editor at another site, OsDEP1, in this embodiment. This site contains five C bases (C1, C9, C13, C16, and C18) in the A strand. NGS analysis of rice protoplasts transformed with different construct combinations targeting the OsDEP1 site showed that the base editing window was mainly located near the 5' region (C9 and C1) of the A strand, while slight editing also occurred at C13 and C16 (e.g., ...). Figure 4A As shown in the figure, this is due to the instantaneous formation of the 3' flap structure after the incision. Importantly, similar to the OsBADH2 site, the OsDEP1 site showed only extremely low levels of insertion / deletion byproducts (such as...). Figure 4B (As shown in the figure). The above results indicate that the novel base editor has the advantage of high product purity.
[0762] exist Figure 4A and Figure 4B The diagram below illustrates the experimental treatments or construct combinations and their associated carriers.
[0763]
[0764] Example 3: The Influence of Exonuclease Digestion Direction and Cleavage Enzyme Single-Strand Preference on Editing Results
[0765] For those with Digestive-directed exonucleases (e.g., rat exonuclease I (mExoI)) expose cytosine residues near the 5' region of the target site on the complementary strand, leading to deamination by deaminases; while for 3' exonucleases, cytosine residues near the 3' region of the target site on the complementary strand are exposed and deaminated by deaminases. To verify that the base editor disclosed in this invention can produce the expected effect for different exonuclease digestive directions, the inventors simultaneously tested the 5' exonuclease mExoI and a 3' exonuclease—human Trex2 exonuclease—at the OsCKX2 target site, and analyzed the editing window of the obtained base editor by NGS. The experimental results showed that for FokI-R nickase In OsCKX2-mediated base editing, when the 5' exonuclease mExoI is used, the editing window is primarily located in the 5' region (C9 and C11) of the A strand at the target site; conversely, when the 3' exonuclease Trex2 is used, the editing window switches to the 3' neighboring region (C11 and C15) of the A strand at the OsCKX2 target site, while cytosine residues of the B strand are not edited (e.g., ...). Figure 5A , Figure 5B (As shown). Furthermore, the inventors evaluated the effect of the single-strand preference of the nicking enzyme used on single strands capable of base editing. FokI-R, which prefers to cleave the B strand... nickaseReplace with FokI-L, which prefers to cut the A chain. nickase As expected, the single strand undergoing base editing switched from strand A to strand B. Figure 5A Meanwhile, for the editing window, when using the 5' exonuclease mExoI, the editing window modifies the 5' adjacent regions (C6 and C8) of the B strand at the OsCKX2 target site. Correspondingly, when using the 3' exonuclease Trex2, the editing window can switch to the 3' adjacent regions (C3 and C6) of the B strand at the OsCKX2 target site, while the cytosine residues of the A strand are not edited. Figure 5A Therefore, the base editor of the present invention can be used with exonucleases of different digestive directions and exert the digestive effect of the corresponding exonucleases, thereby selectively editing target sites.
[0766] Different construct combinations were transformed into rice protoplasts targeting the OsCKX2 site. The C>T base editing efficiency and indel byproduct frequency were detected by NSG. Figure 5A and Figure 5B The diagram below illustrates the experimental treatments or construct combinations and their associated carriers.
[0767]
[0768] Example 4: Effect of Cytosine Deaminase Type
[0769] The novel base editor of this invention is independent of the type of deaminase and is compatible with different types of deaminases. To eliminate the dependence of the editing capability of the novel base editor on the deaminase hAPOBEC3A (A3A), the inventors tested another cytosine deaminase, rAPOBEC1 (APOBEC1), in this embodiment. NGS analysis results showed that in the co-existence of exonucleases, such as mExoI (e.g., Figure 6A (as shown) and Trex2 (as shown) Figure 6B In the case shown, after replacing hAPOBEC3A with rAPOBEC1 at the OsBADH2 site, targeted base editing with high product purity was also achieved, indicating that different types of deaminases are suitable for the base editor of this invention.
[0770] exist Figure 6A In this study, different construct combinations were transformed into rice protoplasts targeting the OsBADH2 site. The C>T base editing efficiency and indel byproduct frequency were detected by NGS. The experimental treatments or construct combinations and their associated vectors are illustrated in the diagram below.
[0771]
[0772] exist Figure 6B In this study, different construct combinations were transformed into rice protoplasts targeting the OsDEP1 site. The C>T base editing efficiency and indel byproduct frequency were detected by NGS. The experimental treatments or construct combinations and their associated vectors are illustrated in the diagram below.
[0773]
[0774] When analyzing the editing windows of these base editors, cytosine residues located near the 5' region of the complementary strand target site of the single strand where cleavage occurs were effectively edited in groups containing mExoI (e.g., Figure 7A (As shown); while cytosine residues located near the 3' region of the target site in the complementary strand are effectively edited in groups containing TREX2 (e.g. Figure 7B As shown in the figure, the results are consistent with those in the above embodiments. These results indicate that the base editing method and base editor disclosed in this invention are compatible with different cytosine deaminases.
[0775] exist Figure 7A In the analysis of the base editing window of the base editor based on the NGS results, the schematic diagram of the experimental treatments or construct combinations and associated vectors involved in the figure is shown below:
[0776]
[0777] exist Figure 7B In the analysis of the base editing window of the base editor based on the NGS results, the schematic diagram of the experimental treatments or construct combinations and associated vectors involved in the figure is shown below:
[0778]
[0779] Example 5: Base Editor Containing Adenine Deaminase
[0780] To expand the range of editable target sequences for the base editor of this invention, this embodiment utilizes an adenine deaminase, TadA-8e, which uses deoxyadenine (A) in single-stranded DNA as a substrate, to target A1, A7, A12, and A13 (e.g., at the OsCKX2 site). Figure 8 (As shown). In this embodiment, UGI is not a necessary component of the base editor under test because it is not required for adenine base editing. Analysis of the adenine base editing window of the base editor based on NGS results showed that an effective A-to-G conversion occurred at the target (as shown). Figure 8This indicates that the base editor described in this invention is compatible with adenine deaminase and can be used for adenine base editing. Combined with Examples 4 and 5, it can be seen that the base editing method and base editor disclosed in this invention are compatible with different deaminases and can perform their corresponding editing functions.
[0781] exist Figure 8 The diagram below illustrates the experimental treatments or construct combinations and their associated carriers.
[0782]
[0783] Example 6: Base editor of a fusion protein containing base editing components
[0784] Having demonstrated the functionality and effectiveness of the base editor of the present invention in the above embodiments, this embodiment verifies whether conversion efficiency (and correspondingly editing efficiency) can be improved by integrating modular components into a single vector. Two examples of the structure of such a base editor containing integrated components are shown below. Figure 9 As shown, the exonuclease is fused to the amino terminus of the deaminase-UGI fusion protein via an XTEN spacer peptide or a 48-amino acid spacer peptide (48aa) to target the OsDEP1 gene, which is to say, the deaminase is fused with the exonuclease.
[0785] Different construct combinations were transformed into rice protoplasts targeting the OsDEP1 site. NGS analysis was used to detect C>T base editing efficiency and the frequency of insertion / deletion (indel) byproducts. NGS analysis showed that targeted base editing could be achieved using a fusion of exonuclease and deaminase, but the efficiency in this vector structure was close to that of expressing exonuclease and deaminase alone (e.g., ...). Figure 10A (As shown). When using this base editor, the editing window is more inclined towards C1 and C9 (as shown). Figure 10B As shown in the figure, this is consistent with the catalytic direction of the mExoI exonuclease.
[0786] Figure 10A and Figure 10B The diagram below illustrates the experimental treatments or construct combinations and their associated carriers.
[0787]
[0788] In addition, the inventors tested other fusion protein structures. The structure of the aforementioned base editor is shown in... Figure 11A and Figure 11B In this process, deaminases (hAPOBEC3A or rAPOBEC1) are separated from TALE-L by a 48-amino acid spacer peptide. Figure 11A ) or TALE-R ( Figure 11BThe amino terminus of the protein is fused, while UGI and exonuclease are expressed by separate vectors, namely deaminase, TALE protein, and cleavage enzyme.
[0789] Integration.
[0790] For deaminase-TALE-FokI-R nickase OsDEP1 was selected as the target gene for characterization (e.g., Figure 12A As shown), while for deaminase-TALE-FokI-L nickase OsCKX2 was selected as the target gene for characterization (e.g., Figure 12B (As shown). NGS analysis showed that the two deaminases -TALE-FokI-L / R nickase Both achieved the C-to-T conversion at the target site. This indicates that deaminases can form fusion complexes with TALE proteins and nickases without interfering with their respective functions. Furthermore, the experimental results further demonstrate that base editing (such as...) can occur using both deaminases hAPOBEC3A and rAPOBEC1. Figure 12A and Figure 12B (As shown).
[0791] exist Figure 12A The diagram below illustrates the experimental treatments or construct combinations and their associated carriers.
[0792]
[0793] exist Figure 12B The diagram below illustrates the experimental treatments or construct combinations and their associated carriers.
[0794]
[0795] To investigate the effects of UGI or exonuclease fusion, a deaminase with the same target specificity as this invention, TALE-FokI-R, was used. nickase In the construct, its base editor is in FokI-L D450A The carboxyl terminus (e.g.) Figure 13A (as shown) or at the amino terminus of deaminase (such as...) Figure 13B (As shown) UGI is linked via a spacer peptide of 48 amino acids or 4 amino acids. NGS analysis indicated that linking UGI to the fusion protein was similar in effect to expressing UGI alone. Figure 14 Furthermore, in deaminase-TALE-FokI-R nickase In the construct, implementations in which the exonuclease is fused to the carboxyl terminus of FokI-R via a spacer peptide of 4, 16, or 48 amino acids also exhibit similar editing efficiency. Figure 14Therefore, both expressing UGI / exonuclease alone and fusing it into a vector for co-expression are technical solutions that can be adopted in this invention.
[0796] exist Figure 14 In this study, different construct combinations were transformed into rice protoplasts targeting the OsDEP1 site. High-throughput sequencing results were used to analyze the DNA strands and editing windows involved in base editing. The experimental treatments or construct combinations and their associated vectors are illustrated in the figure below:
[0797]
[0798] Based on the above results, each modular component of the base editor of the present invention can be expressed individually or can form one or more fusion proteins with each other.
[0799] Example 7: Base Editing of Plant Nuclear Genome
[0800] The above embodiments verified the function and characteristics of the base editor of the present invention, namely, that the composition containing modular components of deaminase, exonuclease, nicking enzyme, and DNA-binding protein TALE can achieve efficient and precise editing of DNA. For ease of description, the above base editor is named DENT (Deaminase-Exonuclease-Nickase-TALE), and named CyDENT (Cytidine Deaminase-Exonuclease-Nickase-TALE) and AdDENT (Adenine Deaminase-Exonuclease-Nickase-TALE) respectively according to the type of deaminase. In this embodiment, the applicable environment and scenarios of the base editor of the present invention are analyzed.
[0801] The inventors selected rice protoplast nuclear genomes to evaluate the editing effect of the base editor of this invention. In this embodiment, four pairs of TALE proteins were designed targeting the rice endogenous gene sites OsDEP1, OsCKX2, OsBADH2, and OsSD1. Exonucleases with 5'→3' (mExol) or 3'→5' (Trex2) restriction enzyme preferences were used to evaluate the effect of combining exonucleases with nicking enzymes to generate ssDNA intermediates. In this embodiment, the highly efficient cytosine deaminase hAPOBEC3A (hA3A) was selected to deaminate the ssDNA intermediates, and a uracil glycosylase inhibitor (UGI) peptide was fused to its C-terminus to further improve editing efficiency by minimizing the impact of DNA base excision repair. Nuclear localization signals (NLS) were fused to the N-terminus of each component, thereby directly targeting the nuclear genome for editing. This combination of base editors targeting the nuclear genome is referred to herein as nuCyDENT, and an exemplary construct diagram is shown below. Figure 18 As shown. nuCyDENT targeting the OsDEP1, OsCKX2, OsBADH2, and OsSD1 sites in rice protoplasts was introduced, and editing efficiency was assessed after 2 days. NGS analysis was used to evaluate targeted cytosine base editing within the 18 bp spacer region between the TALE binding sites at all four nuclear genomic sites. We observed editing efficiencies of 3–18% and lower indel frequencies compared to the corresponding wild-type TALEN systems. Figure 19A and Figure 19B These results demonstrate that the base editor of this invention can achieve efficient base editing in the nuclear genome while inducing only low-level indel byproducts.
[0802] Regarding single-strand editing performance, the inventors used nuCyDENT-L (containing FokI-L) at rice genome sites OsCKX2 and OsSD1, respectively. nickase nuCyDENT (structure) and nuCyDENT-R (containing Foki-R) nickase Base editing was performed using nuCyDENT (with a specific structure). The results showed that when nuCyDENT-R was used, the top strand of DNA was edited; when nuCyDENT-L was used, the bottom strand of DNA was edited. Figure 20 This conclusion is the same as in Example 2, also demonstrating the single-strand editing performance of CyDENT in the nuclear genome.
[0803] exist Figure 19A , Figure 19B , Figure 20 The experimental treatments or construct combinations involved in the figure are shown below:
[0804]
[0805] Example 8: Base Editing of Animal Nuclear Genome
[0806] This embodiment compares the base editing efficiency of CyDENT and DdCBE at the target site of the human SIRT6 gene. The inventors designed the TALE protein targeting SIRT6 and obtained nuCyDENT-L according to the method in Example 7. They also designed the DddA-dependent DdCBE according to existing methods (Nakazato, I. et al. Targeted base editing in the mitochondrial genome of Arabidopsis thaliana. Proc. Natl. Acad. Sci. USA. 119, e2121177119 (2022).). Experimental results show that nuCyDENT-L has a higher base editing efficiency at the target site than DdCBE. Figure 21 This indicates that the base editing system of the present invention has good base editing performance in animal cell nuclear genomes.
[0807] exist Figure 21 The experimental treatments or construct combinations involved in the figure are shown below:
[0808]
[0809] Example 9: Base Editing of DNA in Organelles—Chloroplasts
[0810] The base editor of this invention can be used for base editing of mitochondrial DNA and chloroplast DNA, offering advantages over CRISPR base editors that require nucleic acid components. The protein components in the base editor of this invention can be transferred to mitochondria and chloroplasts, respectively, via mitochondrial targeting sequences (MTS) and chloroplast transport peptides (CTPs). In these embodiments, MTS or CTP can be selected to replace NLS depending on the type of target organelle.
[0811] First, we attempted to edit the bases of plant chloroplast DNA using the CyDENT base editing strategy. Plant chloroplast DNA is an important organelle unique to plants, possessing its own genomic DNA (cpDNA), which cannot be edited using CRISPR-derived base editors. The inventors replaced NLS with chloroplast transport peptides (CTPs) in nuCyDENT, designed according to the method in Example 7 (Kang, BC et al. Chloroplast and mitochondrial DNA editing in plants. Nat Plants 7, 899-905 (2021).). Figure 22AThe inventors named it cpCyDENT. cpCyDENT-L (containing FokI-L) utilizes the TALE protein, which targets the large subunit gene (rbcL) of endogenous ribulose-1,5-bisphosphate carboxylase / oxygenase (RuBisCO), as its precursor. nickase ) and cpCyDENT-R (including Foki-R) nickase Transformation of rice protoplasts. Base editing at the rbcL target site was detected in cpCyDENT-L treatment. Figure 22B It is worth noting that by regulating the type and orientation of the nicking enzymes and exonucleases in cpCyDENT, precise editing of specific bases can be achieved. For example, for the G1 base (with the nucleotide at the 5'th nucleotide in the spacer region counted as position 1, see...), precise editing of specific bases can be achieved. Figure 22B Only when using the cpCyDENT-L(mExol) tool, which contains FokI-Lnickase and the 5'→3' mExol exonuclease, can this base be edited efficiently, with an editing efficiency of approximately 1.67%. This result is consistent with the conclusions of the aforementioned examples. These results demonstrate that cpCyDENT can selectively and precisely edit the DNA strands in the chloroplast genome.
[0812] exist Figure 22B The experimental treatments or construct combinations involved in the figure are shown below:
[0813]
[0814] Example 10: Base Editing of DNA in Organelles—Mitochondria
[0815] In this embodiment, the inventors evaluated the effect of CyDENT base editing on human mitochondrial DNA (mtDNA) base editing. Mitochondrial targeting sequences (MTS) were used instead of NLS, and suitable promoters and terminators for expression in HEK293T cells were selected, resulting in a base editor for mtDNA named mtCyDENT. The mtCyDENT construct generated in this embodiment is shown below. Figure 15A As shown (TALE-FokI-R) nickase and TALE-FokI-L nickase ).
[0816] First, a target site in the ND6 gene of human mitochondrial DNA was selected to construct TALE-FokI-R. nickase and TALE-FokI-L nickaseAn expression vector containing a TALE protein modified to target this site was transfected into HEK293T cells along with vectors expressing deaminases (hAPOBEC3A or C57), exonucleases (mExoI or Trex2), and UGI. The mitochondrial targeting sequence (MTS) was fused to the protein terminus. NGS was used to determine the base editing frequency after transfection with the base editor. Results showed that targeted cytosine base editing was achieved at an efficiency of approximately 6.0% in human mitochondrial DNA targets. Figure 15C The results show that the base editor of this invention can be used for base editing of organelle genomes.
[0817] exist Figure 15C In this study, different construct combinations were transfected into HEK293T cells targeting the mitochondrial ND6 site. High-throughput sequencing results were used to analyze the DNA strands and editing windows involved in base editing. The experimental treatments or construct combinations and their associated vectors are illustrated in the figure below:
[0818]
[0819] Example 11: The effect of base editor fusion state on mitochondrial DNA editing
[0820] Next, the inventors verified the effects of separately expressed deaminases, exonucleases, UGI, and TALE-FokI nickases on the base editing efficiency of mtDNA.
[0821] To this end, the inventors utilized a small peptide called γb fused to the N-terminus of the structural domain of one or more modular components in mtCyDENT to drive the recruitment of individual protein components. Figure 23A γb is an RNA silencing repressor derived from barley stripe mosaic virus (BSMV) and exhibits self-interaction (Jiang, Z., Yang, M., Zhang, Y., Jackson, AO & Li, D. in Encyclopedia of Virology 420-429 (2021).). In this experiment, the inventors chose Trex2 as the exonuclease. The inventors designed various fusion schemes of γb with different components to screen for the base editor composition with the best editing effect. Figure 23B Considering the size of the protein components entering the mitochondria, this embodiment adopts the following approach: Figure 23AThe five protein / fusion proteins shown were expressed using constructs: a fusion protein of TALE-L and FokI-L (abbreviated as TALE-L-FokI-L, TALEL-FL, or TALEL-FokI-L), a fusion protein of TALE-R and FokI-R (abbreviated as TALE-R-FokI-R, TALEL-FR, or TALER-FokI-R), an hA3A deaminase protein, a Trex2 exonuclease protein, and a UGI protein. The suffix D450A indicates the mutant, and WT indicates the wild type.
[0822] Experimental results show that γb has a higher editing effect when it is fused only with UGI and Trex2. The base editor composition with γb fusion structure with UGI and Trex2 respectively is named mtCyDENT1b.
[0823] Next, we evaluated mtCyDENT and mtCyDENT1b at seven additional endogenous mtDNA genomic loci. We observed that the average editing frequency of mtCyDENT was 1.16–11.7%, while mtCyDENT1b further improved the average editing efficiency by 2.42–6.18 times, reaching 4.55–39.3%. Figure 24 Furthermore, mtCyDENT1b exhibited higher editing efficiency than DdCBE at target sites with the same TALE sequence, specifically ND1.2, ND1.3, ND3, and ND6.2. Additionally, we noted a lower indel frequency compared to DdCBE when using CyDENT for base editing at mtDNA target sites. Figure 25 In summary, both mtCyDENT and mtCyDENT1b can achieve efficient base editing in human mitochondrial DNA.
[0824] exist Figure 23B The experimental treatments or construct combinations involved in the figure are shown below (from top to bottom):
[0825]
[0826] exist Figure 24-27 The experimental treatments or construct combinations involved in the figure are shown below:
[0827]
[0828]
[0829] Example 12: Improving CyDENT's editing efficiency and accuracy
[0830] As described in Example 4 above, the base editor of the present invention can be formed by the self-assembly of multiple functional modules and is compatible with different types of deaminases. Therefore, the deaminase domain in the base editor can be replaced with known deaminases in the prior art to utilize the unique characteristics of each deaminase, thereby enhancing activity or further improving the precision of the edited strand. A newly discovered single-stranded DNA (ssDNA)-specific cytidine deaminase, Sdd7, has been found to have higher editing activity than other deaminases (Huang, J. et al. Discovery of new deaminase functions by structure-based protein clustering. bioRxiv (2023).). In this example, the inventors used the mtCyDENT1b composition as an example, using Sdd7 as the deaminase of the editor, and evaluated the editing efficiency at mtDNA target sites ND5.1, ND6, and ND1.3. We observed that 87.5% of base edits induced by Sdd7-mtCyDENT1b-L occurred on only one DNA strand, and 93.0% of base edits induced by Sdd7-mtCyDENT1b-R occurred on only one DNA strand. These results further demonstrate that CyDENT possesses superior strand-specificity for base editing. Figure 26 The average editing efficiency of these two editors on the target strand at the bottom of the DNA layer ranged from 4.88% to 9.13%. Figure 27 These results further validate the ability of the base editor of this invention to replace deaminase domains during modular assembly.
[0831] Example 13: Improvements to the base editor
[0832] In the foregoing embodiments, the inventors experimentally verified that the base editor composition of the present invention possesses technical advantages such as single-strand editing specificity, modular assembly, efficient, precise, and controllable base editing, and low indel frequency. In subsequent embodiments, the inventors further optimized the base editor to obtain a base editor composition with even better functionality.
[0833] In this embodiment, the inventors fused the deaminase and exonuclease domains to the N-terminus of TALE-L and TALE-R using a 48-amino acid spacer peptide (flexible linker), and fused UGI to the C-terminus and N-terminus of FokI-L and FokI-R, respectively. This construct is referred to herein as mtCyDENT2. Figure 28A Use mtCyDENT2-L (containing Foki-L) nickase ) Detecting base editing effects on ND6 ( Figure 28BFurthermore, 94.5% of the base editing occurred only on the top strand, demonstrating the excellent single-strand specific editing capability of the CyDENT system.
[0834] exist Figures 28A-28B The experimental treatments or construct combinations involved in the figure are shown below:
[0835]
[0836] Example 14 For G C mtCyDNET base editing of motifs
[0837] Because the DddA-dependent DdCBE system has a strict T-condition for cytidine deamination. C The contextual constraints of motifs, and studies have found that in G C Editing in the sequence context occurs at a low frequency (Nakazato, I. et al. Targeted baseediting in the mitochondrial genome of Arabidopsis thaliana. Proc. Natl. Acad. Sci. USA. 119, e2121177119 (2022).). A phage-assisted discontinuous and continuous evolution was used to evolve “wild-type” DddA (Mok, BY et al. CRISPR-free base editors with enhanced activity and expanded targeting scope in mitochondrial and nuclear DNA. Nat. Biotechnol. 40, 1378-1387 (2022).), and the resulting DddA11 variant has better A C and C C Sequence motif compatibility is achieved, but editing DddA11 on GC sequence motifs remains challenging. This embodiment utilizes modular substitution of the deaminase domain of CyDENT to achieve efficient and chain-selective G... C Sequence motif editing.
[0838] The inventor introduced a G C A single-stranded DNA-specific cytidine deaminase with sequence motif editing activity was developed, thereby creating a G... C A compatible mtCyDENT base editor. Recently, a newly discovered single-stranded DNA-specific, G... C and A C Compatible cytidine deaminase Sdd3, which in G CIt exhibits higher editing activity on sequence motifs than other deaminases (Huang, J. et al. Discovery of new deaminase functions by structure-based protein clustering. bioRxiv (2023).).
[0839] Therefore, this invention designs a TALE array targeting ND1.2 and ND6.2 sites in HEK293T cells. Figure 29 This study aimed to assess the editing preference for sequence motifs that are difficult to edit using existing techniques. Notably, G at ND1.2 and ND6.2 sites... C The chain-specific cytosine base editing efficiencies observed at the sequence motifs reached 21.0% and 20.0%, respectively, which are editing efficiencies that existing DdCBE technologies cannot achieve at the same target sites; at the ND1.2 site, 96.9% of the editing selectively occurred on the top DNA strand, while at the ND6.2 site, 92.0% of the editing selectively occurred on the bottom DNA strand. Figure 29 ).
[0840] Subsequently, the inventors adjusted the TALE binding site and observed that Sdd3-mtCyDENT had an editing efficiency of 2.06% at the ND6.2 site. Figure 30 This specific mutation (m.14453G>A) has been reported to be directly associated with the development of Leigh syndrome, and existing DdCBE technology cannot achieve editing in this same target sequence background. Therefore, mtCyDENT and future optimized products could be used as a superior base editing method for precise editing of pathogenic mutations on mtDNA.
[0841] exist Figure 29 , 30 The experimental treatments or construct combinations involved in the figure are shown below:
[0842]
[0843] Example 15: Off-target analysis of mtCyDENT
[0844] Existing DdCBE technology for mitochondrial editing can induce a large number of nuclear off-target edits. To evaluate the off-target rate of CyDENT in the whole nuclear and mitochondrial genomes, this embodiment obtained 2.25 Tb of pure bases, with an average base count of 281.13 Gb per sample. The average depth of mitochondrial genome sequencing was approximately 6362-fold, and the human reference genome used was hg19.
[0845] In this embodiment, DdCBE plasmid and mtCyDENT1b-R (hA3A) plasmid targeting ND3, and mtCyDENT2-L (Sdd3) plasmid targeting ND6.2 were designed and transfected into HEK293T cells. Whole-genome sequencing (WGS) and NGS analysis showed that these plasmids could be transfected into HEK293T cells. C Editing on the upper sequence motif ( Figure 31A The off-target rates of the whole mitochondrial and nuclear genomes were then analyzed. The results showed that the average C·G-to-T·A and G·C-to-A·T base transition frequencies in the whole mitochondrial genome of the untreated negative control group, the DdCBE, mtCyDENT1b-R (hA3A), and the mtCyDENT2-L (Sdd3) treatment groups were 4.8%, 6.9%, 16.5%, and 5.9%, respectively. Compared with the control group, we found an average of 32, 678, and 16 single nucleotide variants (SNVs) in the mitochondrial genomes of the DdCBE, mtCyDENT1b-R (hA3A), and mtCyDENT2-L (Sdd3) treatment groups, respectively. By analyzing the upstream and downstream 5 bp regions of each potential off-target SNV, conserved TC motifs were found in the DdCBE and mtCyDENT1b-R(hA3A) groups, while conserved GC / AC motifs were found in the mtCyDENT2-L (Sdd3) group. Figure 31B ).
[0846] In the nuclear genome, the inventors analyzed TALE-dependent off-target effects. A total of 74,963 potential off-target regions were identified (including 0-3 regions that do not match the ND3 and ND6.2 TALE binding sites). We observed no difference in allelic SNV frequencies and indel frequencies at the ND3 or ND6.2 sites among the control group, DdCBE, mtCyDENT1b-R (hA3A), and mtCyDENT2-L (Sdd3) treatment groups. Figure 31C These results demonstrate that modular assembly and optimization of CyDENT can minimize off-target effects in both the mitochondrial and nuclear genomes. mtCyDENT is a valuable tool for mitochondrial genome editing.
[0847] exist Figures 31A-31C The experimental treatments or construct combinations involved in the figure are shown below:
[0848]
[0849] The above description of the embodiments is only for the purpose of helping to understand the method and core ideas of the present invention. It should be noted that those skilled in the art can make several improvements and modifications to the present invention without departing from the principles of the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.
Claims
1. A nucleic acid base editor for organelles, characterized in that, The nucleic acid base editor includes the following components: a) Sequence-specific DNA-binding proteins; b) Cutting enzyme; c) Exonucleases; d) Base-specific deaminases; and e) Organelle-targeting signaling sequences; The sequence-specific DNA-binding protein is the TALE protein; The cleavage enzyme is a dimer of a FokI cleavage domain monomer or a mutant thereof, wherein the dimer or mutant thereof consists of a pair of interacting FokI cleavage domain monomers, and only one of the FokI cleavage domain monomers in the dimer or mutant thereof has DNA endonuclease activity. The cleavage domain monomers of FokI with DNA endonuclease activity are selected from FokI-L protein with amino acid sequence as shown in SEQ ID No. 87 and FokI-R protein with amino acid sequence as shown in SEQ ID No. 88; The cleavage domain monomer of FokI, which has lost its DNA endonuclease activity, is selected from FokI-L with the amino acid sequence shown in SEQ ID No.
60. D450A Protein, selected from FokI-L with an amino acid sequence as shown in SEQ ID No.
61. D467A Protein, selected from FokI-R with an amino acid sequence as shown in SEQ ID No.
62. D450A Proteins or amino acid sequences as shown in SEQ ID No. 63, FokI-R D467A protein; The base-specific deaminase is selected from cytosine-specific deaminase or adenine-specific deaminase; The exonuclease is selected from 5' exonuclease or 3' exonuclease; The organelle targeting signal sequence is selected from the mitochondrial targeting sequence MTS or the chloroplast transport peptide CTP.
2. The nucleic acid base editor according to claim 1, characterized in that, The components of the nucleic acid base editor exist independently or form one or more fusion proteins.
3. The nucleic acid base editor according to claim 1 or 2, characterized in that, The amino acid sequence of the base-specific deaminase is selected from SEQ ID NO. 36-59, 80-86.
4. The nucleic acid base editor according to claim 1 or 2, characterized in that, The base-specific deaminase is a cytosine-specific deaminase.
5. The nucleic acid base editor according to claim 4, characterized in that, The cytosine-specific deaminase is one or more of hAPOBEC3A, rAPOBEC1, hAID, pmCDA1, or Sdd deaminase.
6. The nucleic acid base editor according to claim 4, characterized in that, The nucleic acid base editor also contains: f) Uracil glycosylation inhibitor (UGI); Furthermore, the uracil glycosylation enzyme inhibitor exists alone or forms at least one fusion protein with other nucleic acid base editor components.
7. The nucleic acid base editor according to claim 1 or 2, characterized in that, The base-specific deaminase is an adenine-specific deaminase.
8. The nucleic acid base editor according to claim 7, characterized in that, The adenine-specific deaminase is TadA-8e.
9. The nucleic acid base editor according to any one of claims 1, 2, 5-6, and 8, characterized in that, The nucleic acid base editor also contains: g) γb; The γb, together with other nucleic acid base editor components, forms at least one fusion protein.
10. The nucleic acid base editor according to claim 1 or 2, characterized in that, The amino acid sequence of the exonuclease is selected from SEQ ID NO.64-67, 153.
11. A fusion protein as a nucleic acid base editor, characterized in that, The fusion protein contains the protein domain of the base editor as described in any one of claims 1-10.
12. A recombinant expression construct for nucleic acid base editing, characterized in that, The recombinant expression construct is used to express the nucleic acid base editor of any one of claims 1-10 or the fusion protein of claim 11.
13. A method for non-therapeutic nucleic acid base editing in cells, characterized in that, The nucleic acid base editor according to any one of claims 1-10 or the recombinant expression construct according to claim 12 is introduced into cells to edit the target gene.
14. The method for nucleic acid base editing according to claim 13, characterized in that, The target gene is selected from mitochondrial genomic DNA or chloroplast genomic DNA.
15. The use of the base editor of any one of claims 1-10, the fusion protein of claim 11, or the recombinant expression construct of claim 12 in the preparation of reagents for base editing of DNA in cells, wherein the cells are mammalian cells, bacteria, protists, fungi, insect cells, or plant cells.
16. The application according to claim 15, characterized in that, The plant cells are derived from the whole plant of a monocotyledonous or dicotyledonous plant.
17. The application according to claim 15, characterized in that, The plant cells are derived from seedlings of monocotyledonous or dicotyledonous plants.
18. The application according to claim 15, characterized in that, The plant cells are derived from the meristematic tissue, ground tissue, vascular tissue, or cortical tissue of monocotyledonous or dicotyledonous plants.
19. The application according to claim 15, characterized in that, The plant cells are derived from the seeds, leaves, roots, buds, stems, flowers, or fruits of monocotyledonous or dicotyledonous plants.
20. The application according to claim 15, characterized in that, The plant cells are derived from the stolons, bulbs, tubers, corms, buds, or asexual terminal branches of monocotyledonous or dicotyledonous plants.
21. The application according to claim 15, characterized in that, The plant cells are derived from tumor tissue of monocotyledonous or dicotyledonous plants.
22. The application according to claim 15, characterized in that, The mammalian cells were selected from human somatic cells.
23. The application according to claim 15, characterized in that, The mammalian cells are selected from human neurons, muscle cells, endocrine / exocrine cells, epithelial cells, hematopoietic cells, and bone cells.
24. The application according to claim 15, characterized in that, The mammalian cells were selected from human induced pluripotent stem cells.
25. The application according to claim 15, characterized in that, The mammalian cells were selected from human tumor cells.
26. The application according to claim 15, characterized in that, The fungus is selected from yeast or unconventional yeast.
27. A pharmaceutical composition for treating a disease in a person in need, characterized in that, The pharmaceutical composition comprises the base editor of any one of claims 1-10, the fusion protein of claim 11, or the recombinant expression construct of claim 12.
28. The pharmaceutical composition according to claim 27, characterized in that, The pharmaceutical composition also includes a pharmaceutically acceptable carrier.
29. A method for producing genetically modified plants, characterized in that, The method includes introducing the base editor of any one of claims 1-10, the fusion protein of any one of claims 11, or the recombinant expression construct of claim 12 into at least one of the plants.
Citation Information
Patent Citations
Delivery, use and therapeutic applications of the crispr-CAS systems and compositions for HBV and viral diseases and disorders
WO2015089465A1
Novel crispr enzymes and systems
WO2016205711A1
Compounds, compositions and methods for cancer treatment
WO2018141835A1
Uses of adenosine base editors
WO2019079347A1
Methods and compositions for editing nucleotide sequences
WO2020191233A1