A DNA glycosylase composition and uses thereof

CN118995665BActive Publication Date: 2026-08-18EIGHTH AFFILIATED HOSPITAL SUN YAT SEN UNIV (SHENZHEN FUTIAN)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410956946.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-16
Publication Date
2026-08-18
Estimated Expiration
2044-07-16

AI Technical Summary

Technical Problem

由于Cas9切口酶的引导RNA(gRNA)难以递送到真核生物的线粒体中,这导致gGBE无法使用到线粒体的碱基编辑中

Benefits of technology

[0052] The DNA glycosylation enzyme composition of this application, when used for guanine base editing of plant or animal mitochondrial genomes, has the following advantages: first, it does not rely on deaminases, avoiding non-target editing caused by deaminases; second, mitochondrial genome G base editing has DNA strand selectivity, increasing the accuracy of base editing; and third, the compact single fusion protein, as a G base editor, reduces the technical difficulty of delivering the base editor to the target cell.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118995665B_ABST
    Figure CN118995665B_ABST
Patent Text Reader

Abstract

The application discloses a DNA glycosylase composition and application thereof. The glycosylase composition is a monomer fusion protein or a double fusion protein. The monomer fusion protein sequentially fuses a transport peptide, an exonuclease, a TALE, a nicking enzyme and a guanine DNA glycosylase. The double fusion protein includes an L fusion protein and an R fusion protein. The L fusion protein sequentially fuses a transport peptide, a TALE-L, a nicking enzyme and an exonuclease. The R fusion protein sequentially fuses a transport peptide, a TALE-R and a guanine DNA glycosylase. All the TALEs are composed of an N-terminal functional domain, an array of repeat units and a C-terminal functional domain. The array of repeat units is a specific binding sequence designed for a target sequence. The glycosylase composition is used for mitochondrial G base editing, does not depend on a deaminase, avoids non-target editing caused by the deaminase, has DNA strand selectivity, increases base editing accuracy, has a compact structure and reduces technical difficulty of delivery to target cells.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of base editor technology, and in particular to a DNA glycosylation enzyme composition and its application. Background Technology

[0002] Mitochondria are the energy and metabolic centers of cells, playing crucial roles in catabolism, respiration, regulation of cell signaling pathways, and apoptosis. Mitochondria are semi-autonomous organelles containing their own genomic DNA (mtDNA). Human mtDNA typically has 2–10 copies per cell, with each mtDNA cell approximately 16.5 kb in length, containing 37 genes encoding 12S rRNA, 16S rRNA, 22 tRNAs, and 13 proteins. Plant mitochondrial genomes are generally between 66 kb and 11 Mb in size, containing 50–60 genes, and are inherited maternally from one parent. Comparatively, plant mtDNA is much larger than that of other eukaryotes and has evolved more rapidly in structure.

[0003] Due to factors such as reactive oxygen species, the mutation rate of human mtDNA is two orders of magnitude higher than that of nuclear DNA. mtDNA mutations lead to impaired mitochondrial ATP synthesis, insufficient cellular energy supply, and metabolic enzyme deficiencies, ultimately resulting in bodily dysfunction. Dozens of genetic diseases, including mitochondrial myopathy, are known to be associated with mtDNA mutations. In plants, mtDNA genetic information influences important agronomic traits, including fertility, plant vigor, chloroplast function, and hybridization compatibility. Editing and modifying mtDNA sequences has significant practical implications, whether for repairing human mtDNA mutations, treating diseases caused by mitochondrial mutations, or modifying plant mtDNA genes to improve important agronomic traits.

[0004] In recent years, genome editing technology has rapidly evolved, with the development of techniques such as base editing and prime editing. Base editors enable highly efficient and precise single nucleotide substitutions and are typically composed of a DNA-binding module and a DNA-modifying enzyme. The DNA-binding module is generally a Cas nickase (e.g., Cas9 D10A nickase) or a transcription activator-like effector (TALE) array. The DNA-binding module carries the DNA-modifying enzyme, which targets the DNA at the target site. The enzyme then modifies the sequence at the target site, ultimately obtaining the genetic variation of interest. Based on the type of DNA-modifying enzyme, two types of base editors have been developed: deaminase-based base editors (dBE) and deaminase-free glycosylation enzyme-based base editors (gBE). dBE utilizes single-stranded or double-stranded DNA deaminases for base editing, including adenine base editors (ABE), cytosine base editors (CBE), and DddA-derived cytosine base editors (DdCBE), as well as their derivative systems (A&C-BEmax, AYBE, AXBE, and CGBE, etc.). gBE utilizes glycosylases for base editing, including thymine base editors (gTBE) based on deaminase-free glycosylases and cytosine base editors (gCBE) based on deaminase-free glycosylases.

[0005] Due to the unique characteristics of mitochondrial membranes and mtDNA repair mechanisms, mitochondrial genome editing technology faces numerous challenges, resulting in a relative scarcity of base editors suitable for the mitochondrial genome. For example, the CRISPR / Cas gene editing system, widely used in nuclear genomes, is difficult to use for mitochondrial genome editing because its guide RNA (gRNA) component is difficult to deliver to eukaryotic mitochondria. Currently, mitochondrial genome base editors consist of transcription activator-like effector (TALE) arrays fused with DNA-modifying enzymes, including DdCBE and CyDENT cytosine base editors, TALED adenine base editors, and mitoBEs mitochondrial DNA base editors. These are all adenine and cytosine base editors based on deaminases. It has been reported that deaminases can cause numerous off-target mutations in cellular DNA and RNA; although deaminases are linked to Cas proteins or TALE arrays, they can also act on chromosomal DNA, inducing non-target base editing. Therefore, developing mitochondrial base editing systems without deaminases is of significant practical importance.

[0006] A recent paper reported a guanine base editor (gGBE) based on a deaminase-free glycosyltransferase that can directly edit G bases. It exhibited high base editing efficiency (up to 81%) in both mouse and cultured human cells, and a G-to-T / C (G-to-Y) conversion rate as high as 95% in embryos. In transgenic rice plants, it achieved G-to-T base conversion with an efficiency of 39.1%. This gGBE was developed by fusing the Cas9 nickase (nCas9) with an engineered N-methylpurine DNA glycosyltransferase protein (MPG) mutant. Because the guide RNA (gRNA) of the Cas9 nickase is difficult to deliver to eukaryotic mitochondria, gGBE cannot be used for mitochondrial base editing. A previous research team reported using the cytosine deaminase DddA... E1347A The mutant and MPG established a mitochondrial genome guanine base editor. However, this mitochondrial editor still used the inactivated cytosine deaminase DddA. E1347A The mutant, and the editor edits both G bases of the target DNA double strand, rather than editing specific single strands.

[0007] In summary, to avoid non-target base editing that may be introduced by deaminases, improve the accuracy of base editing, and reduce the technical difficulty of vector delivery, the development of a deaminase-independent, DNA strand-selective, and compact mitochondrial genome guanine base editor remains a key research focus and challenge in the field of base editor technology. Summary of the Invention

[0008] The purpose of this application is to provide a novel DNA glycosylation enzyme composition, an expression vector for expressing the DNA glycosylation enzyme composition, and the application of the DNA glycosylation enzyme composition and the expression vector.

[0009] The following technical solution is adopted in this application:

[0010] The first aspect of this application discloses a DNA glycosylation enzyme composition, which is a monomeric fusion protein or a dimeric fusion protein. The monomeric fusion protein sequentially fuses a transport peptide, a DNA exonuclease, a TALE array, a DNA nicking enzyme, and a guanine DNA glycosylation enzyme. The dimeric fusion protein includes an L fusion protein and an R fusion protein, wherein the L fusion protein sequentially fuses a transport peptide, TALE-L, a DNA nicking enzyme, and a DNA exonuclease, and the R fusion protein sequentially fuses a transport peptide, TALE-R, and a guanine DNA glycosylation enzyme. The transport peptide is used to guide the guanine base editor to the corresponding organelle to achieve G base editing. The TALE array, TALE-L, and TALE-R are all composed of an N-terminal functional domain, a repeat unit array, and a C-terminal functional domain, wherein the repeat unit array is a specific binding sequence designed for the target sequence. The TALE array in the monomeric fusion protein is typically the TALE-L in the L fusion protein of the dimeric fusion protein.

[0011] It should be noted that the DNA glycosylation enzyme composition of this application creatively combines a transport peptide, a TALE array, and engineered guanine DNA glycosylation enzyme (GDG), DNA nicking enzyme, and DNA exonuclease, enabling the constructed glycosylation enzyme composition to be used for guanine base editing in the mitochondrial genomes of various animals and plants. Using the DNA glycosylation enzyme composition of this application for mitochondrial genome guanine editing offers several advantages: first, it is independent of deaminases, avoiding non-target editing that may be caused by deaminases; second, its G base editing exhibits DNA strand selectivity, increasing the accuracy of base editing; and third, the compact single fusion protein provided in this application, acting as a G base editor, reduces the technical difficulty of delivering the base editor to the target cell.

[0012] It should also be noted that the DNA glycosylation enzyme composition of this application does not depend on deaminase and has high editing efficiency and low off-target effect. It is especially suitable for guanine base editing systems of eukaryotic organelle genomes such as mitochondria, realizing base editing of organelle genomes from G to Y, i.e., G to T or G to C, avoiding off-target effects that may be caused by deaminase and improving the accuracy of G base editing.

[0013] It is understood that the DNA glycosylation enzyme composition of this application, firstly, utilizes transport peptides and TALE arrays to guide the base editor into the organelle genome target site; secondly, it avoids the use of deaminases, reducing potential off-target effects caused by deaminases; and thirdly, it utilizes DNA nicking enzymes and engineered guanine DNA glycosylation enzymes to achieve G-to-Y base editing of organelle genomes, enriching the base editing techniques for organelle genomes in higher organisms. When used as a guanine base editor system, the glycosylation enzyme composition established in this application can be applied to organelle gene function research in animal, plant, and human cells, and can also be applied to agricultural bio-breeding scenarios such as modifying crop organelle genes to improve agronomic traits.

[0014] In one implementation of this application, the transport peptides of the monomeric fusion protein and the dimorphic fusion protein are plant mitochondrial transport peptides or animal mitochondrial transport peptides.

[0015] It can be understood that when the transport peptide is a plant mitochondrial transport peptide, the DNA glycosylation enzyme composition of this application is a guanine base editor targeting the plant mitochondrial genome; when the transport peptide is an animal mitochondrial transport peptide, the DNA glycosylation enzyme composition of this application is a guanine base editor targeting the animal mitochondrial genome.

[0016] In one implementation of this application, the plant mitochondrial transport peptide is derived from the rice Rf1b cytoplasmic male sterility restoration gene transport peptide.

[0017] In one implementation of this application, the animal mitochondrial transport peptide is a human mitochondrial transport peptide, which is a mitochondrial transport peptide derived from the human COX8A gene.

[0018] In one implementation of this application, the plant mitochondrial transport peptide is the sequence shown in SEQ ID NO.1.

[0019] In one implementation of this application, the animal mitochondrial transport peptide is the sequence shown in SEQ ID NO.2.

[0020] It should be noted that the key to this application lies in the creative combination of transport peptides, TALE arrays, and engineered GDG, DNA nickases, and DNA exonucleases to construct a base editing system suitable for guanine in the genomes of higher biological organelles. The specific transport peptides shown in SEQ ID NO.1 and SEQ ID NO.2 are merely one implementation of this application. It is understood that other specific transport peptide sequences may be used within the inventive concept of this application, and no specific limitations are made here.

[0021] In one implementation of this application, the N-terminal functional domain of the monomeric fusion protein and the dimeric fusion protein is the sequence shown in SEQ ID NO.3.

[0022] In one implementation of this application, the C-terminal functional domains of the monomeric fusion protein and the dimeric fusion protein are the sequences shown in SEQ ID NO.4.

[0023] It should be noted that in the TALE array, the N-terminal functional domain (NTD) and C-terminal functional domain (CTD) are fixed amino acid sequences, as shown in SEQ ID NO.3 and SEQ ID NO.4. The repeat unit array is designed and assembled based on the sequence of the target editing site. The repeat unit array is assembled using the Golden Gate method, and the assembly of TALE repeat units can be outsourced to a commercial company. It is understood that the sequences shown in SEQ ID NO.3 and SEQ ID NO.4 are only specific N-terminal and C-terminal functional domains used in one implementation of this application. Those skilled in the art can adjust the specific sequences with reference to existing technologies, and no specific limitations are made here.

[0024] In one implementation of this application, the DNA nickase for the monomeric fusion protein and the dimeric fusion protein is BspD6I nickase, MutH nickase, or Fok1-Fok1 nickase. D450A Cutting enzyme.

[0025] In one implementation of this application, the BspD6I nickase is the sequence shown in SEQ ID NO.5.

[0026] In one implementation of this application, the MutH nickase is the sequence shown in SEQ ID NO.6.

[0027] In one implementation of this application, Fok1-Fok1 D450A The nicking enzyme has the sequence shown in SEQ ID NO.7.

[0028] In one implementation of this application, the DNA exonuclease for the monomeric fusion protein and the dimeric fusion protein is either Trex2 exonuclease or mExoI exonuclease.

[0029] In one implementation of this application, the Trex2 exonuclease is the sequence shown in SEQ ID NO.8.

[0030] In one implementation of this application, the mExoI exonuclease is the sequence shown in SEQ ID NO.9.

[0031] In one implementation of this application, the DNA exonuclease of the monomeric fusion protein is connected to the TALE array via a 48aa adapter.

[0032] In one implementation of this application, the 48aa connector is the sequence shown in SEQ ID NO.10.

[0033] In one implementation of this application, the DNA nickase of the monomeric fusion protein and the guanine DNA glycosylation enzyme are connected by a 16aa linker.

[0034] In one implementation of this application, the DNA cleavage enzyme and the DNA exonuclease of the L fusion protein are connected via a 16aa adapter.

[0035] In one implementation of this application, the 16aa connector is the sequence shown in SEQ ID NO.11.

[0036] It should be noted that BspD6I nickase, MutH nickase, and Fok1-Fok1 are mentioned. D450A The cleavage enzyme, Trex2 exonuclease, and mExoI exonuclease are just some of the DNA cleavage enzymes and exonucleases with high editing efficiency used in one implementation of this application; it is understood that other known DNA cleavage enzymes and exonucleases may also be used. Similarly, the linker sequence between the DNA cleavage enzyme and the DNA exonuclease can be modified with reference to existing technologies, and is not limited to the 48aa linker shown in SEQ ID NO. 10 or the 16aa linker shown in SEQ ID NO. 11.

[0037] In one implementation of this application, the guanine DNA glycosylation enzyme is an engineered mutant of gMPGV6.3, with mutation information of G163R, N169G, D175R, C178N, S198A, K202A, G203A, S206A, K210A, and K294R.

[0038] In one implementation of this application, the guanine DNA glycosylation enzyme is the sequence shown in SEQ ID NO.12.

[0039] It should be noted that the engineered mutant of gMPGV6.3 can be found in the reference: Programmable deaminase-free base editors for G-to-Y conversion by engineered glycosylase https: / / doi.org / 10.1093 / nsr / nwad143. The key to this application lies in the inventive combination of engineered guanine DNA glycosylase with transport peptides, TALE arrays, DNA nicking enzymes, and DNA exonucleases to construct the glycosylase composition of this application for guanine base editing. It is understood that, within the inventive concept of this application, other engineered guanine DNA glycosylases with the same or similar functions may also be used, and no specific limitations are made here.

[0040] The second aspect of this application discloses an expression vector for expressing the DNA glycosylation enzyme composition of this application.

[0041] It is understood that when the DNA glycosylation enzyme composition of this application is a dimeric fusion protein, the two fusion proteins, L fusion protein and R fusion protein, are expressed separately, i.e., a vector expressing the L fusion protein and a vector expressing the R fusion protein. Of course, other methods of expressing the DNA glycosylation enzyme composition of this application are not excluded.

[0042] In one implementation of this application, when the DNA glycosylase composition is used for base editing of plant organelles, the transport peptide is a plant mitochondrial transport peptide, and the expression vector is driven by the rice Ubi promoter (rUbi) and / or the maize Ubi promoter (ZmUbi); when the DNA glycosylase composition is used for base editing of animal organelles, the transport peptide is an animal mitochondrial transport peptide, and the expression vector is driven by the human elongation factor EF-1α core promoter.

[0043] It should be noted that the DNA glycosylation enzyme composition of this application, depending on the specific transport peptide used, can be used for mitochondrial genome editing in plants or animals. It is understood that the specific promoters used for the plant and animal expression vectors are different. The rice Ubi promoter and the maize Ubi promoter are only specific promoters used for plant expression vectors in one implementation of this application; other promoters suitable for plant expression vectors may also be used. Similarly, the human elongation factor EF-1α core promoter is only a specific promoter used for animal expression vectors in one implementation of this application; other promoters suitable for animal expression vectors may also be used.

[0044] In one implementation of this application, the expression vector has a replicon that reproduces in Escherichia coli and / or Agrobacterium.

[0045] In one implementation of this application, the expression carrier has a screening marker.

[0046] In one implementation of this application, the screening marker is a kanamycin resistance gene or an ampicillin resistance gene.

[0047] The third aspect of this application discloses the use of the DNA glycosylation enzyme composition of this application or the expression vector of this application in the preparation of reagents for mitochondrial genome editing.

[0048] The fourth aspect of this application discloses the use of the DNA glycosylation enzyme composition of this application or the expression vector of this application in the preparation of breeding reagents or gene therapy reagents.

[0049] It should be noted that the DNA glycosylation enzyme composition or expression vector of this application can achieve guanine base editing of the genome of organelles such as plant mitochondria or animal mitochondria, that is, G to T or G to C base editing; therefore, it can not only be used to prepare mitochondrial genome editing reagents, such as for use as a mitochondrial genome editor; but also for the study of organelle gene function in animal, plant and human cells, for agricultural biological breeding such as modifying crop organelle genes to improve agronomic traits, and for gene therapy based on base editing.

[0050] It should also be noted that the guanine base editing system based on the DNA glycosylation enzyme composition of this application provides technical support for the clinical treatment of related mitochondrial diseases and repairs mtDNA mutations that cause mitochondrial diseases, such as (1) Leigh syndrome, which mostly occurs in childhood, with symptoms including motor developmental delay, symptoms and signs of brainstem and / or basal ganglia involvement, and is prone to respiratory failure; it can also occur later, and may include cerebellar ataxia, dementia, hearing loss, vision loss, cardiomyopathy, which is caused by the mtDNA mutation m.8993T>G. (2) Neurogenic myasthenia gravis, ataxia, retinitis pigmentosa (NARP), which often begins in childhood to early adulthood, causing myopathy, peripheral neuropathy, ataxia, retinitis pigmentosa, epilepsy, dementia, and may also include migraine and mental retardation, which is caused by the mtDNA mutation m.8993T>G / C. (3) Myoclonic epilepsy with fragmented red fibers (MERRF) is a disease phenotype associated with epilepsy, often accompanied by ataxia. It exhibits both myoclonic epilepsy and fragmented red fiber phenotypes and is usually caused by mutations in the mtDNA gene encoding lysine-containing tRNAs. Approximately 90% of MERRF cases are caused by mutations in the MT-TK gene m.8344A>G, m.8356T>C, m.8363G>A, and m.8361G>A.

[0051] The beneficial effects of this application are as follows:

[0052] The DNA glycosylation enzyme composition of this application, when used for guanine base editing of plant or animal mitochondrial genomes, has the following advantages: first, it does not rely on deaminases, avoiding non-target editing caused by deaminases; second, mitochondrial genome G base editing has DNA strand selectivity, increasing the accuracy of base editing; and third, the compact single fusion protein, as a G base editor, reduces the technical difficulty of delivering the base editor to the target cell. Attached Figure Description

[0053] Figure 1 This is a schematic diagram illustrating the glycosylation enzyme composition as an organelle genome guanine base editor in the embodiments of this application and its working principle.

[0054] Figure 2It is a schematic structural diagram of the supporting expression vector of the plant gMtGBE system in the embodiments of the present application;

[0055] Figure 3 It is a schematic structural diagram of the supporting expression vector of the human gMtGBE system in the embodiments of the present application. Detailed implementation manners

[0056] To avoid non-target base editing that may be introduced by deaminases, improve the accuracy of base editing, and reduce the technical difficulty of vector delivery technology, the purpose of the present application is to establish a mitochondrial genome guanine base editor with high editing efficiency and low off-target effects, which is applicable to eukaryotic mitochondrial genomes, does not rely on deaminases, has DNA strand selectivity, and is structurally compact, so as to achieve base editing of G to Y (i.e., G to T or G to C) in organelle genomes. This base editing system does not rely on deaminases, avoids non-target editing that may be caused by deaminases, and at the same time has DNA strand selectivity for base editing, improving the accuracy of base editing. (1) A functional complex is formed by using a DNA nickase, an engineered guanine-DNA glycosylase (GDG), a transit peptide, a TALE array polypeptide, etc., to achieve G to Y base editing of the mitochondrial genome, avoiding the use of deaminases and reducing the possible off-target effects caused by deaminases; (2) A DNA nickase and an exonuclease are used to generate strand-selective DNA nicks and naked single strands, to achieve strand-selective G to Y base editing of DNA, abbreviate the editing interval of the target site, and increase the accuracy of base editing; (3) A structurally compact single fusion protein G base editor is established to reduce the technical difficulty of delivering the base editor to target cells.

[0057] The present application has established a guanine base editor applicable to eukaryotic mitochondrial genomes, which does not rely on deaminases and has DNA strand selectivity, that is, the DNA glycosylase composition of the present application. The DNA glycosylase composition of the present application, that is, the guanine base editor, is named the glycosylase-based Mitochondia Guanine Base Editor (gmtGBE). gmtGBE contains the following key components: a transit peptide, a TALE array, a nickase, an exonuclease, and a guanine-DNA glycosylase (GDG), such as Figure 1As shown. Based on the combination patterns of the above components, glycosylation enzyme compositions can be divided into dimeric fusion proteins and monomeric fusion proteins. Dimeric fusion proteins (i.e., dimeric gmtGBE) are composed of an L fusion protein and an R fusion protein, while monomeric fusion proteins (monomeric gmtGBE) are single fusion proteins. Dimeric gmtGBE consists of an L fusion protein sequentially fused with a transport peptide, a TALE-L array, a nicking enzyme, and an exonuclease component, and an R fusion protein sequentially fused with a transport peptide, a TALE-R array, and a GDG component. Monomeric gmtGBE sequentially fused with a transport peptide, an exonuclease, a TALE array, a nicking enzyme, and a GDG component, as shown... Figure 1 As shown.

[0058] The basic working principle of gmtGBE described in this application is as follows: a transport peptide and a TALE array carry the nicking enzyme, exonuclease, and GDG components into the mitochondria and locate them at the base editing target site; the nicking enzyme fuses with the TALE, recognizes one strand of the target site's DNA double helix, cuts that strand to create a nick, exhibiting strand selectivity; the exonuclease recognizes the nick region and digests the nicked DNA strand, thereby forming an exposed short single-stranded DNA (ssDNA) fragment on the other strand; GDG uses ssDNA as a substrate to excise the G base at the target editing site, generating apyrimidine- or apurine-free (AP) site, which is repaired through trans-damage synthesis (TLS) and / or DNA replication, resulting in G-to-Y base editing, such as... Figure 1 As shown.

[0059] The transport peptide in the gmtGBE fusion protein functions to guide the base editor into the mitochondria. This application provides two transport peptides, selected based on two application scenarios: targeting plant mitochondria and human mitochondrial genomes. Specifically, the plant mitochondrial transport peptide is used for targeted editing of the plant mitochondrial genome, while the human mitochondrial transport peptide is used for targeted editing of the human and animal mitochondrial genomes. The plant mitochondrial transport peptide is derived from the rice Rf1b cytoplasmic male sterility restoration gene transport peptide, and the human mitochondrial transport peptide is derived from the human COX8A gene. It is worth noting that the selection of transport peptides is not limited to the two provided above; other transport peptides with mitochondrial localization functions can also be used.

[0060] The TALE arrangement in the gmtGBE fusion protein functions to precisely deliver components such as nickases, exonucleases, and GDG to the target genome editing site. Since the gRNA components of CRISPR / Cas9 cannot enter organelles, while TALEs, guided by signal peptides, can enter organelles and precisely locate the target genome editing site, this application selects the TALE arrangement as the DNA-binding domain component. The TALE consists of three parts: an N-terminal functional domain (NTD), a repeat unit arrangement, and a C-terminal functional domain (CTD). The N-terminal and C-terminal functional domains (NTD) are fixed amino acid sequences. The repeat unit arrangement is designed and assembled according to the sequence of the target editing site. The repeat units are assembled using the Golden Gate method, and the assembly of the TALE repeat units can be outsourced to a commercial company.

[0061] The L fusion protein of gmtGBE has a nicking enzyme fused to the C-terminus of TALE, such as Figures 1 to 3 As shown. This application uses a portion of the functional domain of a nickase derived from Bacillus subtilis Nt.BspD6I. Nt.BspD6I can form a heterodimer with BspD6I (small subunit, 20 kDa) and function as a restriction endonuclease called R.BspD6I 25. The Nt.BspD6I(C) used in this application for fusion with TALE is only the C-terminal cleavage domain, i.e., 382-604 amino acids, which has very low chain bias and can cleave at the target DNA editing site to create a nick. The nickase used is Nt.BspD6I(C), but not limited to Nt.BspD6I(C). This application also provides a variety of nickases, such as MutH nickase and Fok1-Fok1. D450A Cutting enzyme.

[0062] An exonuclease is also fused to the C-terminus of the L fusion protein nickase. The DNA exonuclease is linked to the aforementioned nickase via a 16aa linker, as shown below. Figure 1 As shown. This application uses Trex2 exonuclease. Trex2 is a 3'→5' digestion-preferred exonuclease. The nicking enzyme and Trex2 are linked via a 16aa linker. The function of Trex2 is to digest the nicked DNA at the DNA nick produced by the nicking enzyme, thereby exposing a short ssDNA fragment at the target DNA editing site. It is worth noting that the choice of exonuclease is not limited to Trex2; other exonucleases, such as mExoI exonuclease, can also be used to achieve a similar effect.

[0063] The DNA glycosylation enzyme GDG is fused to the C-terminus of the R fusion protein TALE, such as... Figure 1As shown. This application configures a guanine DNA glycosylase (GDG). The function of GDG is to excise the G base of ssDNA exposed at the target editing site, generating an apurinine (AP) site. The AP site is then repaired through trans-damage synthesis (TLS) and / or DNA replication, resulting in G-to-Y base editing, such as... Figure 1 As shown. The GDG configured in this application is gMPGV6.3, containing mutations in G163R, N169G, D175R, C178N, S198A, K202A, G203A, S206A, K210A, and K294R. gMPGV6.3 has the function of removing typical G bases, while wild-type MPG can recognize and remove a variety of damaged purine bases, but the rate of recognizing and removing normal G is very low. At the same time, its G base removal activity on single-stranded DNA substrates is higher than that on double-stranded DNA substrates.

[0064] Monomeric gmtGBE is a single fusion protein in which the above-mentioned transport peptides, DNA exonuclease, TALE array, DNA nicking enzyme, and GDG are sequentially fused. The DNA exonuclease and TALE array are linked via a 48aa linker, and the nicking enzyme and GDG are linked via a 16aa linker.

[0065] This application provides binary expression vectors for the gmtGBE system. The pEF1a-Puro vector is used for the expression of human gmtGBE. Complementary vectors for the human gmtGBE system include pgmtGBE-hL, pgmtGBE-hR, and pgmtGBE-hm, such as... Figure 3 As shown. The pgGBEp vector is used for the expression of plant gmtGBE. The complementary vectors for the plant gmtGBE system also include pgmtGBE-L, pgmtGBE-R, and pgmtGBE-m, as shown below. Figure 2 As shown.

[0066] Human gmtGBE expression is driven by the human elongation factor EF-1α core promoter. Plant gmtGBE expression is driven by the rice Ubi promoter (rUbi) and the maize Ubi promoter (ZmUbi), with the Nos terminator (NosT). The pEF1a-Puro plasmid vector contains a replicon that can reproduce in *E. coli*, and includes the selection marker ampicillin resistance gene (ampicillin resistance provided by the β-lactamase gene, AmpR) and the eukaryotic cell selection marker PuroR. The pgGBEp plasmid vector contains a replicon that can reproduce in *E. coli* and *Agrobacterium*, and includes the selection marker kanamycin resistance gene (Kanamycin resistance provided by the aminoglycoside phosphotransferase gene, KanR).

[0067] This application provides a method for constructing the aforementioned base editing system.

[0068] Construction method of the human gmtGBE system: TALE repeat unit arrays were designed based on the sequences flanking the target editing site of interest in the mitochondrial genome. The TALE repeat units can be assembled by a commercial company. The assembled TALE repeat unit arrays were inserted into the pgmtGBE-hL and pgmtGBE-hR vectors via the BsaI site. The pgmtGBE-hL vector released the L fusion protein gene fragment via KpnI-SpeI restriction endonuclease, and then cloned into the pEF1a-Puro vector via KpnI-XbaI restriction endonuclease. The R fusion protein expression cassette of the pgmtGBE-hR vector (driven by the EF1a promoter) was released via EcoRI-XbaI restriction endonuclease, and then cloned into the pEF1a-Puro vector via the EcoRI-NheI restriction endonuclease site. Finally, the pEF1a-Puro vector integrated both the L and R fusion protein expression cassettes for cell infection and transformation. For the human monomeric gmtGBE, a TALE repeat unit array is designed based on the sequence on one side of the target editing site in the mitochondrial genome of interest, and the TALE repeat units can be assembled by a commercial company. The assembled TALE repeat unit array is inserted into pgmtGBE-hm via the BsaI site. The pgmtGBE-hm monomeric fusion protein is released via KpnI-SpeI restriction endonuclease, and then cloned into the pEF1a-Puro vector via KpnI-XbaI restriction endonuclease. Positive clones are used for cell infection and transformation.

[0069] Construction method of plant gmtGBE system: TALE repeat unit arrays are designed based on the sequences flanking the target editing site in the mitochondrial genome of interest, and assembly of the TALE repeat units can be outsourced to commercial companies. The assembled TALE repeat unit arrays are inserted into the pgmtGBE-L and pgmtGBE-R vectors via the BsaI site. In pgmtGBE-L, the L fusion protein gene fragment is cloned into the ZmUbi-NosT expression cassette of the pgGBEp vector via the KpnI-SacI restriction endonuclease site; the R fusion protein expression cassette (driven by rUbi) of pgmtGBE-R replaces the ccdB toxic protein marker expression cassette of the pgGBEp vector via the LR Gateway. Finally, the pgGBEp vector integrates two expression cassettes, L and R, for plant genetic transformation. For plant monoclonal gmtGBE, TALE repeat unit arrays are designed based on the sequences flanking the target editing site in the mitochondrial genome of interest, and assembly of the TALE repeat units can be outsourced to commercial companies. The assembled TALE repeat unit array was inserted into pgmtGBE-m via the BsaI site. The pgmtGBE-m fusion protein expression cassette (driven by rUbi) replaced the toxic protein marker ccdB of the pgGBEp vector via LR Gateway, and positive clones were used for plant transformation.

[0070] This application also provides a method for applying the aforementioned base editing system. After selecting mitochondrial DNA target site sequences from plants such as rice, and human mitochondrial DNA target site sequences, TALE repeat units are assembled, and corresponding vectors are constructed. The vectors can be transformed into eukaryotic cells or plants can be transformed using Agrobacterium / gene gun to achieve GY base editing at their target sites.

[0071] The application method of the base editing system of this application further includes: extracting organelle DNA from transformed plants or humans, then performing PCR amplification using target-specific primers, constructing a high-throughput next-generation DNA sequencing library using the PCR products of the organelle DNA, and performing high-throughput next-generation DNA sequencing. The organelle DNA targets of a certain species include, but are not limited to, at least one of the Cob and Cox2 targets.

[0072] The experimental method of this application includes the following steps:

[0073] 1) Construction of the expression vector. TALE target sequences are designed flanking the target base editing site. The base preceding the 5' end of the TALE target sequence must be a T base, and the length of the TALE target sequence should be in the range of 12-24 bp. Simultaneously, a distance of 3-20 bp between the 3' end of the TALE target sequence and the target base editing site results in higher editing efficiency. After designing and determining the TALE target sequence, a TALE repeat unit array is assembled according to the TALE target sequence. The TALE repeat unit array is then cloned into the corresponding expression vector provided in this application via the BsaI restriction site.

[0074] 2) Plant expression vectors use Agrobacterium to transform plant tissues; while human cell expression vectors transform the HEK293T cell line.

[0075] 3) Genomic DNA extraction and amplification. Genomic DNA is extracted from the transformed cells and then amplified by PCR using specific primers targeting the editing site.

[0076] 4) High-throughput next-generation DNA sequencing of genomic DNA PCR products. High-throughput next-generation DNA sequencing libraries were constructed using genomic DNA PCR products, and high-throughput next-generation DNA sequencing was performed.

[0077] The present application will be further described in detail below through specific embodiments. These embodiments are for illustrative purposes only and should not be construed as limiting the scope of the application. Unless otherwise specified, the methods for plasmid or strain construction in this application refer to Molecular Cloning: A Laboratory Manual (Fourth Edition).

[0078] Example

[0079] The experimental method includes the following steps:

[0080] 1) Construction of the expression vector: TALE target sequences are designed on both sides of the target base editing site. The base preceding the 5' end of the TALE target sequence must be a T base, and the length of the TALE target sequence should be in the range of 12-24 bp. Simultaneously, a distance of 3-20 bp between the 3' end of the TALE target sequence and the target base editing site results in higher editing efficiency. After designing and determining the TALE target sequence, a TALE repeat unit array is assembled according to the TALE target sequence. Then, the TALE repeat unit array is cloned into the corresponding expression vector provided in this application via the BsaI restriction site.

[0081] 2) Plant expression vectors use Agrobacterium to transform plant tissues; while human cell expression vectors transform the HEK293T cell line.

[0082] 3) Genomic DNA extraction and amplification: Genomic DNA is extracted from the transformed cells and then amplified by PCR using specific primers targeting the editing site.

[0083] 4) High-throughput next-generation DNA sequencing of genomic DNA PCR products; constructing a high-throughput next-generation DNA sequencing library using genomic DNA PCR products and performing high-throughput next-generation DNA sequencing.

[0084] I. Base editing of the rice Cob gene using the gmtGBE system for dimorphic plants

[0085] 1) Reagents

[0086] Primers were synthesized using standard methods by Sangon Biotech (Shanghai) Co., Ltd.; plasmid sequencing was commissioned to Sangon Biotech (Shanghai) Co., Ltd.; KpnI, SacI, BsaI restriction endonucleases, DNA ligase (T4 DNA ligase), and high-fidelity DNA polymerase were purchased from NEB; LR Gateway enzyme was used. Mix was purchased from Thermo Fisher Scientific; kanamycin sulfate and rifampicin were purchased from Beijing Solarbio Science & Technology Co., Ltd.; plasmid miniprep kit, DNA gel extraction kit, and high-efficiency plant genomic DNA extraction kit were purchased from Tiangen Biotech Co., Ltd.; DB3.1 and Trans1-T1 Escherichia coli competent cells and EHA105 Agrobacterium competent cells were purchased from Shanghai Weidi Biotechnology Co., Ltd.; vector plasmids pUC-KANA and pCAMBIA1300 were purchased from Sangon Biotech (Shanghai) Co., Ltd.; TALE repeat unit array was assembled by Guangzhou Yijin Biotechnology Co., Ltd.; gene synthesis was completed by Yunzhou Biotechnology (Guangzhou) Co., Ltd.; rice genetic transformation experiments were completed by Wuhan Boyuan Biotechnology Co., Ltd.; QuickExtract™ genomic DNA extraction reagent and high-throughput DNA next-generation sequencing library construction kit TmSeqChIPLibrary Preparation Kit were purchased from Illumina.

[0087] 2) Carrier construction

[0088] The pgmtGBE-L, pgmtGBE-R, and pgmtGBE-m vectors used pUC-KANA plasmid as the starting backbone, while the pgGBEp vector used pCAMBIA1300 plasmid as the starting backbone. The synthesis and cloning of elements such as the rUbi promoter, Nos terminator, aatL, and ccdB were all performed by commercial companies.

[0089] 3) TALE target sequence design and assembly

[0090] Left and right TALE repeat unit arrays were designed based on the sequences flanking the target editing site of the Cob gene of interest in the mitochondrial gene. Following the rule that NI, NG, NN, and HD TALE repeat units recognize nucleotides "A", "T", "G", and "C" respectively, the TALE repeat units were assembled by a commercial company. The assembled TALE repeat unit arrays were inserted into pgmtGBE-L and pgmtGBE-R vectors via the BsaI site, respectively, and transformed into Trans1-T1 *E. coli*, cultured on kanamycin sulfate plates. The TALE repeat unit arrays of both pgmtGBE-L and pgmtGBE-R were confirmed by sequencing using NTD-F and CTD-R primers.

[0091] The left TALE repeat unit array is the sequence shown in SEQ ID NO.13, the right TALE repeat unit array is the sequence shown in SEQ ID NO.14, the sequencing primer NTD-F is the sequence shown in SEQ ID NO.15, and the sequencing primer CTD-R is the sequence shown in SEQ ID NO.16.

[0092] SEQ ID NO.13: 5'-TGTTTAAGAAGAGAGAATCG-3'

[0093] SEQ ID NO.14: 5'-CACGATAGAAAAGAGAAATG-3'

[0094] SEQ ID NO.15: 5'-CACCACTCCAGTTGGACACAG-3'

[0095] SEQ ID NO.16: 5'-TAGGAACGCGGTGGGATGTG-3'

[0096] The Cob target site editing chain is the sequence shown in SEQ ID NO.25.

[0097] SEQ ID NO.25: 5'-ACTATAAGGAACCAA-3'

[0098] The pgmtGBE-L vector cloned the L fusion protein gene fragment into the ZmUbi-NosT expression cassette of the pgGBEp vector via the KpnI-SacI restriction endonuclease site, and transformed it into DB3.1 *E. coli*, cultured on kanamycin sulfate plates. Next, the pgmtGBE-R vector's R fusion protein expression cassette, driven by rUbi, replaced the toxic protein marker ccdB expression cassette of the pgGBEp vector via the LR Gateway, and was transformed into Trans1-T1 *E. coli*, cultured on kanamycin sulfate plates. Finally, the pgGBEp vector integrated two expression cassettes, the L and R fusion proteins, and was designated pgGBEp-Cob, for plant genetic transformation.

[0099] Among them, the transport peptide is a plant mitochondrial transport peptide, namely the sequence shown in SEQ ID NO.1; the N-terminal functional domain of TALE-L and TALE-R is the sequence shown in SEQ ID NO.3; the C-terminal functional domain of TALE-L and TALE-R is the sequence shown in SEQ ID NO.4; the BspD6I nickase is the sequence shown in SEQ ID NO.5; the Trex2 exonuclease is the sequence shown in SEQ ID NO.8; the BspD6I nickase and the Trex2 exonuclease are linked by a 16aa linker, the 16aa linker being the sequence shown in SEQ ID NO.11.

[0100] The DNA glycosylation enzyme in this example is an engineered mutant of gMPGV6.3, with the sequence shown in SEQ ID NO.12.

[0101] In addition to BspD6I nickase, MutH nickase or Fok1-Fok1 nickase can also be used. D450A The nicking enzyme, MutH nicking enzyme, has the sequence shown in SEQ ID NO.6, Fok1-Fok1. D450A The cleavage enzyme has the sequence shown in SEQ ID NO.7. In addition to the Trex2 exonuclease, the mExoI exonuclease can also be used, and the mExoI exonuclease has the sequence shown in SEQ ID NO.9.

[0102] SEQ ID NO.1:

[0103] ARRVAARARARFGGVPRSEGTIQDRARVGSGGAEDALDVFDELLRRGIGAPIRSLNGALADVARDNPAAAVSRFN

[0104] SEQ ID NO.3:

[0105] DLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVTYQDIIRALPEATHEDIVGVGKQWSGARALEALLTEAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVH

[0106] SEQ ID NO.4:RPALDAVKKGLPHAPELIRRINRRIPERTSHRVP

[0107] SEQ ID NO.5:

[0108] RQLEEVIDLLEVYHEKKNVIEEKIKARFIANKNTVFEWLTWNGFIILGNALEYKNNFVIDEELQPVTHAAGNQPDMEIIYEDFIVLGEVTTSKGATQFKMESEPVTRHYLNKKKELEKQGVEKELYCLFIAPEINKNTFEEFMKYNIVQNTRIIPLSLKQFNMLLMVQKKLIEKGRRLSSYDIKNLMVSLYRTTIECERKYTQIKAGLEETLNNWVVDKEVRF

[0109] SEQ ID NO.6:

[0110] MSQPRPLLSPPETEEQLLAQAQQLSGYTLGELAALVGLVTPENLKRDKGWIGVLLEIWLGASAGSKPEQDFAALGVELKTIPVDSLGRPLETTFVCVAPLTGNSGVTWETSHVRHKLKRVLWIPVEGERSIPLAQRRVGSPLLWSPNEEEDRQLREDWEELMDMIVLGQVERITARHGEYLQIRPKAANAKALTEAIGARGERILTLPRGFYLKKNFTSALLARHFLIQ

[0111] SEQ ID NO.7:

[0112] FKQLVKSELEEKKSELRHKLKYVPHEYIELIEIARNSTQDRILEMKVMEFFMKVYGYRGKHLGGSRKPDGAIYTVGSPIDYGVIVDTKAYSGGYNLPIGQADEMQRYVEENQTRNKHINPNEWWKVYPSSVTEFKLFVSGHFKGNYKAQLTRLNHITNCNGAVLSVEELLIGMEKIGAGTLTLEEVRRKFNNGEINFSGGSSGGSSGSETPGTSESATPESSGGSSGGSSGSETPGTSESATPESQLVKSELEEKKSELRHKLKYVPHEYIELIEIARNSTQDRILEMKVMEFFMKVYGYRGKHLGGSRKPAGAIYTVGSPIDYGVIVDTKAYSGGYNLPIGQADEMQRYVEENQTRNKHINPNEWWKVYPSSVTEFKLFVSGHFKGNYKAQLTRLNHITNCNGAVLSVEELLIGMEKIGAGTLTLEEVRRKFNNGEINF

[0113] SEQ ID NO.8:

[0114] MSEPPRAETFVFLDLEATGLPNMDPEIAEISLFAVHRSSLENPERDDSGSLVLPRWLDKLTLCMCPERPFTAKASEITGLSSESLMHCGKAGFNGAVVRTLQGFLSRQEGPICLVAHNGFDY DFPLLCTELQRLGAHLPQDTVCLDTLPLARGLDRAHSHGTRAQGRKSYSLASLFHRYFQAEPSAAHSAEGDVHTLLLIFLHRAPELLAWADEQARSWAHIEPMYVPPDGPSLEA

[0115] SEQ ID NO.9:

[0116] MGIQGLLQFIQEASEPVNVKKYKGQAVAVDTYCWLHKGAIACAEKLAKGEPTDRYVGFCMKFVNMLLSYGVKPILIFDGCTLPSKKEVERSRRERRQSNLLKGKQLLREGKVSEARDCFARSINITHAMAHKVIKAARALGVDCLVAPYEADAQLAYLNKAGIVQAVITEDSDLLAFGCKKVILKMDQFGNGLEVDQARLGMCKQLGDVFTEEKFRYMCILSGCDYLASLRGIGLAKACKVLRLANNPDIVKVIKKIGHYLRMNITVPEDYITGFIRANNTFLYQLVFDPIQRKLVPLNAYGDDVNPETLTYAGQYVGDSVALQIALGNRDVNTFEQIDDYSPDTMPAHSRSHSWNEKAGQKPPGTNSIWHKNYCPRLEVNSVSHAPQLKEKPSTLGLKQVISTKGLNLPRKSCVLKRPRNEALAEDDLLSQYSSVSKKIKENGCGDGTSPNSSKMSKSCPDSGTAHKTDAHTPSKMRNKFATFLQRRNEESGAVVVPGTRSRFFCSSQDFDNFIPKKESGQPLNETVATGKATTSLLGALDCPDTEGHKPVDANGTHNLSSQIPGNAAVSPEDEAQSSETSKLLGAMSPPSLGTLRSCFSWSGTLREFSRTPSPSASTTLQQFRRKSDPPACLPEASAVVTDRCDSKSEMLGETSQPLHELGCSSRSQESMDSSCGLNTSSLSQPSSRDSGSEESDCNNKSLDNQGEQNSKQHLPHFSKKDGLRRNKVPGLCRSSSMDSFSTTKIKPLVPARVSGLSKKSGSMQTRKHHDVENKPGLQTKISELWKNFGFKKDSEKLPSCKKPLSPVKDNIQLTPETEDEIFNKPECVRAQRAIFH

[0117] SEQ ID NO.11:SGSETPGTSESATPES

[0118] SEQ ID NO.12:

[0119] VTPALQMKKPKQFCRRMGQKKQRPARAGQPHSSSDAAQAPAEQPHSSSDAAQAPCPRERCLGPPTTPGPYRSIYFSSPKGHLTRLGLEFFDQPAVPLARAFLGQVLVRRLPNGTELRGRIVETEAYLGPEDEAAHSRGGRQTPRNRGM FMKPGTLYVYIIYRMYFCMGISSQGRGANVLLRALEPLEGLETMRQLRATLRAATAARVLADRELCSGPSKLCQALAINKSFDQRDLAQDEAVWLERGPLEPSEPAVVAAARVGVGHAGEWARKPLRFYVRGSPWVSVVDRVAERDTQA

[0120] 4) Rice genetic transformation

[0121] pgGBEp-Cob positive clone plasmid was extracted, and 50 ng was used to transform EHA105 Agrobacterium competent cells. The cells were then plated on LB agar plates containing 50 μg / mL kanamycin sulfate and 50 μg / mL rifampin, and cultured at 28°C for 36 h to obtain single colonies. The positive Agrobacterium colonies were then used by Wuhan Boyuan Biotechnology Co., Ltd. for rice callus infection, subsequent resistant callus (screened with hygromycin as the antibiotic), and callus regeneration to obtain transgenic plants.

[0122] 5) Identification of transgenic plants and analysis of gene editing efficiency at target sites

[0123] Total genomic DNA was extracted from leaves of transgenic plants using a high-efficiency plant genomic DNA extraction kit. Primers ZmUbi-F and NTD-R were used to identify whether the transgenic plants were positive, i.e., containing the pgGBEp-Cob transgenic fragment. Mitochondrial DNA fragments containing the target site sequence were amplified using primers Cob-F and Cob-R to analyze the gene editing efficiency at the target site.

[0124] Among them, ZmUbi-F is the sequence shown in SEQ ID NO.17, NTD-R is the sequence shown in SEQ ID NO.18, PsbA-F is the sequence shown in SEQ ID NO.19, and PsbA-R is the sequence shown in SEQ ID NO.20.

[0125] SEQ ID NO.17: 5'-TTTAGCCCTGCCTTCATACGC-3'

[0126] SEQ ID NO.18: 5'-CTGGCTGAGAGCAACGATGTG-3'

[0127] SEQ ID NO.19: 5'-ACGCTGTTGAAAGCTAGATCC-3'

[0128] SEQ ID NO.20: 5'-AGCTTAAGAAACCGGCAGTT-3'

[0129] The PCR system is as follows:

[0130] Phusion (NEB, M0535L) 1 μL; 5× buffer (NEB, B0518S) 4 μL; Forward Primer (5 μM) 1 μL; Reverse Primer (5 μM) 1 μL; Template 50ng; RNase-free water was added to 50 μL.

[0131] The PCR program was as follows: 95℃, 2 min; then 33 cycles: 95℃, 15 sec, 60℃, 30 sec, 72℃, 1 min; after the cycles, 72℃, 3 min.

[0132] High-throughput DNA-seq libraries were prepared from the PCR products of Cob-F and Cob-R primers. Specifically, the concentration and purity of the PCR products were first checked using a Qubit Fluorometer and an Aglient Bioanalyzer 2100, respectively. For samples that passed the quality checks, different biotinylated barcodes were added to different PCR products using the Illumina TruSeq ChIP Sample Preparation Kit for PCR amplification and enrichment, thus completing the preparation of the sequencing libraries. The concentration and fragment size of the constructed libraries were then checked. Finally, high-throughput sequencing was performed using an Illumina HiSeq 2500 sequencer.

[0133] High-throughput sequencing results showed that the conversion efficiencies of the two target G bases at the Cob target site to T bases reached 71% and 74%, respectively, and the conversion efficiencies to C bases reached 23% and 19%, respectively.

[0134] II. Base editing of the rice Cob gene using the monoplant gmtGBE system

[0135] 1) Reagents

[0136] Same as "I. Base editing of rice mitochondrial Cob gene using the gmtGBE system of dimorphic plants".

[0137] 2) Carrier construction

[0138] Same as "I. Base editing of rice mitochondrial Cob gene using the gmtGBE system of dimorphic plants".

[0139] 3) TALE target sequence design and assembly

[0140] A TALE repeat unit array was designed based on the sequence on one side of the target editing site of the Cob gene of interest in the mitochondrial gene. In this example, the specific design is the left-side TALE repeat unit array with the same sequence as shown in SEQ ID NO.13, serving as the TALE array for the monomeric fusion protein. According to the rule that the four types of TALE repeat units (NI, NG, NN, and HD) recognize nucleotides "A", "T", "G", and "C" respectively, the TALE repeat units were assembled by a commercial company. The assembled TALE repeat unit arrays were then inserted into the pgmtGBE-m vector (transformed into Trans1-T1 E. coli and cultured on kanamycin sulfate plates) via the BsaI site. The TALE repeat unit array of pgmtGBE-m was confirmed by sequencing using NTD-F and CTD-R primers.

[0141] The monomeric fusion proteins are, in sequence, the transport peptide shown in SEQ ID NO.1, the Trex2 exonuclease shown in SEQ ID NO.8, the N-terminal functional domain shown in SEQ ID NO.3, the left-side TALE repeat unit array shown in SEQ ID NO.13, the C-terminal functional domain shown in SEQ ID NO.4, the BspD6I nickase shown in SEQ ID NO.5, and an engineered mutant of gMPGV6.3 shown in SEQ ID NO.12. Furthermore, the Trex2 exonuclease and the N-terminal functional domain are linked by a 48aa linker, and the BspD6I nickase and the engineered mutant of gMPGV6.3 are linked by a 16aa linker. The 48aa linker is the sequence shown in SEQ ID NO.10, and the 16aa linker is the sequence shown in SEQ ID NO.11.

[0142] SEQ ID NO.10:

[0143] SGGSSGGSSGSETPGTSESATPESSGGSSGGSSGSETPGTSESATPES

[0144] The monomeric fusion protein expression cassette of pgmtGBE-m (driven by rUbi) replaced the toxic protein marker ccdB expression cassette of the pgGBEp vector via LR Gateway (transformed into Trans1-T1 E. coli, cultured on kanamycin sulfate plates). Positive clones were designated as pgGBEp-Cob-m and used for plant genetic transformation.

[0145] 4) Rice genetic transformation

[0146] The genetic transformation of rice using pgGBEp-Cob-m is the same as "I. Base editing of rice mitochondrial Cob gene using the gmtGBE system of dimorphic plants".

[0147] 5) Identification of transgenic plants and analysis of gene editing efficiency at target sites

[0148] The methods for identifying transgenic plants and analyzing the efficiency of gene editing at target sites are the same as those in "I. Base editing of rice mitochondrial Cob gene using the gmtGBE system of dioecious plants".

[0149] High-throughput sequencing results showed that the conversion efficiencies of the two target G bases at the Cob target site to T bases reached 68% and 70%, respectively, and the conversion efficiencies to C bases reached 20% and 17%, respectively.

[0150] III. Base Editing of the Human Mitochondrial Cox2 Gene Using the Disosomal Human gmtGBE System

[0151] 1. Experimental Materials

[0152] 1) Reagents

[0153] Primers were synthesized using standard methods by Sangon Biotech (Shanghai) Co., Ltd.; plasmid sequencing was performed by Sangon Biotech (Shanghai) Co., Ltd.; TALE repeat unit arrays were assembled by Guangzhou Yijin Biotechnology Co., Ltd.; vector plasmid pUC-KANA was purchased from Sangon Biotech (Shanghai) Co., Ltd., and vector plasmid pLV3-CMV was purchased from Yunzhou Biotechnology (Guangzhou) Co., Ltd.; TALE repeat unit arrays were assembled by Guangzhou Yijin Biotechnology Co., Ltd.; gene synthesis was performed by Yunzhou Biotechnology (Guangzhou) Co., Ltd.; EcoRI, SpeI, XbaI, NheI, BsaI restriction endonucleases, DNA ligase (T4 DNA ligase), and high-fidelity DNA polymerase were purchased from NEB; DB3.1 and Trans1-T1 Escherichia coli competent cells were purchased from Shanghai Weidi Biotechnology Co., Ltd.; ampicillin and kanamycin sulfate were purchased from Beijing Solarbio Science & Technology Co., Ltd.; EndoFree plasmid kits were purchased from Qiagen; GeneJET gel extraction kits were purchased from Thermo Scientific; and mycoplasma detection kits were purchased from TransGen. Biotech Inc.; 48-well poly-d-lysine coated plates were purchased from Corning Incorporated; Polyplus Transfection Cell Transfection Kit was purchased from Annoron Inc.; Proteinase K and 1×Dulbecco's PBS were purchased from Thermo Fisher Scientific; QuickExtract™ Genomic DNA Extraction Reagent and TmSeqChIP Library Preparation Kit for Next-Generation Sequencing were purchased from Illumina.

[0154] 2) Cell line: HEK293T (purchased from the Cell Bank of the Chinese Academy of Sciences)

[0155] 2. Experimental Methods

[0156] 1) Carrier construction

[0157] The pgmtGBE-hL, pgmtGBE-hR, and pgmtGBE-hm vectors used pUC-KANA plasmid as the starting backbone, while the pEF1a-Puro vector used pLV3-CMV plasmid as the starting backbone. The synthesis and cloning of elements such as the EF-1α promoter, Cox8 transport peptide, and ccdB were all performed by commercial companies.

[0158] 2) TALE target sequence design and assembly

[0159] Left and right TALE repeat unit arrays were designed based on the sequences flanking the target editing site of the Cox2 mitochondrial gene of interest. Following the rules that NI, NG, NN, and HD TALE repeat units recognize nucleotides “A,” “T,” “G,” and “C,” respectively, the TALE repeat units were assembled by a commercial company. The pgmtGBE-hL and pgmtGBE-hR vectors were propagated in DB3.1 *E. coli*, and plasmid DNA was extracted using a plasmid kit. The assembled TALE repeat unit arrays were inserted into the pgmtGBE-hL and pgmtGBE-hR vectors via the BsaI site (transformed into Trans1-T1 *E. coli*, cultured on kanamycin sulfate plates). Both pgmtGBE-hL and pgmtGBE-hR TALE repeat unit arrays were confirmed by sequencing using the sequencing primers NTD-F and CTD-R. The sequencing primers NTD-F and CTD-R are the same as those used in “I. Base Editing of the Rice Cob Gene Using the Dioecious Plant gmtGBE System.”

[0160] pgmtGBE-hL releases the L fusion protein gene fragment via KpnI-SpeI restriction endonuclease, and then clones it into the pEF1a-Puro vector via KpnI-XbaI restriction endonuclease. The R fusion protein expression cassette of pgmtGBE-hR (driven by the EF1a promoter) is released via EcoRI-XbaI restriction endonuclease, and then cloned into the pEF1a-Puro vector via the EcoRI-NheI restriction endonuclease site. Finally, the pEF1a-Puro vector integrates both the L and R fusion protein expression cassettes, denoted as pEF1a-Cox2, for cell infection and transformation.

[0161] The left TALE repeating unit array is the sequence shown in SEQ ID NO.21, and the right TALE repeating unit array is the sequence shown in SEQ ID NO.22.

[0162] SEQ ID NO.21: 5'-CCTATATATCTTAATGGCAC-3'

[0163] SEQ ID NO.22: 5'-AGGGGAAGTAGCGTCTTGTA-3'

[0164] The Cox2 target site editing chain is the sequence shown in SEQ ID NO.26.

[0165] SEQ ID NO.26: 5'-ACCTACTTGCGCTGCAT-3'

[0166] The Cox8 transport peptide is the sequence shown in SEQ ID NO.2. The N-terminal and C-terminal functional domains of TALE-L and TALE-R, as well as the BspD6I nickase and Trex2 exonuclease, and the linkers of the BspD6I nickase and Trex2 exonuclease are all the same as "I. Base editing of rice mitochondrial Cob gene using the dimorphic plant gmtGBE system".

[0167] SEQ ID NO.2:SVLTPLLLRGLTGSARRLPVPRAKIHSL

[0168] 3) Eukaryotic cell transfection

[0169] pEF1a-Cox2 plasmid for transfection of HEK293T cell line was extracted and purified using the EndoFree Plasmid Kit (Qiagen).

[0170] HEK293T cells were cultured in DMEM (Gibco) supplemented with 10% (vol / vol) fetal bovine serum (FBS, Gibco) and 1% (vol / vol) penicillin-streptomycin (Gibco) in a humidified incubator at 37°C with 5% CO2. All cells were routinely tested for mycoplasma contamination using a mycoplasma detection kit (TransGen Biotech). A total of 40,000 cells per well were seeded into antibiotic-free 48-well poly-d-lysine-coated plates. After 16–24 hours, each well was transfected with 1 μL of jetPRIME transfection reagent (Polyplus Transfection kit) and 400 ng of pEF1a-Cox2 plasmid, with a cell confluence of 60–80%. Cells were washed with PBS and DNA was extracted 72 hours post-transfection.

[0171] 4) Genomic DNA extraction and PCR amplification

[0172] After washing cells once with 1×Dulbecco's PBS, genomic DNA was extracted by adding 100 μL of freshly prepared lysis buffer (10 mM Tris-HCl (pH 8.0), 0.05% SDS, and 25 μg / mL proteinase K) directly to 48-well cultures. The mixture was incubated at 37°C for 60 min, followed by incubation at 80°C for 20 min.

[0173] The extracted genomic DNA was amplified by PCR using the primers Cox-hF and Cox-hR. The PCR products were then excised and recovered using the GeneJET gel extraction kit according to the reagent instructions for subsequent construction of high-throughput next-generation DNA sequencing libraries. Cox-hF is the sequence shown in SEQ ID NO. 23, and Cox-hR is the sequence shown in SEQ ID NO. 24.

[0174] SEQ ID NO.23: 5'-CCTGGAGTGACTATATGGATGC-3'

[0175] SEQ ID NO.24: 5'-GCGATGAGGACTAGGATGATG-3'

[0176] The PCR system used to amplify genomic DNA is as follows:

[0177] Phusion (NEB, M0535L) 1μL; 5×buffer (NEB, B0518S) 4μL; Forward Primer (5μM) 1μL; Reverse Primer (5μM) 1μL; Template 10ng; RNase-free water is added to 50μL.

[0178] The PCR procedure is as follows:

[0179] 95℃, 2 min; then 30 cycles: 95℃, 15 sec, 55℃, 15 sec, 72℃, 30 sec; after the cycle, 72℃, 3 min.

[0180] 5) High-throughput next-generation DNA sequencing

[0181] High-throughput DNA-seq libraries were prepared from the PCR products. Specifically, the concentration and purity of the PCR products were first checked using a Qubit Fluorometer and an Aglient Bioanalyzer 2100, respectively. Samples that passed the quality checks were then fragmented using a Covaris S220 ultrasonic DNA disruptor. Next, using the Illumina TruSeq ChIP Sample Preparation Kit, different biotinylated barcodes were added to different PCR products for PCR amplification and enrichment, completing the preparation of the sequencing libraries. The concentration and fragment size of the constructed libraries were then checked. Finally, high-throughput sequencing was performed using an Illumina HiSeq 2500 sequencer.

[0182] High-throughput sequencing results showed that the conversion efficiencies of the three target G bases at the Cox2 target site to T bases reached 62%, 70%, and 77%, respectively, and the conversion efficiencies to C bases reached 23%, 25%, and 17%, respectively.

[0183] IV. Base editing of the human mitochondrial Cox2 gene using the single-cell human gmtGBE system

[0184] 1) Reagents and cell lines

[0185] Same as "III. Base editing of human cell mitochondrial Cox2 gene using the dual-body human gmtGBE system".

[0186] 2) Carrier construction

[0187] Same as "III. Base editing of human cell mitochondrial Cox2 gene using the dual-body human gmtGBE system".

[0188] 3) TALE target sequence design and assembly

[0189] A TALE repeat unit array was designed based on the sequence on one side of the target editing site of the Cox2 mitochondrial gene of interest. In this example, the specific design is the left-side TALE repeat unit array identical to the sequence shown in SEQ ID NO. 21, serving as the TALE array for the monomeric fusion protein. Following the rule that the four types of TALE repeat units (NI, NG, NN, and HD) recognize nucleotides "A", "T", "G", and "C" respectively, the TALE repeat units were assembled by a commercial company. The assembled TALE repeat unit arrays were then inserted into the pgmtGBE-hm vector (transformed into Trans1-T1 E. coli and cultured on kanamycin sulfate plates) via the BsaI site. The TALE repeat unit array of pgmtGBE-hm was confirmed by sequencing using NTD-F and CTD-R primers.

[0190] The monomeric fusion proteins are, in sequence, the transport peptide shown in SEQ ID NO.2, the Trex2 exonuclease shown in SEQ ID NO.8, the N-terminal functional domain shown in SEQ ID NO.3, the left-side TALE repeat unit array shown in SEQ ID NO.21, the C-terminal functional domain shown in SEQ ID NO.4, the BspD6I nickase shown in SEQ ID NO.5, and an engineered mutant of gMPGV6.3 shown in SEQ ID NO.12. Furthermore, the Trex2 exonuclease and the N-terminal functional domain are linked by a 48aa linker, and the BspD6I nickase and the engineered mutant of gMPGV6.3 are linked by a 16aa linker. The 48aa linker is the sequence shown in SEQ ID NO.10, and the 16aa linker is the sequence shown in SEQ ID NO.11.

[0191] The pgmtGBE-hm monomeric fusion protein is released via KpnI-SpeI restriction endonuclease and then cloned into the pEF1a-Puro vector via KpnI-XbaI restriction endonuclease. The positive clone is designated pEF1a-Cox2m and used for cell infection and transformation.

[0192] 3) Eukaryotic cell transfection

[0193] Same as "III. Base editing of human cell mitochondrial Cox2 gene using the dual-body human gmtGBE system".

[0194] 4) Genomic DNA extraction and PCR amplification

[0195] Same as "III. Base editing of human cell mitochondrial Cox2 gene using the dual-body human gmtGBE system".

[0196] 5) High-throughput next-generation DNA sequencing

[0197] Same as "III. Base editing of human cell mitochondrial Cox2 gene using the dual-body human gmtGBE system".

[0198] High-throughput sequencing results showed that the conversion efficiencies of the three target G bases at the Cox2 target site to T bases reached 60%, 67%, and 73%, respectively, and the conversion efficiencies to C bases reached 22%, 21%, and 15%, respectively.

[0199] The above experimental results show that the guanine base editor of this application, without relying on deaminase, can edit the guanine bases of the genome in plant and animal mitochondria by G to Y, i.e., G to T or G to C. This fills the gap in guanine base editing technology and tools for eukaryotic organelle genomes and expands the target range of base editors.

[0200] The above description, in conjunction with specific embodiments, provides a further detailed explanation of this application and should not be construed as limiting the specific implementation of this application to these descriptions. Those skilled in the art to which this application pertains can make several simple deductions or substitutions without departing from the concept of this application.

Claims

1. A DNA glycosylation enzyme composition, characterized in that: The DNA glycosylation enzyme composition is a monomeric fusion protein or a dimeric fusion protein; The monomeric fusion protein sequentially fuses a transport peptide, a DNA exonuclease, a TALE array, a DNA cleavage enzyme, and a guanine DNA glycosylation enzyme. The dual-body fusion protein includes an L fusion protein and an R fusion protein, wherein the L fusion protein sequentially fuses a transport peptide, TALE-L, a DNA nicking enzyme, and a DNA exonuclease, and the R fusion protein sequentially fuses a transport peptide, TALE-R, and a guanine DNA glycosylation enzyme. The transport peptide is used to guide the guanine base editor to the corresponding organelle to achieve G base editing; The TALE array, TALE-L, and TALE-R are all composed of an N-terminal functional domain, a repeating unit array, and a C-terminal functional domain. The repeating unit array is a specific binding sequence designed for the target sequence. The transport peptide is a plant mitochondrial transport peptide or an animal mitochondrial transport peptide. The plant mitochondrial transport peptide is derived from the rice Rf1b cytoplasmic male sterility restoration gene transport peptide. The animal mitochondrial transport peptide is a human mitochondrial transport peptide, which is a mitochondrial transport peptide derived from the human COX8A gene. The N-terminal functional domain is the sequence shown in SEQ ID NO.3; The C-terminal functional domain is the sequence shown in SEQ ID NO.4; The DNA cleavage enzyme is BspD6I cleavage enzyme, MutH cleavage enzyme, or Fok1-Fok1. D450A Cutting enzyme; The DNA exonuclease is either Trex2 exonuclease or mExoI exonuclease; The guanine DNA glycosylation enzyme is an engineered mutant of gMPGV6.3, with the mutation information being G163R, N169G, D175R, C178N, S198A, K202A, G203A, S206A, K210A, and K294R.

2. The DNA glycosylation enzyme composition according to claim 1, characterized in that: The plant mitochondrial transport peptide has the sequence shown in SEQ ID NO.

1.

3. The DNA glycosylation enzyme composition according to claim 1, characterized in that: The animal mitochondrial transport peptide has the sequence shown in SEQ ID NO.

2.

4. The DNA glycosylation enzyme composition according to claim 1, characterized in that: The BspD6I nickase has the sequence shown in SEQ ID NO.

5.

5. The DNA glycosylation enzyme composition according to claim 1, characterized in that: The MutH nickase has the sequence shown in SEQ ID NO.

6.

6. The DNA glycosylation enzyme composition according to claim 1, characterized in that: The Fok1-Fok1 D450A The nicking enzyme has the sequence shown in SEQ ID NO.

7.

7. The DNA glycosylation enzyme composition according to claim 1, characterized in that: The Trex2 exonuclease has the sequence shown in SEQ ID NO.

8.

8. The DNA glycosylation enzyme composition according to claim 1, characterized in that: The mExoI exonuclease has the sequence shown in SEQ ID NO.

9.

9. The DNA glycosylation enzyme composition according to claim 1, characterized in that: In the monomeric fusion protein, the DNA exonuclease is connected to the TALE array via a 48aa adapter.

10. The DNA glycosylation enzyme composition according to claim 9, characterized in that: The 48aa connector is the sequence shown in SEQ ID NO.

10.

11. The DNA glycosylation enzyme composition according to claim 1, characterized in that: In the monomeric fusion protein, the DNA nicking enzyme and the guanine DNA glycosylation enzyme are linked by a 16aa linker.

12. The DNA glycosylation enzyme composition according to claim 11, characterized in that: In the L fusion protein, the DNA cleavage enzyme and the DNA exonuclease are linked by a 16aa adapter.

13. The DNA glycosylation enzyme composition according to claim 12, characterized in that: The 16aa connector is the sequence shown in SEQ ID NO.

11.

14. The DNA glycosylation enzyme composition according to claim 1, characterized in that: The guanine DNA glycosylation enzyme has the sequence shown in SEQ ID NO.

12.

15. An expression vector for expressing the DNA glycosylation enzyme composition according to any one of claims 1-14.

16. The expression vector according to claim 15, characterized in that: When the transport peptide of the DNA glycosylation enzyme composition is a plant mitochondrial transport peptide, the expression vector is driven by the rice Ubi promoter and / or the maize Ubi promoter; when the transport peptide of the DNA glycosylation enzyme composition is an animal mitochondrial transport peptide, the expression vector is driven by the human elongation factor EF-1α core promoter.

17. The expression vector according to claim 15, characterized in that: The expression vector has a replicon that reproduces in Escherichia coli and / or Agrobacterium.

18. The expression vector according to claim 15, characterized in that: The expression vector has a screening marker.

19. The expression vector according to claim 18, characterized in that: The screening markers are kanamycin resistance genes or ampicillin resistance genes.

20. The use of the DNA glycosylation enzyme composition according to any one of claims 1-14 or the expression vector according to any one of claims 15-19 in the preparation of reagents for editing the rice mitochondrial Cob gene or the human mitochondrial Cox2 gene.

Citation Information

Patent Citations

  • Base editor and application thereof

    CN117384885A

  • Recombinant protein capable of realizing wide-range adenine base editing on organelle DNA, nucleic acid, vector and application

    CN118165122A