Base editor ACGBEmax for simultaneously realizing purine and pyrimidine substitution
By developing the ACGBEmax base editor, which consists of a variety of proteins and glycosylases, can simultaneously realize the replacement of purine and pyrimidine, solving the problems of mutation diversity and low editing efficiency in the prior art, achieving efficient base editing and reducing the Index rate.
Patent Information
- Application Number
- CN202510270599.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-03-07
AI Technical Summary
Existing base editors cannot achieve the replacement of purine and pyrimidine at the same time, which limits the diversity and editing efficiency of mutations, and produces high Indels, affecting the effectiveness of mutants.
A base editor called ACGBEmax is developed, which consists of the HMCES protein, the dual-function deaminase TadDual, the nCas9(D10A) protein and the engineered N-methylpurine DNA glycosylase eMPG, which enables the replacement of purine and pyrimidine at the same time.
ACGBEmax significantly increases the diversity of targeted editing mutants, has high editing efficiency and low Indels rate, and is suitable for protein mutation screening and identification of oncogenic amino acid mutations in vitro and in vitro.
Smart Images

Figure CN120060241A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of gene editing technology, and particularly relates to a base editor ACGBEmax that can simultaneously achieve purine and pyrimidine substitutions. Background Art
[0002] With the development of high-throughput genomic sequencing technology, a large number of human genetic diseases have been proven to be related to gene mutations. Similarly, the production traits of livestock and crops are also affected by gene mutations. However, the relationship between most mutations and their related phenotypes remains poorly understood. Importantly, using gene editing technology to induce mutations in situ in target genes has become a powerful method to clarify the relationship between these genotypes and phenotypes. In addition, in protein directed evolution, introducing mutations into target proteins is a key step in creating a protein mutant library, from which better variants can be screened. Using gene editing tools can efficiently introduce mutations into target proteins in mammalian cells, making it possible to achieve protein directed evolution in mammalian cells.
[0003] CRISPR / Cas9-based gene editing technologies usually induce insertions or deletions, which often lead to loss of gene function, so they are not very suitable for studying the effects of specific nucleotide mutations. Although homologous recombination (by combining Cas9 and donor DNA) can theoretically introduce the desired mutations, the low efficiency of this method limits its practical application in inducing precise nucleotide mutations. Recently, base editors (BEs) and Prime editors (PEs) emerged as powerful tools that can induce base substitutions at target genomic sites in mammalian cells, and they are very valuable for generating saturated mutation libraries of SNVs and MNVs.
[0004] Cytosine base editors (CBEs) and adenine base editors (ABEs) are two main single-base editing technologies. CBEs convert C-G to T-A, while ABEs convert A-T to G-C. These editors directly induce base conversions through deamination reactions without causing double-strand breaks (DSBs). However, CBEs and ABEs are limited to single-base conversions, which restricts their application in site-directed saturation mutagenesis. To expand the diversity of base mutations, by fusing cytosine and adenine deaminases with nCas9(D10A), dual-base editors were developed that can simultaneously achieve the conversion of C to T and A to G through a single sgRNA. These editors, such as A&C-BEmax, SPACE, and STEME, expand the scope of base editing. However, whether single-base editors or dual-base editors, they are still limited to conversion mutations and cannot induce transversion mutations, which further restricts the diversity of mutations. To overcome this limitation, engineering glycosylases (such as uracil DNA glycosylase, UNG, and N-methylpurine DNA glycosylase, MPG) are integrated into base editors, which can induce transversion mutations such as C-G or A-Y, thus expanding the mutation spectrum. In addition, researchers have also developed deaminase-independent base editors. For example, an engineered UNG mutant is fused with nCas9(D10A) to induce substitution and transversion mutations at target T and C sites. Another deaminase-independent guanine base editor (gGBE) has also been developed by fusing nCas9(D10A) with an engineered MPG variant to achieve editing of G at the target site. Despite these advances, existing base editors still cannot meet the needs of saturation mutagenesis. Summary of the Invention
[0005] To solve the above technical problems, the object of the present invention is to provide a base editor ACGBEmax that can simultaneously achieve purine and pyrimidine substitutions. This base editor can simultaneously achieve purine and pyrimidine substitutions, significantly increasing the diversity of mutants after targeted editing, and having the advantages of high editing efficiency and a low Indels rate, and has great application value in aspects such as in vitro and in vivo protein mutation screening and identification of carcinogenic amino acid mutations.
[0006] The technical solution of the present invention to solve the above technical problems is as follows: Provide a base editor ACGBEmax that can simultaneously achieve purine and pyrimidine substitutions. The base editor ACGBEmax sequentially includes an HMCES protein, a bifunctional deaminase TadDual, an nCas9(D10A) protein, and an engineered N-methylpurine DNA glycosylase eMPG from the N-terminus to the C-terminus.
[0007] Furthermore, the nucleotide sequence encoding the HMCES protein is as shown in SEQ ID NO.1.
[0008] Furthermore, the nucleotide sequence encoding the bifunctional deaminase TadDual is as shown in SEQ ID NO.2.
[0009] Furthermore, the nucleotide sequence encoding the nCas9(D10A) protein is as shown in SEQ ID NO.3.
[0010] Furthermore, the nucleotide sequence encoding the engineered N-methylpurine DNA glycosylase eMPG is as shown in SEQ ID NO.4.
[0011] The present invention also provides a gene encoding the above base editor ACGBEmax, and its nucleotide sequence is as shown in SEQ ID NO.5.
[0012] The present invention also provides an expression vector including the above gene.
[0013] The present invention also provides a host cell comprising the above expression vector.
[0014] The present invention also provides the use of the above base editor ACGBEmax, gene, expression vector or host cell in gene editing.
[0015] The present invention has the following beneficial effects: 1. The base editor of the present invention can efficiently induce A, C, and G mutations in vivo and in vitro, and has the ability to generate diverse base substitution mutations at the editing site.
[0016] 2. The concept of ACGBEmax as an effective tool can effectively mediate random mutations of endogenous genes, generate diverse protein mutant libraries in situ for functional or phenotypic screening, and has great potential in the analysis of carcinogenic amino acid mutations. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is a schematic diagram for the design and optimization of adenine / cytosine / guanine base editor (ACGBE); Figure 2 is a distribution diagram of base editing results at genomic loci HPRT, PDCD1, CIITA, VEGFA, and DNMT3A; Figure 3 is a distribution diagram of base editing results at genomic loci B2M, ACE2, and PCSK9; Figure 4 is a statistical result diagram of the editing efficiency of ACGBEv1, ACGBEv2, and ACGBEv3; Figure 5 is a statistical result diagram of the Indels rate of ACGBEv1, ACGBEv2, and ACGBEv3; Figure 6 is a statistical result diagram of the editing efficiency of ACGBEv3, ACGBEv4, and ACGBEv5; Figure 7 is a statistical result diagram of the Indels rate of ACGBEv3, ACGBEv4, and ACGBEv5; Figure 8 is a characterization result diagram of the base editing activity of ACGBEmax; Figure 9 is a statistical result diagram of the number of mutations at the HPRT-site2 mutation site at the DNA level; Figure 10 is a statistical result diagram of the number of mutations at the HPRT-site2 mutation site at the amino acid level; Figure 11Results graph for orthogonal R-loop analysis to evaluate Cas9-independent off-target effects; Figure 12 Results graph for deep RNA sequencing to evaluate the off-target effects of ACGBEmax; Figure 13 Results graph for detecting the fold change of amino acid mutations; Figure 14 Results graph for statistics of the incidence of liver tumors in experimental mice induced by ACGBEmax; Figure 15 Results graph for immunohistochemistry (IHC) analysis; Figure 16 Results graph for next-generation sequencing (NGS) of mouse liver tumor tissues. Detailed implementation mode
[0018] The principles and features of the present invention are described below. The examples given are only used to explain the present invention and are not intended to limit the scope of the present invention. For those not specified in the examples, they are carried out according to conventional conditions or the conditions recommended by the manufacturer. For reagents or instruments not specified for the manufacturer, they are all conventional products that can be obtained through commercial purchase.
[0019] Example 1 Design and optimization of ACGBEs Schematic diagram for the design and optimization of adenine / cytosine / guanine base editor (ACGBE) is as Figure 1 shown.
[0020] (1) To achieve simultaneous editing of A, C, and G bases, first, the cytosine deaminase hA3A (Y130F) and the adenine deaminase TadA8e were fused to the N-terminus of nCas9 (D10A), and at the same time, the engineered uracil glycosylase (eMPG) was linked to the C-terminus of nCas9 (D10A) to construct ACGBEv1. Subsequently, hA3A (Y130F) was replaced with TadA-derived TadCDd to construct ACGBEv2. To obtain a more compact protein structure, the cytosine and adenine deaminases were replaced with the bifunctional deaminase TadDual derived from the TadA variant to form ACGBEv3.
[0021] To compare these three ACGBE versions, their editing activities at multiple endogenous genomic loci were evaluated in HEK293T cells. First, the plasmids encoding ACGBEs and the sgRNA expression vectors were transiently transfected into the cells, and the base mutations at the target sites were analyzed 48 hours after transfection. The results are as Figure 2 and 3As shown, among the eight genomic loci tested (HPRT, PDCD1, CIITA, VEGFA, DNMT3A, B2M, ACE2, and PCSK9), all three base editors efficiently edited A, C, and G bases within the editing window. Notably, since none of these three ACGBEs contain uracil DNA glycosylase inhibitor (UGI), the uracil generated by cytosine deamination was subsequently removed by endogenous UNG, forming abasic sites (AP sites). These AP sites are repaired through the base excision repair (BER) pathway, resulting in cytosine-to-thymine (C-to-T) mutations, and occasionally cytosine-to-guanine (C-G) and a small number of cytosine-to-adenine (C-A) conversions. Similarly, eMPG-mediated guanine (G) removal activates the BER pathway, inducing G-to-other-base mutations. Unexpectedly, eMPG also exhibits a certain hypoxanthine (I) glycosylase activity, leading to A-to-other-base mutations.
[0022] In addition, all three versions of ACGBEs generated a certain proportion of insertions and deletions (Indels), which may be due to the inevitable simultaneous generation of AP sites and nicking sites by glycosylase and Cas9 (D10A) nickase in the non-target and target DNA strands, respectively, leading to the formation of double-strand breaks (DSBs). It can be observed that these three ACGBEs showed different editing efficiencies and patterns for different bases within the editing window.
[0023] The statistical results of the editing efficiencies and Indels rates of the three ACGBEs are shown respectively in Figure 4 and 5 As shown, when considering the overall editing of the eight endogenous genomic loci, the editing window sizes of ACGBEv1, ACGBEv2, and ACGBEv3 were similar, but the total base editing efficiency of ACGBEv1 (average 70.81%) was significantly lower than that of ACGBEv2 (average 75.07%) and ACGBEv3 (average 73.26%). In addition, the Indels rate of ACGBEv1 (average 11.76%) was significantly higher than that of ACGBEv2 (average 6.93%) and ACGBEv3 (7.76%). The base editing efficiencies and Indels rates of ACGBEv2 and ACGBEv3 were comparable. Given its more compact structure, we selected ACGBEv3 for further optimization.
[0024] (2) Although the above ACGBEs can achieve various base substitution mutations, the Indels they generate may lead to frameshift mutations, hindering the generation of effective variants and thus affecting the screening effect. To reduce Indels, nCas9 (D10A) in ACGBEv3 was replaced with inactivated Cas9 (dCas9) to generate ACGBEv4. The AP-site protecting protein HMCES was also introduced into ACGBEv3 to generate ACGBEv5 (i.e., ACGBEmax).
[0025] Then, the base editing capabilities of ACGBEv4 and ACGBEv5 at the same eight endogenous genomic loci were further tested in HEK293T cells, and the statistical results of the editing efficiency and Indels rate are shown in Figure 6 and 7 respectively. The results showed that ACGBEv4 effectively induced A, C, and G mutations, but the Indels rate was significantly reduced (1.81%), at the cost of a decrease in editing efficiency (45.16%). In contrast, the editing efficiency of ACGBEv5 was comparable to that of ACGBEv3 (70.06%), while significantly reducing the Indels rate (5.81%). Therefore, ACGBEmax was selected for further characterization.
[0026] (3) The base editing activity of ACGBEmax was characterized at ten additional endogenous genomic loci in HEK293T cells, and the results are shown in Figure 8 respectively. The results demonstrated its high efficiency in generating diverse base substitution mutations at different loci. For A bases, A-to-G mutations were the most frequent (0.60% to 69.53%), followed by A-to-C (0.17% to 22.92%) and A-to-T mutations (0.11% to 7.84%). For C bases, C-to-T mutations were dominant (1.85% to 36.35%), followed by C-to-G (0.26% to 19.03%), and C-to-A mutations were less common (0.11% to 6.02%). For G bases, G-to-C (0.20% to 32.98%) and G-to-T (0.35% to 15.05%) were the most common mutations, while G-to-A mutations were less frequent (0.21% to 6.22%).
[0027] In addition, ACGBEmax was also evaluated in two human cancer cell lines, HeLa and Hu7, and the results also showed that it could effectively induce A, C, and G mutations. The above results indicate that ACGBEmax can generate diverse base substitution mutations at multiple endogenous genomic loci and has a certain preference for certain specific mutation types.
[0028] Example 2 ACGBEmax generates more diverse mutations than other base editors Select the HPRT-site2 mutation site in HEK293T cells and compare the base editing results of ACGBEmax with those of traditional CBE, ABE, and ACBE. The statistical results of the number of mutations at the DNA and amino acid levels are shown respectively as Figure 9 and 10 shown.
[0029] The results showed that at the DNA level, the number of mutation types induced by ACGBEmax (25 - 42 mutation types per sgRNA) was higher than that of ABE (10 - 18 types), CBE (7 - 25 types), and ACBE (7 - 32 types); similarly, at the amino acid level, ACGBEmax produced 10 - 35 mutation types per sgRNA, far exceeding ABE (5 - 18 types), CBE (2 - 24 types), and ACBE (2 - 12 types). This indicates that ACGBEmax can generate more abundant amino acid mutations in endogenous protein-coding genes, which has advantages for protein mutation screening.
[0030] Example 3 ACGBEmax induces low levels of Cas9-dependent and -independent off-target editing To evaluate the potential off-target effects of ACGBEmax, we first used the Cas-OFFinder software to predict the possible Cas9-dependent off-target sites of three sgRNAs targeting PDCD1, VEGFA, and HPRT. Through next-generation sequencing (NGS) analysis, we found that except for the HPRT-OT4 site, ACGBEmax did not detect off-target effects, and its off-target effects were significantly lower than those of CBE, ABE, and ACBE. We also used orthogonal R-loop analysis to evaluate the Cas9-independent off-target effects. The results are shown as Figure 11 shown. Among the six R-loops generated by inactivated SaCas9, the off-target effects of ACGBEmax were significantly lower than those of ABE, CBE, and ACBE. In addition, we also evaluated the off-target effects of ACGBEmax at the RNA level by deep RNA sequencing. The results are shown as Figure 12 shown. The results showed that the number of A-to-I mutations induced by ACGBEmax was slightly higher than that of the control group, but the number of C-to-U mutations was comparable to that of the control group.
[0031] These results indicate that while achieving efficient targeted editing, ACGBEmax minimizes off-target effects at the DNA and RNA levels.
[0032] Example 4 ACGBEmax mediates random mutations of endogenous genes for phenotypic screening To verify the potential of ACGBEmax in generating a saturated protein mutant library for functional screening studies, we utilized ACGBEmax to introduce and screen for in situ mutations at the HPRT gene locus in human HHL-5 cells. Some single amino acid mutations in the HPRT protein can confer drug resistance to 6-thioguanine (6TG) in cells. Therefore, we established a proof-of-concept mutation screening experiment with 6TG resistance as the readout to evaluate the HPRT mutant library generated by ACGBEmax.
[0033] First, we evaluated the sensitivity of HHL-5 cells to 6TG, and the results showed that 100 μM of 6TG was sufficient to induce cell death. Subsequently, we designed 33 sgRNAs to target exons 3-8 of the HPRT gene for evaluation. Due to the read length limitation of next-generation sequencing and to eliminate the potential linkage effect of mutations between different exons, we co-transfected cells with sgRNAs (forward and reverse) targeting different exons and ACGBEmax, and then mixed the cells. After screening for successfully transfected cells with puromycin, 100 μM of 6TG was added for further screening. Compared with the control group, monoclonal formation was observed in the HPRT mutant cell population. After 6TG screening, exons 3, 4, 6, 7, and 8 of the HPRT gene in the surviving cells were analyzed by next-generation sequencing (NGS) to identify the corresponding HPRT mutants. These 6TG-screened cells were compared with cells treated with DMSO to calculate the fold change of HPRT protein mutants generated by ACGBEmax. The results are as Figure 13 shown that compared with the untreated group, the 6TG-screened cells were enriched in mutations in Exon3 and Exon8, among which p.C66R, p.G70K, p.G70W, p.G190R, and p.G190A were recurrent high-frequency mutations.
[0034] To verify the truly enriched mutations, we constructed vectors expressing mutant and wild-type HPRT and transfected them into the HPRT gene knockout (KO) HHL-5 cell line for 6TG sensitivity analysis. The results showed that HPRT-KO cells and cells supplemented with the corresponding mutants were completely resistant to 6TG treatment, while cells expressing wild-type HPRT died in 6TG medium. These findings demonstrated the concept of ACGBEmax as an effective tool capable of generating diverse protein mutant libraries in situ for functional screening.
[0035] Example 5 ACGBEmax Achieves In Vivo Mutation of Endogenous Genes in Mouse Liver to Identify Carcinogenic Amino Acid Mutations in CTNNB1 Protein Next, aiming to verify the effectiveness of ACGBEmax in in vivo protein mutagenesis and screening, CTNNB1 was selected as the target gene for in situ mutagenesis by ACGBEmax at the single-base level, with the aim of identifying new oncogenic mutations in CTNNB1 that might promote liver tumor formation in mice. For this purpose, 9 forward and 11 reverse sgRNAs were designed to target the key functional domains in exon 3 of CTNNB1. To induce liver tumor formation, ACGBEmax and the sgRNA plasmid targeting CTNNB1 were hydrodynamically injected via the tail vein into wild-type mice, along with the c-Myc overexpression plasmid. To avoid the generation of double-strand breaks (DSBs), the sgRNAs targeting CTNNB1 were divided into two groups: one group was the forward sgRNAs, and the other group was the reverse sgRNAs. In addition, to control the potential interference of background c-Myc expression, we also set up non-targeting sgRNAs as a control group.
[0036] Thirty days after injection, the mice were dissected. No liver tumors were observed in the control group. In contrast, liver tumors appeared in 2 out of 10 mice injected with forward sgRNAs and in 3 out of 10 mice injected with reverse sgRNAs. Extending the observation time (up to 60 days) showed that no tumor formation occurred in the control group, indicating that overexpressing c-Myc alone was not sufficient to induce liver tumor formation within this time range. In contrast, as Figure 14 shown, the CTNNB1 mutations induced by ACGBEmax significantly increased the incidence of liver tumors in experimental mice.
[0037] Fluorescence microscopy analysis showed that strong and uniform green fluorescence was presented in the tumor tissues of experimental mice, while only weak fluorescence was observed in the livers of control group mice, indicating that the CTNNB1 mutations induced by ACGBEmax were crucial for tumor formation. In addition, the results of immunohistochemistry (IHC) analysis were as Figure 15 shown, the expression of CTNNB1 in tumor tissues was significantly increased, and CTNNB1 accumulated in the nucleus, which is characteristic of activating CTNNB1 mutations.
[0038] Furthermore, the next-generation sequencing (NGS) results of the liver tumor tissues of experimental mice were as Figure 16As shown. The results showed that several mutations of CTNNB1 were enriched in the forward sgRNA group and the reverse sgRNA group. The two most common single amino acid mutations were p.D32G (caused by the mutation from A to G) and p.L63V (caused by the mutation from G to C). CTNNB1 p.D32G is a known oncogenic mutation, while p.L63V has not been reported previously. In addition, complex mutant alleles involving point mutations and deletions were also enriched in tumors. These results highlight the effectiveness of ACGBEmax in inducing in vivo mutations of endogenous genes and further confirm its great application value in identifying oncogenic mutations.
[0039] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A base editor ACGBEmax that simultaneously achieves purine and pyrimidine substitution, characterized in that: The base editor ACGBEmax includes HMCES protein, bifunctional deaminase TadDual, nCas9 (D10A) protein and engineered N-methylpurine DNA glycosylase eMPG from N-terminus to C-terminus.
2. The base editor ACGBEmax for simultaneously replacing purine and pyrimidine according to claim 1, characterized in that: The nucleotide sequence encoding the HMCES protein is shown in SEQ ID NO.
1.
3. The base editor ACGBEmax for simultaneously replacing purine and pyrimidine according to claim 1, characterized in that: The nucleotide sequence encoding the bifunctional deaminase TadDual is shown in SEQ ID NO.
2.
4. The base editor ACGBEmax for simultaneously replacing purine and pyrimidine according to claim 1, characterized in that: The nucleotide sequence encoding the nCas9 (D10A) protein is shown in SEQ ID NO.
3.
5. The base editor ACGBEmax for simultaneously replacing purine and pyrimidine according to claim 1, characterized in that: The nucleotide sequence encoding the engineered N-methylpurine DNA glycosylase eMPG is shown in SEQ ID NO.
4.
6. A gene encoding the base editor ACGBEmax according to claim 1, characterized in that Its nucleotide sequence is shown in SEQ ID NO.
5.
7. An expression vector comprising the gene according to claim 6.
8. A host cell comprising the expression vector of claim 7.
9. Use of the base editor ACGBEmax described in any one of claims 1 to 5, the gene described in claim 6, the expression vector described in claim 7 or the host cell described in claim 8 in gene editing.
Citation Information
Patent Citations
Base editing system for realizing C to A and C to G base mutation and application thereof
CN111763686A
High-efficiency and high-precision base editor for conversion from cytosine C to guanine G
CN115703842A
Targeted mutagenesis system based on adenine and cytosine double-base editor
CN115704015A
Fusion protein, uracil-N-glycosylase mutant-mediated base editing system and application of fusion protein and uracil-N-glycosylase mutant-mediated base editing system
CN117126827A
A:t to c:g base editors and uses thereof
WO2020181180A1