C-to-G double-enzyme synergistic base editor with high efficiency and wide targeting range and application of C-to-G double-enzyme synergistic base editor
By designing a dual-enzyme synergistic C-to-G base editor that integrates cytosine deaminase and cytosine DNA glycosylase, the problems of low efficiency and limited targeting range of existing C-to-G base editors are solved, achieving efficient and broad-spectrum base editing effects.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SANYA INSTITUTE OF NANJING AGRICULTURAL UNIVERSITY
- Filing Date
- 2025-12-30
- Publication Date
- 2026-05-01
AI Technical Summary
Existing C-to-G base editors have low editing efficiency and limited targeting range, making it difficult to meet the needs of a wide range of scientific research and applications.
A dual-enzyme synergistic C-to-G base editor was designed, which integrates cytosine deaminase and cytosine DNA glycosylase. Guided to specific targets by sgRNA, it achieves efficient and broad-spectrum base editing through synergistic action.
It improves the efficiency and targeting flexibility of C-to-G base editing, expands the editing window, and meets a wider range of scientific research and application needs.
Smart Images

Figure CN121950753A_ABST
Abstract
Description
A highly efficient and broad-targeting C-to-G dual-enzyme synergistic base editor and its applications Technical Field
[0001] This invention belongs to the field of gene editing technology, and particularly relates to base editing technology based on the CRISPR system. More specifically, it relates to a novel gene editing tool capable of achieving cytosine-to-guanine (C-to-G) base transversion, its construction method, and its application. Background Technology
[0002] The CRISPR-Cas system, as a revolutionary gene-editing tool, has greatly advanced the development of life sciences and biotechnology. Traditional CRISPR-Cas9 technology introduces double-strand breaks (DSBs) at the target DNA site, relying on the cell's own repair mechanisms (such as NHEJ or HDR) to achieve gene editing. However, the introduction of DSBs is often accompanied by unpredictable insertions or deletions (indels), and efficiency is difficult to guarantee when relying on the inefficient HDR pathway for precise repair.
[0003] To overcome these shortcomings, base editing technology has emerged. Base editors (BEs) consist of an inactivated or nicked form of a Cas protein (dCas9 or nCas9) fused with a nucleic acid deaminase, enabling precise conversion of specific bases, such as C-to-T or A-to-G, without generating DSBs. These conversion editors have shown great potential in basic research, disease treatment, and crop breeding.
[0004] However, the existing base editing toolkit remains incomplete, especially for base transversions—the conversion between purines and pyrimidines—which present significant technical challenges. C-to-G transversions, in particular, are crucial for correcting approximately 10% of pathogenic point mutations and for introducing specific modifications in areas such as protein engineering and crop improvement.
[0005] Currently, reported C-to-G base editors (CGBEs) are mainly achieved through two different technical approaches. One mainstream strategy involves constructing deaminase-dependent CGBEs, typically composed of nCas9, cytosine deaminases (such as APOBEC1 and CDA1), and uracil DNA glycosyltransferase (UNG). The mechanism involves the deaminase converting target C to U, followed by UNG recognizing and cleaving the U to form an abase (AP) site. Cells then have a certain probability of inserting G when repairing this site. However, the core drawback of this type of editor lies in its generally low editing efficiency, narrow editing window, and strict positional preference (e.g., APOBEC1-CGBE exhibits the highest activity at C6, while CDA1-CGBE prefers C3), which severely limits its targeting range and application scenarios.
[0006] To address some of the issues, another class of glycosylation-dependent CGBEs has been developed. This system involves fusing nCas9 with an engineered cytosine DNA glycosylation enzyme (CDG) or thymine DNA glycosylation enzyme (TDG), which directly cleaves the target C to form an AP site, thereby inducing a C-to-G conversion. Although this strategy reduces deaminase byproducts (C-to-T) to some extent, its editing efficiency and targeting flexibility still need improvement.
[0007] In conclusion, there is an urgent need for a new type of CGBE system that should have higher editing efficiency, a wider or more flexible editing window, and maintain high specificity to meet a wider range of research and application needs. Summary of the Invention
[0008] The purpose of this invention is to solve the technical problems of low efficiency and limited targeting range of existing CGBE, and to provide a novel system for efficient and broad-targeted C-to-G base editing through the synergistic effect of two enzymes, as well as its construction method and application.
[0009] To achieve the above objectives, this invention proposes a dual-enzyme synergistic C-to-G base editor (hereinafter referred to as "dual-enzyme CGBE"). This editor is a fusion protein, and its core design concept is to synergistically integrate the functions of cytosine deaminase (CDA) and cytosine DNA glycosylation enzyme in a single editor molecule.
[0010] Specifically, the dual-enzyme CGBE fusion protein consists of three core functional parts: a modified Cas protein, a cytosine deaminase, and a cytosine DNA glycosylase. It uses a Cas protein with DNA-targeting ability but reduced or lost endonuclease activity as its localization module, preferably a Cas9 nickase (nCas9), such as the SpCas9(D10A) mutant, or a Cas9 inactivating protein (dCas9). Fused with this Cas protein is a cytosine deaminase domain, whose function is to deaminate cytosine (C) to generate uracil (U). The domain has broad selectivity in terms of origin and type, including PmCDA1 and its truncated mutants (such as CDA1Δ194) from moray eels, APOBEC1 and its mutants (such as A1(R33A)) from rats, and APOBEC3A and its mutants (such as eA3A) from humans. In addition, the protein integrates a cytosine DNA glycosylation enzyme domain, which can directly recognize and remove cytosine on the DNA strand to generate AP sites. Preferably, it is an engineered uracil DNA glycosylation enzyme, such as CDG4.
[0011] These three functionally distinct domains are linked by a suitable linker to form a single fusion protein. In use, this fusion protein works in conjunction with a single guide RNA (sgRNA). The sgRNA guides the fusion protein to a specific target site in the genome, whereupon the deaminase and glycosylation domains work synergistically to target C within the editing window.
[0012] The first objective of this invention is to provide a dual-enzyme CGBE editor comprising a cytosine DNA glycosylase CDG4, a cytosine deaminase, an SpCas9 nickase nCas9(D10A), and a nuclear localization signal NLS; wherein the cytosine DNA glycosylase CDG4 is fused to the N-terminus of the cytosine deaminase, and the cytosine deaminase is fused to the N-terminus of nCas9(D10A); wherein the cytosine deaminase is selected from any one or more of CDA1Δ194, A1(R33A), eA3A, mini-Sdd7, and TadA8e(N46L); wherein the amino acid sequence of CDA1Δ194 is as shown in SEQ ID NO.1, the amino acid sequence of A1(R33A) is as shown in SEQ ID NO.5, the amino acid sequence of eA3A is as shown in SEQ ID NO.6, the amino acid sequence of mini-Sdd7 is as shown in SEQ ID NO.7, and the amino acid sequence of TadA8e(N46L) is as shown in SEQ ID NO.1. As shown in NO.8; the amino acid sequence of the cytosine DNA glycosylase CDG4 is shown in SEQ ID NO.2; the amino acid sequence of the nCas9(D10A) is shown in SEQ ID NO.3.
[0013] Furthermore, the dual-enzyme CGBE editor also includes a nuclear localization signal NLS, with nCas9 (D10A) directly or indirectly fused to the N-terminus of the nuclear localization signal NLS; the amino acid sequence of the nuclear localization signal NLS is shown in SEQ ID NO.4.
[0014] Further, the dual-enzyme CGBE editor is selected from any one of the following (A1)-(A5): (A1) CDG4-CDA1Δ194-miniCGBE: composed of SEQ ID NO.2, SEQ ID NO.1, SEQ ID NO.3 and SEQ ID NO.4 in sequence; (A2) CDG4-A1(R33A)-miniCGBE: composed of SEQ ID NO.2, SEQ ID NO.5, SEQ ID NO.3 and SEQ ID NO.4 in sequence; (A3) CDG4-eA3A-miniCGBE: composed of SEQ ID NO.2, SEQ ID NO.6, SEQ ID NO.3 and SEQ ID NO.4 in sequence; (A4) CDG4-miniSdd7-miniCGBE: composed of SEQ ID NO.2, SEQ ID NO.7, SEQ ID NO.3 and SEQ ID NO.4 in sequence; (A5) CDG4-TadA8e(N46L)-miniCGBE: composed of SEQ ID NO.2, SEQ ID NO.8, SEQ ID NO.4 and SEQ ID NO.4 in sequence. The segments (A1)-(A5) formed by ID NO.3 and SEQ ID NO.4 may or may not include the linker sequence.
[0015] In a particular embodiment, the segments (A1)-(A5) include linker sequences, specifically as follows: (A1) CDG4-CDA1Δ194-miniCGBE: composed of SEQ ID NO.2, SEQ ID NO.14, SEQ ID NO.1, SEQ ID NO.3, SGGS linker, and SEQ ID NO.4 connected sequentially; (A2) CDG4-A1(R33A)-miniCGBE: composed of SEQ ID NO.2, SEQ ID NO.14, SEQ ID NO.5, SEQ ID NO.14, SEQ ID NO.3, SGGS linker, and SEQ ID NO.4 connected sequentially; (A3) CDG4-eA3A-miniCGBE: composed of SEQ ID NO.2, SEQ ID NO.14, SEQ ID NO.6, SEQ ID NO.14, SEQ ID NO.3, SGGS linker, and SEQ ID NO.4 connected sequentially; (A4) CDG4-miniSdd7-miniCGBE: composed of SEQ ID NO.2, SEQ ID NO.14, SEQ ID NO.14, SEQ ID NO.3, SGGS linker, and SEQ ID NO.4 connected sequentially. Composed of NO.14, SEQ ID NO.7, SEQ ID NO.14, SEQ ID NO.3, SGGS linker and SEQ ID NO.4; (A5)CDG4-TadA8e(N46L)-miniCGBE: Composed of SEQ ID NO.2, SEQ ID NO.14, SEQ ID NO.8, SEQ ID NO.14, SEQ ID NO.3, SGGS linker and SEQ ID NO.4 connected in sequence.
[0016] The dual-enzyme CGBE editor (A1)-(A5) is suitable for yeast expression vectors.
[0017] Preferably, the dual-enzyme CGBE editor is selected from any one of (A1), (A2), and (A3).
[0018] Furthermore, the N-terminus of the cytosine DNA glycosylase CDG4 of the dual-enzyme CGBE editor is also connected to a nuclear localization signal NLS, and nCas9(D10A) is not directly fused to the N-terminus of the nuclear localization signal NLS. There is also a nucleoplasmic protein NLS between nCas9(D10A) and the nuclear localization signal NLS; the amino acid sequence of the nucleoplasmic protein NLS is shown in SEQ ID NO.13.
[0019] Furthermore, the dual-enzyme CGBE editor is selected from any one of the following (B1)-(B2): (B1) P-CDG4-CDA1∆194-miniCGBE: composed of SEQ ID NO.4, SEQ ID NO.2, SEQ ID NO.1, SEQ ID NO.3, SEQ ID NO.13 and SEQ ID NO.4 in sequence; (B2) P-CDG4-eA3A-miniCGBE: composed of SEQ ID NO.4, SEQ ID NO.2, SEQ ID NO.6, SEQ ID NO.3, SEQ ID NO.13 and SEQ ID NO.4 in sequence; the fragments of (B1) and (B2) may or may not include a linker sequence.
[0020] In a particular embodiment, the segments of (B1) and (B2) include a linker sequence, as follows: (B1) P-CDG4-CDA1∆194-miniCGBE: composed of SEQ ID NO.4, SEQ ID NO.2, SEQ ID NO.16, SEQ ID NO.1, SEQ ID NO.3, SEQ ID NO.13 and SEQ ID NO.4 connected in sequence; (B2) P-CDG4-eA3A-miniCGBE: composed of SEQ ID NO.4, SEQ ID NO.2, SEQ ID NO.16, SEQ ID NO.6, SEQ ID NO.16, SEQ ID NO.3, SEQ ID NO.13 and SEQ ID NO.4 connected in sequence.
[0021] The dual-enzyme CGBE editor (B1)-(B2) is suitable for rice expression vectors.
[0022] Preferably, the dual-enzyme CGBE editor system is selected from either (B1) or (B2).
[0023] A second objective of this invention is to provide a dual-enzyme CGBE editor system, comprising the aforementioned dual-enzyme CGBE editor, and further comprising sgRNA or an sgRNA expression vector. By designing the dual-enzyme CGBE editor to form a complex with sgRNA, it is possible to target a target sequence and perform base editing.
[0024] A third objective of this invention is to provide a nucleic acid molecule encoding the aforementioned dual-enzyme CGBE editor or the aforementioned dual-enzyme CGBE editor system.
[0025] A fourth object of the present invention is to provide biological materials related to the aforementioned nucleic acid molecules, said biological materials being any of the following: (C1) an expression cassette containing the aforementioned nucleic acid molecules; (C2) a recombinant vector containing the aforementioned nucleic acid molecules, or a recombinant vector containing the expression cassette of (C1); (C3) a recombinant microorganism containing the aforementioned nucleic acid molecules, or a recombinant microorganism containing the expression cassette of (C1), or a recombinant microorganism containing the recombinant vector of (C2).
[0026] A fifth object of the present invention is to provide the use of the aforementioned dual-enzyme CGBE editor, the aforementioned dual-enzyme CGBE editor system, the aforementioned nucleic acid molecule, and the aforementioned biological material in any of the following (D1)-(D2): (D1) in the preparation of gene editing products; (D2) in improving the scope and / or efficiency of gene editing for non-disease treatment purposes.
[0027] Furthermore, the gene editing is C-to-G base editing.
[0028] A sixth objective of this invention is to provide a method for improving the scope and / or efficiency of gene editing for non-disease treatment purposes, using the aforementioned dual-enzyme CGBE editor system.
[0029] The dual-enzyme CGBE editor fused with cytosine deaminase and cytosine DNA glycosylase provided by this invention comprises cytosine deaminase from different sources, an engineered variant CDG4 derived from human uracil DNA glycosylase UNG, an SpCas9 nickase (nCas9(D10A)) derived from Streptococcus pyogenes, and a nuclear localization signal NLS fusion protein, collectively referred to as CDG4-Deaminase-miniCGBE.
[0030] In a specific implementation, the cytosine deaminase in the dual-enzyme CGBE editor is selected from any one of the following (E1)-(E5): (E1) PmCDA1 (hereinafter referred to as CDA1Δ194), a C-terminus truncated to 194 amino acids from petrmyzon marinus, and a dual-enzyme CGBE editor containing this deaminase is called CDG4-CDA1Δ194-miniCGBE; (E2) A variant A1 (R33A) from rat APOBEC1, and a dual-enzyme CGBE editor containing this deaminase is called CDG4-A1(R33A)-miniCGBE; (E3) A variant eA3A from human APOBEC3A, and a dual-enzyme CGBE editor containing this deaminase is called CDG4-eA3A-miniCGBE; (E4) Sdd7 from Actinosynnema mirum. The N-terminal truncated variant mini-Sdd7, containing this deaminase, is called CDG4-miniSdd7-miniCGBE; (E5) is derived from the TadA variant TadA8e(N46L) of Escherichia coli, containing this deaminase, and is called CDG4-TadA8e(N46L)-miniCGBE.
[0031] The CDA1Δ194 amino acid sequence includes the sequence shown in SEQ ID NO.1.
[0032] The CDG4 amino acid sequence described in SEQ ID NO.1:MTDAEYVRIHEKLDIYTFKKQFFNNKKSVSHRCYVLFELKRRGERRACFWGYAVNKPQSGTERGIHAEIFSIRKVEEYLRDNPGQFTINWYSSWSPCADCAEKILEWYNQELRGNGHTLKIWACKLYYEKNARNQIGLWNLRDNGVGLNVMVSEHYQCCRKIFIQSSHNQLNENRWLEKTLKRAEKRRSELSIM includes the sequence shown in SEQ ID NO.2.
[0033] The nCas9 (D10A) amino acid sequence described in SEQ ID NO.2:MFGESWKKHLSGEFGKPYFIKLMEFVAEERKHYTVYPPPHQVFTWTQMCDIKDVKVVILGQDPYHGPNQAHGLCFSVQRPVPPPPSLENIYEELSTDIEGFVHPGHGDLSGWAKQGVLLLDAVLTVRAHQANSHKEQGWEQFTDAVVSWLNQNSNGLVFLLWGSHAQKKGSAIDRKRHHVLQAAHPSPLSAHRGFFGCRHFSKTNELLQKSGKKPIDWTEL includes the sequence shown in SEQ ID NO.3.
[0034]
[0035] The A1 (R33A) amino acid sequence described in SEQ ID NO.4:PKKKRKV includes the sequence shown in SEQ ID NO.5.
[0036] SEQ ID NO.5: The eA3A amino acid sequence described in MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELAKETCLLYEINWGGRHSIWRHTSQNTNKHVEVNFIEKFTTERYFCPNTRCSITWFLSWSPCGECSRAITEFLSRYPHVTLFIYIARLYHHADPRNRQGLRDLISSGVTIQIMTEQESGYCWRNFVNYSPSNEAHWPRYPHLWVRLYVLELYCIILGLPPCLNILRRKQPQLTFFTIALQSCHYQRLPPHILWATGLK includes the sequence shown in SEQ ID NO.6.
[0037] The mini-Sdd7 amino acid sequence described in SEQ ID NO.6: MEASPASGPRHLMDPHIFTSNFNNGIGRHKTYLCYEVERLDNGTSVKMDQHRGFLHGQAKNLLCGFYGRHAELRFLDLVPSLQLDPAQIYRVTWFISWSPCFSWGCAGEVRAFLQENTHVRLRIFAARIYDYDPLYKEALQMLRDAGAQVSIMTYDEFKHCWDTFVDHQGCPFQPWDGLDEHSQALSGRLRAILQNQGN includes the sequence shown in SEQ ID NO.7.
[0038] The TadA8e(N46L) amino acid sequence described in SEQ ID NO.7: MEGGGPGAVPEGGDGPPAVPAEEVERLRGELPPPVVPGTGQKTHGRWIGPDGRVRAIVSGRDEDAALVHAQLAAKGIPDEPTRNSDVEQKLAAHMVANGIRHVTLVINHRPCRGFDDSCDTLVPIILPEGCTLTVHGQTDKGMRVRVRYTGGARPWWS includes the sequence shown in SEQ ID NO.8.
[0039] SEQ ID NO.8: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWLRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN The embedded structure described (i.e., the editor whose name contains CE) contains nCas9 1-1046 The amino acid sequence includes the sequence shown in SEQ ID NO.9.
[0040] 1063-1367 The amino acid sequence includes the sequence shown in SEQ ID NO.10.
[0041] SEQ ID NO.10: ETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQ KGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD The hCDG amino acid sequence includes SEQ The sequence shown in ID NO.11.
[0042] SEQ ID NO.11: The CDA1 amino acid sequence described herein includes the sequence shown in SEQ ID NO.12.
[0043] SEQ ID NO.12: The nucleoplasmic protein NLS amino acid sequence described herein includes the sequence shown in SEQ ID NO.13.
[0044] The 16 aa XTEN linker amino acid sequence described in SEQ ID NO.13: KRPAATKKAGQAKKKK includes the sequence shown in SEQ ID NO.14.
[0045] The SGGS linker amino acid sequence described in SEQ ID NO.14:SGSETPGTSESATPES includes the sequence shown below.
[0046] The GSSGS linker amino acid sequence described in SGGS includes the sequence shown in SEQ ID NO.15.
[0047] The 18 aa XTEN linker amino acid sequence described in SEQ ID NO.15:GSSGS includes the sequence shown in SEQ ID NO.16.
[0048] The TadA8e amino acid sequence described in SEQ ID NO.16:SGSETPGTSESATPESLK includes the sequence shown in SEQ ID NO.17.
[0049] The nucleotide sequence of the yeast expression vector encoding the CDA1Δ194 amino acid sequence shown in SEQ ID NO.17:MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN is shown in SEQ ID NO.18.
[0050] SEQ ID NO.18: ATGACCGACGCTGAGTACGTGAGAATCCATGAGAAGTTGGACATCTACACGTTTAAGAAACAGTTTTTCAACAACAAAAAATCCGTGTCGCATAGATGCTACGTTCTCTTTGAATTAAAACGACGGGGTGAACGTAGAGCGTGTTTTTGGGGCTATGCTGTGAATAAACCACAGAGCGGGACAGAACGTGGCATTCACGCCGAAATCTTTAGCATTAGAAAAGTCGAAGAATACCTGCGCGACAACCCCGGACAATTCACGATAAATTGGTACTCATCCTGGAGTCCTTGTGCAGATTGCGCTGAAAAGATCTTAGAATGGTATAACCAGGAGCTGCGGGGGAACGGCCACACTTTGAAAATCTGGGCTTGCAAACTCTATTACGAGAAAAATGCGAGGAATCAAATTGGGCTGTGGAACCTCAGAGATAACGGGGTTGGGTTGAATGTAATGGTAAGTGAACACTACCAATGTTGCAGGAAAATATTCATCCAATCGTCGCACAATCAATTGAATGAGAATAGATGGCTTGAGAAGACTTTGAAGCGAGCTGAAAAACGACGGAGCGAGTTGTCCATTATG. The nucleotide sequence of the rice expression vector encoding the CDA1Δ194 amino acid sequence shown in SEQ ID NO.1 is as shown in SEQ ID NO.19.
[0051] SEQ ID NO.19: ATGACCGATGCGGAATATGTTCGTATTCATGAAAAGCTTGATATCTACACATTCAAGAAGCAATTCTTCAACAATAAGAAATCCGTCTCGCACCGCTGCTACGTCTTGTTTGAGCTCAAAAGGCGCGGCGAGCGGCGCGCATGTTTCTGGGGCTATGCCGTGAACAAGCCGCAAAGTGGGACTGAACGCGGCATCCATGCTGAGATATTTTCCATTAGGAAGGTGGAGGAGTATCTACGGGACAACCCAGGGCAGTTCACCATAAACTGGTATTCAAGCTGGAGCCCCTGTGCCGACTGCGCCGAAAAAATTCTTGAGTGGTACAACCAGGAGCTTCGGGGTAATGGTCATACTCTCAAGATCTGGGCGTGCAAGCTGTACTACGAGAAGAATGCGAGGAATCAGATCGGACTCTGGAACCTGAGAGATAATGGAGTTGGCTTGAATGTCATGGTGTCTGAGCACTACCAGTGCTGCAGAAAAATATTTATCCAATCTTCGCACAACCAATTGAATGAGAACAGATGGCTCGAAAAAACGCTGAAACGAGCAGAAAAGAGGAGGAGTGAGCTGTCCATCATG The nucleotide sequence of the yeast expression vector encoding the CDG4 amino acid sequence shown in SEQ ID NO.2 is shown in SEQ ID NO.20.
[0052] SEQ ID NO.20: ATGTTCGGTGAATCTTGGAAGAAACATTTGTCCGGTGAATTCGGTAAACCTTATTTCATCAAGTTGATGGAATTTGTTGCTGAAGAAAGAAAGCACTACACCGTCTACCCACCACCACACCAAGTTTTCACCTGGACCCAAATGTGTGATATCAAGGATGTCAAGGTTGTTATTTTGGGTCAAGACCCATACCACGGTCCAAACCAAGCCCATGGCTTGTGTTTCTCCGTCCAAAGACCAGTTCCACCTCCACCATCCTTGGAAAACATTTACGAAGAATTATCTACTGACATTGAAGGTTTCGTTCACCCAGGTCACGGTGACTTGTCTGGTTGGGCTAAGCAAGGTGTTTTGTTATTAGATGCTGTCTTGACTGTCAGAGCTCACCAAGCTAATTCCCACAAAGAACAAGGTTGGGAACAATTCACTGATGCCGTCGTTTCCTGGTTGAACCAAAACTCTAACGGTTTGGTCTTCTTGTTATGGGGTTCTCACGCTCAAAAGAAGGGTTCTGCTATCGACAGAAAGCGTCACCACGTTCTACAAGCTGCTCATCCATCTCCATTGTCAGCCCACAGAGGTTTCTTTGGTTGTAGACATTTCAGTAAGACCAACGAATTGTTGCAAAAGTCTGGTAAGAAGCCAATCGACTGGACTGAATTG. The nucleotide sequence of the rice expression vector encoding the CDG4 amino acid sequence shown in SEQ ID NO.2 is shown in SEQ ID NO.21.
[0053] SEQ ID NO.21: ATGTTCGGCGAGTCTTGGAAGAAGCACCTCTCAGGTGAGTTTGGTAAGCCGTACTTCATTAAGCTGATGGAGTTCGTTGCGGAAGAGCGGAAGCATTACACTGTGTATCCGCCACCGCACCAGGTCTTTACTTGGACACAGATGTGCGACATCAAAGACGTGAAGGTGGTGATTCTGGGTCAAGATCCGTACCACGGTCCAAATCAGGCGCATGGCCTCTGCTTCAGCGTCCAGCGGCCAGTGCCGCCACCGCCTAGCCTGGAGAATATCTACGAAGAACTCAGCACCGACATCGAGGGTTTCGTTCATCCGGGTCACGGCGACCTGTCGGGATGGGCGAAGCAGGGCGTGCTCTTGCTTGATGCAGTGCTCACCGTGAGGGCGCATCAGGCCAACTCTCATAAAGAGCAGGGATGGGAACAGTTCACCGACGCTGTGGTGTCTTGGCTGAACCAGAACAGCAACGGCCTGGTGTTTCTGCTCTGGGGTAGCCATGCTCAGAAGAAGGGCAGCGCCATTGATCGGAAGAGACACCACGTGCTCCAAGCTGCGCATCCATCACCACTGTCAGCTCACAGAGGCTTCTTCGGTTGCAGGCATTTCAGCAAGACTAATGAACTGCTCCAGAAATCTGGCAAGAAGCCGATTGATTGGACGGAGTTG. The nucleotide sequence of the yeast expression vector encoding the amino acid sequence of nCas9 (D10A) shown in SEQ ID NO.3 is shown in SEQ ID NO.22.
[0054]
[0055]
[0056] The nucleotide sequence of the rice expression vector encoding the SV40 NLS amino acid sequence shown in SEQ ID NO. 4, SEQ ID NO. 25, is shown in SEQ ID NO. 25. SEQ ID NO. 24: CCCAAGAAGAAGAGGAAGGTG
[0057] SEQ ID NO.25: CCAAAGAAGAAGCGGAAGGTG encodes the nucleotide sequence of the yeast expression vector with the amino acid sequence A1 (R33A) shown in SEQ ID NO.5, as shown in SEQ ID NO.26.
[0058] SEQ ID NO.26: ATGAGTTCCGAGACAGGCCCTGTAGCTGTTGATCCCACTCTGAGGAGAAGAATTGAGCCCCACGAGTTTGAAGTCTTCTTTGACCCCCGGGAACTTGCGAAAGAGACCTGTCTGCTGTATGAGATCAACTGGGGAGGAAGGCACAGCATCTGGCGACACACGAGCCAAAACACCAACAAACACGTTGAAGTCAATTTCATAGAAAAATTTACTACAGAAAGATACTTTTGTCCAAACACCAGATGCTCCATTACCTGGTTCCTGTCCTGGAGTCCCTGTGGGGAGTGCTCCAGGGCCATTACAGAATTTTTGAGCCGATACCCCCATGTAACTCTGTTTATTTATATAGCACGGCTTTATCACCACGCAGATCCTCGAAATCGGCAAGGACTCAGGGACCTTATTAGCAGCGGTGTTACTATCCAGATCATGACGGAGCAAGAGTCTGGCTACTGCTGGAGGAATTTTGTCAACTACTCCCCTTCGAATGAAGCTCATTGGCCAAGGTACCCCCATCTGTGGGTGAGGCTGTACGTACTGGAACTCTACTGCATCATTTTAGGACTTCCACCCTGTTTAAATATTTTAAGAAGAAAACAACCTCAACTCACGTTTTTCACGATTGCTCTTCAAAGCTGCCATTACCAAAGGCTACCACCCCACATCCTGTGGGCCACAGGGTTGAAA. The nucleotide sequence of the yeast expression vector encoding the eA3A amino acid sequence shown in SEQ ID NO.6 is shown in SEQ ID NO.27.
[0059] SEQ ID NO.27: ATGGAAGCCAGCCCAGCATCCGGGCCCAGACACTTGATGGATCCACACATATTCACTTCCAACTTTAACAATGGCATTGGAAGGCATAAGACCTACCTGTGCTACGAAGTGGAGCGCCTGGACAATGGCACCTCGGTCAAGATGGACCAGCACAGGGGCTTTCTACACGGCCAGGCTAAGAATCTTCTCTGTGGCTTTTACGGCCGCCATGCGGAGCTGCGCTTCTTGGACCTGGTTCCTTCTTTGCAGTTGGACCCGGCCCAGATCTACAGGGTCACTTGGTTCATCTCCTGGAGCCCCTGCTTCTCCTGGGGCTGTGCCGGGGAAGTGCGTGCGTTCCTTCAGGAGAACACACACGTGAGACTGCGTATCTTCGCTGCCCGCATCTATGATTACGACCCCCTATATAAGGAGGCACTGCAAATGCTGCGGGATGCTGGGGCCCAAGTCTCCATCATGACCTACGATGAATTTAAGCACTGCTGGGACACCTTTGTGGACCACCAGGGATGTCCCTTCCAGCCCTGGGATGGACTAGATGAGCACAGCCAAGCCCTGAGTGGGAGGCTGCGGGCCATTCTCCAGAATCAGGGAAAC. The nucleotide sequence of the rice expression vector encoding the eA3A amino acid sequence shown in SEQ ID NO.6 is shown in SEQ ID NO.28.
[0060] SEQ ID NO.28: ATGGAGGCGTCTCCTGCTAGTGGACCGAGGCATCTCATGGACCCCCACATCTTCACCAGCAATTTCAACAATGGTATTGGGCGGCATAAAACATATCTCTGCTACGAGGTGGAAAGACTCGACAACGGTACTTCAGTTAAGATGGATCAGCATCGTGGCTTTCTCCACGGTCAAGCTAAGAACCTTCTTTGTGGTTTCTACGGCCGCCACGCGGAGCTGAGGTTTTTAGATTTGGTACCGTCGCTGCAATTGGATCCCGCACAAATATACAGGGTCACATGGTTTATTAGCTGGTCACCATGCTTCTCCTGGGGCTGCGCCGGGGAAGTCCGCGCCTTCTTGCAGGAGAACACGCACGTGAGGCTGCGGATATTTGCAGCTCGCATCTACGACTATGATCCTCTCTACAAAGAGGCCCTACAAATGCTACGGGACGCGGGCGCCCAGGTGTCCATCATGACCTATGACGAGTTCAAGCACTGCTGGGACACTTTTGTTGATCATCAAGGATGTCCATTCCAGCCGTGGGATGGACTTGATGAACATTCTCAGGCGCTGTCGGGCAGATTACGAGCAATTCTTCAGAATCAGGGGAAC. The nucleotide sequence of the yeast expression vector encoding the mini-Sdd7 amino acid sequence shown in SEQ ID NO.7 is shown in SEQ ID NO.29.
[0061] SEQ ID NO.29: ATGGAAGGTGGTGGTCCAGGTGCTGTCCCAGAAGGTGGTGACGGTCCTCCAGCTGTTCCAGCTGAAGAAGTTGAAAGATTGAGAGGTGAATTGCCACCTCCAGTTGTTCCAGGTACCGGTCAAAAGACTCACGGTAGATGGATTGGTCCAGACGGCCGTGTCAGAGCTATTGTCTCCGGTAGAGACGAAGATGCTGCTTTGGTTCACGCTCAATTGGCCGCTAAGGGTATTCCAGATGAACCAACCAGAAACTCCGATGTGGAACAAAAATTGGCTGCTCACATGGTTGCCAACGGTATCAGACACGTCACTTTGGTCATCAACCACAGACCATGTCGTGGTTTCGATGACTCTTGTGACACTCTAGTTCCAATCATCTTACCAGAAGGTTGTACTTTGACCGTCCACGGTCAAACTGACAAGGGTATGCGTGTTAGAGTCAGATACACTGGTGGTGCCAGACCTTGGTGGTCT The nucleotide sequence of the yeast expression vector encoding the TadA8e(N46L) amino acid sequence shown in SEQ ID NO.8 is shown in SEQ ID NO.30.
[0062] SEQ ID NO.30: ATGTCCGAAGTCGAGTTTTCCCATGAGTACTGGATGAGACACGCATTGACTCTCGCAAAGAGGGCTCGAGATGAACGCGAGGTGCCCGTGGGGGCAGTACTCGTGCTCAACAATCGCGTAATCGGCGAAGGTTGGCTGAGGGCAATCGGACTCCACGACCCCACTGCACATGCGGAAATCATGGCCCTTCGACAGGGAGGGCTTGTGATGCAGAATTATCGACTTATCGATGCGACGCTGTACGTCACGTTTGAACCTTGCGTAATGTGCGCGGGAGCTATGATTCACTCCCGCATTGGACGAGTTGTATTCGGTGTTCGCAACTCAAAGAGAGGTGCCGCAGGTTCACTGATGAACGTGCTGAACTACCCAGGCATGAACCACCGGGTAGAAATCACAGAAGGCATATTGGCGGACGAATGTGCGGCGCTGTTGTGTGATTTTTATCGCATGCCCAGGCAGGTCTTTAACGCCCAGAAAAAAGCACAATCCTCTATCAAC encodes the nCas9 shown in SEQ ID NO.9 1-1046 The nucleotide sequence of the yeast expression vector for the amino acid sequence is shown in SEQ ID NO.31.
[0063] 1063-1367 The nucleotide sequence of the yeast expression vector with the amino acid sequence is shown in SEQ ID NO.32.
[0064] SEQ ID NO.32: GAAACAAACGGAGAAACAGGAGAAATCGTGTGGGACAAGGGTAGGGATTTCGCGACAGTCCGGAAGGTCCTGTCCATGCCGCAGGTGAACATCGTTAAAAAGACCGAAGTACAGACCGGAGGCTTCTCCAAGGAAAGTATCCTCCCGAAAAGGAACAGCGACAAGCTGATCGCACGCAAAAAAGATTGGGACCCCAAGAAATACGGCGGATTCGATTCTCCTACAGTCGCTTACAGTGTACTGGTTGTGGCCAAAGTGGAGAAAGGGAAGTCTAAAAAACTCAAAAGCGTCAAGGAACTGCTGGGCATCACAATCATGGAGCGATCAAGCTTCGAAAAAAACCCCATCGACTTTCTCGAGGCGAAAGGATATAAAGAGGTCAAAAAAGACCTCATCATTAAGCTTCCCAAGTACTCTCTCTTTGAGCTTGAAAACGGCCGGAAACGAATGCTCGCTAGTGCGGGCGAGCTGCAGAAAGGTAACGAGCTGGCACTGCCCTCTAAATACGTTAATTTCTTGTATCTGGCCAGCCACTATGAAAAGCTCAAAGGGTCTCCCGAAGATAATGAGCAGAAGCAGCTGTTCGTGGAACAACACAAACACTACCTTGATGAGATCATCGAGCAAATAAGCGAATTCTCCAAAAGAGTGATCCTCGCCGACGCTAACCTCGATAAGGTGCTTTCTGCTTACAATAAGCACAGGGATAAGCCCATCAGGGAGCAGGCAGAAAACATTATCCACTTGTTTACTCTGACCAACTTGGGCGCGCCTGCAGCCTTCAAGTACTTCGACACCACCATAGACAGAAAGCGGTACACCTCTACAAAGGAGGTCCTGGACGCCACACTGATTCATCAGTCAATTACGGGGCTCTATGAAACAAGAATCGACCTCTCTCAGCTCGGTGGAGAC The nucleotide sequence of the yeast expression vector encoding the hCDG amino acid sequence shown in SEQ ID NO.11 is as shown in SEQ ID NO. 33.
[0065] SEQ ID NO.33: ATGTTCTTTTCTCCATCACCTGCTCGTAAACGTCACGCTCCAAGTCCAGAACCAGCTGTTCAAGGTACCGGTGTTGCCGGTGTTCCAGAAGAATCCGGTGATGCTGCCGCCATTCCAGCCAAGAAGGCCCCAGCTGGTCAAGAAGAACCAGGTACTCCACCATCTTCTCCATTATCCGCTGAGCAATTGGACAGAATCCAAAGAAACAAGGCTGCTGCTTTGTTGAGATTGGCTGCTAGAAACGTTCCAGTTGGTTTCGGTGAATCTTGGAAAAAACATTTGTCTGGTGAATTCGGTAAGCCTTACTTCATCAAGTTGATGGGTTTCGTTGCTGAAGAAAGAAAGCACTACACCGTCTACCCACCACCACACCAAGTTTTCACCTGGACCCAAATGTGTGATATCAAGGATGTTAAGGTAGTTATCTTGGGTCAAGACCCATACCACGGTCCAAACCAAGCTCATGGTCTCTGTTTCTCTGTCCAAAGACCAGTCCCACCACCTCCATCTTTGGAAAACATTTACAAGGAATTGTCTACTGACATTGAAGACTTTGTCCATCCGGGTCACGGTGACTTGTCCGGTTGGGCTAAGCAAGGTGTCTTGTTGTTGGATGCGGTCTTGACTGTTAGAGCTCACCAAGCTAATTCTCACAAGGAAAGAGGCTGGGAACAATTCACTGATGCTGTCGTCTCCTGGTTGAACCAAAACTCCAACGGTTTAGTCTTCTTGCTATGGGGTTCCTATGCTCAAAAAAAGGGTTCTGCCATCGACAGAAAGAGACACCACGTTTTGCAAACTGCTCACCCATCTCCATTATCCGTTTACAGAGGTTTTTTCGGTTGTAGACACTTCTCCAAGACCAACGAATTACTTCAAAAGTCTGGTAAGAAGCCAATTGACTGGAAGGAATTG. The nucleotide sequence of the yeast expression vector encoding the CDA1 amino acid sequence shown in SEQ ID NO.12 is as shown in SEQ ID NO.34.
[0066] SEQ ID NO.34: The nucleotide sequence of the rice expression vector encoding the NLS amino acid sequence of the nucleoplasmic protein shown in SEQ ID NO.13 is shown in SEQ ID NO.35.
[0067] SEQ ID NO.35: AAGCGGCCAGCGGCGACGAAGAAGGCGGGGCAGGCGAAGAAGAAGAAG encodes the nucleotide sequence of the yeast expression vector containing the 16 aa XTEN linker amino acid sequence shown in SEQ ID NO.14, as shown in SEQ ID NO.36.
[0068] The nucleotide sequence of the yeast expression vector encoding the amino acid sequence of the SGGS linker, SEQ ID NO.36:TCTGGTTCTGAAACTCCAGGTACTTCTGAATCTGCTACTCCAGAATCT, is shown in SEQ ID NO.37.
[0069] SEQ ID NO.37:TCTGGTGGTTCA encodes the nucleotide sequence of the yeast expression vector containing the GSSGS linker amino acid sequence shown in SEQ ID NO.15, as shown in SEQ ID NO.38.
[0070] The nucleotide sequence of the rice expression vector encoding the 18 aa XTEN linker amino acid sequence shown in SEQ ID NO.16 (SEQ ID NO.38: GGTTCTTCTGGTTCC) is shown in SEQ ID NO.39.
[0071] SEQ ID NO.39: TCCGGCAGCGAGACGCCAGGCACGTCCGAGAGCGCTACGCCAGAGTCCCTTAAG encodes the TadA8e amino acid sequence shown in SEQ ID NO.17. The nucleotide sequence of the yeast expression vector is shown in SEQ ID NO.40.
[0072] SEQ ID NO.40: ATGTCCGAAGTCGAGTTTTCCCATGAGTACTGGATGAGACACGCATTGACTCTCGCAAAGAGGGCTCGAGATGAACGCGAGGTGCCCGTGGGGGCAGTACTCGTGCTCAACAATCGCGTAATCGGCGAAGGTTGGAATAGGGCAATCG GACTCCACGACCCACTGCACATGCGGAAATCATGGCCCTTCGACAGGGAGGGCTTGTGATGCAGAATTATCGACTTATCGATGCGACGCTGTACGTCACGTTTGAACCTTGCGTAATGTGCGCGGGAGCTATGATTCACTCCCGCATTGGACG The technical solution of the present invention has the following beneficial effects: The technical solution of the present invention achieves synergistic effects through three core functional parts: Cas protein, a cytosine deaminase and a cytosine DNA glycosylase. Its mechanism lies in the efficient generation of a common intermediate product—AP site—through two independent biochemical pathways.
[0073] Specifically, when the editor binds to the target, its deaminase domain catalyzes the deamination of C to U, while the glycosylation domain simultaneously and competitively cleaves cytosine directly. This "two-pronged" strategy significantly increases the generation rate of AP sites, potentially surpassing the processing capacity of the conventional intracellular base excision repair (BER) pathway. Consequently, it tends to repair via alternative pathways such as transdamage synthesis (TLS) and preferentially incorporates guanine (G), ultimately achieving efficient C-to-G transversion.
[0074] In a preferred embodiment, the domains of the dual-enzyme CGBE are arranged from the N-terminus to the C-terminus as follows: cytosine DNA glycosylase CDG4, adapter, cytosine deaminase, adapter, nCas9 (D10A), and nuclear localization signal (NLS). Placing CDG4 at the N-terminus and fusing it with the deaminase-based editor is a key structural design for achieving optimal editing performance.
[0075] Compared with existing technologies, the dual-enzyme synergistic CGBE proposed in this invention has several significant advantages.
[0076] 1. Significantly improved editing efficiency: Through the dual action of CDA and CDG, the editor of this invention can synergistically enhance the editing efficiency of C-to-G. Experimental data show that compared with the traditional CGBE containing only a single enzyme, the dual-enzyme system of this invention can improve the C-to-G editing efficiency by an average of 1.7 times.
[0077] 2. The targeting range and flexibility of this invention are greatly expanded: On the one hand, it can expand the editing window. For example, fusing CDG4 with certain deaminases (such as A1(R33A)) can extend the originally narrow C5-C6 editing window to a wider C5-C8 range, thus making previously uneditable sites editable. On the other hand, it can also shift the optimal editing location. For example, fusing CDG4 with deaminases such as eA3A can significantly change its editing location preference, shifting the optimal editing window from the traditional C5-C6 region to the more downstream C7-C10 region, thereby enabling the targeting of previously inaccessible genomic sites. This feature greatly enhances the targeting flexibility and practicality of the editor.
[0078] 3. This invention maintains high genome specificity while improving performance: Comprehensive whole-genome sequencing analysis shows that the dual-enzyme system of this invention significantly improves on-target efficiency and flexibility without introducing additional off-target mutations across the entire genome. Its off-target effects are comparable to those of corresponding single-enzyme editors, demonstrating the safety and reliability of this strategy.
[0079] 4. The invention possesses immense versatility and application potential: This dual-enzyme CGBE editor system not only performs exceptionally well in model organisms such as yeast, but also exhibits powerful editing capabilities and window reshaping effects in higher plant cells (such as rice), demonstrating its good universality across different species. Furthermore, the "CDG4 fusion" strategy disclosed in this invention is modular and can be applied to various cytosine deaminases, universally improving or modifying their performance, providing a new design paradigm for developing more customized advanced base editors. Attached Figure Description
[0080] Figure 1: Schematic diagram of the hypothetical mechanism of action of the dual-enzyme CGBE (containing deaminase CDA and glycosylation enzyme CDG) of the present invention.
[0081] Figure 2: Schematic diagram of CDG variant selection and fusion strategy carrier structure.
[0082] Figure 3: Evaluation results of CDG variant selection and fusion strategy optimization.
[0083] Figure 4: Schematic diagram of the dual-enzyme CGBE editor vector integrating CDG4 and CDA1 deaminase and the control vector.
[0084] Figure 5: Evaluation results of the dual-enzyme CGBE editor integrating CDG4 and CDA1 deaminases.
[0085] Figure 6: Optimized C-to-G editing efficiency of the dual-enzyme CGBE editor at six genomic targets in yeast.
[0086] Figure 7: Summary of C-to-G editing efficiency of the optimized dual-enzyme CGBE editor.
[0087] Figure 8: The regulatory effect of CDG4 fusion on the editing window and position preference of various deaminases (A1(R33A), eA3A, mini-Sdd7, TadA8e(N46L)). Figure A shows the schematic diagram of the dual-enzyme CGBE editor fused with different deaminases and the control vector. Figure B shows the effect evaluation of the dual-enzyme CGBE editor fused with different deaminases at the PolyC-1 target. Figure C shows the effect evaluation of the dual-enzyme CGBE editor fused with different deaminases at the PolyC-2 target.
[0088] Figure 9: Comprehensive validation of the CDG4-mediated editing window expansion effect on six yeast genome targets.
[0089] Figure 10: Off-target effects analysis of yeast whole-genome DNA, where Figure A shows genome-wide insertions and deletions, Figure B shows genome-wide single nucleotide variants, and Figure C shows the specific categories of single nucleotide variants.
[0090] Figure 11: Schematic diagram of the structure of the dual-enzyme CGBE editor vector and control vector in plant cells (rice).
[0091] Figure 12: Evaluation of C-to-G editing efficiency of the dual-enzyme CGBE editor in plant cells (rice). Detailed Implementation
[0092] The present invention will be further described in detail below with reference to the embodiments. The embodiments shown in this invention are merely preferred embodiments and are not intended to limit the invention to other forms. The invention can be implemented in different forms. Unless otherwise specified, the experimental methods in the following embodiments are conventional methods, performed according to the techniques, experimental conditions, reagents described in the literature in the art or according to the product instructions. Unless otherwise specified, the instruments, materials, reagents, etc. used below can be obtained commercially. Unless otherwise specified, all sequences described in this invention are in the sequence listing where the DNA / RNA nucleotide sequence is 5' end → 3' end, and the protein sequence is N-terminus → C-terminus.
[0093] The YPDA medium in the following examples was prepared from 20 g / L peptone, 10 g / L yeast extract, 20 g / L glucose, 0.12 g / L adenine hemisulfate and water. For solid medium, an additional 15 g / L agarose was added.
[0094] The deficient culture medium in the following examples is prepared from 6.7 g / L YNB, 20 g / L glucose, an appropriate amount of SC-LU mixture of uridine and leucine deficient amino acids (SC-LU), and sterile water. Solid culture medium requires the addition of 15 g / L agarose.
[0095] Example 1: Vector Construction 1. Vector construction method: (1) Design the primers required for constructing the vector, as shown in Table 1; (2) Perform PCR amplification using Phanta Max Super-Fidelity DNA Polymerase (Vazyme), corresponding primer pairs and DNA template; (3) Digest the plasmid vector as the backbone using restriction endonuclease, reaction conditions: 37℃, 4h; (4) Identify the fragment size of the PCR amplification product and the digestion product by agarose gel electrophoresis, and then recover the fragments by gel extraction; (5) Perform seamless cloning and ligation of the purified linear vector and fragment using OK Clon DNA Ligation Kit II (Accurate); (6) Transform into Escherichia coli DH5α competent cells, plate onto resistant LB medium, pick single colonies for sequencing verification; (7) Transfer the correctly sequenced colonies to 5mL liquid resistant LB medium, and culture at 37℃ and 225r / min for 12-18 hours; (8) Extract the plasmid using the OMEGA plasmid mini extraction kit.
[0096] 2. Construction of all gRNA expression vectors using the GoldenGate method: (1) Design the primers required for constructing the vectors, as shown in Table 2; (2) Mix the corresponding primer pairs in a 1:1 ratio, heat to 95°C and then gradually anneal to 25°C to obtain the primer annealing products; (3) Prepare a working system using Class IIs restriction endonuclease and T4 DNA ligase (NEB), mix them, add the primer annealing products and the vector backbone to the system for GoldenGate ligation, and the reaction conditions are: 37°C, 2 min; 16°C, 30 s, 50 cycles, then 37°C, 5 min, 85°C, 5 min. (4) Transform the GoldenGate reaction product into E. coli DH5α competent cells, spread it on resistant LB medium, pick single clones for sequencing verification; (5) Transfer the correctly sequenced colonies to 5 mL of liquid resistant LB medium, and culture at 37℃ and 225 r / min for 12-18 hours; (6) Extract plasmids using the OMEGA plasmid mini-extraction kit.
[0097] Table 1 Primers for Editor Vector Construction
[0098]
[0099]
[0100]
[0101]
[0102]
[0103] Table 2 Primers for constructing sgRNA expression vectors
[0104]
[0105] Table 3. sgRNA Target Sequences Target Site Name Sequence (5' → 3') PAM Analysis Method PolyC-1 C1 C2 C3 C4 C5 C6 C7 C8 ATGTTCCGAGATCGG High-throughput Sequencing PolyC-2 AC2 C3 C4 C5 C6 C7 C8 TAATATATTCAAAGG High-throughput Sequencing Site 1 GC2 C3 C4 C5 C6 C7 C8 TCCCCAAAAAAATGG High-throughput Sequencing Site 2 TC2 C3 C4 C5 C6 C7 C8 C9 ATGGGATTATACGG High-throughput Sequencing Site 3 AGC3 C4 C5 C6 C7 C8 C9 TCTAAAACAAGTGG High-throughput Sequencing Site 5 TTC3 C4 C5 C6 C7 C8 C9 AATGTTGGAAACGG High-throughput Sequencing Site 7 GAC3 C4 C5 C6 C7 C8 C9 TATACAAGTGTTGG High-throughput Sequencing Site 8CTC3C4C5C6C7C8C9TAATAGAACGTTGG high-throughput sequencing site 9TC2C3AATAAC9GGAATC 15 C 16 AACTGGG high-throughput sequencing site 11TGC3C4TC6AATGTC 12 TC 14 TTCTATCGG high-throughput sequencing site 12GTGTGTGC8C9AC 11 TC13 C 14 C 15 GGCCCCGG high-throughput sequencing site 16C1ATGC5AGC8GAATAAAC 16 ATC 19 AGGG high-throughput sequencing site 25GTTAC5ATGTATTGGTTTTCTTGGL-canavanine resistance selection T1TGC3GGC6AGC9TGC 12 GTTTTCCTTGG high-throughput sequencing T2C1C2C3AC5C6GC8AATATGCCATTCAGG high-throughput sequencing T3GC2C3C4C5C6AC8C9C 10 GGCCTCGAGCGGG high-throughput sequencing T4GC2GC4C5C6C7C8AC 10 TTGGGATCATAGG high-throughput sequencing T5C1C2AC4C5C6AC8C9C 10 C 11 GCGCGTGGATGG High-throughput sequencing T6TC2C3C4TC6C7C8C9C 10 C 11 TTC 14 C 15 TGGCCCGG high-throughput sequencing Table 3. The specific construction process of each expression vector is as follows: (1) Construction process of yeast expression vector: (F1) Construction of CDA1Δ194 / A1(R33A) / eA3A-miniCGBE vector: Using the primer pairs miniCGBE-1F / 1R and 2F / 2R in Table 1, pJT84_GalL_nCDA1Δ194-BE3 (Addgene, Plasmid #145047), pJT25_GalL_BE3 (Addgene, Plasmid #143736), and pJT176_GalL_eA3A-BE3 (Addgene, Plasmid #145125) were used as templates to amplify the fragments 1 and 2 of each vector, and then amplified the fragments with the restriction endonuclease AscI / MluI. The digested CDA1Δ194-BE3 was ligated to obtain the vectors CDA1Δ194-miniCGBE (as shown in Figure 4), A1-miniCGBE, and eA3A-miniCGBE (as shown in Figure 8A). Using the point mutation-containing primer pairs A1(R33A)-1F / 1R and 2F / 2R from Table 1, A1-miniCGBE was amplified as a template. After obtaining the fragment, it was ligated with A1-miniCGBE digested with restriction endonucleases SpeI / SbfI to obtain the vector A1(R33A)-miniCGBE, as shown in Figure 8A.
[0106] (F2) Construction of the fusion vector of glycosylation enzyme CDG and nCas9: The yeast codon-optimized CDG4 shown in SEQ ID NO.20 and the yeast codon-optimized hCDG shown in SEQ ID NO.33 were artificially synthesized by Qingke Biotechnology Co., Ltd. (Beijing). The synthesized fragment CDG4 was amplified using the primer pair DAF-CBE-1F / 1R in Table 1 as a template to obtain fragment 1. The CDA1Δ194-miniCGBE constructed in (F1) was amplified using the primer pair DAF-CBE-2F / 2R in Table 1 as a template to obtain fragment 2. The two fragments were ligated with CDA1Δ194-miniCGBE digested with restriction endonucleases SpeI / SbfI to obtain the vector DAF-CBE, as shown in Figure 2.
[0107] Using primer pairs CE-CDG-1F / 1R and CE-CDG4-1F / 1R from Table 1, amplification was performed using pJT46_GalL_cCDA1-BE3 (Addgene, Plasmid #145039) as a template to obtain fragment 1. Using primer pairs CE-CDG-2F / 2R and CE-CDG4-2F / 2R from Table 1, amplification was performed using the synthesized fragments hCDG and CDG4 as templates to obtain fragment 2. Using primer pairs CE-CDG-3F / 3R and CE-CDG4-3F / 3R from Table 1, amplification was performed using the CDA1Δ194-miniCGBE constructed by (F1) as a template to obtain fragment 3. The three fragments were ligated with CDA1Δ194-miniCGBE digested with restriction endonucleases SpeI / AscI to obtain vectors CE-CDG and CE-CDG4, as shown in Figure 2.
[0108] Using the primer pairs nCas9-CDG4-1F / 1R in Table 1, amplification was performed with pJT46_GalL_cCDA1-BE3 (Addgene, Plasmid #145039) as a template to obtain fragments. These fragments were then ligated with CDA1Δ194-miniCGBE ((F1) construction) digested with restriction endonucleases SpeI / SbfI to obtain an intermediate vector. Using the primer pairs nCas9-CDG4-2F / 2R and 4F / 4R in Table 1, amplification was performed with the intermediate vector as a template to obtain fragment 1 and fragment 3, respectively. Using the primer pairs nCas9-CDG4-3F / 3R in Table 1, amplification was performed with the synthesized fragment CDG4 as a template to obtain fragment 2. The three fragments were then ligated with the intermediate vector digested with restriction endonucleases AscI / MluI to obtain the vector nCas9-CDG4, as shown in Figure 2.
[0109] Construction of the (F3)CDG4-CDA1Δ194-miniCGBE vector: Using the primer pair CDG4-CDA1∆-1F / 1R in Table 1, the synthesized fragment CDG4 was amplified to obtain fragment 1; using the primer pair CDG4-CDA1∆-2F / 2R in Table 1, the CDA1Δ194-miniCGBE constructed in (F1) was amplified to obtain fragment 2; the two fragments were ligated with CDA1Δ194-miniCGBE digested with restriction endonucleases SpeI / SbfI to obtain the vector CDG4-CDA1Δ194-miniCGBE, as shown in Figure 4.
[0110] Construction of (F4)CDA1Δ194-DAF-CBE and DAF-CBE-CDA1 vectors: Using the primer pair CDG4-CDA1∆-1F / 1R in Table 1, CDA1Δ194-miniCGBE constructed in (F1) was amplified as a template to obtain fragment 1; using the primer pair CDG4-CDA1∆-2F / 2R in Table 1, CDG4 was amplified as a template to obtain fragment 2; the two fragments were ligated with CDA1Δ194-miniCGBE digested with restriction endonucleases SpeI / SbfI to obtain the vector CDA1Δ194-DAF-CBE, as shown in Figure 4.
[0111] Using the primer pair DAF-CBE-CDA1-1F / 1R in Table 1, amplification was performed with pJT46_GalL_cCDA1-BE3 (Addgene, Plasmid #145039) as a template to obtain fragment 1; using the primer pair DAF-CBE-CDA1-2F / 2R in Table 1, amplification was performed with DAF-CBE as a template to obtain fragment 2; the two fragments were ligated with DAF-CBE after digestion with restriction endonucleases AscI / MluI to obtain the vector DAF-CBE-CDA1, as shown in Figure 4.
[0112] Construction of (F5)CDA1Δ194-CE-CDG4 and CE-CDG4-CDA1 vectors: Using primers in Table 1, CDA1Δ194-miniCGBE constructed in (F1) was amplified with CDA1Δ194-miniCGBE as a template to obtain fragment 1 and fragment 2. These two fragments were then ligated with CE-CDG4 digested with restriction endonucleases SpeI / SbfI (constructed in (F2)) to obtain the vector CDA1Δ194-CE-CDG4, as shown in Figure 4. Using primers in Table 1, CE-CDG4-CDA1-F / R was amplified with DAF-CBE-CDA1 as a template to obtain a fragment, which was then ligated with CE-CDG4 digested with restriction endonucleases AscI / MluI to obtain the vector CE-CDG4-CDA1, as shown in Figure 4.
[0113] Construction of the (F6)miniSdd7 / TadA8e(N46L)-miniCGBE vector: The mini-Sdd7 vector with yeast codon optimized as shown in SEQ ID NO.29 and SEQ ID NO.29 were artificially synthesized by Qingke Biotechnology Co., Ltd. (Beijing). The yeast codon-optimized TadA8e shown in NO.40 was amplified using primer pairs CDG4-miniSdd7-1F / 1R and TadA8e-1F / 1R in Table 1, with the synthesized fragments mini-Sdd7 and TadA8e as templates, respectively, to obtain their respective fragments 1. CDG4-miniSdd7-2F / 2R and TadA8e-2F / 2R were also amplified using the (F1) constructed CDA1Δ194-miniCGBE as a template, with their respective fragments 2. The two fragments were then ligated with CDA1Δ194-miniCGBE digested with restriction endonucleases SpeI / SbfI to obtain the vectors miniSdd7-miniCGBE (as shown in Figure 8A) and TadA8e-miniCGBE. Using the primer pairs TadA8e(N46L)-1F / 1R and 2F / 2R with point mutation pairs in Table 1, fragments 1 and 2 were amplified using TadA8e-miniCGBE as a template. These fragments were then ligated with TadA8e-miniCGBE digested with restriction endonucleases SpeI / SphI to obtain the TadA8e(N46L)-miniCGBE vector, as shown in Figure 8A.
[0114] Construction of the (F7)CDG4-A1(R33A) / eA3A / miniSdd7 / TadA8e(N46L)-miniCGBE vector: Using the F1 / R1 primer pairs required for constructing the corresponding vectors in Table 1, the DAF-CBE constructed in (F2) was used as a template to amplify fragment 1 containing CDG4 and XTEN linker. The corresponding F2 / R2 primer pairs were used to amplify fragment 2 containing deaminase. The two fragments were ligated with CDA1Δ194-miniCGBE (constructed in (F1)) digested with restriction endonucleases SpeI / SbfI to obtain the corresponding four vectors CDG4-A1(R33A)-miniCGBE, CDG4-eA3A-miniCGBE, CDG4-miniSdd7-miniCGBE, and CDG4-TadA8e(N46L)-miniCGBE, as shown in Figure 8A.
[0115] (2) Construction process of rice expression vector: (G1) Construction of P-eA3A-miniCGBE vector: Using the primer pair P-A3A-1F / 1R in Table 1, the fragment was amplified with pH-A3A-PBE (Addgene, Plasmid #119774) as a template. It was then ligated with pH-A3A-PBE digested with restriction endonucleases MluI / SacI to obtain the vector P-A3A-miniCGBE. Using the primer pair P-eA3A-1F / 1R and 2F / 2R with point mutations in Table 1, fragments 1 and 2 were amplified with P-A3A-miniCGBE as a template. They were then ligated with pH-A3A-PBE digested with restriction endonucleases AvrII / SbfI to obtain the vector P-eA3A-miniCGBE, as shown in Figure 11.
[0116] Construction of (G2)P-DAF-CBE and P-CDA1Δ194-miniCGBE vectors: The rice codon-optimized CDG4 and P-CDA1Δ194-miniCGBE vectors shown in SEQ ID NO. 21 were artificially synthesized by Qingke Biotechnology Co., Ltd. (Beijing). The rice codon-optimized CDA1Δ194 shown in NO.19 was amplified using primer pairs P-DAF-CBE-1F / 1R and P-CDA1∆194-1F / 1R from Table 1, with the synthesized fragments CDG4 and CDA1Δ194 as templates, to obtain their respective fragments 1. Using primer pairs P-DAF-CBE-2F / 2R and P-CDA1∆194-2F / 2R from Table 1, the P-eA3A-miniCGBE constructed with (G1) was amplified using (G1) constructed as a template to obtain their respective fragments 2. The two fragments were then ligated with P-eA3A-miniCGBE digested with restriction endonucleases AvrII / SbfI to obtain the vectors P-DAF-CBE and P-CDA1Δ194-miniCGBE, as shown in Figure 11.
[0117] Construction of the (G3)P-CDG4-CDA1Δ194 / eA3A-miniCGBE vector: Using primer pairs P-CDG4-CDA1∆-1F / 1R and P-CDG4-eA3A-1F / 1R in Table 1, the P-DAF-CBE constructed in (G2) was amplified using the corresponding fragment 1; using primer pairs P-CDG4-CDA1∆-2F / 2R and P-CDG4-eA3A-2F / 2R in Table 1 respectively... Using P-CDA1Δ194-miniCGBE constructed with (G2) and P-eA3A-miniCGBE constructed with (G1) as templates, fragment 2 was obtained. The two fragments were then ligated with P-eA3A-miniCGBE digested with restriction endonucleases AvrII / SbfI to obtain vectors P-CDG4-CDA1Δ194-miniCGBE and P-CDG4- / eA3A-miniCGBE, as shown in Figure 11.
[0118] (3) Construction of gRNA expression vector: (H1) Using the primer set for constructing the gRNA expression vector backbone in Table 1, sgRNA-scaffold-F1 / R1 was used to amplify with ccdB expression frame as template to obtain fragment 1. sgRNA-scaffold-F2 / R2 primers were used to amplify with pJT303_SNR52_sgRNA_Can1-3 (Addgene, Plasmid #145066) as template to obtain fragment 2. The two fragments were ligated with pJT303_SNR52_sgRNA_Can1-3 digested with restriction endonuclease AatII / KpnI to obtain the gRNA expression vector backbone.
[0119] (H2) Using the primer sets gRNA-F1 / R1~gRNA-F13 / R13 in Table 2 for constructing gRNA expression vectors, all yeast gRNA expression vectors are used with PolyC-1 as an example. The primers gRNA-F1 / R1 in Table 2 are used for annealing. The annealed product is mixed with the gRNA expression vector backbone and ligated with restriction endonuclease PaqCI and T4 DNA ligase GoldenGate to obtain the corresponding gRNA expression vector.
[0120] (H3) Using the primer sets gRNA-F14 / R14~gRNA-F19 / R19 in Table 2 for constructing rice target series vectors, all rice gRNA expression cassettes are in the rice editor. Taking the construction of the vector P-DAF-CBE targeting the rice T1 target (sequence shown in Table 3) as an example, annealing is performed using the primer gRNA-F14 / R14 in Table 2. The annealed product is mixed with P-DAF-CBE and ligated with restriction endonuclease BsaI and T4 DNA ligase using GoldenGate to obtain the corresponding P-DAF-CBE vector targeting the rice T1 target.
[0121] Example 2 Transformation and induction of inducible Saccharomyces cerevisiae BY4743 and high-throughput sequencing analysis In order to detect the editing status of the editor in yeast, this example will use the vectors in (F1)-(F7) above and the sgRNA vector in (H2) to co-transform and induce yeast. After DNA extraction, the target fragments will be sequenced by high-throughput sequencing, and the editing efficiency will be obtained by analysis and visualization. The specific method is as follows: 1. Yeast transformation (1) Streak Saccharomyces cerevisiae BY4743 on YPDA medium and culture at 28℃ for 2-3 days; (2) After washing with sterile ddH2O and collecting yeast cells, add 100mM LiAc and incubate at 28℃ for 10 minutes; (3) After centrifugation for 5 seconds, remove the supernatant and add the cells with 0.5-1μg plasmid DNA obtained in Example 1, 240μL 50%PEG3350, 36μL 1M LiAc, and 50μL 2 mg / mL salmon sperm DNA and 20 μL sterile water were mixed in a centrifuge tube and incubated at 42°C for 1.5 h; (4) After centrifugation for 5 seconds, the supernatant was removed and cultured at 28°C for 2-3 days on SC-LU (a yeast synthesis medium lacking uracil and leucine); (5) Single clones were picked and cultured at 28°C and 225 r / min for 18-20 hours on SC-LU (a liquid medium lacking uracil and leucine), and the positive result was confirmed by PCR amplification.
[0122] 2. Yeast induction: (1) Pick 3-5 positive colonies and incubate them in 3 mL of Deficit Liquid Medium (SC-LU) containing 2% glucose at 28°C for 18-20 hours; (2) Take 0.8 mL of bacterial solution, centrifuge and discard the supernatant, wash 3 times with sterile water to remove residual glucose, and then resuspend in 5 mL of SC-LU liquid induction medium containing 2% galactose and 1% raffinose, and incubate at 28°C and 225 r / min for 21 hours; (3) Take 0.5 mL of bacterial solution, briefly centrifuge and discard the supernatant to obtain the induced bacterial cells.
[0123] 3. Crude extraction of yeast DNA: (1) Add yeast cell lysis buffer containing 1% SDS and 0.2M LiAc to the induced cells, resuspend the yeast cells, and incubate at 70 °C for 8 min; (2) Add three volumes of anhydrous ethanol (Shanghai test) to the treated resuspension, shake to mix, briefly centrifuge and discard the supernatant to obtain a precipitate containing yeast DNA; (3) Wash the precipitate with 70% ethanol and air dry at room temperature; (4) Add 100 μL of sterile water to the dried precipitate to obtain yeast DNA.
[0124] 4. High-throughput sequencing library construction and analysis: (1) Phanta Max Super-Fidelity DNA Polymerase, corresponding primer pairs, and yeast genomic DNA obtained in the above examples were used as templates for PCR amplification. The primers were the high-throughput sequencing primers shown in Table 4. (2) The PCR products were purified using the Cycle Pure kit (OMEGA). (3) The purified products were subjected to PCR-free library construction, high-throughput sequencing (Kingwiz Biotechnology, Suzhou, China), and data analysis. The sequencing was performed using the Illumina NovaSeq 6000 platform. (4) On average, more than 100,000 reads were obtained per sample. After data filtering, the FASTQ files were analyzed using the script https: / / github.com / zfcarpe / Cas9Sequencing. (5) Vector graphics were drawn using the vector drawing tools GraphPad Prism 8 and Adobe Illustrator.
[0125] The editing efficiency tests in Examples 3 and 4 were conducted using the methods described above, and the specific analysis results are shown in the corresponding examples.
[0126] Example 3: Design, Construction, and Optimization of a Dual-Enzyme CGBE Editor. Based on the working mechanism of C-to-G generation, this study proposes a novel dual-enzyme CGBE editor, the core of which is a fusion protein called CDG4-Deaminase-miniCGBE. Figure 1 illustrates the hypothetical working mechanism of the fusion protein, which achieves efficient C-to-G base editing by binding the activities of cytosine deaminase (CDA) and cytosine DNA glycosyltransferase (CDG). After the editor binds to the target site, it efficiently generates AP sites through two pathways (C→U deamination and direct C excision), ultimately promoting the C-to-G transversion process. The construction and optimization process of the dual-enzyme CGBE editor is as follows: 1. Screening of Glycosyltransferases and Fusion Sites: In order to achieve the optimal working effect of the dual-enzyme CGBE editor, the reported cytosine DNA glycosyltransferases (human CDG (hCDG) and engineered CDG4) and their optimal fusion sites with nCas9 (D10A) were first screened. Four CGBE base editors fused with nCas9(D10A) were constructed, namely (I1)DAF-CBE, (I2)CE-CDG, (I3)CE-CDG4, and (I4)nCas9-CDG4 (as shown in Figure 2), wherein: (I1) DAF-CBE is composed of SEQ ID NO.2, SEQ ID NO.14, SEQ ID NO.3, SGGSlinker, and SEQ ID NO.4 connected in sequence; (I2) CE-CDG is composed of SEQ ID NO.9, SEQ ID NO.15, SEQ ID NO.11, SEQ ID NO.15, SEQ ID NO.10, SGGS linker, and SEQ ID NO.4 connected in sequence; (I3) CE-CDG4 is composed of SEQ ID NO.9, SEQ ID NO.15, SEQ ID NO.2, SEQ ID NO.15, SEQ ID NO.10, SGGS linker, and SEQ ID NO.4 connected in sequence. NO.4 is composed of (I4) nCas9-CDG4 consisting of SEQ ID NO.3, SEQ ID NO.14, SEQ ID NO.2, SGGS linker and SEQ ID NO.4 connected in sequence; and the editing efficiency was detected at two target sites, PolyC-1 and PolyC-2 in yeast (sequences are shown in Table 3).
[0127] The specific construction methods of each editor and the experimental process of yeast editing efficiency test are shown in Example 1 and Example 2. The results are shown in Figure 3: (1) The editing effect test of CE-CDG and CE-CDG4 shows that the editing efficiency of CDG4 is significantly higher than that of hCDG.
[0128] (2) Compare the effects of CDG4 on the N-terminal fusion (DAF-CBE), C-terminal fusion (nCas9-CDG4) and internal embedding (CE-CDG4) of nCas9, and determine that the N-terminal fusion (DAF-CBE configuration) is the optimal strategy, followed by the internal embedding (CE-CDG4 configuration).
[0129] 2. Introduction and conformation optimization of deaminase: Further, the C-terminal truncated variant CDA1Δ194 of the highly efficient PmCDA1 deaminase was selected as the deaminase module and fused into the CDG4-nCas9 backbone (i.e., DAF-CBE conformation) and embedded internally with nCas9. 1 -1046 -CDG4-nCas9 1063-1367 The editing efficiency was detected in the backbone (i.e., the CE-CDG4 configuration) and at the yeast PolyC-2 target site.
[0130] Specifically, the following editors were constructed: (J1)CDA1∆194-miniCGBE, (J2)CDA1∆194-DAF-CBE, (J3)DAF-CBE-CDA1, (J4)CDA1∆194-CE-CDG4, (J5)CE-CDG4-CDA1, and (A1)CDG4-CDA1∆194-miniCGBE mentioned above. Among them, (J1)CDA1∆194-miniCGBE, along with (I1)DAF-CBE and (I3)CE-CDG4 constructed above, served as control editors. (A1) and (J2)-(J3) show the dual-enzyme editors constructed by fusing the deaminase module into the DAF-CBE configuration, while (J4)-(J5) show the dual-enzyme editors constructed by fusing the deaminase module into the CE-CDG4 configuration. The reason for choosing full-length CDA1 as the deaminase in (J3) is that previous studies have shown that C-to-G editing efficiency of full-length CDA1 fused to the C-terminus of nCas9(D10A) is better than that of truncated variants.
[0131] Specifically, (J1) CDA1∆194-miniCGBE is composed of SEQ ID NO.1, SEQ ID NO.3, SGGS linker, and SEQ ID NO.4 connected in sequence; (J2) CDA1∆194-DAF-CBE is composed of SEQ ID NO.1, SEQ ID NO.2, SEQ ID NO.14, SEQ ID NO.3, SGGS linker, and SEQ ID NO.4 connected in sequence; (J3) DAF-CBE-CDA1 is composed of SEQ ID NO.2, SEQ ID NO.14, SEQ ID NO.3, SEQ ID NO.14, SEQ ID NO.12, SGGS linker, and SEQ ID NO.4 connected in sequence; (J4) CDA1∆194-CE-CDG4 is composed of SEQ ID NO.1, SEQ ID NO.9, SEQ ID NO.15, SEQ ID NO.2, SEQ ID NO.15, SEQ ID NO.10, SGGS linker, and SEQ ID NO.4 connected in sequence. NO.4 constitutes (J5) the CE-CDG4-CDA1 consisting of SEQ ID NO.9, SEQ ID NO.15, SEQ ID NO.2, SEQ ID NO.15, SEQ ID NO.10, SEQ ID NO.14, SEQ ID NO.12, SGGS linker and SEQ ID NO.4 connected in sequence. NO.4 Composition; The specific construction methods of each editor and the yeast editing efficiency test process are shown in Example 1 and Example 2. The results are shown in Figure 5: (1) The CDG4-CDA1Δ194-miniCGBE editor showed the highest editing efficiency compared with other dual-enzyme fusion editors, reaching 60.2%. Among them, the C-to-G editing efficiency reached 25%, which was higher than the C-to-G editing mediated by the monodeaminase CDA1Δ194 (20.2%) and slightly lower than the C-to-G editing mediated by the monosaccharidase CDG4 (30.1%). Its fusion structure is N-terminal-CDG4-CDA1Δ194-nCas9(D10A)-NLS-C-terminal.
[0132] (2) CDA1Δ194-CE-CDG4 performed second best, reaching 35.2%, of which the editing efficiency of C-to-G reached 17%. Its fusion structure is N-end-CDA1Δ194-nCas9 1-1046 -CDG4-nCas9 1063-1367 -NLS-C terminal.
[0133] In summary, CDG4-CDA1Δ194-miniCGBE, which has the highest efficiency, and CDA1Δ194-CE-CDG4, which has the second highest efficiency, were selected for comprehensive verification of C-to-G editing efficiency at multiple yeast targets.
[0134] 3. Performance verification: CDG4-CDA1Δ194-miniCGBE and CDA1Δ194-CE-CDG4, along with three control groups (including monosaccharide saccharidase editors DAF-CBE and CE-CDG4, and monodeaminase editor CDA1Δ194-miniCGBE), were tested at six endogenous gene target sites in yeast (sequences are shown in Table 3). The yeast editing efficiency test process is shown in Example 2, and the results are shown in Figures 6 and 7: (1) CDG4-CDA1Δ194-miniCGBE showed higher C-to-G editing efficiency than DAF-CBE and CDA1Δ194-miniCGBE at multiple target sites, with an average C-to-G editing efficiency of 24.9%, while the average C-to-G editing efficiency of DAF-CBE and CDA1Δ194-miniCGBE was only 16.9% and 14.8%, respectively. At site 1, C-to-G editing of CDG4-CDA1Δ194-miniCGBE reached 43.73%, significantly higher than single-enzyme-mediated C-to-G editing (21.9% and 19.5%), demonstrating the significant synergistic effect of the two enzymes.
[0135] (2) CDA1Δ194-CE-CDG4 performed better than CE-CDG4 at some target sites, with an average C-to-G editing efficiency of 13.4%, but significantly lower than CDG4-CDA1Δ194-miniCGBE, indicating that CDG4-Deaminase-miniCGBE is the optimal dual-enzyme CGBE editor fusion method.
[0136] Based on the above experimental results, in order to develop the optimal dual-enzyme CGBE editor, this invention screened the types of cytosine DNA glycosylases and their fusion positions with nCas9(D10A). The optimal fusion was determined to be the N-terminus fusion of cytosine DNA glycosylase CDG4. Furthermore, the highly efficient cytosine deaminase CDA1Δ194 (a C-terminal truncated variant of PmCDA1) was fused to screen for the best dual-enzyme fusion strategy. Through editing efficiency testing in yeast, CDG4-Deaminase-miniCGBE was determined to be the optimal dual-enzyme CGBE editor fusion method, i.e., glycosylase CDG4 is fused to the N-terminus of the deaminase, and the deaminase is fused to the N-terminus of nCas9(D10A).
[0137] Example 4: Test of CDG4 fusion strategy on the editing window of different deaminases. To verify the universality of the strategy of the present invention, CDG4 was fused to four other commonly used deaminase editors, including miniCGBE based on A1(R33A), eA3A, mini-Sdd7, and TadA8e(N46L). The fusion mode was that CDG4 was fused to the N-terminus of the deaminase, and a series of dual-enzyme editors such as CDG4-A1(R33A)-miniCGBE and CDG4-eA3A-miniCGBE were constructed.
[0138] Specifically, the following editors were constructed: (K1)A1(R33A)-miniCGBE, (K2)eA3A-miniCGBE, (K3)miniSdd7-miniCGBE, (K4)TadA8e(N46L)-miniCGBE, and the aforementioned (A2)CDG4-A1(R33A)-miniCGBE, (A3)CDG4-eA3A-miniCGBE, (A4)CDG4-miniSdd7-miniCGBE, and (A5)CDG4-TadA8e(N46L)-miniCGBE. Among them, (K1)-(K4) are single-enzyme CGBE editors for the control group, and (A2)-(A5) are dual-enzyme CGBE editors.
[0139] Specifically, (K1) the A1(R33A)-miniCGBE consists of SEQ ID NO.5, SEQ ID NO.14, SEQ ID NO.3, SGGS linker, and SEQ ID NO.4 connected in sequence; (K2) the eA3A-miniCGBE consists of SEQ ID NO.6, SEQ ID NO.14, SEQ ID NO.3, SGGS linker, and SEQ ID NO.4 connected in sequence; (K3) the miniSdd7-miniCGBE consists of SEQ ID NO.7, SEQ ID NO.14, SEQ ID NO.3, SGGS linker, and SEQ ID NO.4 connected in sequence; and (K4) the TadA8e(N46L)-miniCGBE consists of SEQ ID NO.8, SEQ ID NO.14, SEQ ID NO.3, SGGS linker, and SEQ ID NO.4 connected in sequence. Testing was performed at two PolyC gene target sites in yeast (sequences are shown in Table 3). The specific construction methods of each editor and the experimental process for testing the editing efficiency of yeast are described in Examples 1 and 2. The results are shown in Figures 8B and 8C: 1. For the A1(R33A) editors (K1) and (A2), after fusing CDG4, the editing efficiency was improved by hundreds or even thousands of times at downstream cytosine sites (such as C7-C9) where the traditional editor has low activity; its effective editing window was expanded from the original C5-C6 to C5-C9. At the previously extremely low activity sites C8 and C9, the editing efficiency was improved by as much as 630 times and 1062 times, respectively.
[0140] 2. For the eA3A editors (K2) and (A3), the integration with CDG4 fundamentally reshaped their editing preferences. Optimal editing activity shifted from C5-C6 to C7-C9, making this editor a highly efficient "downstream" editor and greatly increasing its targeting flexibility.
[0141] 3. For the mini-Sdd7 editor (K3) and (A4), after merging with CDG4, the editing window remains unchanged, but the editing efficiency of C6-C8 is slightly improved.
[0142] 4. For the TadA8e(N46L) editor (K4) and (A5), after fusion with CDG4, its effective editing window is expanded from the original C5-C6 to C5-C8, but the overall C-to-G editing efficiency is lower than that of the dual-enzyme CGBE editor fused with A1(R33A) and eA3A.
[0143] Furthermore, the two dual-enzyme editors with the highest editing efficiency (CDG4-A1(R33A)-miniCGBE and CDG4-eA3A-miniCGBE) and their control groups (DAF-CBE, A1(R33A)-miniCGBE and eA3A-miniCGBE) were selected for validation at six genomic targets in yeast (sequences are shown in Table 3). The results are shown in Figure 9. Consistent with the performance of the two targets in Figures 8B and 8C, the dual-enzyme CGBE editor fused with CDG4 and deaminase effectively expanded the editing range and transformed previously marginal or inaccessible C7-C9 editing sites into highly efficient targets, fundamentally expanding the utility and targeting range of these CGBE platforms.
[0144] Example 5: Off-target assay of yeast whole genome DNA. The off-target detection mainly involves screening positive bacteria with L-canavanine, performing whole genome sequencing, and detecting off-target effects. Using the primer pair gRNA-F13 / R13 shown in Table 2, an sgRNA expression vector was constructed to target the site25 target of the CAN1 gene (sequence shown in Table 3) (the specific construction process of the vector is shown in Example 1). The induced edited bacterial culture was cultured in a medium containing canavanine. Colonies that grow normally in the medium are the successfully edited positive bacteria. The detailed steps are as follows: 1. Similar to Example 2, yeast was transformed and cultured using the expression vector containing the editor and the sgRNA vector, while a blank control without any expression vector was cultured simultaneously; 2. Subsequently, similar to Example 2, the bacterial cells were transferred to a liquid induction medium containing 2% galactose and 1% raffinose, and cultured at 28°C with shaking at 225 rpm for 20 hours; 3. The bacterial culture was diluted 10,000 times, and the blank control was plated on YPDA medium, while the remaining bacterial culture was plated on SC-Arg solid medium containing 60 μg / mL L-canavanine, and cultured at 28°C for 2-3 days; 4. Colonies were picked from each culture dish and cultured in YPDA liquid medium with shaking at 28°C at 225 rpm for 20 hours; 5. 0.5-1 mL of bacterial culture was taken and yeast genomic DNA was extracted using a yeast genomic DNA extraction reagent (Solarbio); 6. The extracted DNA samples were quality assessed, and library construction, whole genome sequencing, and bioinformatics analysis were performed.
[0145] This invention investigated the off-target effects of the fusion proteins (I1) DAF-CBE, (J1) CDA1Δ194-miniCGBE, (A1) CDG4-CDA1Δ194-miniCGBE, (K1) A1(R33A)-miniCGBE, (K2) eA3A-miniCGBE, (A2) CDG4-A1(R33A)-miniCGBE, and (A3) CDG4-eA3A-miniCGBE at site 25 of the CAN1 gene. As shown in Figure 10: 1. Compared with the untreated control group, none of the CGBE editors (whether single-enzyme or dual-enzyme) caused a significant increase in the frequency of insertions and deletions (Figure 10A).
[0146] 2. The total number of single nucleotide variants (SNVs) induced by the dual-enzyme CGBE shown in (A1)-(A3) is comparable to that of the corresponding single-enzyme CGBE shown in (I1), (J1), (K1), and (K2), without introducing an additional mutation burden (Figure 10B).
[0147] 3. In terms of mutation types, the dual-enzyme system did not change the inherent mutation spectrum of deaminases (mainly C-to-T byproducts), and even slightly reduced it in some cases (Figure 10C).
[0148] Conclusion: The efficient dual-enzyme system of this invention maintains the same high genome specificity as existing single-enzyme editors.
[0149] Example 6: Application of Dual Enzyme CGBE in Higher Plants To verify the application potential of this invention in agricultural biotechnology, the optimized dual enzyme editor CDG4-CDA1Δ194-miniCGBE and CDG4-eA3A-miniCGBE with the most significant window remodeling effect were applied to rice. T-DNA vectors and control vectors suitable for rice expression were constructed, containing corresponding editors and sgRNA expression cassettes. The specific construction process of the editor vector and sgRNA expression cassette is described in Example 1. The T1-T6 target sequences are shown in Table 3.
[0150] Specifically, the following single-enzyme control editors were constructed: (L1) P-DAF-CBE, (L2) P-CDA1∆194-miniCGBE, (L3) P-eA3A-miniCGBE, and the aforementioned dual-enzyme editors: (B1) P-CDG4-CDA1∆194-miniCGBE and (B2) P-CDG4-eA3A-miniCGBE.
[0151] The following components are described: (L1) P-DAF-CBE: composed of SEQ ID NO.4, SEQ ID NO.2, SEQ ID NO.16, SEQ ID NO.3, SEQ ID NO.13 and SEQ ID NO.4 connected in sequence; (L2) P-CDA1∆194-miniCGBE: composed of SEQ ID NO.4, SEQ ID NO.1, SEQ ID NO.3, SEQ ID NO.13 and SEQ ID NO.4 connected in sequence; (L3) P-eA3A-miniCGBE: composed of SEQ ID NO.4, SEQ ID NO.6, SEQ ID NO.16, SEQ ID NO.3, SEQ ID NO.13 and SEQ ID NO.4 connected in sequence; rice callus tissue was transformed by Agrobacterium-mediated transformation, and DNA was extracted. The editing efficiency was obtained by high-throughput sequencing of the target fragment, followed by analysis and visualization. Except for replacing the template with the rice callus DNA mentioned above, the library preparation and analysis steps for high-throughput sequencing are the same as in Example 2. The high-throughput sequencing primers used are the primer sets T1-NGS-F / R~T6-NGS-F / R in Table 4. The specific methods for Agrobacterium transformation, rice callus transformation, and DNA extraction are as follows: 1. Agrobacterium transformation: (1) Place Agrobacterium competent cells in a 37℃ water bath and culture for 5 min; (2) Add the vector plasmid DNA corresponding to Figure 11 to Agrobacterium competent cells, and incubate in ice bath for 30 min, liquid nitrogen for 5 min, 37℃ water bath for 5 min, and ice bath for 5 min in sequence; (3) Add 700 μL LB liquid medium (antibiotic-free), and culture at 28℃ and 225 rpm for 2-3 h with shaking; (4) Take out the bacterial solution, centrifuge at 6000 rpm for 1 min, and keep 100 μL of supernatant to resuspend the bacterial block; (5) Spread the bacterial solution on a plate containing 50 mg / L kanamycin and 20 mg / L rifampicin resistant medium, and incubate upside down at 28℃ for 2-3 days. Perform colony PCR on the colonies and sequence the amplified products. Agrobacterium with correct sequencing results is recombinant Agrobacterium containing the corresponding vector.
[0152] 2. Rice callus transformation: (1) Soak mature seeds of Ningjing No. 7 in NaCl for 2 hours and rinse with sterile water 3-4 times; (2) Place the treated seeds on the callus induction medium and culture in the dark for 15-20 days; (3) In a liquid resistance medium containing 50 mg / L kanamycin and 20 mg / L rifampicin, culture at 28°C and 225 rpm until the OD value is in the range of 0.6-1.0; (4) Transform into the induced embryogenic callus; (5) 3 days after transformation, transfer the callus to N6 medium containing 2,4-dichlorophenoxyacetic acid, tyrosine, gel stone, sucrose, carbenicillin and hygromycin for 30 days.
[0153] 3. Extraction of DNA from rice callus: (1) Take new rice callus tissue, grind it with liquid nitrogen, 25 Hz, 90s, to obtain the ground sample; (2) Extract genomic DNA from the callus tissue using the TPS method.
[0154] The results are shown in Figure 12: Compared with its single-enzyme control group, P-CDG4-CDA1Δ194-miniCGBE significantly improved C-to-G editing efficiency at all tested sites, with a maximum improvement of 157-fold, confirming that the synergistic effect also exists in plants. P-CDG4-eA3A-miniCGBE also successfully replicated its editing window reshaping effect in yeast in plants, transferring the optimal editing activity to the C8-C10 region, becoming a high-performance plant CGBE tool with a unique targeting range. These results demonstrate that the dual-enzyme synergistic strategy of this invention is also applicable and effective in higher plants, providing a powerful new tool for precision crop breeding.
[0155] The scope of protection of this invention is not limited to the above embodiments. Variations and advantages that can be conceived by those skilled in the art without departing from the spirit and scope of the inventive concept are included in this invention and are protected by the appended claims.
Claims
1. A dual-enzyme CGBE editor, characterized in that, The dual-enzyme CGBE editor comprises cytosine DNA glycosylase CDG4, cytosine deaminase, and SpCas9 nickase nCas9(D10A); the cytosine DNA glycosylase CDG4 is fused to the N-terminus of the cytosine deaminase, and the cytosine deaminase is fused to the N-terminus of nCas9(D10A); the cytosine deaminase is selected from any one or more of CDA1Δ194, A1(R33A), eA3A, mini-Sdd7, and TadA8e(N46L); the amino acid sequence of CDA1Δ194 is shown in SEQ ID NO.1, the amino acid sequence of A1(R33A) is shown in SEQ ID NO.5, the amino acid sequence of eA3A is shown in SEQ ID NO.6, the amino acid sequence of mini-Sdd7 is shown in SEQ ID NO.7, and the amino acid sequence of TadA8e(N46L) is shown in SEQ ID NO.8; the amino acid sequence of the cytosine DNA glycosylase CDG4 is shown in SEQ ID NO.
1. As shown in NO.2; the amino acid sequence of the nCas9(D10A) is shown in SEQ ID NO.
3.
2. The dual-enzyme CGBE editor according to claim 1, characterized in that, The dual-enzyme CGBE editor also includes a nuclear localization signal NLS, with nCas9(D10A) directly or indirectly fused to the N-terminus of the nuclear localization signal NLS; the amino acid sequence of the nuclear localization signal NLS is shown in SEQ ID NO.
4.
3. The dual-enzyme CGBE editor according to claim 1, characterized in that, The dual-enzyme CGBE editor is selected from any one of the following (A1)-(A5): (A1) CDG4-CDA1Δ194-miniCGBE: composed of SEQ ID NO.2, SEQ ID NO.1, SEQ ID NO.3 and SEQ ID NO.4 in sequence; (A2) CDG4-A1(R33A)-miniCGBE: composed of SEQ ID NO.2, SEQ ID NO.5, SEQ ID NO.3 and SEQ ID NO.4 in sequence; (A3) CDG4-eA3A-miniCGBE: composed of SEQ ID NO.2, SEQ ID NO.6, SEQ ID NO.3 and SEQ ID NO.4 in sequence; (A4) CDG4-miniSdd7-miniCGBE: composed of SEQ ID NO.2, SEQ ID NO.7, SEQ ID NO.3 and SEQ ID NO.4 in sequence; (A5) CDG4-TadA8e(N46L)-miniCGBE: composed of SEQ ID NO.2, SEQ ID NO.8, SEQ ID NO.9 and SEQ ID NO.1 in sequence; (A5) CDG4-TadA8e(N46L)-miniCGBE: composed of SEQ ID NO.2, SEQ ID NO.9 and SEQ ID NO.1 in sequence; (A1) CDG4-CDA1Δ194-miniCGBE: composed of SEQ ID NO.2, SEQ ID NO.1, SEQ ID NO.2 and SEQ ID NO.1 in sequence; (A2) CDG4-A1(R33A)-miniCGBE: composed of SEQ ID NO.2, SEQ ID NO.9 ...3) CDG4-eA3A-miniCGBE: composed of SEQ ID NO.2, SEQ ID NO.9 The sequence consists of NO.3 and SEQ ID NO.4; the segments (A1)-(A5) may or may not include the linker sequence.
4. The dual-enzyme CGBE editor according to claim 1, characterized in that, The N-terminus of the cytosine DNA glycosylase CDG4 in the dual-enzyme CGBE editor is also connected to the nuclear localization signal NLS. nCas9(D10A) is not directly fused to the N-terminus of the nuclear localization signal NLS. There is also a nucleoplasmic protein NLS between nCas9(D10A) and the nuclear localization signal NLS. The amino acid sequence of the nucleoplasmic protein NLS is shown in SEQ ID NO.
13.
5. The dual-enzyme CGBE editor according to claim 4, characterized in that, The dual-enzyme CGBE editor is selected from any one of the following (B1)-(B2): (B1) P-CDG4-CDA1∆194-miniCGBE: composed of SEQ ID NO.4, SEQ ID NO.2, SEQ ID NO.1, SEQ ID NO.3, SEQ ID NO.13 and SEQ ID NO.4 in sequence; (B2) P-CDG4-eA3A-miniCGBE: composed of SEQ ID NO.4, SEQ ID NO.2, SEQ ID NO.6, SEQ ID NO.3, SEQ ID NO.13 and SEQ ID NO.4 in sequence; the fragments of (B1) and (B2) may or may not include a linker sequence.
6. A dual-enzyme CGBE editor system, characterized in that, The dual-enzyme CGBE editor system includes the dual-enzyme CGBE editor as described in any one of claims 1 to 5, and further includes sgRNA or sgRNA expression vector.
7. Encoding a nucleic acid molecule of the dual-enzyme CGBE editor of any one of claims 1 to 5 or the dual-enzyme CGBE editor system of claim 6.
8. A biological material relating to the nucleic acid molecule of claim 7, wherein the biological material is any one of the following: (C1) an expression cassette containing the nucleic acid molecule of claim 7; (C2) a recombinant vector containing the nucleic acid molecule of claim 7, or a recombinant vector containing the expression cassette of (C1); (C3) a recombinant microorganism containing the nucleic acid molecule of claim 7, or a recombinant microorganism containing the expression cassette of (C1), or a recombinant microorganism containing the recombinant vector of (C2).
9. The use of the dual-enzyme CGBE editor of any one of claims 1 to 5, the dual-enzyme CGBE editor system of claim 6, the nucleic acid molecule of claim 7, and the biomaterial of claim 8 in any of the following (D1)-(D2): (D1) in the preparation of gene editing products; (D2) in improving the scope and / or efficiency of gene editing for non-disease treatment purposes.
10. The application according to claim 9, characterized in that, The gene editing described is C-to-G base editing.
11. A method for improving the scope and / or efficiency of gene editing for non-disease treatment purposes, characterized in that, Editing is performed using the dual-enzyme CGBE editor system described in claim 6.
Citation Information
Cited By
Fusion protein, uracil-N-glycosylase mutant-mediated base editing system and application of fusion protein and uracil-N-glycosylase mutant-mediated base editing system
CN117126827A