Fusion protein having function of editing three bases and use thereof

By fusing adenine deaminase, N-methylpurine DNA glycosylase and Cas9n, a three-base editing tool that can edit three bases A, C, and G at the same time was developed, which solved the problem that the existing technology could not edit three bases at the same time, and achieved efficient gene editing and expanded application fields.

WO2025119385A1PCT designated stage expired Publication Date: 2025-06-12SUZHOU INST OF SYST MEDICINE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/137649
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-06
Filing Date
2024-12-06
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

The prior art has not yet developed a three-base editing technology that can edit three base substrates simultaneously, limiting the application of gene therapy, animal models of disease and crop genetic breeding.

Method used

By fusing adenine deaminase and N-methylpurine DNA glycosylase with Cas9n, a three-base editing tool can be developed that can simultaneously edit the three base substrates of A, C, and G.

Benefits of technology

The three bases A, C and G are simultaneously edited on alleles, which improves the flexibility and efficiency of gene editing and expands the application prospects of gene therapy and genetic breeding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024137649_12062025_PF_FP_ABST
    Figure CN2024137649_12062025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of base editing. Provided are a fusion protein having the function of editing three bases and the use thereof. The fusion protein comprises adenine deaminase, N-methylpurine DNA glycosylase, and nuclease. With regard to the problem of existing base editing tools not being able to simultaneously use three types of bases as substrates for editing, the present invention provides a new multifunctional base editor which is named A,C&G-BE. The A,C&G-BE can separately use A, C and G as a substrate for editing to realize A-G / T / C, C-T / G / A and G-T / A / C editing activities, and can also simultaneously use A, C and G as substrates on an allele for editing.
Need to check novelty before this filing date? Find Prior Art

Description

A fusion protein with three-base editing function and its application

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on December 6, 2023, with application number "202311666216.7" and invention name "A fusion protein with three-base editing function and its application", the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of base editing technology, and in particular to a fusion protein with three-base editing function and its application. Background Art

[0003] At present, base editing technologies mainly include cytosine base editing technology (CBE), adenine base editing technology (ABE) and guanine base editing technology (gGBE). They can use cytosine (C), adenine (A) and guanine (G) as substrates to achieve efficient single-substrate base editing, and have broad application prospects in disease animal models, gene therapy, genetic screening and crop genetic breeding.

[0004] In addition, dual-base editing technologies have been developed by integrating two deaminases with the CRISPR system to edit two bases. However, there is currently no triple-base editing technology that can edit three bases simultaneously, and basic and applied research urgently needs such new technologies. Summary of the Invention

[0005] To solve the above problems, the applicant, based on the working principle of base editing technology, fused adenine deaminase that can edit two base substrates (A and C) and N-methylpurine DNA glycosylase that can edit guanine bases with Cas9n, and developed a three-base editing tool that can simultaneously edit three base substrates A, C, and G.

[0006] On the one hand, the present application provides a fusion protein for multi-base editing, the fusion protein comprising: adenine deaminase, N-methylpurine DNA glycosylase, nuclease; the adenine deaminase is selected from TadA dual, T AD AC 3.1, T AD AC 3.155, TadA-8e; the N-methyl purine DNA glycosylase is selected from one or more of N-methyl purine DNA glycosylase MPG v6.3, N-methyl purine DNA glycosylase MPG v3 from humans, alkyl adenosine DNA glycosylase mAAG from mice, and adenosine DNA glycosylase from rat or Bacillus subtilis.

[0007] The multi-base editing is to edit 1-3 types of bases; optionally, the multi-base editing is to edit 3 types of bases; more optionally, the 3 types of bases include A, C, and G.

[0008] Optionally, the multi-base editing includes base conversion from A to G or T or C, C to T or G or A, G to T or A or C.

[0009] Optionally, the multi-base editing is simultaneous editing of multiple bases of alleles.

[0010] Furthermore, the adenine deaminase is TadA dual; optionally, the amino acid sequence of the adenine deaminase includes the amino acid sequence shown in SEQ ID No.1 or an amino acid sequence having at least 98% sequence identity with SEQ ID NO.1; more optionally, the coding sequence of the adenine deaminase includes the nucleotide sequence shown in SEQ ID No.2 or a nucleotide sequence having at least 98% sequence identity with SEQ ID NO.2.

[0011] Furthermore, the amino acid sequence of the adenine deaminase includes the amino acid sequence shown in SEQ ID No. 1 or an amino acid sequence having 98%, 98.1%, 98.2%, 98.3%, 98.4%, 98.5%, 98.6%, 98.7%, 98.8%, 98.9%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9% sequence identity with SEQ ID NO. 1; the coding sequence of the adenine deaminase includes the nucleotide sequence shown in SEQ ID No. 2 or an amino acid sequence having 98%, 98.1%, 98.2%, 98.3%, 98.4%, 98.5%, 98.6%, 98.7%, 98.8%, 98.9%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9% sequence identity with SEQ ID NO.2 has a nucleotide sequence with 98%, 98.1%, 98.2%, 98.3%, 98.4%, 98.5%, 98.6%, 98.7%, 98.8%, 98.9%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9% sequence identity

[0012] More optionally, the amino acid sequence of the adenine deaminase is shown as SEQ ID No. 1; the nucleotide sequence encoding the adenine deaminase is shown as SEQ ID No. 2.

[0013] Furthermore, the N-methylpurine DNA glycosylase is N-methylpurine DNA glycosylase MPG v6.3 derived from humans; optionally, the amino acid sequence of the N-methylpurine DNA glycosylase includes the amino acid sequence shown in SEQ ID No.3 or an amino acid sequence having at least 98% sequence identity with SEQ ID NO.3; more optionally, the coding sequence of the N-methylpurine DNA glycosylase includes the nucleotide sequence shown in SEQ ID No.4 or a nucleotide sequence having at least 98% sequence identity with SEQ ID NO.4.

[0014] Furthermore, the amino acid sequence of the N-methylpurine DNA glycosylase includes the amino acid sequence shown in SEQ ID No. 3 or an amino acid sequence having 98%, 98.1%, 98.2%, 98.3%, 98.4%, 98.5%, 98.6%, 98.7%, 98.8%, 98.9%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9% sequence identity with SEQ ID NO. 3; the coding sequence of the N-methylpurine DNA glycosylase includes the nucleotide sequence shown in SEQ ID No. 4 or an amino acid sequence having 98%, 98.1%, 98.2%, 98.3%, 98.4%, 98.5%, 98.6%, 98.7%, 98.8%, 98.9%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9% sequence identity with SEQ ID NO.4 has a nucleotide sequence with 98%, 98.1%, 98.2%, 98.3%, 98.4%, 98.5%, 98.6%, 98.7%, 98.8%, 98.9%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, and 99.9% sequence identity.

[0015] More optionally, the amino acid sequence of the N-methylpurine DNA glycosylase is shown as SEQ ID No. 3; the nucleotide sequence encoding the N-methylpurine DNA glycosylase is shown as SEQ ID No. 4.

[0016] Furthermore, the nuclease is selected from one or any several of Cas9, Cas3, Cas8a, Cas8b, Cas10d, Cse1, Csy1, Csn2, Cas4, Cas10, Csm2, Cmr5, Fok1, and Cpf1; optionally, the nuclease is Cas9; more optionally, the amino acid sequence of Cas9 includes the amino acid sequence shown in SEQ ID No.5 or an amino acid sequence having at least 98% sequence identity with SEQ ID NO.5; more optionally, the coding sequence of Cas9 includes the nucleotide sequence shown in SEQ ID No.6 or a nucleotide sequence having at least 98% sequence identity with SEQ ID NO.6.

[0017] Furthermore, the amino acid sequence of the Cas9 includes the amino acid sequence shown in SEQ ID No. 5 or an amino acid sequence having 98%, 98.1%, 98.2%, 98.3%, 98.4%, 98.5%, 98.6%, 98.7%, 98.8%, 98.9%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9% sequence identity to SEQ ID NO. 5; the coding sequence of the Cas9 includes the nucleotide sequence shown in SEQ ID No. 6 or an amino acid sequence having 98%, 98.1%, 98.2%, 98.3%, 98.4%, 98.5%, 98.6%, 98.7%, 98.8%, 98.9%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9% sequence identity to SEQ ID NO.6 has a nucleotide sequence with 98%, 98.1%, 98.2%, 98.3%, 98.4%, 98.5%, 98.6%, 98.7%, 98.8%, 98.9%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, and 99.9% sequence identity.

[0018] More optionally, the amino acid sequence of the Cas9 is shown as SEQ ID No.5; the nucleotide sequence encoding the Cas9 is shown as SEQ ID No.6.

[0019] Furthermore, the fusion protein further comprises an NLS (nuclear localization sequence); optionally, the NLS is located at at least one end of the fusion protein; more optionally, the NLS is located at both ends of the fusion protein; more optionally, the amino acid sequence of the NLS comprises the amino acid sequence shown in SEQ ID No.7 or an amino acid sequence having at least 98% sequence identity with SEQ ID NO.7; more optionally, the coding sequence of the NLS comprises the nucleotide sequence shown in SEQ ID No.8 or a nucleotide sequence having at least 98% sequence identity with SEQ ID NO.8.

[0020] Furthermore, the amino acid sequence of the NLS includes the amino acid sequence shown in SEQ ID No. 7 or an amino acid sequence having 98%, 98.1%, 98.2%, 98.3%, 98.4%, 98.5%, 98.6%, 98.7%, 98.8%, 98.9%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9% sequence identity to SEQ ID NO. 7; the coding sequence of the NLS includes the nucleotide sequence shown in SEQ ID No. 8 or an amino acid sequence having 98%, 98.1%, 98.2%, 98.3%, 98.4%, 98.5%, 98.6%, 98.7%, 98.8%, 98.9%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9% sequence identity to SEQ ID NO.8 has a nucleotide sequence with 98%, 98.1%, 98.2%, 98.3%, 98.4%, 98.5%, 98.6%, 98.7%, 98.8%, 98.9%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, and 99.9% sequence identity.

[0021] More optionally, the amino acid sequence of the NLS is shown as SEQ ID No.7; the nucleotide sequence encoding the NLS is shown as SEQ ID No.8.

[0022] Furthermore, in the fusion protein, adenine deaminase, nuclease, and N-methylpurine DNA glycosylase are sequentially connected starting from the N (nitrogen) end of the sequence toward the C (carbon) end.

[0023] In an optional embodiment, starting from the N-terminus of the sequence toward the C-terminus, the NLS, adenine deaminase, nuclease, N-methylpurine DNA glycosylase, and NLS in the fusion protein are connected in sequence.

[0024] In an optional embodiment, the amino acid sequence of the fusion protein includes the amino acid sequence shown in SEQ ID No.29 or an amino acid sequence having at least 98% sequence identity with SEQ ID NO.29; the sequence encoding the fusion protein includes the nucleotide sequence shown in SEQ ID No.30 or a nucleotide sequence having at least 98% sequence identity with SEQ ID NO.30.

[0025] Furthermore, the amino acid sequence of the fusion protein includes the amino acid sequence shown in SEQ ID No. 29 or an amino acid sequence having 98%, 98.1%, 98.2%, 98.3%, 98.4%, 98.5%, 98.6%, 98.7%, 98.8%, 98.9%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9% sequence identity to SEQ ID NO. NO.30 has a nucleotide sequence with 98%, 98.1%, 98.2%, 98.3%, 98.4%, 98.5%, 98.6%, 98.7%, 98.8%, 98.9%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, and 99.9% sequence identity.

[0026] It will be understood by those skilled in the art that a known linker can be added between the NLS, adenine deaminase, nuclease and N-methylpurine DNA glycosylase for connection, and that the fusion protein can also be subjected to known modifications, including phosphorylation, acetylation, ubiquitination, glycosylation, etc., without affecting the function of the fusion protein itself.

[0027] On the other hand, the present application also provides a biomaterial, which comprises any one of the following A)-D):

[0028] A) a gene encoding the fusion protein;

[0029] B) an expression cassette comprising the gene described in A);

[0030] C) a recombinant vector comprising A) the gene and / or B) the expression cassette;

[0031] D) A recombinant cell or recombinant bacterium containing the fusion protein, A) the gene, B) the expression cassette or C) the recombinant vector.

[0032] The expression cassette described herein may further include functional elements such as a promoter, a terminator, and a marker gene. Those skilled in the art may make routine selections based on actual conditions, as long as the expression of the gene described in A) can be achieved. The structure and composition of the expression cassette are not further restricted herein.

[0033] The vector described herein refers to a vector that can carry exogenous DNA or target gene into host cells for amplification and expression. The vector can be a cloning vector or an expression vector. Those skilled in the art can choose according to actual circumstances, and no excessive restrictions are imposed here.

[0034] It is understandable that those skilled in the art can select appropriate gene editing systems and gene editing methods according to actual conditions to complete the construction of the above-mentioned biomaterials.

[0035] On the other hand, the present application also provides a multi-base gene editing system, which includes the fusion protein; optionally, the multi-base gene editing system also includes sgRNA, and the sgRNA guides the fusion protein to perform gene editing on the target gene in the target cell; more optionally, the target sequence of the sgRNA includes at least one of SEQ ID No.37-SEQ ID No.73; optionally, the multi-base gene editing system can edit 1-3 types of bases; more optionally, the three-base gene editing system can edit three types of bases; more optionally, the three types of bases include A, C, and G; optionally, the multi-base gene editing system is a dual-base editing system or a three-base editing system.

[0036] Optionally, the three-base gene editing system can simultaneously edit three types of bases of the allele, and can achieve base conversion from A to G or T or C, C to T or G or A, and G to T or A or C in the editing target.

[0037] In an optional embodiment, the editing targets are PLS3-AS1-sg1, PPP1R12C site 3, EGFR-sg23, GJB2-sg1, EGFR-sg4, DMD-sg1, EMX1-sg1, USP46-sg1, EMX1-sg2p, PTEN-sg1, MAGEA1-Mb, PD-1-sg10, EMX1-sg6p, HFE-sg1, ABE site30, FANCF-sg2, FGF6-sg3, FANCF-sg3, EGFR-sg46, PD-1-sg3, TTR-sg1, Lag3-sg1, RPS24-sg1, RUNX1-sg3, HBG-sg9, AHCY-sg1, ABE site17, CCR5-sg1, ABE One or more of site13, PPP1R12C site 15, HEK site2, FANCF-sg1, CTLA4-sg2, FANCF-Mb, GANAB-sg1, RELA-sg1, and TIM3-sg2.

[0038] The inventors of this application combined the known adenine deaminase mutant (TadA dual) with two bases (A / C) editing activity and the N-methylpurine DNA glycosylase mutant (MPG v6.3) with guanine base editing activity. Through fusion and screening, the editing efficiency of different constructions for the three bases A, C, and G was compared. After verification, TadA-dual-MPG v6.3 was constructed and named A, C&G-BE7. In addition, this application verified that A, C&G-BE7 can use the three bases A, C, and G as substrates at the same time to achieve editing activities of AG / T / C, CT / G / A, and GT / A / C. Most importantly, it can achieve simultaneous editing of the three bases on the same DNA (i.e., allele).

[0039] On the other hand, the present application also provides the use of the fusion protein, the biomaterial, or the multi-base gene editing system in gene editing for non-disease diagnosis and treatment purposes and / or preparation of gene editing products.

[0040] Furthermore, the application is to edit 1-3 types of bases; optionally, the application is to edit 3 types of bases; more optionally, the 3 types of bases include A, C, and G.

[0041] The application is to simultaneously edit two types of bases of the allele, which can achieve base conversion from A to C, A to G, and C to G in the editing target.

[0042] The application is to simultaneously edit three types of bases of the allele, which can achieve base conversion from A to G or T or C, C to T or G or A, and G to T or A or C in the editing target.

[0043] Preferably, the editing oligos from A to G or T or C in the editing target are A4-A8; the editing oligos from C to T or G or A in the editing target are C4-C8; and the editing oligos from G to T or A or C are G5-G12.

[0044] More preferably, adenine located in YAN (Y=C / T, N=A / T / C / G) and GAT is easily edited; guanine located in NGR region (R=A / G) is easily edited.

[0045] More preferably, targets containing As and C at positions 4-8 and G at positions 5-12 are susceptible to simultaneous editing of A, C, or G bases.

[0046] The multi-base gene editing system also has the characteristics of high DNA mutation efficiency and low indel rate.

[0047] The present invention has the following beneficial effects:

[0048] To address the problem that existing base editing tools cannot simultaneously edit three bases, the present invention provides a new multifunctional base editor, named A,C&G-BE. A,C&G-BE can edit with A, C, and G as substrates individually, achieving editing activity for AG / T / C, CT / G / A, and GT / A / C. It can also simultaneously edit with A, C, and G as substrates on alleles. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0050] Figure 1 is a schematic diagram of the A, C & G-BE base editing principle;

[0051] Figure 2 is a schematic diagram of different plasmid designs and constructions;

[0052] Figure 3 is a statistical diagram of representative base editing efficiencies of different construction groups A, C & G-BE at the endogenous human target site PLS3-AS1-sg1;

[0053] Figure 4 is a statistical comparison of the base editing efficiency of A, C&G-BE7 at the endogenous human target site PLS3-AS1-sg1;

[0054] Figure 5 is a statistical diagram of representative base editing efficiencies of different construction groups A, C & G-BE at the endogenous human target site PLS3-AS1-sg1;

[0055] Figure 6 is a statistical graph of representative base editing efficiencies of different construction groups A, C&G-BE at the endogenous target PPP1R12C site 3 of humans;

[0056] Figure 7 is a statistical diagram of representative base editing efficiencies of different construction groups A, C & G-BE at the endogenous human target EGFR-sg23;

[0057] Figure 8 is a statistical diagram of representative base editing efficiencies of different construction groups A, C & G-BE at the endogenous human target site GJB2-sg1;

[0058] Figure 9 is a statistical diagram of representative base editing efficiencies of different construction groups A, C & G-BE at the endogenous human target EGFR-sg4;

[0059] Figure 10 is a statistical diagram of representative base editing efficiencies of different construction groups A, C & G-BE at the endogenous human target DMD-sg1;

[0060] Figure 11 is a statistical diagram of representative base editing efficiencies of different construction groups A, C & G-BE at the endogenous human target EMX1-sg1;

[0061] Figure 12 is a statistical diagram of representative base editing efficiencies of different construction groups A, C & G-BE at the endogenous human target USP46-sg1;

[0062] Figure 13 is a statistical diagram of representative base editing efficiencies of different construction groups A, C & G-BE at the endogenous human target EMX1-sg2p;

[0063] Figure 14 is a statistical diagram of representative base editing efficiencies of different construction groups A, C & G-BE at the endogenous human target PTEN-sg1;

[0064] Figure 15 is a statistical diagram of representative base editing efficiencies of different construction groups A, C & G-BE at the endogenous human target site MAGEA1-Mb;

[0065] Figure 16 is a statistical diagram of representative base editing efficiencies of different construction groups A, C & G-BE at the endogenous human target PD-1-sg10;

[0066] Figure 17 is a statistical diagram of representative base editing efficiencies of different construction groups A, C & G-BE at the endogenous human target EMX1-sg6p;

[0067] Figure 18 is a statistical diagram of representative base editing efficiencies of different construction groups A, C & G-BE at the endogenous human target HFE-sg1;

[0068] Figure 19 is a statistical diagram of representative base editing efficiencies of different construction groups A, C & G-BE at the endogenous human target site ABE site 30;

[0069] Figure 20 is a statistical diagram of representative base editing efficiencies of different construction groups A, C & G-BE at the endogenous human target FANCF-sg2;

[0070] Figure 21 is a statistical diagram of representative base editing efficiencies of different construction groups A, C & G-BE at the endogenous human target FGF6-sg3;

[0071] Figure 22 is a statistical diagram of representative base editing efficiencies of different construction groups A, C & G-BE at the endogenous human target FANCF-sg3;

[0072] Figure 23 is a statistical diagram of representative base editing efficiencies of different construction groups A, C & G-BE at the endogenous human target EGFR-sg46;

[0073] Figure 24 is a statistical diagram of representative base editing efficiencies of different construction groups A, C & G-BE at the endogenous human target PD-1-sg3;

[0074] Figure 25 is a statistical diagram of representative base editing efficiencies of different construction groups A, C & G-BE at the endogenous human target TTR-sg1;

[0075] Figure 26 is a statistical graph showing representative base editing efficiencies of different construction groups A, C & G-BE at the endogenous human target Lag3-sg1;

[0076] Figure 27 is a statistical diagram of representative base editing efficiencies of different construction groups A, C & G-BE at the endogenous human target RPS24-sg1;

[0077] Figure 28 is a statistical diagram of representative base editing efficiencies of different construction groups A, C & G-BE at the endogenous human target RUNX1-sg3;

[0078] Figure 29 is a statistical diagram of representative base editing efficiencies of different construction groups A, C & G-BE at the endogenous human target HBG-sg9;

[0079] Figure 30 is a statistical graph showing representative base editing efficiencies of different construction groups A, C & G-BE at the endogenous human target AHCY-sg1;

[0080] Figure 31 is a statistical diagram of representative base editing efficiencies of different construction groups A, C & G-BE at the endogenous human target site ABE site 17;

[0081] Figure 32 is a statistical diagram of representative base editing efficiencies of different construction groups A, C&G-BE at the endogenous human target CCR5-sg1;

[0082] Figure 33 is a statistical diagram of representative base editing efficiencies of different construction groups A, C & G-BE at the endogenous human target site ABE site 13;

[0083] Figure 34 is a statistical graph of representative base editing efficiencies of different construction groups A, C & G-BE at the endogenous human target PPP1R12C site 15;

[0084] Figure 35 is a statistical diagram of representative base editing efficiencies of different construction groups A, C & G-BE at the endogenous target site HEK site 2 in humans;

[0085] Figure 36 is a statistical graph showing representative base editing efficiencies of different construction groups A, C & G-BE at the endogenous human target FANCF-sg1;

[0086] Figure 37 is a statistical graph showing representative base editing efficiencies of different construction groups A, C & G-BE at the endogenous human target CTLA4-sg2;

[0087] Figure 38 is a statistical graph showing representative base editing efficiencies of different construction groups A, C & G-BE at the endogenous human target FANCF-Mb;

[0088] Figure 39 is a statistical graph showing representative base editing efficiencies of different construction groups A, C & G-BE at the endogenous human target GANAB-sg1;

[0089] Figure 40 is a statistical graph showing representative base editing efficiencies of different construction groups A, C & G-BE at the endogenous human target RELA-sg1;

[0090] Figure 41 is a statistical graph showing representative base editing efficiencies of different construction groups A, C & G-BE at the endogenous human target site TIM3-sg2;

[0091] FIG42 is a statistical graph showing the catalytic efficiency of A,C&G-BE5, A,C&G-BE6, and A,C&G-BE7 for A;

[0092] FIG43 is a statistical graph showing the catalytic efficiency of A,C&G-BE5, A,C&G-BE6, and A,C&G-BE7 for C;

[0093] FIG44 is a statistical graph showing the catalytic efficiency of A,C&G-BE5, A,C&G-BE6, and A,C&G-BE7 for G;

[0094] Figure 45 is a graph showing the efficiency of simultaneous multi-base editing of target sites detected by A, C&G-BE5, A, C&G-BE6, and A, C&G-BE7;

[0095] Figure 46 is a statistical diagram of the main A to G / C / T editing windows of A,C&G-BE5, A,C&G-BE6, and A,C&G-BE7;

[0096] Figure 47 is a statistical diagram of the main C to T / G / A editing windows of A,C&G-BE5, A,C&G-BE6, and A,C&G-BE7;

[0097] Figure 48 is a statistical diagram of the main G to C / T / A editing windows of A,C&G-BE5, A,C&G-BE6, and A,C&G-BE7;

[0098] Figure 49 is a statistical graph of the A / C / G simultaneous editing efficiency of A,C&G-BE5, A,C&G-BE6, and A,C&G-BE7;

[0099] Figure 50 is a statistical diagram of DNA mutation types in A, C&G-BE5, A, C&G-BE6, and A, C&G-BE7;

[0100] Figure 51 is a statistical graph of indel rates for A,C&G-BE5, A,C&G-BE6, and A,C&G-BE7;

[0101] Figure 52 is a statistical diagram of mutant allele types of the endogenous target site PLS3-AS1-sg1 DNA in humans for different construction groups A, C & G-BE;

[0102] Figure 53 is a statistical graph of adenine motif preferences for A,C&G-BE5, A,C&G-BE6, and A,C&G-BE7;

[0103] Figure 54 is a statistical graph of cytosine motif preferences for A,C&G-BE5, A,C&G-BE6, and A,C&G-BE7;

[0104] Figure 55 is a statistical graph showing the guanine motif preference of A,C&G-BE5, A,C&G-BE6, and A,C&G-BE7;

[0105] Figure 56 is a statistical graph showing the editing efficiency of PPP1R12Csite 3 of a mixture of A, C&G-BE7, AYBE v3, TadA-dual (-UGI), and gGBE v 6.3;

[0106] Figure 57 is a statistical graph showing the editing efficiency of the EGFR-sg23 site of a mixture of A, C & G-BE7, AYBE v3, TadA-dual (-UGI), and gGBE v 6.3;

[0107] Figure 58 is a statistical graph showing the editing efficiency of the GJB2-sg1 site of a mixture of A, C & G-BE7 with AYBE v3, TadA-dual (-UGI), and gGBE v 6.3;

[0108] Figure 59 is a statistical graph showing the editing efficiency of the EGFR-sg4 site of a mixture of A, C&G-BE7, AYBE v3, TadA-dual (-UGI), and gGBE v 6.3;

[0109] Figure 60 is a statistical graph showing the DMD-sg1 site editing efficiency of a mixture of A, C & G-BE7 with AYBE v3, TadA-dual (-UGI), and gGBE v 6.3;

[0110] Figure 61 is a statistical graph showing the editing efficiency of the EMX1-sg1 site of a mixture of A, C&G-BE7 with AYBE v3, TadA-dual (-UGI), and gGBE v 6.3;

[0111] Figure 62 is a statistical graph showing the editing efficiency of the USP46-sg1 site of a mixture of A, C&G-BE7 with AYBE v3, TadA-dual (-UGI), and gGBE v 6.3;

[0112] Figure 63 is a statistical graph showing the editing efficiency of the EMX1-sg2p site of a mixture of A, C&G-BE7 with AYBE v3, TadA-dual (-UGI), and gGBE v 6.3;

[0113] Figure 64 is a statistical graph showing the editing efficiency of the PTEN-sg1 site of a mixture of A, C & G-BE7 with AYBE v3, TadA-dual (-UGI), and gGBE v 6.3;

[0114] Figure 65 is a statistical graph showing the editing efficiency of the MAGEA1-Mb site of a mixture of A, C & G-BE7 with AYBE v3, TadA-dual (-UGI) and gGBE v 6.3;

[0115] Figure 66 is a statistical graph showing the editing efficiency of the PD-1-sg10 site of a mixture of A, C&G-BE7, AYBE v3, TadA-dual (-UGI), and gGBE v 6.3;

[0116] Figure 67 is a statistical graph showing the editing efficiency of the EMX1-sg6p site of a mixture of A, C & G-BE7 with AYBE v3, TadA-dual (-UGI), and gGBE v 6.3;

[0117] Figure 68 is a statistical graph showing the HFE-sg1 site editing efficiency of a mixture of A, C & G-BE7 with AYBE v3, TadA-dual (-UGI), and gGBE v 6.3;

[0118] Figure 69 is a statistical graph showing the ABE site 30 editing efficiency of a mixture of A, C&G-BE7 with AYBE v3, TadA-dual (-UGI), and gGBE v 6.3;

[0119] Figure 70 is a statistical graph showing the FANCF-sg2 site editing efficiency of a mixture of A, C & G-BE7 with AYBE v3, TadA-dual (-UGI), and gGBE v 6.3;

[0120] Figure 71 is a statistical graph showing the editing efficiency of the FGF6-sg3 site of a mixture of A, C&G-BE7, AYBE v3, TadA-dual (-UGI), and gGBE v 6.3;

[0121] Figure 72 is a statistical graph showing the FANCF-sg3 site editing efficiency of a mixture of A, C & G-BE7 with AYBE v3, TadA-dual (-UGI), and gGBE v 6.3;

[0122] Figure 73 is a statistical graph showing the editing efficiency of the EGFR-sg46 site of a mixture of A, C & G-BE7, AYBE v3, TadA-dual (-UGI), and gGBE v 6.3;

[0123] Figure 74 is a graph showing the multi-base simultaneous editing efficiency of a mixture of A, C & G-BE7 with AYBE v3, TadA-dual (-UGI), and gGBE v 6.3;

[0124] Figure 75 is a statistical graph of the simultaneous editing efficiency of A, C&G-BE7, AYBE v3, TadA-dual (-UGI), and gGBE v 6.3A / C / G;

[0125] Figure 76 is a statistical graph of DNA mutation types in A, C & G-BE7, AYBE v3, TadA-dual (-UGI), and gGBE v 6.3;

[0126] Figure 77 is a statistical graph of indel rates among A, C&G-BE7, AYBE v3, TadA-dual (-UGI), and gGBE v 6.3. DETAILED DESCRIPTION

[0127] Technical terms:

[0128] Identity: refers to the degree of similarity between the nucleotide sequences of two nucleic acid molecules or the amino acid sequences of two protein molecules in molecular evolution research.

[0129] Recombination: In a broad sense, any gene exchange process that causes genotype changes is called recombination.

[0130] Expression cassette: An expression cassette is a set of DNA sequences consisting of a promoter, target gene, and reporter gene, which can be expressed in specific tissues and easily detected.

[0131] Recombinant vector: A recombinant vector is a vector that transfers the target gene into the basic skeleton of a cloning vector, thereby enabling the target gene to be expressed.

[0132] Recombinant bacteria: A fungal cell line in which exogenous genes are efficiently expressed using genetic engineering methods.

[0133] Recombinant cell: The term "recombinant cell" means any cell type that is susceptible to transformation, transfection, transduction, etc. with a nucleic acid construct or expression vector comprising a polynucleotide of the present invention. The term "recombinant cell" encompasses any progeny of a parent cell that is not completely identical to the parent cell due to mutations that occur during replication.

[0134] In order to more clearly illustrate the overall concept of the present application, the following is a detailed description of the embodiments in conjunction with the accompanying drawings. In the following description, a large number of specific details are provided to provide a more thorough understanding of the present invention. However, it is obvious to those skilled in the art that the present invention can be implemented without one or more of these details. In other examples, in order to avoid confusion with the present invention, some technical features well known in the art are not described.

[0135] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present application belongs.

[0136] Unless otherwise specified, in the following embodiments, the reagents or instruments used without indicating the manufacturer are all conventional products that can be purchased from the market.

[0137] The procedures, conditions, reagents, and experimental methods used in the present invention, except those specifically mentioned below, are generally common knowledge and general understanding in the art and are not particularly limited herein. These procedures may be performed in accordance with Sambrook et al., Molecular Cloning, A Laboratory Manual (New York: Cold Spring Harbor Laboratory Press, 1989), or according to manufacturer's recommendations.

[0138] 1.1 Plasmid design and construction

[0139] 1.1.1. Based on the functional evaluation of existing base editing proteins, seven fusion plasmids were constructed (Figure 2): A,C&G-BE1, A,C&G-BE2, A,C&G-BE3, A,C&G-BE4, A,C&G-BE5, A,C&G-BE6, and A,C&G-BE7.

[0140] The principle diagram is shown in Figure 1. TadA deaminates adenine (A) and cytosine (C), and MPG cuts guanine (G) as well as hypoxanthine (I) and uracil (U) to produce abasic sites (AP). After DNA repair or replication, base conversion from A to G or T or C, C to T or G or A, and G to T or A or C in the editing target is achieved.

[0141] Among them, the amino acid sequence of TadA dual is shown in SEQ ID No.1, and the nucleotide sequence is shown in SEQ ID No.2, the amino acid sequence of TadA-8e is shown in SEQ ID No.9, and the nucleotide sequence is shown in SEQ ID No.10, ADThe amino acid sequence of AC-3.1 is shown in SEQ ID No. 11, and the nucleotide sequence is shown in SEQ ID No. 12. AD The amino acid sequence of AC-3.155 is shown in SEQ ID No.13, and the nucleotide sequence is shown in SEQ ID No.14, the amino acid sequence of MPG v6.3 is shown in SEQ ID No.3, and the nucleotide sequence is shown in SEQ ID No.4, the amino acid sequence of MPG v3 is shown in SEQ ID No.15, and the nucleotide sequence is shown in SEQ ID No.16, the amino acid sequence of spCas9n is shown in SEQ ID No.5, and the nucleotide sequence is shown in SEQ ID No.6, the amino acid sequence of bNLS is shown in SEQ ID No.7, and the nucleotide sequence is shown in SEQ ID No.8.

[0142] Specifically, SpCas9 TadDE (addgene #193837) and ABE8e (addgene #138489) were used as vectors to synthesize MPG v3 and MPG v6.3, and then cloned and assembled to construct plasmids A, C&G-BE1, A, C&G-BE2, A, C&G-BE3, A, C&G-BE4, A, C&G-BE5, A, C&G-BE6, A, C&G-BE7, and the amino acid and nucleotide sequences of the control plasmids TadA-dual v3 (without UGI), gGBE v6.3, and AYBE v3 are shown in Table 1.

[0143] Table 1. Construction group plasmid sequences

[0144] 1.1.2. 37 endogenous human targets were designed

[0145] Endogenous targets include PLS3-AS1-sg1, PPP1R12C site 3. EGFR-sg23, GJB2-sg1, EGFR-sg4, DMD-sg1, EMX1-sg1, USP46-sg1, EMX1-sg2p, PTEN-sg1, MAGEA1-Mb, PD-1-sg10, EMX1-sg6p, HFE-sg1, ABE site30, FANCF-sg2, FGF6-sg3, FANCF-sg3, EGFR-sg46, PD-1-sg3, TTR-sg1, Lag3-sg1, RPS24-sg1, RUNX1-sg3, HBG-sg9, AHCY-sg1, ABE site17, CCR5-sg1, ABE site13, PPP1R12C site 15. HEK site2, FANCF-sg1, CTLA4-sg2, FANCF-Mb, GANAB-sg1, RELA-sg1, and TIM3-sg2 (sequences are shown in Table 2 ) were screened. Synthetic target PLS3-AS1-sg1, PPP1R12C site 3. EGFR-sg23, GJB2-sg1, EGFR-sg4, DMD-sg1, EMX1-sg1, USP46-sg1, EMX1-sg2p, PTEN-sg1, MAGEA1-Mb, PD-1-sg10, EMX1-sg6p, HFE-sg1, ABE site30, FANCF-sg2, FGF6-sg3, FANCF-sg3, EGFR-sg46, PD-1-sg3, TTR-sg1, Lag3-sg1, RPS24-sg1, RUNX1-sg3, HBG-sg9, AHCY-sg1, ABE site17, CCR5-sg1, ABE site13, PPP1R12C site 15. The DNA of HEK site2, FANCF-sg1, CTLA4-sg2, FANCF-Mb, GANAB-sg1, RELA-sg1, and TIM3-sg2 were ligated into the BbsⅠ site of the sgRNA expression plasmid U6-sgRNA-EF1α-GFP (used to express the corresponding target sgRNA) by Golden Gate cloning to obtain the recombinant target plasmid. The target sites and sequences are shown in Table 2.

[0146] Table 2. Targets and sequences used

[0147] 1.1.3. Perform Sanger sequencing on the plasmids constructed in 1.1.1 and 1.1.2 to ensure they are completely correct.

[0148] 1.2 Cell Transfection

[0149] HEK293T 2×10 5 The cells were plated in 24-well plates. When the cells grew to 70%-80%, the plasmid combination was transfected according to the working plasmid:sgRNA=750ng:250ng. Each plasmid combination was transfected in 3 replicates, with 2×10 cells per well. 5 cells. A blank control without any plasmid transfection was also established. 24 and 48 hours after transfection, the supernatant was discarded and 500 μL of complete culture medium (89% DMEM (Gibco) + 10% FBS (BioVision Technology) + 1% PS (Gibco)) was added until the cell genome was extracted.

[0150] Working plasmids: plasmids A,C&G-BE1, A,C&G-BE2, A,C&G-BE3, A,C&G-BE4, A,C&G-BE5, A,C&G-BE6, A,C&G-BE7, with plasmids gGBE v6.3, AYBE v3, TadA-dual v3 (without UGI) as controls.

[0151] 1.3. Genome Extraction and Amplicon Library Preparation

[0152] 72h after transfection, BIOSEARCH QuickExtract TM Genomic DNA was extracted using a DNA extraction kit (QE09050). Following the Hitom kit protocol, corresponding identification primers were designed (Table 3). The forward identification primer was terminated with a bridging sequence (5'-ggagtgagtacggtgtgc-3') and the reverse identification primer was terminated with a bridging sequence (5'-gagttggatgctggatgg-3'). This yielded a first-round PCR product. This product was then used as a template for a second round of PCR. The resulting product was then mixed, purified, and harvested for gel cleavage and sent to a sequencing company for next-generation sequencing (NGS).

[0153] Table 3. Target identification primers used

[0154] 1.4 Data Analysis

[0155] NGS results were processed using www.rgenome.net. A, C, and G bases upstream of the PAM sequence (usually 2-14 nt) were counted and the data were summarized and plotted using GraphPad to obtain the final editing efficiency results. The editing window is numbered starting from the first base (nt) away from the PAM, and so on, up to the 20th base.

[0156] in conclusion:

[0157] Preliminary statistical analysis of high-throughput sequencing data at one target site (PLS3-AS1-sg1) revealed that high-throughput sequencing (HTS) showed that, as shown in Figure 3, MPG v3-based triple base editors (A, C&G-BE1, A, C&G-BE2, A, C&G-BE3) exhibited high editing efficiencies for A (6.67%-41.73%) and C (37.6-65.06%), but low editing efficiencies for G (0.17-2.1%). However, triple base editors based on MPG v6.3 (A, C&G-BE4, A, C&G-BE5, A, C&G-BE6, A, C&G-BE7) induced high editing efficiencies for A, C, and G, ranging from 5.43%-75.7%, 0.3%-51.03%, and 2.1%-43.7%, respectively. As shown in Figures 4 and 5, after further analysis of the two or three bases edited simultaneously in the same allele, we found that the efficiency of A / C / G simultaneous editing induced by the MPG v6.3-based triple base editor was significantly higher than that of the corresponding MPG v3-based triple base editor, among which A, C&G-BE7 had the highest efficiency (20.1%). At the same time, the MPG v6.3-based triple base editor also produced A / C, A / G and C / G simultaneous editing. The MPG v3-based triple base editor mainly induced A / C simultaneous editing, among which A, C&G-BE3 had the highest efficiency (63.3%). In addition, the MPG v6.3-based triple base editor produced more DNA mutation allele types than the corresponding MPG v3-based triple base editor (Figure 52). Therefore, we selected the MPG v6.3-based triple base editor (hereinafter named A, C&G-BE5-7) for further study.

[0158] To objectively analyze the properties of these optimized three-base editors, we further tested 36 endogenous targets containing multiple As, Cs, or Gs in HEK293T cells. As shown in Figures 6-41, after HTS analysis of the 36 targets, we found that all three-base editors can act on three base substrates (As, Cs, and Gs), effectively editing almost all tested targets. As shown in Figure 42, the catalytic efficiency of A for A,C&G-BE5, A,C&G-BE6, and A,C&G-BE7 ranged from 0.4-63.1% (median 17.5%), 0.4-67.4% (median 21.6%), and 2.3-80.5% (median 40.5%), respectively. As shown in Figure 43, the C catalytic efficiency of A, C & G-BE5, A, C & G-BE6 and A, C & G-BE7 were 4.0-72.1% (median 29.7%), 0.6-65.6% (median 21.1%) and 2.2-75.8% (median 32.7%), respectively. As shown in Figure 44, the G catalytic efficiency of A, C & G-BE5, A, C & G-BE6 and A, C & G-BE7 were 0.4-61.1% (median 11%), 0.3-63.4% (median 12.4%) and 0.5-61.4% (median 17.7%), respectively. It is worth noting that compared with the corresponding single-base editors, the A, C or G editing activity of A, C & G-BE was weakened, which may be due to the competitive binding between adenosine deaminase and N-methylpurine DNA glycosylase to the base substrate. The major A to G / C / T editing window of A, C&G-BE5, A, C&G-BE6 and A, C&G-BE7 is A4-A8 (the end away from the protospacer adjacent motif (PAM) is counted as position 1), which is slightly narrower than the editing window of AYBE v3 (A3-A9) ( Figure 46 ); the major C to T / G / A editing window of A, C&G-BE5, A, C&G-BE6 and A, C&G-BE7 is C4-C8, which is similar to TadA-dual(-UGI) ( Figure 47 ; the major G to T / C / A editing window of A, C&G-BE5, A, C&G-BE6 and A, C&G-BE7 is G5-G12, which is similar to gGBE v6.3 (3-12) is similar (Figure 48). To further analyze the motif preferences of these three-base editors in all tested targets, we found that adenines that are easily edited by A, C&G-BE5, A, C&G-BE6, and A, C&G-BE7 (average> 10%) are located in YAN (Y = C / T, N = A / T / C / G) and GAT. In addition, A, C&G-BE5 and A, C&G-BE6 prefer the AAG motif, but A, C&G-BE7 does not (Figure 53). Guanines that are easily edited (average> 10%) are located near the NGR region (R = A / G) (Figure 55). However, no motif preference was observed for cytosine (Figure 54).Further analysis of A / C / G co-editing on the same allele (>1%) revealed that A,C&G-BE5, A,C&G-BE6, and A,C&G-BE7 exhibited A / C / G co-editing at 19, 17, and 24 of the 36 targets tested, respectively (Figure 45). The A / C / G co-editing efficiency of A,C&G-BE7 ranged from 1.6% to 31.3%, similar to that of A,C&G-BE5 (1.1% to 23.6%) and higher than that of A,C&G-BE6 (1.1% to 14.0%) (Figure 49). Higher A / C / G co-editing efficiencies were observed at targets containing As and C at positions 4-8 and G at positions 5-12. If the A / C / G editing window contained the corresponding preferential motifs, the A / C / G co-editing efficiency would be further improved. We also compared the A / C / G simultaneous editing efficiency and DNA mutation types of A,C&G-BE7 with those of the AYBE v3, TadA-dual (-UGI), and gGBE v 6.3 mixtures. The results showed that A,C&G-BE7 exhibited higher A / C / G simultaneous editing efficiency than the mixture, and there was no significant difference in DNA mutation types between the two groups (Figures 56-77). In addition, A,C&G-BEs can also simultaneously edit two base substrates (A / C, A / G, C / G) (>1%) (Figure 45). The A / C editing efficiencies of A,C&G-BE5, A,C&G-BE6, and A,C&G-BE7 were 1.7-55.0%, 1.1-39.3%, and 1.0-46.9%, respectively (Figure 45); the A / G editing efficiencies of A,C&G-BE5, A,C&G-BE6, and A,C&G-BE7 were 1.0-52.5%, 1.0-55.0%, and 1.1-66.5%, respectively (Figure 45); the C / G editing efficiencies of A,C&G-BE5, A,C&G-BE6, and A,C&G-BE7 were 1.2-36.3%, 1.2-23.5%, and 1.2-40.3%, respectively (Figure 45). Therefore, A,C&G-BEs can also be used as dual base editors, such as A&G, C&G, or A&C base editors, for editing certain targets containing only A / C, A / G, and C / G within the editing window. We also compared the DNA mutations and indel profiles induced by A,C&G-BEs and found that A,C&G-BE7 induced the most DNA mutations and the lowest indel rate compared to A,C&G-BE5 and A,C&G-BE6 (Figures 50 and 51).

[0159] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application should be included within the scope of the claims of the present application.

Claims

1. A fusion protein for multi-base editing, characterized in that: The fusion protein comprises: adenine deaminase, N-methylpurine DNA glycosylase, and nuclease; the adenine deaminase is selected from TadA dual, T AD AC 3.1, T AD AC 3.155, TadA-8e; the alkyladenine DNA glycosylase is selected from one or more of N-methylpurine DNA glycosylase MPG v6.3, N-methylpurine DNA glycosylase MPG v3 derived from humans, alkyladenosine DNA glycosylase mAAG from mice, and adenosine DNA glycosylase from rats or Bacillus subtilis.

2. The fusion protein according to claim 1, characterized in that The adenine deaminase is TadA dual; optionally, the amino acid sequence of the adenine deaminase includes the amino acid sequence shown in SEQ ID No.1 or an amino acid sequence having at least 98% sequence identity with SEQ ID NO.1; more optionally, the coding sequence of the adenine deaminase includes the nucleotide sequence shown in SEQ ID No.2 or a nucleotide sequence having at least 98% sequence identity with SEQ ID NO.

2.

3. The fusion protein according to claim 1, characterized in that The N-methyl purine DNA glycosylase is N-methyl purine DNA glycosylase MPG v6.3 derived from humans; optionally, the amino acid sequence of the N-methyl purine DNA glycosylase includes the amino acid sequence shown in SEQ ID No.3 or an amino acid sequence having at least 98% sequence identity with SEQ ID NO.3; more optionally, the coding sequence of the N-methyl purine DNA glycosylase includes the nucleotide sequence shown in SEQ ID No.4 or a nucleotide sequence having at least 98% sequence identity with SEQ ID NO.

4.

4. The fusion protein according to claim 1, characterized in that The nuclease is selected from one or any several of Cas9, Cas3, Cas8a, Cas8b, Cas10d, Cse1, Csy1, Csn2, Cas4, Cas10, Csm2, Cmr5, Fok1, and Cpf1; optionally, the nuclease is Cas9; more optionally, the amino acid sequence of Cas9 includes the amino acid sequence shown in SEQ ID No.5 or an amino acid sequence having at least 98% sequence identity with SEQ ID NO.5; more optionally, the coding sequence of Cas9 includes the nucleotide sequence shown in SEQ ID No.6 or a nucleotide sequence having at least 98% sequence identity with SEQ ID NO.

6.

5. The fusion protein according to claim 1, characterized in that The fusion protein also includes an NLS; optionally, the NLS is located at at least one end of the fusion protein; more optionally, the NLS is located at both ends of the fusion protein; more optionally, the amino acid sequence of the NLS includes the amino acid sequence shown in SEQ ID No.7 or an amino acid sequence having at least 98% sequence identity with SEQ ID NO.7; more optionally, the coding sequence of the NLS includes the nucleotide sequence shown in SEQ ID No.8 or a nucleotide sequence having at least 98% sequence identity with SEQ ID NO.

8.

6. The fusion protein according to claim 1, characterized in that In the fusion protein, adenine deaminase, nuclease, and N-methylpurine DNA glycosylase are sequentially connected from the N-terminal to the C-terminal.

7. Biomaterial, characterized in that The biological material includes any one of the following A)-D): A) a gene encoding the fusion protein according to any one of claims 1 to 6; B) an expression cassette, the expression cassette comprising the gene described in A); C) a recombinant vector, the recombinant vector containing A) the gene and / or B) the expression cassette; D) A recombinant cell or recombinant bacterium, comprising the fusion protein according to any one of claims 1 to 6, the gene according to A), the expression cassette according to B or the recombinant vector according to C).

8. A multi-base gene editing system, characterized in that: The polybasic gene editing system comprises the fusion protein according to any one of claims 1 to 6; optionally, the polybasic gene editing system further comprises sgRNA, wherein the sgRNA guides the fusion protein to perform gene editing on the target gene in the target cell; More optionally, the target sequence of the sgRNA includes at least one of SEQ ID No.37-SEQ ID No.73; optionally, the multi-base gene editing system is capable of editing 1-3 types of bases; more optionally, the multi-base gene editing system is capable of editing 3 types of bases; more optionally, the 3 types of bases include A, C, and G; optionally, the multi-base gene editing system is a dual-base editing system or a tri-base editing system.

9. Use of the fusion protein as described in any one of claims 1 to 6, or the biomaterial as described in claim 7, or the multi-base gene editing system as described in claim 8 in gene editing for non-disease diagnosis and treatment purposes and / or in preparing gene editing products.

10. The use according to claim 9, characterized in that: The application is to edit 1-3 types of bases; optionally, the application is to edit 3 types of bases; more optionally, the 3 types of bases include A, C, and G.

Citation Information

Patent Citations

  • Alkyl adenine DNA glycosylase mutant, fusion protein and application

    CN116694605A

  • Programmable adenine base editor and uses thereof

    WO2023217280A1