Novel base editor and use thereof

By mutating the amino acid sequence of double-stranded DNA deaminase and fusing it with a programmable DNA-binding protein, a highly efficient base editing system was formed, solving the problem of low mitochondrial DNA editing efficiency in existing technologies and achieving efficient editing of DNA fragments such as "AC" and "GC".

CN121931087APending Publication Date: 2026-04-28LINGANG LAB
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
LINGANG LAB
Filing Date
2024-10-25
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing gene editing methods are inefficient when editing mitochondrial DNA, especially when editing DNA fragments such as "AC" and "GC", which limits their application in the treatment of mitochondrial genetic diseases.

Method used

Develop a double-stranded DNA deaminase that improves the editing efficiency of DNA fragments such as "AC" and "GC" by making specific mutations or deletions in the amino acid sequence of wild-type double-stranded DNA deaminase, and fuse it with a programmable DNA-binding protein to form a base editing system.

Benefits of technology

It improved the efficiency of editing mitochondrial DNA, especially the editing efficiency of DNA fragments such as "AC" and "GC", reduced cytotoxicity, and expanded the editing range.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005103195220000271
    Figure BDA0005103195220000271
  • Figure BDA0005103195220000281
    Figure BDA0005103195220000281
  • Figure HDA0005103195230000011
    Figure HDA0005103195230000011
Patent Text Reader

Abstract

The invention provides a novel base editor and application thereof. Specifically, the invention provides a novel double-stranded DNA deaminase, and a fusion protein (novel base editor), a base editing system, a base editing method and the like based on the novel double-stranded DNA deaminase, so as to improve the DNA editing efficiency of eukaryotic cell genes (especially mitochondrial genes), and especially improve the editing efficiency of non-TC DNA fragments such as' AC ',' GC 'and the like. According to the method, the types of the DNA fragments which can be effectively edited are expanded, and a wider prospect is provided for clinical application of DNA editing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of gene editing technology, and more specifically to a novel base editor and its uses. Background Technology

[0002] In the biomedical field, gene editing technology has always been a research hotspot. Especially in the treatment of mitochondrial genetic diseases, the search for efficient and cell-free gene editing tools is crucial. Currently, various gene editing technologies exist, including CRISPR-Cas-based editing methods and base editing methods that utilize single-stranded DNA cytosine deaminase to catalyze the conversion of cytosine to uracil.

[0003] However, these methods still have limitations when editing mitochondrial DNA. Since naturally occurring cytosine deaminases typically only catalyze the deamination of cytosine within single-stranded DNA, existing gene editing methods primarily focus on single-stranded DNA editing, with limited research on double-stranded DNA editing. This is especially true in mitochondria, a unique organelle whose genome encodes proteins of the respiratory chain, making gene editing there particularly important. However, due to the protective double-stranded structure of the mitochondrial genome and its difficulty in entry, existing gene editing methods have limited effectiveness.

[0004] A recently discovered double-stranded DNA deaminase, DddA (BruceDddA), from Burkholderia, can recognize and deaminate cytosine in double-stranded DNA, showing potential for the treatment of mitochondrial genetic diseases. However, its efficiency in editing mitochondrial DNA still needs improvement. Furthermore, BruceDddA exhibits a high preference for "TC" DNA fragments, but its editing efficiency on "AC" and "GC" DNA fragments is lower. This limits the effectiveness and applicability of this enzyme in clinical applications.

[0005] Therefore, there is an urgent need in this field to develop a new base editor with higher editing efficiency and a wider range of editable DNA fragments, as well as its applications. Summary of the Invention

[0006] The purpose of this invention is to provide a novel double-stranded DNA deaminase to improve the editing efficiency of eukaryotic gene (especially mitochondrial gene) DNA, particularly for editing DNA fragments such as "AC" and "GC".

[0007] In a first aspect of the invention, a double-stranded DNA deaminase is provided, wherein the amino acid sequence of the double-stranded DNA deaminase corresponding to the wild-type double-stranded DNA deaminase has one or more mutant amino acids selected from the group consisting of alanine (A), serine (S), valine (V), glycine (G), or combinations thereof; and the amino acid sequence of the wild-type double-stranded DNA deaminase is SEQ ID NO:1, SEQ ID NO:2, or SEQ ID NO:3.

[0008] In another preferred embodiment, the double-stranded DNA deaminase also has a deletion mutation in the amino acid sequence corresponding to the wild-type double-stranded DNA deaminase.

[0009] In another preferred embodiment, the double-stranded DNA deaminase is mutated at one or more sites selected from the group consisting of the amino acid sequence of the wild-type double-stranded DNA deaminase:

[0010] 37th, 44th, 45th, 58th, 59th, 63rd, 64th, 73rd, 93rd, 94th, 95th, 109th, 114th, 115th, 129th, 134th.

[0011] In another preferred embodiment, the double-stranded DNA deaminase has an amino acid sequence corresponding to the wild-type double-stranded DNA deaminase having one or more mutations selected from the group consisting of:

[0012] E44A, E58A, E63A; S37G, D115G, S129G, P134G; Q45S, G59S, G64S, G73S; A95V, A109V, A114V; deletion mutation at position 93, deletion mutation at position 94.

[0013] In another preferred embodiment, the double-stranded DNA deaminase has an E58A mutation corresponding to the amino acid sequence of the wild-type double-stranded DNA deaminase, and the amino acid sequence of the wild-type double-stranded DNA deaminase is SEQ ID NO:1.

[0014] In another preferred embodiment, the double-stranded DNA deaminase has an E44A mutation corresponding to the amino acid sequence of the wild-type double-stranded DNA deaminase, and the amino acid sequence of the wild-type double-stranded DNA deaminase is SEQ ID NO:2.

[0015] In another preferred embodiment, the double-stranded DNA deaminase has an E63A mutation corresponding to the amino acid sequence of the wild-type double-stranded DNA deaminase, and the amino acid sequence of the wild-type double-stranded DNA deaminase is SEQ ID NO:3.

[0016] In another preferred embodiment, the double-stranded DNA deaminase has S37G, G59S, G73S, A109V and S129G mutations corresponding to the amino acid sequence of the wild-type double-stranded DNA deaminase, and the amino acid sequence of the wild-type double-stranded DNA deaminase is SEQ ID NO:1.

[0017] In another preferred embodiment, the double-stranded DNA deaminase corresponds to the amino acid sequence of the wild-type double-stranded DNA deaminase having Q45S, deletion mutation at position 93, deletion mutation at position 94, A95V, and D115G mutations, and the amino acid sequence of the wild-type double-stranded DNA deaminase is SEQ ID NO:2.

[0018] In another preferred embodiment, the double-stranded DNA deaminase has G64S, A114V and P134G mutations in the amino acid sequence corresponding to the wild-type double-stranded DNA deaminase, and the amino acid sequence of the wild-type double-stranded DNA deaminase is SEQ ID NO:3.

[0019] In another preferred embodiment, the wild-type double-stranded DNA deaminase is derived from Burkholderiacepacia, Simiaoa Sunni, Roseburia intestinalis, or a combination thereof.

[0020] In another preferred embodiment, the amino acid sequence of the double-stranded DNA deaminase is selected from the group consisting of:

[0021] (a) SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ IDNO:8 or SEQ IDNO:9;

[0022] (b) Derived sequences having at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, or SEQ ID NO:9;

[0023] (c) Derivative sequences obtained by adding, deleting, modifying and / or substituting at least one (e.g., 1-20, 1-10, 1-5) amino acids based on SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8 or SEQ ID NO:9.

[0024] In another preferred embodiment, the amino acid sequence of the double-stranded DNA deaminase is SEQ ID NO:4, SEQ ID NO:5, or SEQ ID NO:6.

[0025] In another preferred embodiment, the amino acid sequence of the double-stranded DNA deaminase is SEQ ID NO:7, SEQ ID NO:8, or SEQ ID NO:9.

[0026] In another preferred embodiment, the double-stranded DNA deaminase has one or more of the following characteristics:

[0027] (i) It exhibits lower cytotoxicity compared to wild-type double-stranded DNA deaminase;

[0028] (ii) It can edit DNA fragments more effectively compared to wild-type double-stranded DNA deaminases.

[0029] In another preferred embodiment, the DNA fragment is a double-stranded DNA fragment.

[0030] In another preferred embodiment, the DNA fragment is a mitochondrial DNA fragment.

[0031] In another preferred embodiment, the DNA fragment comprises an "AC" DNA fragment.

[0032] In another preferred embodiment, the DNA fragment is an "AC" DNA fragment.

[0033] In another preferred embodiment, the DNA fragment comprises a “CC” DNA fragment.

[0034] In another preferred embodiment, the DNA fragment is a “CC” DNA fragment.

[0035] In another preferred embodiment, the DNA fragment comprises a “GC” DNA fragment.

[0036] In another preferred embodiment, the DNA fragment is a “GC” DNA fragment.

[0037] In another preferred embodiment, the DNA fragment comprises a “TC” DNA fragment.

[0038] In another preferred embodiment, the DNA fragment is a “TC” DNA fragment.

[0039] In another preferred embodiment, the double-stranded DNA deaminase has one or more of the following characteristics:

[0040] (i) It exhibits lower cytotoxicity compared to wild-type double-stranded DNA deaminase;

[0041] (ii) Compared with wild-type double-stranded DNA deaminase, it can edit “AC” DNA fragments more effectively;

[0042] (iii) Compared with wild-type double-stranded DNA deaminase, it can edit the “CC” DNA fragment more effectively;

[0043] (iv) Compared with wild-type double-stranded DNA deaminase, it can edit “GC” DNA fragments more effectively;

[0044] (v) Compared with wild-type double-stranded DNA deaminase, it can edit the “TC” DNA fragment more effectively.

[0045] In another preferred embodiment, the double-stranded DNA deaminase is capable of catalyzing cytosine deaminase activity in double-stranded DNA, and is preferably capable of serving as a core component of a base editor to precisely induce C-to-T conversion in double-stranded DNA at the target sequence and its complementary strand.

[0046] In another preferred embodiment, a C-to-T conversion occurs on the "AC" DNA segment of the double-stranded DNA.

[0047] In another preferred embodiment, a C-to-T conversion occurs on the "CC" DNA fragment of the double-stranded DNA.

[0048] In another preferred embodiment, a C-to-T conversion occurs on the "GC" DNA segment of the double-stranded DNA.

[0049] In another preferred embodiment, a C-to-T conversion occurs on the “TC” DNA fragment of the double-stranded DNA.

[0050] In a second aspect of the invention, a fusion protein is provided, the fusion protein comprising:

[0051] (a) a double-stranded DNA deaminase as described in the first aspect of the invention; and

[0052] (b) One or more elements or functional domains selected from the group consisting of: signal peptides, localization signals, reporter proteins, programmable DNA-binding proteins, DNA-binding domains, epitope tags, transcription activation domains, transcription repression domains, nucleases, methyltransferases, transcription release factors, deacetylases, cleavage active peptides, linkers, ligases, integrases, transposases, recombinases, polymerases, base excision repair inhibitors, or combinations thereof.

[0053] In another preferred embodiment, the fusion protein comprises:

[0054] (i) a double-stranded DNA deaminase as described in the first aspect of the invention; and

[0055] (ii) Programmable DNA-binding proteins.

[0056] In another preferred embodiment, the programmable DNA-binding protein is selected from the group consisting of: Cas protein, Argonaute (Ago) protein, TALE protein, zinc finger protein, or a combination thereof.

[0057] In another preferred embodiment, the programmable DNA-binding protein is a Cas protein or a TALE protein.

[0058] In another preferred embodiment, the Cas protein includes, but is not limited to: type II CRISPR-Cas peptide, type I CRISPR-Cas peptide, type III CRISPR-Cas peptide, type IV CRISPR-Cas peptide, type V CRISPR-Cas peptide, type VI CRISPR-Cas peptide, type VII CRISPR-Cas peptide, IscB peptide, TnpB peptide, IsrB peptide, or combinations thereof.

[0059] In another preferred embodiment, the type I CRISPR-Cas polypeptide includes I-A, I-B, I-C, I-D, I-E, and I-FCRISPR-Cas proteins.

[0060] In another preferred embodiment, the V-type CRISPR-Cas polypeptide includes the Cas12 protein.

[0061] In another preferred embodiment, the type VI CRISPR-Cas polypeptide includes the Cas13 protein.

[0062] In some embodiments, the Cas protein is selected from the group consisting of Cas9, CasX, CasY, Cas12a (Cpf1), Cas12b (C2c1), Cas13a (C2c2), Cas12c (C2c3), Cas12g, Cas12h, Cas12i, Cas13b, Cas13c, Cas13d, Cas14, Csn2, or combinations thereof.

[0063] In another preferred embodiment, the Cas protein is selected from the group consisting of: Cas9 proteins (e.g., SpCas9, SaCas9, GeoCas9, CjCas9, Cas9-KKH, cyclically arranged Cas9, SmacCas9, Spy-macCas9, xCas9, SpCas9-NG, SaCas9-NG); Cas12 proteins (e.g., Cas12a, AsCas12a, LbCas12a, Cas12b, Cas12c, Cas12d, Cas12e, Cas12f (Cas14), Cas12g, Cas12h, Cas12i, xCa s12i, Cas12Max, hfCas12Max, Cas12j, Cas12k, Cas12l, Cas12m, Cas12n, Cas12o, Cas12p, Cas12q, Cas12r, Cas12s, Cas12t, Cas12u, Cas12v, Cas12w, Cas12x, Cas12y, Cas12z); Cas13 proteins (e.g., Cas13a, Cas13b, Cas13c, Cas13d, Cas13e, Cas13f, Cas13x, Cas13y); Csn2; or mutants thereof, or combinations thereof.

[0064] In another preferred embodiment, the Cas protein is a nicking enzyme.

[0065] In another preferred embodiment, the Cas protein is the dCas protein.

[0066] In another preferred embodiment, the Cas protein is SpCas9 or its mutant SpCas9-SpRY.

[0067] In another preferred embodiment, the SpCas9 mutant is selected from the group consisting of SpCas9-SpRY, eSpCas9, SpCas9-HF1, SpCas9-NG, VQR-SpCas9, VRER-SpCas9, or combinations thereof.

[0068] In another preferred embodiment, the amino acid sequence of SpCas9-SpRY is shown in SEQ ID NO:13.

[0069] In another preferred embodiment, the Ago protein is selected from the group consisting of pAgo, eAgo, Ago1, Ago2, Ago3, Ago4, or combinations thereof.

[0070] In another preferred embodiment, when the programmable DNA-binding protein is a Cas protein, the fusion protein further comprises one or more elements selected from the group consisting of: a promoter, a localization signal (nuclear localization signal), a linker, a base excision repair inhibitor (preferably a uracil glycosylase inhibitor (UGI)), a reporter protein, or a combination thereof.

[0071] In another preferred embodiment, when the programmable DNA-binding protein is a TALE protein, the fusion protein further contains: a signal peptide (preferably a mitochondrial-targeting signal peptide (MTS)) and / or a base excision repair inhibitor (preferably a uracil glycosylase inhibitor (UGI)).

[0072] In another preferred embodiment, when the programmable DNA-binding protein is a zinc finger protein, the fusion protein further contains: a nuclease (preferably an endonuclease) and / or a base excision repair inhibitor (preferably a uracil glycosylase inhibitor (UGI)).

[0073] In another preferred embodiment, the DNA binding domain includes: methylation-binding protein, LexADBD, and Gal4DBD.

[0074] In another preferred embodiment, the epitope tag includes: histidine tag, V5 tag, FLAG tag, influenza virus hemagglutinin tag, Myc tag, VSV-G tag, thioredoxin tag, and streptavidin tag.

[0075] In another preferred embodiment, the epitope tag is a tag sequence that assists in expression and / or purification.

[0076] In another preferred embodiment, the fusion protein comprises:

[0077] (i) a double-stranded DNA deaminase as described in the first aspect of the invention; and

[0078] (ii) Optional tag sequences to assist in expression and / or purification.

[0079] In another preferred embodiment, the tag sequence includes: 6His tag, HA tag, Flag tag (including 3×Flag tag), Fc tag, or a combination thereof.

[0080] In another preferred embodiment, the transcriptional activation domain includes VP64 and / or VPR.

[0081] In another preferred embodiment, the transcriptional activation domain includes VP64 and / or VPR.

[0082] In another preferred embodiment, the nuclease comprises FokI.

[0083] In another preferred embodiment, the deacetylase is a histone deacetylase (HDAC).

[0084] In another preferred embodiment, the cleavage-active polypeptide includes: a polypeptide having single-stranded RNA cleavage activity, a polypeptide having double-stranded RNA cleavage activity, a polypeptide having single-stranded DNA cleavage activity, or a polypeptide having double-stranded DNA cleavage activity.

[0085] In another preferred embodiment, the ligase comprises DNA ligase and / or RNA ligase.

[0086] In another preferred embodiment, the base excision repair inhibitor is a uracil glycosylation inhibitor (UGI).

[0087] In another preferred embodiment, the fusion protein contains any of the following structures from the N-terminus to the C-terminus:

[0088] Z1-Z2(A);

[0089] Z2-Z1(B);

[0090] Z3-Z1-Z4(C);

[0091] In the formula,

[0092] Z1 is a double-stranded DNA deaminase as described in the first aspect of the present invention;

[0093] Z2 is a programmable DNA-binding protein;

[0094] Z3 is the N-terminal fragment of a programmable DNA-binding protein;

[0095] Z4 is the C-terminal fragment of a programmable DNA-binding protein;

[0096] The hyphen "-" indicates a linking peptide or peptide bond.

[0097] In another preferred embodiment, the fusion protein contains the structure shown in Formula I, Formula II, or Formula III from the N-terminus to the C-terminus:

[0098] NLS1-L1-DddA-L2-Cas-(L3-U) x -NLS2(I)

[0099] NLS1-L1-DddA-L2-Cas-(L3-U) x -NLS2-F(II)

[0100] NLS1-L1-DddA-L2-Cas-(L3-U) x -NLS2-2A-R(III)

[0101] In the formula,

[0102] NLS1 is the nuclear positioning signal;

[0103] DddA is a double-stranded DNA deaminase as described in the first aspect of the present invention;

[0104] Cas is the Cas protein;

[0105] U is a base excision repair inhibitor;

[0106] x is a positive integer from 1 to 5;

[0107] NLS2 is the nuclear positioning signal;

[0108] F represents the label sequence;

[0109] 2A is a self-cleaving 2A peptide;

[0110] R stands for reporter protein;

[0111] L1, L2, and L3 are each independently either empty or connected.

[0112] The hyphen "-" indicates a linking peptide or peptide bond.

[0113] In another preferred embodiment, the NLS1 is a bipartite NLS (bpNLS).

[0114] In another preferred embodiment, the amino acid sequence of the NLS1 is shown in SEQ ID NO:10 or SEQ ID NO:11.

[0115] In another preferred embodiment, L2 is an XTEN element.

[0116] In another preferred embodiment, the amino acid sequence of L2 is shown in SEQ ID NO:12.

[0117] In another preferred embodiment, the amino acid sequence of the double-stranded DNA deaminase is selected from the group consisting of: SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, or a combination thereof.

[0118] In another preferred embodiment, the Cas protein is SpCas9-SpRY.

[0119] In another preferred embodiment, the amino acid sequence of SpCas9-SpRY is shown in SEQ ID NO:13.

[0120] In another preferred embodiment, L3 is a flexible peptide having (GGGGS) n The structure is , where n is a positive integer from 1 to 5.

[0121] In another preferred embodiment, the amino acid sequence of L3 is shown in SEQ ID NO:14.

[0122] In another preferred embodiment, the U is a uracil glycosylase inhibitor (UGI).

[0123] In another preferred embodiment, the amino acid sequence of U is shown in SEQ ID NO:17.

[0124] In another preferred embodiment, x is 2.

[0125] In another preferred embodiment, the NLS2 is an SV40 NLS.

[0126] In another preferred embodiment, the amino acid sequence of the NLS2 is shown in SEQ ID NO:18.

[0127] In another preferred embodiment, the amino acid sequence of the linker peptide is as shown in SEQ ID NO:15 or SEQ ID NO:16.

[0128] In another preferred embodiment, the amino acid sequence of "-NLS2-" in Formula I, Formula II or Formula III is composed of SEQ ID NO:15, SEQ ID NO:18 and SEQ ID NO:16 respectively.

[0129] In another preferred embodiment, F is a Flag tag sequence (including a 3×Flag tag sequence).

[0130] In another preferred embodiment, the amino acid sequence of F is shown in SEQ ID NO:19.

[0131] In another preferred embodiment, the self-cleaving 2A peptide is selected from the group consisting of P2A, T2A, E2A, F2A, or combinations thereof; preferably, the self-cleaving 2A peptide is T2A.

[0132] In another preferred embodiment, the reporter protein is selected from the group consisting of glutathione S-transferase (GST), horseradish peroxidase (HRP), chloramphenicol acetyltransferase (CAT), β-galactosidase, β-glucuronidase, fluorescent protein, or combinations thereof.

[0133] In another preferred embodiment, the reporter protein is a fluorescent protein.

[0134] In another preferred embodiment, the fluorescent protein is selected from the group consisting of green fluorescent protein (GFP), yellow fluorescent protein (YFP), red fluorescent protein (RFP), blue fluorescent protein (BFP), or combinations thereof.

[0135] In another preferred embodiment, the reporter protein is mNeonGreen (mNG) protein.

[0136] In another preferred embodiment, the fusion protein contains one or more structures from the N-terminus to the C-terminus as shown in Formula 1 or Formula 2:

[0137] S-TALE-DddA-U(1)

[0138] U-DddA-TALE-S(2)

[0139] In the formula,

[0140] S represents either no signal peptide or a signal peptide;

[0141] TALE is the TALE protein;

[0142] DddA is a double-stranded DNA deaminase as described in the first aspect of the present invention;

[0143] U is a base excision repair inhibitor;

[0144] The hyphen "-" indicates a linking peptide or peptide bond.

[0145] In another preferred embodiment, S is a mitochondrial targeting signal peptide (MTS).

[0146] In another preferred embodiment, the TALE contains a nuclear localization signal (NLS).

[0147] In another preferred embodiment, the amino acid sequence of DddA is selected from the group consisting of: SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, or a combination thereof.

[0148] In another preferred embodiment, the U is a uracil glycosylase inhibitor (UGI).

[0149] In another preferred embodiment, the amino acid sequence of U is shown in SEQ ID NO:17.

[0150] In another preferred embodiment, the fusion protein contains the structures shown in both Formula 1 and Formula 2.

[0151] In another preferred embodiment, the amino acid sequence of the fusion protein is selected from the group consisting of:

[0152] (i) SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24 or SEQ ID NO:25;

[0153] (ii) Derived sequences having at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, or SEQ ID NO:25;

[0154] (iii) Derived sequences obtained by adding, deleting, modifying and / or substituting at least one (e.g., 1-20, 1-10, 1-5) amino acids based on SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24 or SEQ ID NO:25.

[0155] In a third aspect of the invention, a base editing system is provided, the base editing system comprising:

[0156] (i) the double-stranded DNA deaminase and programmable DNA-binding protein as described in the first aspect of the present invention; or

[0157] (ii) The fusion protein as described in the second aspect of the present invention.

[0158] In another preferred embodiment, the base editing system further comprises guide RNA (gRNA).

[0159] In another preferred embodiment, the programmable DNA-binding protein or the fusion protein forms a ribonucleoprotein complex with the guide RNA and binds to the target nucleic acid under the guidance of the guide RNA.

[0160] In another preferred embodiment, the target nucleic acid is DNA.

[0161] In another preferred embodiment, the target nucleic acid is double-stranded DNA.

[0162] In another preferred embodiment, the target nucleic acid is mitochondrial DNA.

[0163] In another preferred embodiment, the programmable DNA-binding protein is selected from the group consisting of: Cas protein, Argonaute (Ago) protein, TALE protein, zinc finger protein, or a combination thereof.

[0164] In another preferred embodiment, the programmable DNA-binding protein is a Cas protein or a TALE protein.

[0165] In another preferred embodiment, the base editing system further comprises a signal peptide.

[0166] In another preferred embodiment, the base editing system contains both the structures shown in Formula 1 and Formula 2.

[0167] In another preferred embodiment, the base editing system further comprises a nuclease (preferably an endonuclease).

[0168] In a fourth aspect of the invention, a separate polynucleotide is provided, the polynucleotide encoding a double-stranded DNA deaminase as described in the first aspect of the invention, or a fusion protein as described in the second aspect of the invention, or a base editing system as described in the third aspect of the invention.

[0169] In another preferred embodiment, the polynucleotide comprises a humanized optimized sequence.

[0170] In another preferred embodiment, the polynucleotide is selected from the group consisting of DNA sequences (including cDNA sequences), RNA sequences (including mRNA sequences), or combinations thereof.

[0171] In another preferred embodiment, the polynucleotide further comprises, flanking the ORF of the double-stranded DNA deaminase, an auxiliary element selected from the group consisting of: signal peptides, secretory peptides, tag sequences (such as 6His), or combinations thereof.

[0172] In another preferred embodiment, the polynucleotide further comprises a promoter operatively linked to the ORF sequence of the double-stranded DNA deaminase.

[0173] In another preferred embodiment, the promoter is a constitutive promoter or an inducible promoter.

[0174] In another preferred embodiment, the promoter is a strong promoter.

[0175] In another preferred embodiment, the promoter is a tissue-specific promoter.

[0176] In another preferred embodiment, the promoter is selected from the group consisting of: CMV promoter, U6 promoter, 35S promoter, T7 phage promoter, Ubiquitin promoter, Actin1 promoter, CsVMV promoter, or a combination thereof; preferably, the promoter is a CMV promoter or a U6 promoter.

[0177] In another preferred embodiment, the polynucleotide also contains the coding sequence of gRNA (i.e., the transcription template of gRNA).

[0178] In another preferred embodiment, the gRNA comprises a direct repeat (DR) sequence and a spacer sequence; the direct repeat sequence forms a gRNA scaffold.

[0179] In another preferred embodiment, the gRNA scaffold includes base modifications (such as loop modifications).

[0180] In another preferred embodiment, the polynucleotide is a polynucleotide whose codons have been optimized according to the codon preferences of the host cell.

[0181] In another preferred embodiment, the host cell includes a prokaryotic cell or a eukaryotic cell.

[0182] In another preferred embodiment, the host cell is a eukaryotic cell, such as a yeast cell, plant cell, or mammalian cell (including human and non-human mammals).

[0183] In another preferred embodiment, the host cell is a prokaryotic cell, such as Escherichia coli.

[0184] In another preferred embodiment, the yeast cells are derived from yeasts selected from the group consisting of Pichia pastoris, Kluyveromyces, or combinations thereof; more preferably, the yeast cells are derived from Kluyveromyces; even more preferably, the Kluyveromyces is Kluyveromyces marxi and / or Kluyveromyces lactis.

[0185] In another preferred embodiment, the host cell is selected from the group consisting of Escherichia coli, wheat germ cells, insect cells, SF9, HeLa, HEK293, CHO, yeast cells, or combinations thereof.

[0186] In a fifth aspect of the invention, a carrier is provided that comprises a polynucleotide as described in the fourth aspect of the invention.

[0187] In another preferred embodiment, the polynucleotide is located on one or more vectors.

[0188] In another preferred embodiment, the vector comprises one or more promoters operatively linked to the polynucleotide, enhancer, transcription termination signal, polyadenylated sequence, origin of replication, selectivity marker, nucleic acid restriction site, and / or homologous recombination site.

[0189] In another preferred embodiment, the promoter is selected from one or more of constitutive promoters, inducible promoters, ubiquitin promoters, cell type-specific promoters, and tissue-specific promoters.

[0190] In another preferred embodiment, the vector includes plasmids and viral vectors.

[0191] In another preferred embodiment, the viral vector is selected from the group consisting of adeno-associated virus (AAV), adenovirus, lentivirus, retrovirus, herpesvirus, SV40, poxvirus, or combinations thereof.

[0192] In another preferred embodiment, the vector includes a cloning vector, a transformation vector, an expression vector, a shuttle vector, an integration vector, and a multifunctional vector.

[0193] In another preferred embodiment, the carrier has a structure as shown in formula (1), formula (2), formula (3) or formula (4):

[0194] P1-NLS1-L1-DddA-L2-Cas-(L3-U) x -NLS2(1)

[0195] P1-NLS1-L1-DddA-L2-Cas-(L3-U) x -NLS2-F(2)

[0196] P1-NLS1-L1-DddA-L2-Cas-(L3-U) x -NLS2-2A-R-pA(3)

[0197] P1-NLS1-L1-DddA-L2-Cas-(L3-U) x -NLS2-2A-R-pA-P2-G(4)

[0198] In the formula,

[0199] P1 and P2 are each independently a no promoter or promoter;

[0200] NLS1 is the nuclear positioning signal;

[0201] DddA is the coding sequence of the double-stranded DNA deaminase as described in the first aspect of the present invention;

[0202] Cas is the coding sequence for the Cas protein;

[0203] U is a base excision repair inhibitor;

[0204] x is a positive integer from 1 to 5;

[0205] NLS2 is the nuclear positioning signal;

[0206] F represents the label sequence;

[0207] 2A is the coding sequence for the self-cleaved 2A peptide;

[0208] R stands for reporter gene (the coding sequence of a reporter protein);

[0209] pA is a polyA element;

[0210] G stands for the guide portion of a programmable DNA-binding protein;

[0211] L1, L2, and L3 are each independently either empty or connected.

[0212] Each "-" represents a bond or nucleotide linking sequence independently.

[0213] In another preferred embodiment, P1 is a constitutive promoter or an inducible promoter.

[0214] In another preferred embodiment, P1 is selected from the group consisting of: CMV promoter, U6 promoter, 35S promoter, T7 phage promoter, Ubiquitin promoter, Actin1 promoter, CsVMV promoter, or combinations thereof; preferably, P1 is a CMV promoter or a U6 promoter.

[0215] In another preferred embodiment, P2 is a constitutive promoter or an inducible promoter.

[0216] In another preferred embodiment, P2 is selected from the group consisting of: CMV promoter, U6 promoter, 35S promoter, T7 phage promoter, Ubiquitin promoter, Actin1 promoter, CsVMV promoter, or combinations thereof; preferably, P2 is a CMV promoter or a U6 promoter.

[0217] In another preferred embodiment, P1 is a CMV promoter and P2 is a U6 promoter.

[0218] In another preferred embodiment, G is the coding sequence of gRNA (i.e., the transcription template of gRNA).

[0219] In another preferred embodiment, the gRNA comprises a direct repeat (DR) sequence and a spacer sequence; the direct repeat sequence forms a gRNA scaffold.

[0220] In another preferred embodiment, the gRNA scaffold includes base modifications (such as loop modifications).

[0221] In a sixth aspect of the invention, a host cell is provided, the host cell comprising a vector as described in the fifth aspect of the invention, or having a genome incorporating a polynucleotide as described in the fourth aspect of the invention.

[0222] In another preferred embodiment, the host cell expresses a double-stranded DNA deaminase as described in the first aspect of the invention, a fusion protein as described in the second aspect of the invention, or a base editing system as described in the third aspect of the invention.

[0223] In another preferred embodiment, the host cell includes a prokaryotic cell or a eukaryotic cell.

[0224] In another preferred embodiment, the host cell is a eukaryotic cell, such as a yeast cell, plant cell, or mammalian cell (including human and non-human mammals).

[0225] In another preferred embodiment, the host cell is a prokaryotic cell, such as Escherichia coli.

[0226] In another preferred embodiment, the yeast cells are derived from yeasts selected from the group consisting of Pichia pastoris, Kluyveromyces, or combinations thereof; more preferably, the yeast cells are derived from Kluyveromyces; even more preferably, the Kluyveromyces is Kluyveromyces marxi and / or Kluyveromyces lactis.

[0227] In another preferred embodiment, the host cell is selected from the group consisting of Escherichia coli, wheat germ cells, insect cells, SF9, HeLa, HEK293, CHO, yeast cells, or combinations thereof.

[0228] In another preferred embodiment, the host cell is HEK293T cell, U2OS cell, NIH3T3 cell, N2A cell, or a combination thereof.

[0229] In a seventh aspect of the invention, there is provided the use of a double-stranded DNA deaminase as described in the first aspect of the invention, a fusion protein as described in the second aspect of the invention, a base editing system as described in the third aspect of the invention, a polynucleotide as described in the fourth aspect of the invention, a vector as described in the fifth aspect of the invention, or a host cell as described in the sixth aspect of the invention, for:

[0230] (1) Preparation of a base editor;

[0231] (2) Editing DNA fragments;

[0232] (3) Preparation of gene editing reagents; and / or

[0233] (4) Preparation of gene editing drugs.

[0234] In another preferred embodiment, the DNA fragment is a double-stranded DNA fragment.

[0235] In another preferred embodiment, the DNA fragment is a mitochondrial DNA fragment (double-stranded DNA fragment).

[0236] In another preferred embodiment, a C-to-T conversion occurs on the "AC" DNA segment of the double-stranded DNA.

[0237] In another preferred embodiment, a C-to-T conversion occurs on the "CC" DNA fragment of the double-stranded DNA.

[0238] In another preferred embodiment, a C-to-T conversion occurs on the "GC" DNA segment of the double-stranded DNA.

[0239] In another preferred embodiment, a C-to-T conversion occurs on the “TC” DNA fragment of the double-stranded DNA.

[0240] In another preferred embodiment, the DNA fragment comprises an "AC" DNA fragment.

[0241] In another preferred embodiment, the DNA fragment is an "AC" DNA fragment.

[0242] In another preferred embodiment, the DNA fragment comprises a “CC” DNA fragment.

[0243] In another preferred embodiment, the DNA fragment is a “CC” DNA fragment.

[0244] In another preferred embodiment, the DNA fragment comprises a “GC” DNA fragment.

[0245] In another preferred embodiment, the DNA fragment is a “GC” DNA fragment.

[0246] In another preferred embodiment, the DNA fragment comprises a “TC” DNA fragment.

[0247] In another preferred embodiment, the DNA fragment is a “TC” DNA fragment.

[0248] In another preferred embodiment, the gene-editing drug is a drug for treating hereditary diseases.

[0249] In another preferred embodiment, the hereditary disease includes: single-gene hereditary disease, polygenic hereditary disease, and chromosomal abnormality disease.

[0250] In another preferred embodiment, the hereditary diseases include, but are not limited to: sickle cell anemia, cystic fibrosis, Duchenne muscular dystrophy, hemophilia, albinism, red-green color blindness, achondroplasia, congenital deafness, hereditary nephritis, asthma, schizophrenia, hypertension, epilepsy, congenital heart disease, cleft lip and palate, Alzheimer's disease, Parkinson's disease, or combinations thereof.

[0251] In another preferred embodiment, the gene-editing drug is a drug for treating tumors or cancer.

[0252] In an eighth aspect of the invention, a base editing method is provided, comprising the steps of: expressing a fusion protein as described in the second aspect of the invention, or a base editing system as described in the third aspect of the invention, in a target cell, thereby causing base editing of the DNA of the target cell.

[0253] In another preferred embodiment, the method includes the steps of introducing a polynucleotide as described in the fourth aspect of the invention, a vector as described in the fifth aspect of the invention, or a host cell as described in the sixth aspect of the invention into a target cell.

[0254] In another preferred embodiment, the base editing method is an in vitro method.

[0255] In another preferred embodiment, the base editing method is for non-disease treatment purposes and non-disease diagnosis purposes.

[0256] In another preferred embodiment, the target cell's DNA is double-stranded DNA.

[0257] In another preferred embodiment, the DNA of the target cell is mitochondrial DNA.

[0258] In another preferred embodiment, the DNA contains an “AC” DNA fragment.

[0259] In another preferred embodiment, the DNA is an “AC” DNA fragment.

[0260] In another preferred embodiment, the DNA comprises a “CC” DNA fragment.

[0261] In another preferred embodiment, the DNA is a “CC” DNA fragment.

[0262] In another preferred embodiment, the DNA comprises a “GC” DNA fragment.

[0263] In another preferred embodiment, the DNA is a “GC” DNA fragment.

[0264] In another preferred embodiment, the DNA comprises a “TC” DNA fragment.

[0265] In another preferred embodiment, the DNA is a “TC” DNA fragment.

[0266] In another preferred embodiment, a C-to-T conversion occurs on the “AC” DNA segment of the DNA.

[0267] In another preferred embodiment, a C-to-T conversion occurs on the “CC” DNA segment of the DNA.

[0268] In another preferred embodiment, a C-to-T conversion occurs on the “GC” DNA fragment of the DNA.

[0269] In another preferred embodiment, a C-to-T conversion occurs on the “TC” DNA segment of the DNA.

[0270] It should be understood that, within the scope of this invention, the above-described technical features of this invention and the technical features specifically described below (such as in the embodiments) can be combined with each other to form new or preferred technical solutions. Due to space limitations, they will not be described in detail here. Attached Figure Description

[0271] Figure 1 A shows a comparative illustration of the protein sequences of Bruce GSVG (also known as Bruce DddA GSVG), DddA1SVG, and DddA8SVG; Figure 1 B shows the integration of toxicity-eliminating mutants BruceDead, DddA1Dead, and DddA8Dead with the CRISPR system; Figure 1 C shows the integration of BruceGSVG, DddA1SVG, and DddA8SVG with the CRISPR system after engineering modifications; Figure 1 D shows the integration of different engineered deaminases (i.e., BruceDead, DddA1Dead, DddA8Dead, BruceGSVG, DddA1SVG, and DddA8SVG) with the TALE system.

[0272] Figure 2 The efficiency of gene editors designed based on different deaminases in base editing of human ROR1 and EMX1 genes was demonstrated. Figure 2 A represents the base editing efficiency of BruceDead, DddA1Dead, and DddA8Dead at site 3 of the EMX1 gene; Figure 2 B represents the base editing efficiency of DddA1SVG and DddA8SVG at site 3 of the EMX1 gene; Figure 2 C represents the base editing efficiency of BruceDead, DddA1Dead, and DddA8Dead at site 1 of the ROR1 gene; Figure 2 D represents the base editing efficiency of BruceGSVG, DddA1SVG, and DddA8SVG at site 1 of the ROR1 gene; Figure 2 E represents the base editing efficiency of BruceDead, DddA1Dead, and DddA8Dead at site 2 of the ROR1 gene; Figure 2 F represents the base editing efficiency of BruceGSVG, DddA1SVG, and DddA8SVG at site 2 of the ROR1 gene; Figure 2 G shows the percentage distribution of base sequence motifs edited by DddA1Dead, DddA8Dead, DddA1SVG, and DddA8SVG.

[0273] Figure 3The efficiency of gene editors designed based on different deaminases in base editing of human mitochondrial ND5 and ATP6 genes was demonstrated. Detailed Implementation

[0274] Through extensive and in-depth research and numerous screenings, the inventors have discovered for the first time a class of double-stranded DNA deaminases. On the one hand, these deaminases exhibit significantly reduced cytotoxicity and can be expressed in large quantities in prokaryotic and eukaryotic cells in vitro. On the other hand, they demonstrate higher single-base editing activity on non-TC DNA sequences, thereby expanding the range of base editing in eukaryotic gene sequences (especially mitochondrial genes) and ensuring that the base area is not limited to TC DNA sequences. This invention is based on these findings.

[0275] the term

[0276] To facilitate a clearer understanding of this disclosure, certain terms are first defined. As used herein, unless otherwise expressly specified herein, each of the following terms shall have the meaning given below.

[0277] The term “about” can refer to a value or composition within an acceptable range of error for a particular value or composition as determined by a person skilled in the art, which will depend in part on how the value or composition is measured or determined.

[0278] DNA deaminase

[0279] DNA deaminases are a class of enzymes that catalyze the deamination of nucleotides in DNA molecules, playing important roles in biochemistry and molecular biology. Based on their substrates and catalytic mechanisms, DNA deaminases can be classified into several types, such as activation-induced cytidine deaminase (AID) and apolipoprotein messenger RNA editing enzyme catalyzed by polypeptides (APOBEC).

[0280] DNA deaminases have potential applications in gene editing. By modifying and optimizing the catalytic activity and specificity of DNA deaminases, more precise and efficient gene editing tools can be developed for applications such as disease treatment and crop improvement.

[0281] The roles of DNA deaminases in base editing mainly include the following two aspects:

[0282] 1. Catalytic deamination reaction: DNA deaminases can catalyze the deamination reaction of specific nucleotides in DNA molecules, converting cytosine (C) into uracil (U) or adenine (A) into inosine (I).

[0283] For cytosine base editors (CBE): cytosine (C) in a certain range on DNA is deaminated to uracil (U), and then U is converted to thymine (T) through DNA replication or repair, ultimately achieving direct substitution of C·G base pairs to T·A base pairs;

[0284] For adenine base editors (ABE): After adenine (A) in a certain range on DNA is deaminated to inosine (I), I is read and copied as G at the DNA level, ultimately achieving direct substitution of A·T base pairs to G·C base pairs.

[0285] 2. Combination with programmable DNA-binding proteins: In base editors (Bes), DNA deaminases typically bind to programmable DNA-binding proteins (such as the Cas9 protein in the CRISPR-Cas system) to form fusion proteins. These fusion proteins, guided by sgRNA (single-stranded guide RNA), can precisely target specific locations in the genome and catalyze deamination reactions.

[0286] For cytosine base editors (CBE): The core components of CBE are nCas9 or dCas9 (Cas9 that has lost its cleavage activity) and cytosine deaminase. When the fusion protein targets genomic DNA under the guidance of sgRNA, cytosine deaminase can bind to the single-stranded DNA (ssDNA) of the R-loop region formed by the Cas9 protein, sgRNA, and genomic DNA, thereby causing subsequent base substitutions;

[0287] For the adenine base editor (ABE): The core components of the ABE are nCas9 (D10A) and an artificially induced adenine deaminase. The mechanism of action of the ABE is similar to that of the CBE, that is, when the fusion protein targets genomic DNA under the guidance of sgRNA, the adenine deaminase can bind to the ssDNA, thereby causing subsequent base substitution.

[0288] In addition to traditional CBE and ABE, researchers have developed a variety of novel base editors. For example, the guanine base editor (gGBE) based on engineered glycosylation enzymes enables direct editing of guanine (G), filling the gap in traditional base editors that cannot directly edit G. Another example is the adenine transversion editor "ACBEs," which achieves precise conversions from A to C and A to T, providing new tools for treating more types of genetic diseases.

[0289] Furthermore, various DNA repair mechanisms exist within cells that can interfere with the base editing process. For example, uracil DNA glycosyltransferase (UDG) can recognize U·G mismatches and cleave the glycosidic bond between uracil and the phosphate backbone, thereby reversing the base editing process. To address this challenge, researchers have developed a base editor version that incorporates a uracil DNA glycosyltransferase inhibitor (UGI) to improve editing efficiency.

[0290] Uracil glycosylation enzyme inhibitor (UGI)

[0291] Uracil glycosylation enzyme inhibitor (UGI), derived from Bacillus subtilis bacteriophage, has a small protein molecular weight (9.5 kDa) and inhibits uracil-DNA glycosylation enzyme (UDG) in Escherichia coli and UDG from other species. UDG is an enzyme responsible for recognizing and cleaving uracil bases in DNA, and UGI can dissociate the UDG-DNA complex. When UDG and UGI reversibly bind in a 1:1 ratio, UGI inhibits UDG, preventing UDG from repairing DNA and thus protecting DNA stability under certain experimental conditions (such as PCR amplification and cloning). Furthermore, UGI can be used to study the enzymatic properties of UDG and its interaction with DNA.

[0292] TALE protein

[0293] TALE (Transcription Activator-Like Effectors) proteins are a class of DNA-binding proteins injected into host cells by the plant pathogen Xanthomonas through a type III secretion system. They possess the ability to specifically recognize DNA sequences, a characteristic that makes TALE proteins promising for broad applications in genome editing and gene regulation.

[0294] The core domain of a TALE protein, its DNA-binding domain, is composed of varying numbers of repeating units, each typically containing 33–35 amino acids. The amino acids at positions 12 and 13 (called RVDs) are highly variable and responsible for specifically recognizing a single DNA base pair. By combining different RVDs, TALE proteins capable of recognizing almost any DNA sequence can be designed. For example, the NI combination recognizes adenine (A), the HD combination recognizes cytosine (C), the NG combination recognizes thymine (T), and the NN combination can recognize either guanine (G) or adenine (A, but with lower preference).

[0295] Due to the specific DNA sequence recognition and flexible assembly capabilities of TALE proteins, scientists can design and assemble arbitrary TALE units to recognize target DNA double helix sequences. This property has been widely applied in gene editing technologies, such as the development of TALEN (TALE Nuclease) technology, for introducing site-directed mutations and site-directed knockouts in the cell genome.

[0296] The binding of the TALE protein to a deaminase forms TALED technology, a novel base editor. TALED technology utilizes the specific DNA-binding ability of the TALE protein to precisely guide the deaminase to the target DNA sequence, thereby achieving base editing at specific sites. In the TALED system, the TALE protein is responsible for recognizing and binding to the target DNA sequence, while the deaminase catalyzes the deamination reaction of the target bases. Depending on the type and properties of the deaminase, TALED can achieve different types of base transitions, such as A·T to G·C, or C·G to T·A.

[0297] The N-terminal domain of TALE proteins typically contains a nuclear localization signal (NLS) linker, which guides the TALE protein into the cell nucleus, a prerequisite for its DNA-binding function. The C-terminal domain of TALE proteins usually contains functional domains, such as a nuclease catalytic domain (in TALEN technology) or other regulatory factors (such as transcription factors). These domains endow TALE proteins with additional functions, such as DNA cleavage or regulation of gene expression. For example, in TALEN technology, the C-terminal domain is often fused with the catalytic domain of the FokI endonuclease, enabling TALEN to cleave target DNA sequences, thereby achieving gene editing. When TALE proteins bind to FokI endonucleases, they typically function as dimers, with each monomer targeting one strand of the DNA double helix and forming a dimer at the target site to cleave the DNA.

[0298] TALE proteins, with their specific DNA sequence recognition capabilities and flexible assembly properties, have shown great potential in gene editing and gene regulation. Particularly in base editing, the combination of TALE proteins with deaminases forms TALED technology, providing new tools and methods for achieving precise base switching. With further research and technological advancements, TALED technology is expected to play an even more important role in areas such as the treatment of genetic diseases and crop breeding.

[0299] Mitochondrial Targeting Signaling Peptide (MTS)

[0300] Mitochondrial targeting signal peptides (MTS) are signal peptides or sequences that guide proteins into the mitochondria. These signal peptides are generally located at the N-terminus of proteins and are cleaved after localization; therefore, they are also called presequences. Their function is to ensure that proteins are correctly transported into the mitochondria.

[0301] Zinc finger protein

[0302] Zinc finger proteins are composed of a ring containing approximately 30 amino acids and a Zn²⁺ molecule that coordinates to either four Cys or two Cys and two His ions on the ring, forming a finger-like structure. This structure allows zinc finger proteins to stably bind to DNA or RNA. Originally used to describe the finger-like appearance of the Xenopus oocyte transcription factor IIIA hypothesis, the name zinc finger now encompasses a variety of different protein structures.

[0303] Zinc finger proteins are mostly involved in the regulation of gene expression and are a type of functional protein. They play a crucial role in gene expression regulation, influencing transcription and translation by binding to DNA, RNA, or proteins. Zinc finger proteins are typically composed of multiple zinc finger units tandemly linked together, each unit recognizing and binding to three consecutive bases on DNA. By combining different zinc finger units, zinc finger proteins capable of recognizing specific DNA sequences can be constructed.

[0304] Zinc finger nucleases (ZFNs) are an important application of zinc finger proteins. They combine the DNA recognition ability of zinc finger proteins with the cleavage function of endonucleases, enabling site-specific cleavage at specific DNA sites to achieve gene editing. Zinc finger nucleases (ZFNs) consist of two parts: a DNA recognition domain (composed of multiple zinc finger proteins (ZFPs) tandemly linked to recognize and bind to specific DNA sequences) and a non-specific endonuclease (such as FokI, which cleaves the DNA double strand at a specific location through dimerization, causing a DNA double-strand break and triggering the cell's DNA repair mechanism to achieve gene editing).

[0305] Zinc finger deaminases (ZFDs) are a novel base editing tool composed of zinc finger DNA-binding proteins, interaminases from cleaving bacteria (such as DddAtox), and uracil glycosylation inhibitors (UGIs). ZFDs recognize and bind to specific DNA sequences via zinc finger proteins, and the deaminases then catalyze C-to-T conversions within those sequences without cleaving the DNA double strand. ZFDs have shown great potential in genome editing in eukaryotic cells and other organisms. They have been used to achieve efficient C-to-T base editing in both nuclear and mitochondrial DNA with minimal induced insertions and deletions. ZFD technology offers new strategies for the treatment of genetic diseases and the optimization of crop traits.

[0306] Main advantages of the invention

[0307] 1. The novel base editor of this invention can be used to effectively edit double-stranded DNA, overcoming the limitation of existing gene editing methods that can often only edit single-stranded DNA.

[0308] 2. The novel base editor of this invention can solve the problem of difficult editing of mitochondrial DNA, effectively improve the editing efficiency of mitochondrial DNA, and expand the types of DNA fragments that can be effectively edited (no longer limited to "TC" DNA fragments), providing a broader prospect for the clinical application of mitochondrial DNA editing.

[0309] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. Experimental methods in the following embodiments, unless otherwise specified, are generally performed under conventional conditions, such as those described in Sambrook et al., Molecular Cloning: A Laboratory Manual (New York: Cold Spring Harbor Laboratory Press, 1989), or as recommended by the manufacturer. Unless otherwise stated, percentages and parts are weight percentages and parts by weight.

[0310] Sequence information

[0311] BruceDddA wild-type sequence (SEQ ID NO:1)

[0312] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC

[0313] DddA1 wild-type sequence (SEQ ID NO: 2)

[0314] MSLPEYDGTTTHGVLVLDDGTQIGFTSGNGDPRYTNYRNNGHVEQKSALYMRENNISNATVYHNNTNGTCGYCNTMTATFLPEGATLTVVPPENAVANNSRAIDYVKTYTGTSNDPKISPRYKGN

[0315] DddA8 wild-type sequence (SEQ ID NO: 3)

[0316] YFVTTTDVWVHNTSPSSSAAPKLPPYDGKTTRGILRLSEGGDDIPLSSGKKVLPNYEGSGHVEGKAALEIRARGSSGGTVWHNNTNGTCGYCNSHTATLLPEGAKLDVVPPSNAVANNSRAVAAPKQYTGNSRPIKSPPK

[0317] BruceDead sequence (SEQ ID NO: 4)

[0318] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHV A GQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC

[0319] DddA1Dead sequence (SEQ ID NO: 5)

[0320] MSLPEYDGTTTHGVLVLDDGTQIGFTSGNGDPRYTNYRNNGHV A QKSALYMRENNISNATVYHNNTNGTCGYCNTMTATFLPEGATLTVVPPENAVANNSRAIDYVKTYTGTSNDPKISPRYKGN

[0321] DddA8Dead sequence (SEQ ID NO: 6)

[0322] YFVTTTDVWVHNTSPSSSAAPKLPPYDGKTTRGILRLSEGGDDIPLSSGKKVLPNYEGSGHV AGKAALEIRARGSSGGTVWHNNTNGTCGYCNSHTATLLPEGAKLDVVPPSNAVANNSRAVAAPKQYTGNSRPIKSPPK

[0323] Bruce GSVG sequence (SEQ ID NO:7)

[0324] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLE G KVFSSGGPTPYPNYANAGHVE S QSALFMRDNGISE S LVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEG V IPVKRGATGETKVFTGNSN G PKSPTKGGC

[0325] DddA1 SVG sequence (SEQ ID NO:8)

[0326] MSLPEYDGTTTHGVLVLDDGTQIGFTSGNGDPRYTNYRNNGHVE S KSALYMRENNISNATVYHNNTNGTCGYCNTMTATFLPEGATLTVVPP V VANNSRAIDYVKTYTGTSN G PKISPRYKGN

[0327] DddA8 SVG sequence (SEQ ID NO:9)

[0328] YFVTTTDVWVHNTSPSSSAAPKLPPYDGKTTRGILRLSEGGDDIPLSSGKKVLPNYEGSGHVEsKAALEIRARGSSGGTVWHNNTNGTCGYCNSHTATLLPEGAKLDVVPPSNvVANNSRAVAAPKQYTGNSRgIKSPPK

[0329] bpNLS sequence (SEQ ID NO:10)

[0330] MKRTADGSEFESPKKKRKV

[0331] bpNLS sequence (SEQ ID NO:11)

[0332] MPKKKRKV

[0333] XTEN Linker Sequence (SEQ ID NO:12)

[0334] SGGSSGGSSGSETPGTSESATPESSGGSSGGS

[0335] SpCas9-SpRY sequence (SEQ ID NO:13)

[0336] DKKYSIGL A

[0337] GS Linker Sequence (SEQ ID NO:14)

[0338] SGGSGGSGGS

[0339] GS Linker Sequence (SEQ ID NO:15)

[0340] SGGS

[0341] GS Linker Sequence (SEQ ID NO:16)

[0342] SSGG

[0343] UGI sequence (SEQ ID NO:17)

[0344] TNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTD ENVMLLTSDAPEYKPWALVIQDSNGENKIKML

[0345] SV40 NLS sequence (SEQ ID NO:18)

[0346] KRTADGSEFEPKKKRKV

[0347] 3×Flag tag sequence (SEQ ID NO:19)

[0348] DYKDHDGDYKDHDIDYKDDDDK

[0349] DddA1-Dead(E44A)-SpRY(D10A)pXZ504 sequence (SEQ ID NO:20)

[0350]

[0351] DddA8-Dead(E63A)-SpRY(D10A)pXZ505Zone(SEQ ID NO:21)

[0352]

[0353] BruceDddA-Dead(E1347A)-SpRY(D10A) pXZ503 sequence (SEQ ID NO:22)

[0354]

[0355] The DddA1(SVGdel2)-SpRY(D10A)pCX2566AA sequence (SEQ ID NO:23)

[0356]

[0357] The DddA8(SVG)-SpRY(D10A)pCX2567 sequence (SEQ ID NO:24)

[0358]

[0359] BruceDddA(GSVG)-SpRY(D10A)pCX2278 sequence (SEQ ID NO:25)

[0360]

[0361] Example 1: Engineering modification of BruceDddA, DddA1, and DddA8

[0362] BruceDddA protein (SEQ ID NO:1) is a double-stranded DNA deaminase DddA from Burkholderiacepacia. This enzyme recognizes and deaminates cytosine in double-stranded DNA. However, this enzyme exhibits strong cytotoxicity in prokaryotic and eukaryotic cells, making it virtually impossible to express in large quantities in vitro.

[0363] Simiaoa Sunni also has wild-type proteins similar to BruceDddA that can recognize and deaminate cytosine in double-stranded DNA. These are named DddA1 (SEQ ID NO:2) and DddA8 (SEQ ID NO:3), respectively. These wild-type proteins also exhibit strong cytotoxicity in prokaryotic and eukaryotic cells.

[0364] After extensive screening and verification, this invention obtained BruceDead (SEQ ID NO:4), DddA1Dead (SEQ ID NO:5), and DddA8Dead (SEQ ID NO:6) proteins by mutation based on wild-type BruceDddA, DddA1, and DddA8 proteins. The cytotoxicity of BruceDead, DddA1Dead, and DddA8Dead proteins was significantly reduced after mutation, enabling them to be expressed in large quantities in prokaryotic and eukaryotic cells in vitro.

[0365] To further improve DNA editing efficiency, this invention further modified the wild-type BruceDddA, DddA1, and DddA8 proteins, obtaining BruceGSVG (SEQ ID NO:7), DddA1SVG (SEQ ID NO:8), and DddA8SVG (SEQ ID NO:9) proteins, respectively. Figure 1 Figure A shows a comparative illustration of the Bruce GSVG (also known as Bruce DddA GSVG), DddA1SVG, and DddA8SVG protein sequences.

[0366] The mutation status of the above proteins is summarized in Table 1 below:

[0367] Table 1

[0368]

[0369]

[0370] Example 2: Verification of the base editing effect of different modified deaminases on non-TC DNA sequences in eukaryotic cells. Rate

[0371] This embodiment designed a CRISPR-based eukaryotic expression vector and obtained the following by fusing SpCas9-RY, which only possesses nickase activity, with different deaminases (i.e., BruceDead, DddA1Dead, DddA8Dead, BruceGSVG, DddA1SVG, and DddA8SVG). Figure 1 B and 1C show gene cytosine (CBE) single-base editing tools. Among them, tools based on modified BruceDead, DddA1Dead, and DddA8Dead proteins were designed as follows: Figure 1 B shows its base editing system / base editor integrated with the CRISPR system; based on the modified BruceGSVG, DddA1SVG, and DddA8SVG proteins, it designed such as Figure 1 C shows its base editing system / base editor integrated with the CRISPR system.

[0372] Different gene editor expression vectors were constructed using sgRNAs targeting human ROR1 and EMX1 genes, transfected into HEK293T cells, and successfully transfected cells were sorted by flow cytometry for sgRNA target site sequence amplification and next-generation sequencing. Base editing efficiency was statistically analyzed as follows: Figure 2 As shown, the results indicate that BruceGSVG, DddA1SVG, and DddA8SVG have higher single-base editing activity in non-"TC" DNA sequences compared to BruceDead, DddA1Dead, and DddA8Dead.

[0373] Example 3: Verification of the base sequence of non-"TC" DNA sequences in eukaryotic mitochondrial genes of different modified deaminases. Editing efficiency

[0374] This embodiment designed a eukaryotic expression vector based on the TALE system. Furthermore, by fusing the TALE component with different deaminases (i.e., BruceDead, DddA1Dead, DddA8Dead, BruceGSVG, DddA1SVG, and DddA8SVG), the following expression was obtained: Figure 1 D shows the gene cytosine (CBE) single-base editing tool (a base editing system / base editor integrated with the TALE system).

[0375] Different gene editor expression vectors were constructed using TALE sequences targeting human ND5 and ATP6 genes, transfected into HEK293T cells, and successfully transfected cells were sorted by flow cytometry for sgRNA target site sequence amplification and next-generation sequencing. Base editing efficiency was statistically analyzed as follows: Figure 3As shown, the results indicate that BruceGSVG, DddA1SVG, and DddA8SVG have higher single-base editing activity in non-"TC" DNA sequences compared to BruceDead, DddA1Dead, and DddA8Dead.

[0376] All documents mentioned in this invention are incorporated herein by reference as if each document were individually incorporated by reference. Furthermore, it should be understood that after reading the foregoing teachings of this invention, those skilled in the art can make various alterations or modifications to this invention, and these equivalent forms also fall within the scope defined by the appended claims.

Claims

1. A double-stranded DNA deaminase, characterized in that, The double-stranded DNA deaminase has an amino acid sequence corresponding to the wild-type double-stranded DNA deaminase containing one or more mutant amino acids selected from the group consisting of alanine (A), serine (S), valine (V), glycine (G), or combinations thereof; the amino acid sequence of the wild-type double-stranded DNA deaminase is SEQ ID NO:1, SEQ ID NO:2, or SEQ ID NO:

3.

2. The double-stranded DNA deaminase as described in claim 1, characterized in that, The amino acid sequence of the double-stranded DNA deaminase is selected from the following group: (a) SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8 or SEQ IDNO:9; (b) Derived sequences having at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, or SEQ ID NO:9; (c) Derived sequences obtained by adding, deleting, modifying and / or substituting at least one (e.g., 1-20, 1-10, 1-5) amino acids based on SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8 or SEQ ID NO:

9.

3. The double-stranded DNA deaminase as described in claim 1, characterized in that, The double-stranded DNA deaminase has one or more of the following characteristics: (i) It exhibits lower cytotoxicity compared to wild-type double-stranded DNA deaminase; (ii) Compared with wild-type double-stranded DNA deaminase, it can edit "AC" DNA fragments more effectively; (iii) Compared with wild-type double-stranded DNA deaminase, it can edit the "CC" DNA fragment more effectively; (iv) Compared with wild-type double-stranded DNA deaminase, it can edit "GC" DNA fragments more effectively; (v) Compared with wild-type double-stranded DNA deaminase, it can edit the "TC" DNA fragment more effectively.

4. A fusion protein, characterized in that, The fusion protein contains: (a) the double-stranded DNA deaminase as described in any one of claims 1-3; and (b) One or more elements or functional domains selected from the group consisting of: signal peptides, localization signals, reporter proteins, programmable DNA-binding proteins, DNA-binding domains, epitope tags, transcription activation domains, transcription repression domains, nucleases, methyltransferases, transcription release factors, deacetylases, cleavage active peptides, linkers, ligases, integrases, transposases, recombinases, polymerases, base excision repair inhibitors, or combinations thereof.

5. A base editing system, characterized in that, The base editing system includes: (i) the double-stranded DNA deaminase and programmable DNA-binding protein as described in any one of claims 1-3; or (ii) The fusion protein as described in claim 4.

6. An isolated polynucleotide, characterized in that, The polynucleotide encodes a double-stranded DNA deaminase as described in any one of claims 1-3, or a fusion protein as described in claim 4, or a base editing system as described in claim 5.

7. A carrier, characterized in that, The vector comprises the polynucleotide as described in claim 6.

8. A host cell, characterized in that, The host cell contains the vector as described in claim 7, or its genome is integrated with the polynucleotide as described in claim 6.

9. Use of a double-stranded DNA deaminase as described in any one of claims 1-3, a fusion protein as described in claim 4, a base editing system as described in claim 5, a polynucleotide as described in claim 6, a vector as described in claim 7, or a host cell as described in claim 8, characterized in that, Used for: (1) Preparation of a base editor; (2) Editing DNA fragments; (3) Preparation of gene editing reagents; and / or (4) Preparation of gene editing drugs.

10. A base editing method, characterized in that, The method includes the step of expressing the fusion protein as described in claim 4 or the base editing system as described in claim 5 in target cells, thereby causing base editing of the DNA of the target cells.