Cytidine deaminases, base editing systems comprising the same, and applications thereof

By using riDddAtox cytosine deaminase derived from gut Rosbairi bacteria in combination with CRISPR/Cas9 and TALEN technologies, a cytosine deaminase base editor for nuclear and mitochondrial genomes was generated. This solved the problems of limited editing windows and sequence bias in existing technologies, achieving efficient nuclear and mitochondrial genome editing, which is suitable for gene therapy.

CN116103271BActive Publication Date: 2026-04-17SHANGHAI TECH UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANGHAI TECH UNIV
Filing Date
2022-11-16
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing technologies lack cytosine base editors with low 5'-TC sequence bias and high efficiency, making it difficult to efficiently edit nuclear and mitochondrial genomes. Furthermore, CRISPR/Cas9 technology faces challenges in delivery within mitochondria.

Method used

Using riDddAtox cytosine deaminase derived from Rosbairi bacteria in the gut, and in conjunction with CRISPR/Cas9 and TALEN technologies, a base editor for nuclear and mitochondrial cytosine deaminase was generated. This expanded the editing window and overcame sequence bias, enabling efficient editing of 5'-TC, 5'-AC, 5'-CC, and 5'-GC.

Benefits of technology

It achieves efficient C-to-T editing of nuclear and mitochondrial genomes, expands the editing window, overcomes sequence bias problems, is suitable for gene editing and gene mutation correction, and has good prospects for gene therapy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0003954384240000151
    Figure BDA0003954384240000151
  • Figure BDA0003954384240000161
    Figure BDA0003954384240000161
  • Figure BDA0003954384240000171
    Figure BDA0003954384240000171
Patent Text Reader

Abstract

This invention provides a cytosine deaminase, a base editing system comprising the same, and their applications. The cytosine deaminase comprises an amino acid sequence having at least 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the amino acid sequence shown in SEQ ID NO:1. The cytosine deaminase and two derived novel cytosine single-base editing tools of this invention are suitable for nuclear genome editing and mitochondrial plastid genome editing, solving problems such as narrow editing windows, transcriptome off-target effects, low editing efficiency, and 5'-TC sequence bias, thus expanding the selection and application range of gene editing and base editing tools. The mitochondrial cytosine base editor provided by this invention features small size, no sequence bias, and no restrictions, making it more suitable for viral vector-based gene therapy, and possessing good prospects for gene therapy and industrialization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of gene editing technology, specifically relating to a cytosine deaminase, a base editing system containing the cytosine deaminase, and their applications. Background Technology

[0002] There are currently three main types of gene editing technologies that rely on programmable nucleases: zinc finger nucleases (ZFNs). [1] Transcription activator-like effector nucleases (TALENs) [2] Clustered regularly interspaced short palindromic repeats (CRISPR) / CRISPR-associated protein 9, CRISPR / Cas9 technology [3] These three types of technologies all achieve gene editing by targeting specific sites in the genome. Nucleases cleave the target DNA, causing double-stranded DNA breaks (DSBs). This induces cells to repair the DNA through non-homologous end-joining (NHEJ), ultimately leading to random insertions or deletions (indels) and frameshift mutations. Alternatively, in the presence of a homologous DNA template, homology-directed repair (HDR) replaces the DNA surrounding the cleavage site, ultimately achieving the goal of gene editing. [4,5] .

[0003] Base editors (BEs) developed based on CRISPR / Cas9 technology can achieve precise point mutations without causing DSBs or introducing a precursor template DNA. Currently, two types of base editors have been developed: cytosine base editors (CBEs). [5] It can perform C-to-G base pair conversion; adenine base editors (ABEs) [6]It enables the conversion of AT base pairs to GC base pairs (A-to-G). The earliest generation of CBE fused rat cytosine deaminase and catalytically inactivated dCas9 (ratAPOBEC1-XTEN-dCas9), catalyzing cytosine C within the editing window to uracil (U). In subsequent DNA repair, U is used as thymine T to pair with adenine A, ultimately achieving the conversion of CG base pairs to TA base pairs. To further improve the editing efficiency of CBE, a uracil DNA glycosyltransferase inhibitor (UGI) was fused to the C-terminus of CBE, and dCas9 was replaced with nickase Cas9 (nCas9 / D10A), resulting in the BE3 editor (rat APOBEC1-XTEN-Cas9(D10A)-UGI). [5] Since Cas9 originates from Streptococcus pyogenes (SpCas9), it is subject to PAM restriction of the NGG sequence and the editing activity window at position 4-8 of the 5' end of the sgRNA target site.

[0004] Mitochondria are double-membrane-bound organelles with a single, multi-copy small circular double-stranded DNA (mtDNA) encoding 13 proteins essential for oxidative phosphorylation (OXPHOS), 22 transfer RNAs (tRNAs), and 2 ribosomal RNAs (rRNAs). [7] mtDNA is susceptible to increased mutation frequency due to exogenous stimuli (such as ultraviolet radiation and radioactive substances) and endogenous factors (such as aerobic respiration byproducts, reactive oxygen species, and reactive nitrogen species). mtDNA defends against damage through repair mechanisms similar to nDNA, such as base excision repair (BER), direct reversal (DR), mismatch repair (MMR), and double strand break repair (DSBR). [8,9] To date, there are two main ways to manipulate mtDNA: one is through the natural input mechanism (TOM-TIM complex).

[10] Restriction endonucleases with mitochondrial targeting sequences (MTS), such as PstI, will be used.

[11] ApaLI

[12] SmaI

[13] 、XmaI

[14] or programmable nuclease ZFN

[15] TALEN[2] Transport into the mitochondria causes DSB (Discharge-Solved DNA), and after DSB occurs, mtDNA tends to degrade rapidly rather than undergo DSBR (Discharge-Solved DNA Bleeding).

[16] Another approach is CRISPR / Cas9 technology, but due to the mitochondrial double membrane barrier, the exact process by which endogenous and exogenous RNA is transported to the mitochondria remains unclear.

[17] Therefore, manipulating mtDNA using CRISPR / Cas9 technology faces the challenge of delivering sgRNA.

[0005] In view of the problems faced by nuclear genome editing and mitochondrial genome editing, David Liu's research team developed DdCBE.

[18] The toxin-like protein (DddA) with double-stranded DNA deaminase activity derived from Burkholderia cenocepacia was used. tox It splits into two non-toxic, inactive parts at G1397, named DddA. tox -N(1264-1397), DddA tox-C(1398-1427) were loaded into dSpCas9 and nSaCas9 editors, respectively. SpCas9 recognizes NGG PAM, while SaKKH-nCas9 (SaCas9 mutant) recognizes NNNRRT PAM. Using SpCas9sgRNA, SaCas9sgRNA guides the two Cas9 proteins to adjacent target sites. The adjacent DddA-N and DddA-C proteins restore cytosine deaminase activity and perform C-to-T editing of nDNA, with an editing window of 17-60 bp, capable of simultaneously editing both template and non-template strands. When DddA-N and DddA-C replace Fok1 and are incorporated into the left side TALE (TALE-L) and right side TALE (TALE-R) plasmids containing MTS to recognize the left and right arms of the target site, respectively, efficient C-to-T editing of mtDNA target sites can also be achieved after transfection into mammalian cells. However, DddA, whether used for nDNA or mtDNA editing, has high requirements for the DNA sequence. The target C to be edited needs to be located in the 5'-TC sequence (the base adjacent to the 5' end of the target cytosine is T). Furthermore, issues such as deaminase cutoff points and deaminase directionality affecting the editing window and sequence bias significantly limit the use of DddA. In addition, phage-assisted continuous evolution (PACE) and phage-assisted non-continuous evolution (PANCE) have been used to introduce mutations into DddA to improve the editing efficiency of DdCBE and reduce 5'-TC sequence bias. This has enabled relatively inefficient editing of 5'-AC and 5'-CC, but effective editing of 5'-GC remains difficult.

[19] . Summary of the Invention

[0006] The technical problem this invention aims to solve is the lack of cytosine base editors with low 5'-TC sequence bias and high efficiency in existing technologies. This invention provides a cytosine deaminase, a base editing system containing the deaminase, and their applications. The cytosine of this invention enables efficient editing of nuclear and mitochondrial genes, overcomes 5'-TC sequence bias, and expands the editing window.

[0007] The inventors used BLAST to search for homologs of DddA, finding a bacterial toxin-like protein derived from the gut bacterium *Roseburia intestinalis*, and optimized its eukaryotic codons, naming it riDddA. tox After homology comparison, it was truncated at amino acid position 2378 and named riDddA. tox-N(2266-2378),riDddA tox -C(2379-2409), based on the Aureus-N orientation for nDNA editing, attempted to add bpNLS to the front end of the plasmid containing SaKKH-nCas9(D10A) and replace the SV40 NLS at the tail with bpNLS to improve nuclear entry efficiency; added the transcription activation domain Rta after the UGI protein to improve editing efficiency; modified the dSpCas9 plasmid to nickase SpCas9 (nSpCas9 / H840A), which can synergize with SaKKH-nCas9(D10A) to generate a "U"-shaped cut on the target DNA to improve efficiency. In the TALE-L and TALE-R plasmids for mtDNA editing, Fok1 was replaced with riDddAtox-N and riDddAtox-C, respectively, and Rta was added after UGI, respectively. Based on this, a riDddA-based... tox The cell nuclear genome base editor and mitochondrial genome editor are two sets of cytosine base editors that can achieve efficient C-to-T editing on double-stranded nDNA and mtDNA, and have efficient editing capabilities for 5'-TC, 5'-AC, 5'-CC, and 5'-GC. This completely solves the sequence bias problem of cytosine base editors based on bacterial toxin-like proteins. The mitochondrial base editor has no sequence restriction, thereby broadening the target range and applicability of single-base editors.

[0008] The present invention solves the above-mentioned technical problems through the following technical solutions.

[0009] A first aspect of the present invention provides a cytosine deaminase comprising an amino acid sequence having at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with the amino acid sequence shown in SEQ ID NO:1.

[0010] In some specific embodiments of the present invention, the amino acid sequence of the cytosine deaminase is shown in SEQ ID NO:1.

[0011] In this invention, the amino acid sequence of the cytosine deaminase also includes an amino acid sequence that has partially the same amino acid sequence as SEQ ID NO:1 and is obtained by amino acid sequence splicing and cutting through intron sequences, etc., to obtain an amino acid sequence with complete fusion protein function.

[0012] In this invention, the cytosine deaminase can catalyze the cytosine deaminase activity of double-stranded DNA, and preferably can serve as a core component of a base editor to precisely induce C-to-T conversion on double-stranded DNA at the target sequence and its complementary strand.

[0013] A second aspect of the present invention provides a polypeptide complex comprising a first polypeptide and a second polypeptide;

[0014] The first polypeptide comprises an amino acid sequence having at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with the amino acid sequence shown in SEQ ID NO:17; the second polypeptide comprises an amino acid sequence having at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with the amino acid sequence shown in SEQ ID NO:30.

[0015] Optionally, the first polypeptide is linked to the second polypeptide.

[0016] A third aspect of the present invention provides a combination of fusion proteins, the combination of fusion proteins comprising a first fusion protein and a second fusion protein; the first fusion protein comprising a first polypeptide and a first nuclease, and the second fusion protein comprising a second polypeptide and a second nuclease;

[0017] Specifically, the first nuclease and the second nuclease break on the DNA strand; the first nuclease and the second nuclease have different recognition sites; and the first polypeptide and the second polypeptide, in their bound state, edit cytosine into thymine.

[0018] In this invention, the first nuclease and the second nuclease are preferably selected from VQR-spCas9, VRER-spCas9, spRY, spNG, SaCas9-KKH and SaCas9-NG.

[0019] In some embodiments of the present invention, the first fusion protein and / or the second fusion protein further include a nuclear localization signal sequence fragment.

[0020] In some embodiments of the present invention, the first fusion protein and / or the second fusion protein further include a uracil DNA glycosylation inhibitor fragment.

[0021] In some embodiments of the present invention, the first fusion protein and / or the second fusion protein further include a transcriptional activation domain fragment.

[0022] In some specific embodiments of the present invention, the first fusion protein further includes a nuclear localization signal sequence fragment, a uracil DNA glycosylation inhibitor fragment, and a transcription activation domain fragment.

[0023] In some specific embodiments of the present invention, the second fusion protein further includes a nuclear localization signal sequence fragment, a uracil DNA glycosylation inhibitor fragment, and a transcription activation domain fragment.

[0024] In some embodiments of the present invention, the nuclear localization signal sequence fragment is located at the N-terminus and / or C-terminus of the first fusion protein and / or the second fusion protein.

[0025] In some embodiments of the present invention, the uracil DNA glycosylation inhibitor fragment is located at the C-terminus of the first fusion protein and / or the second fusion protein.

[0026] In some embodiments of the present invention, the transcriptional activation domain fragment is located at the C-terminus of the first fusion protein and / or the second fusion protein.

[0027] In some embodiments of the present invention, the nuclear localization signal sequence fragment, the uracil DNA glycosylation inhibitor fragment, and the transcription activation domain fragment are linked to the first polypeptide, the second polypeptide, the first nuclease, and / or the second nuclease via linkers.

[0028] In some specific embodiments of the present invention, the first fusion protein comprises, from the N-terminus to the C-terminus, the following components in sequence: a nuclear localization signal sequence fragment, a first polypeptide, a first nuclease, a uracil DNA glycosyltransferase inhibitor fragment, a nuclear localization signal sequence fragment, and a transcription activation domain fragment.

[0029] In some specific embodiments of the present invention, the second fusion protein comprises, from the N-terminus to the C-terminus, the following components in sequence: a nuclear localization signal sequence fragment, a second polypeptide, a second nuclease, a uracil DNA glycosyltransferase inhibitor fragment, a nuclear localization signal sequence fragment, and a transcription activation domain fragment.

[0030] In this invention, the first polypeptide preferably comprises an amino acid sequence having at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with the amino acid sequence shown in SEQ ID NO:17; the second polypeptide preferably comprises an amino acid sequence having at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with the amino acid sequence shown in SEQ ID NO:30.

[0031] In this invention, the first nuclease preferably comprises an amino acid sequence having at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with the amino acid sequence shown in SEQ ID NO:21; the second nuclease preferably comprises an amino acid sequence having at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with the amino acid sequence shown in SEQ ID NO:32.

[0032] Alternatively, the first nuclease preferably comprises an amino acid sequence having at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with the amino acid sequence shown in SEQ ID NO:32; and the second nuclease preferably comprises an amino acid sequence having at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with the amino acid sequence shown in SEQ ID NO:21.

[0033] In some specific embodiments of the present invention, the nuclear localization signal sequence fragment has an amino acid sequence as shown in SEQ ID NO:13.

[0034] In some specific embodiments of the present invention, the uracil DNA glycosylation inhibitor fragment has an amino acid sequence as shown in SEQ ID NO:25.

[0035] In some specific embodiments of the present invention, the transcription activation domain fragment has an amino acid sequence as shown in SEQ ID NO:15.

[0036] In some specific embodiments of the present invention, the first polypeptide and the first nuclease are linked by a linker with an amino acid sequence as shown in SEQ ID NO:19.

[0037] In some specific embodiments of the present invention, adjacent uracil DNA glycosylation inhibitor fragments are connected by linkers with amino acid sequences as shown in SEQ ID NO:23.

[0038] In some specific embodiments of the present invention, the nuclear localization signal sequence fragment is connected to the adjacent transcription activation domain fragment via a linker with an amino acid sequence as shown in SEQ ID NO:27.

[0039] A fourth aspect of the present invention provides a combination of fusion proteins, the combination of fusion proteins comprising a first fusion protein and a second fusion protein; the first fusion protein comprising a first polypeptide and a nuclease, and the second fusion protein comprising a second polypeptide and a nuclease;

[0040] The nuclease causes breaks in the DNA strand; the first polypeptide and the second polypeptide, in a bound state, edit cytosine into thymine.

[0041] In some specific embodiments of the present invention, the nuclease is a transcription activator-like effector nuclease (TALEN).

[0042] In some embodiments of the present invention, the first fusion protein and / or the second fusion protein further include a mitochondrial targeting sequence fragment.

[0043] In some embodiments of the present invention, the first fusion protein and / or the second fusion protein further include a uracil DNA glycosylation inhibitor fragment.

[0044] In some embodiments of the present invention, the first fusion protein and / or the second fusion protein further include a transcriptional activation domain fragment.

[0045] In some specific embodiments of the present invention, the mitochondrial targeting sequence fragment is located at the N-terminus and / or C-terminus of the first fusion protein and / or the second fusion protein.

[0046] In some specific embodiments of the present invention, the uracil DNA glycosylation inhibitor fragment is located at the C-terminus of the first fusion protein and / or the second fusion protein.

[0047] In some specific embodiments of the present invention, the transcriptional activation domain fragment is located at the C-terminus of the first fusion protein and / or the second fusion protein.

[0048] In some specific embodiments of the present invention, the mitochondrial targeting sequence fragment, the uracil DNA glycosylation inhibitor fragment, and the transcription activation domain fragment are linked to the first polypeptide, the second polypeptide, and / or the nuclease via linkers.

[0049] In some specific embodiments of the present invention, the first fusion protein further includes the mitochondrial targeting sequence fragment, the uracil DNA glycosylation inhibitor fragment, and the transcription activation domain fragment.

[0050] In some specific embodiments of the present invention, the second fusion protein further includes the mitochondrial targeting sequence fragment, the uracil DNA glycosylation inhibitor fragment, and the transcription activation domain fragment.

[0051] In some specific embodiments of the present invention, the first fusion protein comprises, from the N-terminus to the C-terminus, the following components in sequence: a mitochondrial targeting sequence fragment, a nuclease, a first polypeptide, a uracil DNA glycosylation inhibitor fragment, and a transcription activation domain fragment.

[0052] In some specific embodiments of the present invention, the second fusion protein comprises, from the N-terminus to the C-terminus, the following components in sequence: a mitochondrial targeting sequence fragment, a nuclease, a second polypeptide, a uracil DNA glycosylation inhibitor fragment, and a transcription activation domain fragment.

[0053] In some embodiments of the present invention, the first polypeptide comprises an amino acid sequence having at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with the amino acid sequence shown in SEQ ID NO:17; the second polypeptide comprises an amino acid sequence having at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with the amino acid sequence shown in SEQ ID NO:30.

[0054] In some embodiments of the present invention, the nuclease comprises a left arm having an amino acid sequence that is at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identical to the amino acid sequence shown in SEQ ID NO:36, and a right arm having an amino acid sequence that is at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identical to the amino acid sequence shown in SEQ ID NO:38.

[0055] In some specific embodiments of the present invention, the mitochondrial targeting sequence fragment has an amino acid sequence as shown in SEQ ID NO:35.

[0056] In some specific embodiments of the present invention, the uracil DNA glycosylation inhibitor fragment has an amino acid sequence as shown in SEQ ID NO:25.

[0057] In some specific embodiments of the present invention, the transcription activation domain fragment has an amino acid sequence as shown in SEQ ID NO:15.

[0058] In some specific embodiments of the present invention, the mitochondrial targeting sequence fragment is linked to the left arm of the nuclease via a linker with an amino acid sequence as shown in SEQ ID NO:42.

[0059] A fifth aspect of the present invention provides a cytosine base editing system comprising a combination as described in the third aspect or the fourth aspect; the combination editing cytosine to thymine at a target site.

[0060] In some embodiments of the present invention, the cytosine base editing system further includes a guide RNA structure for guiding the combination to a target site.

[0061] A sixth aspect of the present invention provides a polynucleotide that encodes a combination as described in the third aspect, a combination as described in the fourth aspect, or a cytosine base editing system as described in the fifth aspect.

[0062] In some embodiments of the present invention, the nucleotide encoding the first polypeptide is shown in SEQ ID NO:18, and the nucleotide encoding the second polypeptide is shown in SEQ ID NO:31.

[0063] In some embodiments of the present invention, the nucleotide encoding the first nuclease is shown in SEQ ID NO:22, and the nucleotide encoding the second nuclease is shown in SEQ ID NO:33.

[0064] In other embodiments of the present invention, the nucleotide encoding the first nuclease is shown in SEQ ID NO:33, and the nucleotide encoding the second nuclease is shown in SEQ ID NO:22.

[0065] In other embodiments of the present invention, the nucleotides encoding the left arm of the nuclease are shown in SEQ ID NO:37, and the nucleotides encoding the right arm of the nuclease are shown in SEQ ID NO:39.

[0066] In some embodiments of the present invention, the nucleotides encoding the nuclear localization signal sequence fragment are shown in SEQ ID NO:29.

[0067] In some embodiments of the present invention, the nucleotide encoding the uracil DNA glycosylation inhibitor fragment is shown in SEQ ID NO:26.

[0068] In some embodiments of the present invention, the nucleotides encoding the transcription activation domain fragment are shown in SEQ ID NO:16.

[0069] A seventh aspect of the invention provides a construct comprising the polynucleotide as described in the sixth aspect.

[0070] In some embodiments of the present invention, the plasmid backbone of the expression vector is a eukaryotic expression vector plasmid backbone.

[0071] In some embodiments of the present invention, the construct comprises a nucleotide sequence as shown in SEQ ID NO:4 and a nucleotide sequence as shown in SEQ ID NO:6.

[0072] In other embodiments of the present invention, the construct comprises a nucleotide sequence as shown in SEQ ID NO:10 and a nucleotide sequence as shown in SEQ ID NO:12.

[0073] In some embodiments of the present invention, the backbone of the guide RNA structure is shown in SEQ ID NO:34.

[0074] In some preferred embodiments of the present invention, the plasmid backbone is selected from one or more of pCMV, pSV2 and pGL3.

[0075] In some preferred embodiments of the present invention, the construct further comprises the nucleotide sequence shown in SEQ ID NO:7 and the nucleotide sequence shown in SEQ ID NO:8.

[0076] An eighth aspect of the present invention provides an expression system comprising the constructs described in the seventh aspect.

[0077] In some embodiments of the present invention, the host cell of the expression system is selected from eukaryotic cells and prokaryotic cells.

[0078] In some preferred embodiments of the present invention, the eukaryotic cells are selected from mouse cells and human cells; preferably selected from mouse neuroma cells and human embryonic renal cell carcinoma cells, such as N2a cells and HEK293T cells.

[0079] A ninth aspect of the present invention provides a composition for nuclear genome base editing or mitochondrial gene base editing, the composition comprising a cytosine deaminase as described in the first aspect, a polypeptide complex as described in the second aspect, a combination as described in the third aspect, a combination as described in the fourth aspect, a cytosine base editing system as described in the fifth aspect, a polynucleotide as described in the sixth aspect, a construct as described in the seventh aspect, or an expression system as described in the eighth aspect.

[0080] The tenth aspect of the present invention provides the use of cytosine deaminase as described in the first aspect, polypeptide complex as described in the second aspect, combination as described in the third aspect, combination as described in the fourth aspect, cytosine base editing system as described in the fifth aspect, polynucleotide as described in the sixth aspect, construct as described in the seventh aspect, or expression system as described in the eighth aspect in the preparation of drugs for nuclear genome base editing or mitochondrial gene base editing, the construction of animal models, or crop breeding.

[0081] The eleventh aspect of the present invention provides a base editing method, the base editing method comprising administering to target cells a cytosine deaminase as described in the first aspect, a polypeptide complex as described in the second aspect, a combination as described in the third aspect, a combination as described in the fourth aspect, a cytosine base editing system as described in the fifth aspect, a polynucleotide as described in the sixth aspect, or a construct as described in the seventh aspect; or comprising contacting the target cells with an expression system as described in the eighth aspect.

[0082] In some embodiments of the present invention, the base editing method is performed in vivo or in vitro.

[0083] In some embodiments of the present invention, the target cell is a eukaryotic cell.

[0084] In some embodiments of the present invention, the base editing method is for non-therapeutic purposes.

[0085] In summary, the beneficial effects of the present invention are as follows:

[0086] Based on common knowledge in the field, the above-mentioned preferred conditions can be combined arbitrarily to obtain various preferred embodiments of the present invention.

[0087] The reagents and raw materials used in this invention are all commercially available.

[0088] The positive and progressive effects of this invention are as follows:

[0089] The cytosine deaminase and its derived nuclear and mitochondrial gene cytosine base editors of this invention feature highly efficient editing and no sequence bias. They are applicable to both in vitro and in vivo gene editing and gene mutation correction, and can be applied to nuclear and mitochondrial genes respectively, expanding the scope of gene editing and base editing applications and providing tool selectivity for base editing and gene correction. The mitochondrial cytosine base editor provided by this invention is characterized by its small size, lack of sequence bias, and lack of restrictions, making it more suitable for viral vector-based gene therapy and showing promising prospects for gene therapy and industrialization. Attached Figure Description

[0090] Figure 1 This is a schematic diagram of the structural domains of the riDddAtox protein.

[0091] Figure 2 This is a schematic diagram of the plasmid structure of plasmid 1 (pCMV-bpNLS-riDddAtox-N-nSpCas9-2*UGI-bpNLS-Rta) for the nuclear gene cytosine base editor in the example.

[0092] Figure 3 This is a schematic diagram of the plasmid structure of plasmid 2 (pCMV-bpNLS-riDddAtox-C-SaKKH-nCas9-UGI-bpNLS-Rta) used in the example.

[0093] Figure 4 This is a schematic diagram of the plasmid structure of the nuclear gene cytosine base editor sgRNA expression plasmid 1 (pGL3-U6-SpCas9sgRNAinsert site-scaffold-hPGK-GFP) in the example.

[0094] Figure 5 This is a schematic diagram of the plasmid structure of the nuclear gene cytosine base editor sgRNA expression plasmid 2 (pGL3-U6-SaCas9sgRNAinsert site-scaffold-hPGK-mCherry) in the example.

[0095] Figure 6 This is a schematic diagram illustrating the editing efficiency of the nuclear gene cytosine base editor in HEK293T cells in this embodiment.

[0096] Figure 7 This is a schematic diagram illustrating the editing efficiency of the nuclear gene cytosine base editor in the 5'-AC, 5'-CC, 5'-GC, and 5'-TC genome sequences of HEK293T cells in the examples.

[0097] Figure 8 This is a schematic diagram of the protein domains of the mitochondrial gene cytosine base editor plasmid 1 (PGL3-MTS-Talen-up-TALEinsert site-ccdB-Talen-down-riDddAtox-N-UGI-Rta) in the example.

[0098] Figure 9 This is a schematic diagram of the protein domains of the mitochondrial gene cytosine base editor plasmid 2 (PGL3-MTS-Talen-up-TALEinsert site-ccdB-Talen-down-riDddAtox-C-UGI-Rta) in the example.

[0099] Figure 10This is a schematic diagram illustrating the editing efficiency of the mitochondrial gene cytosine base editor in HEK293T cells in this embodiment.

[0100] Figure 11 This is a schematic diagram illustrating the editing efficiency of the mitochondrial gene cytosine base editor in the 5'-AC, 5'-CC, 5'-GC, and 5'-TC genome sequences of HEK293T cells in this embodiment.

[0101] Figure 12 This is a schematic diagram illustrating the editing efficiency of the mitochondrial gene cytosine base editor in N2a cells, as shown in the example. Detailed Implementation

[0102] This invention provides a novel double-stranded DNA cytosine deaminase optimized with eukaryotic codons, derived from a bacterial toxin of *Roseburia intestinalis*, and named riDddA. tox And truncate it to riDddA tox -N(2266-2378),riDddA tox -C(2379-2409), subsequently coupled with CRISPR / Cas9 technology to generate a nuclear genome cytosine deaminase base editor, and coupled with TALEN technology to generate a mitochondrial genome cytosine deaminase base editor. The nuclear genome cytosine deaminase base editor components nSpCas9 and SaKKH-nCas9 recognize NGG and NNNRRT PAM sequences, respectively, with the editing activity window (17-60 bp) in the middle region of the nuclear DNA sequences targeted by the two sgRNAs, simultaneously enabling C-to-T editing of both template and non-template strands. The mitochondrial genome cytosine deaminase base editor uses two sets of TALEs to recognize target sequences, with the editing activity window (14-18 bp) in the middle region of the mitochondrial DNA sequences targeted by the two sets of TALEs, simultaneously enabling C-to-T editing of both template and non-template strands. This system has no sequence restrictions or sequence bias and can edit 5'-AC, 5'-CC, 5'-GC, and 5'-TC.

[0103] In some embodiments of the present invention, the nuclear genome (nDNA) cytosine base editor consists of two cytosine editor plasmids and two sgRNA expression plasmids.

[0104] The nuclear genome cytosine base editor plasmid 1 (pCMV-bpNLS-riDddA) tox -N-nSpCas9-2*UGI-bpNLS-Rta) includes bpNLS peptide and riDddA from the N-terminus to the C-terminus. toxThe following components are sequentially fused together: a polypeptide fragment consisting of 1–113 amino acids at the N-terminus, a linker consisting of 32 amino acids, SpCas9 (H840A), a linker consisting of 10 amino acids, a 2*UGI polypeptide, a linker consisting of 4 amino acids, bpNLS, a linker consisting of 12 amino acids, and an Rta polypeptide fragment consisting of 1–190 amino acids. The above components can achieve cytosine base editing functionality through rearrangement or addition / removal of components.

[0105] The nuclear genome cytosine base editor plasmid 2 (pCMV-bpNLS-riDddA) tox -C-SaKKH-nCas9-UGI-bpNLS-Rta) consists of bpNLS peptide and riDddA from the N-terminus to the C-terminus. tox The following components are sequentially fused together: a C-terminal polypeptide fragment of 114–144 amino acids, a 32-amino acid linker, a SaKKH-nCas9(D10A) linker, a 10-amino acid linker, a UGI polypeptide, a 4-amino acid linker, bpNLS, a 12-amino acid linker, and an Rta polypeptide fragment of 1–190 amino acids. These components can achieve cytosine base editing functionality through rearrangement or addition / removal of components.

[0106] In some embodiments of the present invention, the nuclear gene cytosine base editing system further includes sgRNA expression plasmid 1 for SpCas9-specific sgRNA scaffold and sgRNA expression plasmid 2 for SaCas9-specific sgRNA scaffold. The DNA sequence of sgRNA expression plasmid 1 (pGL3-U6-SpCas9 sgRNA insert site-scaffold-hPGK-GFP) is shown in SEQ ID NO:7, and its expression elements from the N-terminus to the C-terminus include, in sequence, a U6 promoter, a gRNA targeting sequence insertion restriction site, a scaffold (SpCas9 specific), an hPGK promoter, a GFP fluorescent protein, and a termination signal. The DNA sequence of the sgRNA expression plasmid 2 (pGL3-U6-SaCas9 sgRNA insert site-scaffold-hPGK-mCherry) is shown in SEQ ID NO:8. Its expression elements, from the N-terminus to the C-terminus, include U6 promoter, gRNA target sequence insertion site, scaffold (SaCas9 specific), hPGK promoter, mCherry fluorescent protein, and termination signal.

[0107] In some embodiments of the present invention, the mitochondrial genome cytosine base editor plasmid 1 (pGL3-MTS-Talen-up-TALE insert site-ccdB-Talen-down-riDddA) tox -N-UGI-Rta) includes, from N-terminus to C-terminus, the MTS peptide, the Taylor-up peptide, the target sequence insertion site, the Taylor-down peptide, and riDddA. tox The cytosine base editor function is achieved by fusing a polypeptide fragment consisting of 1–113 amino acids at the N-terminus, a UGI polypeptide, and an Rta polypeptide fragment consisting of 1–190 amino acids. These components can be rearranged or have additional components added or removed to achieve the desired cytosine base editor function.

[0108] In some embodiments of the present invention, the mitochondrial genome cytosine base editor plasmid 2 (pGL3-MTS-Talen-up-TALE insert site-ccdB-Talen-down-riDddA) tox -C-UGI-Rta) includes, from N-terminus to C-terminus, the MTS peptide, the Taylor-up peptide, the target sequence insertion site, the Taylor-down peptide, and riDddA. tox The cytosine base editor function is achieved by fusing a polypeptide fragment consisting of 114–144 amino acids at the C-terminus, a UGI polypeptide, and an Rta polypeptide fragment consisting of 1–190 amino acids. These components can be rearranged or have additional components added or removed to achieve the desired cytosine base editor function.

[0109] In some embodiments of the present invention, the nuclear gene cytosine base editing system includes two fusion proteins and two sgRNA expression vectors. The fusion proteins include nuclear genome cytosine base editor plasmid 1 (pCMV-bpNLS-riDddA). tox -N-nSpCas9-2*UGI-bpNLS-Rta), nuclear genome cytosine base editor plasmid 2 (pCMV-bpNLS-riDddA) tox-C-SaKKH-nCas9-UGI-bpNLS-Rta); the sgRNA expression vectors include sgRNA expression plasmid 1 (pGL3-U6-SpCas9 sgRNA insert site-scaffold-hPGK-GFP) and sgRNA expression plasmid 2 (pGL3-U6-SaCas9 sgRNA insert site-scaffold-hPGK-mCherry). The two fusion proteins and the two sgRNA expression vectors work together to target nuclear gene DNA sequences under the guidance of SpCas9 sgRNA (recognizing NGG PAM) and SaCas9 sgRNA (recognizing NNNRRT PAM), respectively. The ununwound DNA double strand between the two sgRNAs forms the editing window of this editing system, with a size of 17-60 bp, where C-to-T conversion can occur on the target strand or its complementary strand.

[0110] The fusion protein is preferably a target sequence capable of editing a PAM recognition sequence with SpCas9 and SaCas9 at both ends of the target sequence.

[0111] In some embodiments of the present invention, the mitochondrial genome cytosine base editing system includes two fusion proteins: mitochondrial cytosine base editor plasmid 1 (PGL3-MTS-Talen-up-TALE insert site-ccdB-Talen-down-riDddA) tox -N-UGI-Rta (abbreviated as TALEN-L) and mitochondrial cytosine base editor plasmid 2 (pGL3-MTS-Talen-up-TALE insert site-ccdB-Talen-down-riDddA) tox -C-UGI-Rta (abbreviated as TALEN-R). The two fusion proteins work together to achieve highly efficient targeted base conversion of mitochondrial DNA. Preferably, the target sequence is one that can theoretically recognize any DNA sequence in the mitochondrial genome; more preferably, it targets the middle 14-18 bp of the sequences recognized by TALEN-L and TALEN-R, which serves as the editing window for this editing system, enabling efficient double-stranded DNA cytosine base editing to thymine (C-to-T).

[0112] In the fusion protein provided by this invention, the substitution, deletion, or addition can be a conserved amino acid substitution. Specifically, "conserved amino acid substitution" refers to the substitution of an amino acid residue by another amino acid residue with a similar side chain. The fusion protein provided by this invention may also include a nuclear localization signal fragment (NLS).

[0113] The constructs described in this invention can typically be obtained by inserting the isolated polynucleotides into a suitable expression vector. Those skilled in the art can select a suitable expression vector, for example, the expression vector may be, but is not limited to, pCMV expression vector, pSV2 expression vector, pGL3 expression vector and other eukaryotic expression vectors.

[0114] The guide RNA described in this invention is, for example, sgRNA in a base editing system known in the art. The sequence of the sgRNA is typically at least partially complementary to the target region, thereby enabling it to interact with the fusion protein and locate the fusion protein to the target region. The ununwound DNA double strand between the two sgRNAs forms the editing window of this editing system, with a size of 17-60 bp, where C-to-T conversion can occur on the target strand or its complementary strand.

[0115] This invention reduces the size of plasmids used for nuclear genome cytosine base editing, making it more suitable for packaging systems including, but not limited to, lentiviruses and adenoviruses. It is applicable to the construction of animal disease models or gene mutation correction therapy for diseases. Furthermore, it enables simultaneous cytosine base conversion in double-stranded DNA, expanding the application and targeting range of base editors. In addition, this invention offers advantages such as high editing precision and low off-target effects, demonstrating promising commercial prospects.

[0116] Under appropriate conditions, the expression system provided in aspect eight of this invention can be expressed, and the protein can be purified; under appropriate conditions, the cytosine base editing system described in aspect five of this invention can be transcribed in vitro, and the corresponding mRNA or sgRNA can be purified. The mRNA corresponding to the protein or the purified protein described in this invention can perform base editing on the target region in the presence of the sgRNA that targets the target region. For mitochondrial genes, this invention can perform base editing on the mitochondrial DNA target region after introduction into cells.

[0117] To better illustrate the purpose, technical solution, and advantages of this invention, the invention will be further described below in conjunction with the accompanying drawings and specific embodiments. Non-essential substitutions or improvements by others are within the scope of protection of this invention. Reagents or instruments not specified in the embodiments can be purchased commercially. Experimental methods not detailed herein should be performed according to conventional conditions or methods recommended by the reagent manufacturer.

[0118] The present invention is further illustrated below by way of embodiments, but the invention is not limited to the scope of the embodiments described herein. Experimental methods in the following embodiments that do not specify specific conditions were performed according to conventional methods and conditions, or as selected according to the product instructions.

[0119] Example 1: riDddA tox Determination of double-stranded DNA cytosine deaminase

[0120] Search the database using BLAST for 1) DddA from Burkholderia. tox 1) Possesses 30-99.9% homology; 2) Contains the SCP1.201 deaminase domain; 3) Contains domains similar to DddA. tox Homologous to this bacterium, it contains the important amino acid E1347, which, when mutated, can inactivate deaminases. Based on this, a bacterial toxin originating from the intestinal bacteria *Roseburia intestinalis* was identified and named riDddA. tox (Amino acids 2266-2409, amino acid sequence as shown in SEQ ID NO:1, nucleotide sequence as shown in SEQ ID NO:2; full-length name riDddA). Reference DddA tox The truncation method is named riDddA tox (divided into riDddA) tox -N(2266-2378), riDddA tox -C(2379-2409). The structure of riDddA is as follows: Figure 1 As shown.

[0121] Example 2: Construction of a nuclear gene cytosine base editing system

[0122] 2.1 Nuclear gene cytosine base editor plasmid 1 (pCMV-bpNLS-riDddA) tox Construction of -N-nSpCas9-2*UGI-bpNLS-Rta)

[0123] riDddA derived from intestinal Rosbyrate bacteria tox Based on the DNA sequence, eukaryotic codon optimization was performed. After optimization, Nanjing GenScript Biotech Co., Ltd. was commissioned to synthesize the full-length DNA fragment for subsequent cloning of riDddA. tox -N(2266-2378), riDddA tox-C(2379-2409) PCR template. Primers were designed using pCMV-AncBE4max plasmid as the backbone (Table 1, PCR for 1; PCR rev 1) to amplify sequences other than AncAPOBEC1. The reaction system was: 1 ng pCMV-AncBE4max plasmid, 1 μL PCR for 1, 1 μL PCR rev 1, 1 μL Phanta Max Super-Fidelity DNA Polymerase (Novaza: P505-d1), 25 μL 2×Phanta Max Buffer, 1 μL dNTP Mix, and ddH2O to make up to 50 μL. The amplification program was 95℃, 5 min; (95℃, 10 s, 60℃, 15 s, 72℃, 5 min) × 18 cycles; 72℃, 3 min; hold at 4℃. Primers were designed (Table 1, PCR for 2; PCR rev 2) to amplify the eukaryotic codon-optimized riDddA tox -N, the amplification reaction system and procedure are the same as above. After confirming the target band by agarose gel electrophoresis, the fragment was purified using a clean-up kit (AxyPrep PCR Clean Kit) and eluted with 20 μL ddH2O. 0.5 μL of DpnI enzyme (NEB:R0176L) was added to the purified pCMV-AncBE4max PCR fragment (without AncAPOBEC) and incubated at 37°C for 30 min. Subsequently, the pCMV-AncBE4max fragment (excluding Anc APOBEC) was recombinated with the riDddAtox-N fragment using a two-fragment recombinase (Norvoza: C112-01 / 02). The recombination ligation system was: 180 ng of pCMV-AncBE4max fragment (excluding Anc APOBEC), riDddAtox-N fragment... tox -N fragment 12ng, 5×CE II Buffer 4μL, Exnase II 2μL, ddH2O to 20μL. After recombination and ligation at 37℃ for 30 min, the recombination and ligation product was transformed into competent E. coli, plated, and incubated overnight at 37℃. Single colonies were picked and sent for identification.

[0124] Using the plasmid with correct sequencing results from the previous step as a template, two pairs of primers (Table 1, PCR for3; PCR rev3; PCR for4; PCR rev4) were designed to introduce the H840A mutation and repair the A10D mutation. The amplification system and procedure were the same as above. The recombinant ligation product was transformed into competent E. coli, plated, and incubated overnight at 37°C. Single colonies were picked and identified, yielding pCMV-bpNLS-riDddA. tox -N-nSpCas9-2*UGI-bpNLS.

[0125] Primers were designed (Table 1, PCR for5; PCR rev5; PCR for6; PCR rev6; PCR for7; PCR rev7; PCR for8; PCR rev8; PCR for9; PCR rev9; PCR for10; PCR rev10; PCR for11; PCR rev11; PCR for12; PCR rev12) to amplify and construct the Rta fragment. The amplification system and procedure were the same as above. Primers were designed (Table 1, PCR for13; PCR rev13) to linearize the plasmid pCMV-bpNLS-riDddAtox-N-nSpCas9-2*UGI-bpNLS obtained above. The amplification system and procedure were the same as above. After confirming the target band by agarose gel electrophoresis, the fragment was purified using a clean-up kit (AxyPrep PCR Clean Kit) and eluted with 20 μL ddH2O. 0.5 μL of DpnI enzyme (NEB:R0176L) was added to the purified linear fragment, and the fragment was incubated at 37°C for 30 min. Subsequently, the cells were recombinantly ligated using two recombinase fragments (Novizan: C112-01 / 02). After ligation at 37°C for 30 min, the ligation product was transformed into competent *E. coli*, plated, and incubated overnight at 37°C. Single colonies were picked for identification, and the correct sequence pCMV-bpNLS-riDddA was finally obtained. tox -N-nSpCas9-2*UGI-bpNLS-Rta (plasmid structure as follows) Figure 2 (As shown).

[0126] Table 1 Primer Design

[0127]

[0128]

[0129]

[0130] 2.2 Nuclear gene cytosine base editor plasmid 2 (pCMV-bpNLS-riDddA) tox Construction of -C-SaKKH-nCas9-UGI-bpNLS-Rta)

[0131] Primers were designed using pCMV-SaKKH-nCas9-BE4max as a template (Table 1, PCR for14: PCR rev14), with one UGI removed. The amplification system and procedure were the same as above. After identifying the target band by lipoglycogel electrophoresis, the fragment was purified using a clean-up kit (AxyPrep PCR Clean Kit) and eluted with 20 μL ddH2O. 0.5 μL of DpnI enzyme (NEB: R0176L) was added to the purified linear fragment, and the mixture was incubated at 37°C for 30 min. Subsequently, the fragment was recombinantly ligated using two-fragment recombinase (Novizan: C112-01 / 02). After ligation at 37°C for 30 min, the ligation product was transformed into competent E. coli, plated, and incubated overnight at 37°C. Single colonies were then selected for identification.

[0132] Using the plasmid identified by the sequencing as a template, primers (except AncAPOBEC1) were designed (Table 1, PCR for 1; PCR rev1) to amplify the backbone. The amplification system and procedure are as described above. Primers (Table 1, PCR for 15; PCR rev15) were designed to amplify the eukaryotic codon-optimized riDddA. tox -C, the amplification reaction system and procedure are the same as above. After identifying the target band by lipoglycolic gel electrophoresis, the fragment was purified using a clean-up kit (AxyPrep PCR Clean Kit) and eluted with 20 μL ddH2O. 0.5 μL of DpnI enzyme (NEB:R0176L) was added to the purified linear backbone fragment, and incubated at 37°C for 30 min. Subsequently, the backbone (except AncAPOBEC1) was recombined with riDddA using a two-fragment recombinase (Novizan: C112-01 / 02). tox -C fragment recombination ligation. After recombination ligation at 37℃ for 30 min, the recombination ligation product was transformed into competent E. coli, plated, and incubated overnight at 37℃. Single colonies were picked and identified, yielding pCMV-bpNLS-riDddA. tox -C-SaKKH-nCas9-UGI-bpNLS.

[0133] Using the plasmid as a template, primers (PCR for13; PCR rev13) were designed to linearize it. The amplification system and procedure were the same as above. After confirming the target band by lipoglycolic acid gel electrophoresis, the fragment was purified using a clean-up kit (AxyPrep PCR Clean Kit) and eluted with 20 μL ddH2O. 0.5 μL of DpnI enzyme (NEB: R0176L) was added to the purified linear fragment, and the mixture was incubated at 37°C for 30 min. Subsequently, the fragment was ligated to the constructed Rta fragment using a two-fragment recombinase (Novizan: C112-01 / 02). After ligation at 37°C for 30 min, the ligation product was transformed into competent *E. coli*, plated, and incubated overnight at 37°C. Single colonies were picked for identification, and pCMV-bpNLS-riDddA was successfully obtained. tox -C-SaKKH-nCas9-UGI-bpNLS-Rta (plasmid structure as follows) Figure 3 (As shown).

[0134] 2.3 sgRNA plasmid 1 (pGL3-U6-SpCas9 sgRNA insert site-scaffold-hPGK-GFP) (plasmid structure as shown) Figure 4 Construction of target plasmids (as shown)

[0135] sgRNAs were designed and two complementary oligos were synthesized. The upstream sequence was 5'-accg-20nt-3', and the downstream sequence was 5'-aaac-20nt-3' (the 20nt substituent is complementary to the upstream sequence, as shown in Table 2). The upstream and downstream sequences were annealed using a program (95℃, 5 min; 95℃-85℃ at-2℃ / s; 85℃-25℃ at-0.1℃ / s; hold at 4℃) and ligated into the pGL3-U6-SpCas9 sgRNA insert site-scaffold-hPGK-GFP vector linearized with BsaI-HFv2 (NEB: R3733). The linearization digestion system was: pGL3-U6-SpCas9 sgRNA insert site-scaffold-hPGK-GFP 3 μg; CutSmart Buffer 5 μL; BsaI-HFv2 2 μL; ddH2O to bring the total to 50 μL. Incubate overnight at 37°C. The ligation mixture consisted of: 1 μL T4 ligation buffer (NEB: M0202L), 20 ng linearized vector, 2 μL annealed oligo fragment (100 μM), 0.5 μL T4 DNA ligase (NEB: M0202L), and ddH2O to a final volume of 10 μL. Incubate overnight at 16°C.

[0136] Table 2 Primer sequences

[0137]

[0138]

[0139]

[0140]

[0141]

[0142] 2.4gRNA plasmid 2 (pGL3-U6-SaCas9 sgRNA insert site-scaffold-hPGK-mCherry) (plasmid structure as shown) Figure 5 Construction of target plasmids (as shown)

[0143] The sgRNA was designed and two complementary oligos were synthesized. The upstream sequence was 5'-accg-20nt-3' and the downstream sequence was 5'-aaac-20nt-3' (the 22 / 23nt substitution sequence is complementary to the upstream sequence, as shown in Table 2). The upstream and downstream sequences were annealed using a program (95℃, 5 min; 95℃-85℃ at-2℃ / s; 85℃-25℃ at-0.1℃ / s; hold at 4℃) and ligated into the pGL3-U6-SaCas9 sgRNA insert site-scaffold-hPGK-mCherry vector linearized with BsaI-HFv2 (NEB: R3733). The linearization digestion system was as follows: pGL3-U6-SaCas9 sgRNA insertsite-scaffold-hPGK-mCherry 3 μg; CutSmart Buffer 5 μL; BsaI-HFv2 2 μL; ddH2O to a final volume of 50 μL. Digestion was performed overnight at 37°C. The ligation system was as follows: T4 ligation buffer (NEB: M0202L) 1 μL, linearized vector 20 ng, annealed oligo fragment (100 μM) 2 μL, T4 DNA ligase (NEB: M0202L) 0.5 μL, ddH2O to a final volume of 10 μL. Ligation was performed overnight at 16°C.

[0144] Example 3

[0145] The nuclear genome editor constructed in Example 2 above was used to transfect HEK293T cells, and the process is as follows:

[0146] 3.1 HEK293T cells (from ATCC) were resuscitated and cultured in 10 cm culture dishes (Corning, 430167) in DMEM (HyClone, SH30243.01) containing 10% fetal bovine serum (HyClone, SV30087). The culture temperature was 37℃ and the CO2 concentration was 5%. After multiple passages, when the cell density reached 90%, the cells were transferred to 24-well plates (Jet Biotech).

[0147] 3.2 After culturing plated cells for 16-18 hours, transfect them when the cell confluence reaches 80%. The transfection system is: editor plasmid 1 (pCMV-bpNLS-riDddA) tox -N-nSpCas9-2*UGI-bpNLS-Rta)562.5ng, editor plasmid 2(pCMV-bpNLS-riDddA tox 562.5 ng of pGL3-U6-SpCas9 sgRNA insert site-scaffold-hPGK-GFP, 187.5 ng of gRNA plasmid 1, 187.5 ng of gRNA plasmid 2, 187.5 ng of gRNA plasmid 2, and 4.5 μL of EZTrans transfection reagent (Liji Biotechnology).

[0148] 3.3 The specific transfection steps are as follows:

[0149] 3.3.1 Preparation of reagent A: For each well of cells, dilute 1.5 μg of plasmid into 40 μL of serum-free, antibiotic-free, high-glucose DMEM medium and mix well.

[0150] 3.3.2 Preparation of reagent B: For each well of cells, dilute 4.5 μL of EZ Trans transfection reagent (EZ Trans: plasmid DNA = 3:1) into 40 μL of serum-free and antibiotic-free high-glucose DMEM medium and vortex to mix.

[0151] 3.3.3 Let reagents A and B stand for 5 minutes, then add reagent B to reagent A and shake to mix.

[0152] 3.3.4 Let stand at room temperature for 15 min to form the EZ Trans-DNA complex. Evenly drop the prepared EZ Trans-DNA transfection complex into the corresponding wells of a 24-well plate, and gently shake the culture dish to disperse the EZ Trans-DNA complex evenly.

[0153] 3.3.5 Incubate at 37℃ in a 5% CO2 incubator for 6 hours, remove the culture medium containing the EZ Trans-DNA complex, replace with new culture medium, and incubate for 3 days.

[0154] 3.4 After culturing the transfected cells for 3 days, the cells were digested with trypsin to obtain cells. Cells were then further sorted by flow cytometry to obtain cells that were positive for both GFP and mCherry (FITC fluorescence intensity, PE-Texas Red-A top 10%). Genomic DNA was extracted from the collected cells using the phenol-chloroform method.

[0155] 3.5 Design primers (Table 1, EMX1-site1-for1; EMX1-site1-rev1; EMX1-site1-for2; EMX1-site1-rev2; EMX1-site1-for3; EMX1-site2-for1; EMX1-site2-rev1; EMX1-site2-for2; EMX1-site2-rev2; EMX1-site2-for3; EMX1-site3-for1; EMX1-site3-rev1; EMX1-site3-for2; EMX1-site3-rev2; EMX1-site3-for3; EMX1-site4-for1; EMX1-site4-rev1; EMX1-site4-for2; EMX1-site4-rev2; EMX1-site4-for3; The target sequence was amplified using the following methods: RNF2-site1-for1; RNF2-site1-rev1; RNF2-site1-for2; RNF2-site1-rev2; RNF2-site1-for3; RNF2-site2-for1; RNF2-site2-rev1; RNF2-site2-for2; RNF2-site2-rev2; RNF2-site2-for3; RNF2-site3-for1; RNF2-site3-rev1; RNF2-site3-for2; RNF2-site3-rev2; RNF2-site3-for3; RNF2-site4-for1; RNF2-site4-rev1; RNF2-site4-for2; RNF2-site4-rev2; RNF2-site4-for3, and then sent for testing. Editing efficiency statistics are as follows: Figure 6 As shown. Furthermore, the editing efficiency of the nuclear gene editor on the 5'-AC, 5'-CC, 5'-GC, and 5'-TC sequences of the HEK293T cell genome was statistically analyzed, such as... Figure 7 As shown.

[0156] Example 4: Construction of a mitochondrial gene cytosine base editing system

[0157] 4.1 Mitochondrial gene cytosine base editor plasmid 1 (pGL3-MTS-Talen-up-TALE insertsite-ccdB-Talen-down-riDddA) tox -N-UGI-Rta)( Figure 8 Construction of a plasmid targeting the human mitochondrial genome

[0158] Using pGL3-MTS-ccdb-DddA(G1397)-UGI-Puro as a template, Puro was replaced with mCherry. First, using pGL3-MTS-ccdb-DddA(G1397)-UGI-Puro as a template, Puro was removed by enzyme digestion, and the plasmid was linearized. The digestion system consisted of 3 μg pGL3-MTS-ccdb-DddA(G1397)-UGI-Puro, 5 μL 10×CutSmart (NEB: B7204S), 2 μL BamHI-HF (NEB: R3136L), and ddH2O to a final volume of 50 μL. Digestion was performed overnight at 37°C. Using pGL3-U6-SaCas9 sgRNA insertsite-scaffold-hPGK-mCherry as a template, primers (PCR for16; PCR rev16) were designed to amplify the mCherry fragment. The amplification system and procedure were as described above. After identifying the target band by lipoglycogel electrophoresis, the digested vector and mCherry fragment were purified using a clean-up kit (AxyPrep PCR Cleaning Kit) and eluted with 20 μL ddH2O. 0.5 μL of DpnI enzyme (NEB:R0176L) was added to the purified linear mCherry fragment, and the mixture was incubated at 37°C for 30 min. Subsequently, the digested vector and mCherry fragment were recombinantly ligated using a two-fragment recombinase (Novizan: C112-01 / 02). After ligation at 37°C for 30 min, the recombinant ligation product was transformed into competent E. coli, plated, and incubated overnight at 37°C. Single colonies were picked and identified, yielding the correct pGL3-MTS-ccdb-DddA(G1397N)-UGI-mCherry.

[0159] Using the correctly identified plasmid as a template, replace DddA(G1397N) with riDddA. tox(G2378N). Using the correctly identified pGL3-MTS-ccdb-DddA(G1397N)-UGI-mCherry as a template, DddA(G1397N) was removed by enzyme digestion. The digestion system consisted of 3 μg pGL3-MTS-ccdb-DddA(G1397N)-UGI-mCherry, 5 μL NEBuffer 3.1 (NEB: B7203S), 2 μL Bgl II (NEB: R0144S), and ddH2O to a final volume of 50 μL. Digestion was performed overnight at 37°C. pCMV-bpNLS-riDddA was then used as a template. tox Using -N-nSpCas9-2*UGI-bpNLS-Rta as a template, primers were designed (PCR for17; PCR rev17) to amplify riDddA tox (G2378N), the amplification system and procedure are as above. After confirming the target band by lipoglycolic acid gel electrophoresis, the digested vector and riDddA were purified using a clean-up kit (AxyPrep PCR Cleaning Kit). tox The (G2378N) fragment was eluted with 20 μL ddH2O. The purified riDddA... tox The (G2378N) fragment was added to 0.5 μL of DpnI enzyme (NEB:R0176L) and incubated at 37°C for 30 min. Subsequently, the digested vector and riDddA were recombinantly incubated using a two-fragment recombinase (Novizan: C112-01 / 02). tox (G2378N) fragment recombination ligation. After recombination ligation at 37℃ for 30 min, the recombination ligation product was transformed into competent E. coli, plated, and incubated overnight at 37℃. Single colonies were picked for identification to obtain the correct pGL3-MTS-ccdb-riDddA fragment. tox (G2378N)-UGI-mCherry.

[0160] Using the correctly identified plasmid as a template, Rta elements were added. The correctly identified pGL3-MTS-ccdb-riDddA... tox Using (G2378N)-UGI-mCherry as a template, the enzyme was linearized by digestion. The digestion system was pGL3-MTS-ccdb-riDddA. tox (G2378)-UGI-mCherry 3ug, 10xCutSmart (NEB: B7204S) 5μL, Pme I (NEB: R0560S) 2μL, ddH2O to make up to 50μL, digested overnight at 37℃. pCMV-bpNLS-riDddA toxUsing -N-nSpCas9-2*UGI-bpNLS-Rta as a template, primers (PCR for18; PCR rev18) were designed to amplify the Rta fragment. The amplification system and procedure were as described above. After confirming the target band by lipoglycolic acid gel electrophoresis, the digested vector and Rta fragment were purified using a clean-up kit (AxyPrep PCR Clean Kit) and eluted with 20 μL ddH2O. 0.5 μL of DpnI enzyme (NEB:R0176L) was added to the purified Rta fragment, and the mixture was incubated at 37°C for 30 min. Subsequently, the digested vector and Rta fragment were recombinantly ligated using a two-fragment recombinase (Novizan: C112-01 / 02). After ligation at 37°C for 30 min, the recombinant ligation product was transformed into competent E. coli, plated, and incubated overnight at 37°C. Single colonies were picked for identification, and pGL3-MTS-ccdb-riDddA was successfully obtained. tox (G2378N)-UGI-GFP-Rta (same as pGL3-PGL3-MTS-Talen-up-TALE insert site-ccdB-Talen-down-riDddA tox -N-UGI-Rta).

[0161] Design the TALE sequence, digest and ligate it to construct pGL3-MTS-Talen-up-TALE, insert site-ccdB-Talen-down-riDddA tox Construction of the -N-UGI-Rta targeting plasmid. The enzyme digestion system was pGL3-MTS-Talen-up-TALE insert site-ccdB-Talen-down-riDddA. tox 100 ng of -N-UGI-Rta, 50 ng of RVD plasmid, 0.8 μL of Bsa I-HFv2 (NEB: R3733), 0.8 μL of T4 DNA Ligase (NEB: M0202M), 1 μL of T4 DNA ligase buffer, and ddH2O to a final volume of 10 μL. The digestion program was 37℃, 30 min; (37℃, 5 min; 16℃, 5 min) × 15 cycles; 50℃, 5 min; 80℃, 5 min; hold at 4℃. The ligation product was transformed into competent *E. coli*, plated, and incubated overnight at 37℃. Single colonies were picked for identification. The plasmid pGL3-PGL3-MTS-Talen-up-TALE inserted-Talen-down-riDddA, which recognizes the mitochondrial DNA target sequence 15 or 17 bp upstream, was successfully obtained. tox-N-UGI-Rta plasmid. The RVD sequence targeting human mitochondrial DNA is shown in Table 2.

[0162] 4.2 Mitochondrial gene cytosine base editor plasmid 2 (pGL3-MTS-Talen-up-TALE-insertsite-ccdB-Talen-down-riDddA) tox -C-UGI-Rta)( Figure 9 Construction of a plasmid targeting the human mitochondrial genome

[0163] Using pGL3-MTS-ccdb-DddA(G1397C)-UGI-Puro as a template, Puro was replaced with GFP. First, using pGL3-MTS-ccdb-DddA(G1397C)-UGI-Puro as a template, Puro was removed by enzyme digestion, and the plasmid was linearized. The digestion system consisted of 3 μg pGL3-MTS-ccdb-DddA(G1397C)-UGI-Puro, 5 μL 10×CutSmart (NEB: B7204S), 2 μL BamHI-HF (NEB: R3136L), and ddH2O to a final volume of 50 μL. Digestion was performed overnight at 37°C. Using pGL3-U6-SpCas9 sgRNA insertsite-scaffold-hPGK-GFP as a template, primers (PCR for19; PCR rev19) were designed to amplify the GFP fragment. The amplification system and procedure were as described above. After identifying the target band by lipoglycolic acid gel electrophoresis, the digested vector and GFP fragment were purified using a clean-up kit (AxyPrep PCR Cleaning Kit) and eluted with 20 μL ddH2O. 0.5 μL of DpnI enzyme (NEB:R0176L) was added to the purified linear mCherry fragment, and the mixture was incubated at 37°C for 30 min. Subsequently, the digested vector and mCherry fragment were recombinantly ligated using a two-fragment recombinase (Novizan: C112-01 / 02). After ligation at 37°C for 30 min, the recombinant ligation product was transformed into competent E. coli, plated, and incubated overnight at 37°C. Single colonies were picked and identified, yielding the correct pGL3-MTS-ccdb-DddA(G1397C)-UGI-GFP.

[0164] Using the correctly identified plasmid as a template, replace DddA(G1397C) with riDddA. tox(G2378C). Using the correctly identified pGL3-MTS-ccdb-DddA(G1397C)-UGI-GFP as a template, DddA(G1397C) was removed by enzyme digestion. The digestion system consisted of 3 μg pGL3-MTS-ccdb-DddA(G1397C)-UGI-GFP, 5 μL NEBuffer 3.1 (NEB: B7203S), 2 μL Bgl II (NEB: R0144S), and ddH2O to a final volume of 50 μL. Digestion was performed overnight at 37°C. pCMV-bpNLS-riDddA was then used as a template. tox Using -C-SaKKH-nCas9-UGI-bpNLS-Rta as a template, primers were designed (PCR for20; PCR rev20) to amplify riDddA tox (G2378C), the amplification system and procedure are as above. After identifying the target band by lipoglycolic acid gel electrophoresis, the digested vector and riDddA were purified using a clean-up kit (AxyPrepPCR Cleaning Kit). tox The (G2378C) fragment was eluted with 20 μL ddH2O. The purified riDddA... tox The (G2378C) fragment was added to 0.5 μL of DpnI enzyme (NEB:R0176L) and incubated at 37°C for 30 min. Subsequently, the digested vector and riDddA were recombined using a two-fragment recombinase (Novizan: C112-01 / 02). tox (G2378C) fragment recombination ligation. After recombination ligation at 37℃ for 30 min, the recombination ligation product was transformed into competent E. coli, plated, and incubated overnight at 37℃. Single colonies were picked for identification, yielding pGL3-MTS-ccdb-riDddA. tox (G2378C)-UGI-GFP.

[0165] Using the correctly identified plasmid as a template, Rta elements were added. The correctly identified pGL3-MTS-ccdb-riDddA... tox Using (G2378C)-UGI-GFP as a template, the enzyme was linearized by digestion. The digestion system was pGL3-MTS-ccdb-riDddA. tox (G2378C)-UGI-GFP 3μg, 10xCutSmart (NEB: B7204S) 5μL, Pme I (NEB: R0560S) 2μL, ddH2O to make up to 50μL, digested overnight at 37℃. pCMV-bpNLS-riDddA toxUsing -N-nSpCas9-2*UGI-bpNLS-Rta as a template, primers (PCR for18; PCR rev18) were designed to amplify the Rta fragment. The amplification system and procedure were as described above. After confirming the target band by lipoglycolic acid gel electrophoresis, the digested vector and Rta fragment were purified using a clean-up kit (AxyPrep PCR Clean Kit) and eluted with 20 μL ddH2O. 0.5 μL of DpnI enzyme (NEB:R0176L) was added to the purified Rta fragment, and the mixture was incubated at 37°C for 30 min. Subsequently, the digested vector and Rta fragment were recombinantly ligated using a two-fragment recombinase (Novizan: C112-01 / 02). After ligation at 37°C for 30 min, the recombinant ligation product was transformed into competent E. coli, plated, and incubated overnight at 37°C. Single colonies were picked for identification, yielding pGL3-MTS-ccdb-riDddA. tox (G2378C)-UGI-GFP-Rta (same as pGL3-MTS-Talen-up-TALE insert site-ccdB-Talen-down-riDddAtox-C-UGI-Rta).

[0166] The TALE sequence was designed, and the target plasmid pGL3-MTS-Talen-up-TALE insert site-ccdB-Talen-down-riDddAtox-C-UGI-Rta was constructed by restriction enzyme digestion and ligation. The restriction enzyme digestion system consisted of 100 ng of pGL3-MTS-Talen-up-TALE insert site-ccdB-Talen-down-riDddAtox-C-UGI-Rta, 50 ng of RVD plasmid, 0.8 μL of Bsa I-HFv2 (NEB: R3733), 0.8 μL of T4 DNA Ligase (NEB: M0202M), 1 μL of T4 DNA ligase buffer, and ddH2O to a final volume of 10 μL. The digestion program was: 37℃, 30 min; (37℃, 5 min; 16℃, 5 min) × 15 cycles; 50℃, 5 min; 80℃, 5 min; hold at 4℃. The ligation product was transformed into competent *E. coli*, plated, and incubated overnight at 37°C. Single colonies were selected for identification, successfully yielding the plasmid pGL3-PGL3-MTS-Talen-up-TALE inserted-Talen-down-riDddA, which recognizes a 15 or 17 bp downstream target sequence of mitochondrial DNA. tox -C-UGI-Rta plasmid. The RVD sequence targeting human mitochondrial DNA is shown in Table 2.

[0167] Example 5

[0168] HEK293T cells were transfected using the mitochondrial genome cytosine base editor constructed in Example 4 above, as follows:

[0169] 5.1 After multiple passages of HEK293T cells, when the cell density reached 90%, the cells were separated into 12-well plates (JET Biotechnology).

[0170] 5.2 After culturing plated cells for 16-18 hours, transfect them when the cell confluence reaches 80%. The transfection system is as follows: 1.5 μg of mitochondrial genome cytosine base editor plasmid 1 (pGL3-MTS-Talen-up-TALE inserted-Talen-down-riDddAtox-N-UGI-Rta) and 1.5 μg of mitochondrial genome cytosine base editor plasmid 2 (pGL3-PGL3-MTS-Talen-up-TALE inserted-Talen-down-riDddA... tox -C-UGI-Rta) 1.5 μg, EZTrans transfection reagent (Liji Biotechnology) 7.5 μL.

[0171] 5.3 The specific transfection steps are as follows:

[0172] 5.3.1 Preparation of Reagent A: For each well of cells, dilute 3 μg of plasmid into 60 μL of serum-free, antibiotic-free, high-glucose DMEM medium and mix well.

[0173] 5.3.2 Preparation of reagent B: For each well of cells, dilute 7.5 μL of EZ Trans transfection reagent (EZ Trans: plasmid DNA = 2.5:1) into 60 μL of serum-free, antibiotic-free, high-glucose DMEM medium and vortex to mix.

[0174] 5.3.3 Let reagents A and B stand for 5 minutes, then add reagent B to reagent A and shake to mix.

[0175] 5.3.4 Let stand at room temperature for 15 min to form the EZ Trans-DNA complex. Evenly drop the prepared EZ Trans-DNA transfection complex into the corresponding wells of a 12-well plate, and gently shake the culture dish to disperse the EZ Trans-DNA complex evenly.

[0176] 5.3.5 Incubate at 37℃ in a 5% CO2 incubator for 6 hours, remove the culture medium containing the EZ Trans-DNA complex, replace with new culture medium, and incubate for 3 days.

[0177] 5.4 After culturing transfected cells for 3 days, cells were digested with trypsin to obtain cells. Further flow cytometry sorting was used to obtain GFP-positive and mCherry-positive cells (FITC fluorescence intensity, PE-Texas Red-A top 10%). On day 3, genomic DNA was extracted from a portion of the collected cells using the phenol-chloroform method. The remaining cells were seeded back into 96-well plates, and genomic DNA was extracted from them on day 6 using the phenol-chloroform method.

[0178] 5.5 Design primers (Table 1, homo-mt-site1-for1; homo-mt-site1-rev1; homo-mt-site1-for2; homo-mt-site1-rev2; homo-mt-site2 / 5 / 6-for1; homo-mt-site2 / 5 / 6-rev1; homo -mt-site2 / 5 / 6-for2;homo-mt-site2 / 5 / 6-rev2;homo-mt-site2 / 5 / 6-for3;homo-mt-site2 / 5 / 6-rev3;homo-mt-site2 / 5 / 6-for4;homo-mt-site3 / 4-for1;homo- The target sequence was amplified using the following methods: mt-site3 / 4-rev1; homo-mt-site3 / 4-for2; homo-mt-site3 / 4-rev2; homo-mt-site7 / 8-for1; homo-mt-site7 / 8-rev1; homo-mt-site7 / 8-for2; homo-mt-site7 / 8-rev2; homo-mt-site7 / 8-for3; homo-mt-site7 / 8-rev3; homo-mt-site11-for1; homo-mt-site11-rev1; homo-mt-site11-for2; homo-mt-site11-rev2) and then sent for analysis. Editing efficiency statistics are as follows: Figure 10 As shown. Furthermore, the editing efficiency of the mitochondrial gene editor on the 5'-AC, 5'-CC, 5'-GC, and 5'-TC sequences of the HEK293T cell genome was statistically analyzed, such as... Figure 11 As shown.

[0179] Example 6

[0180] 6.1 Following the method described in Example 4, construct a plasmid targeting mouse mitochondrial DNA, including pGL3-PGL3-MTS-Talen-up-TALE inserted-Talen-down-riDddA tox-N-UGI-Rta,PGL3-MTS-Talen-up-TALEinserted-Talen-down-riDddA tox -C-UGI-Rta. The RVD sequence targeting mouse mitochondrial DNA is shown in Table 2.

[0181] 6.2 N2a cells were transfected using the mitochondrial genome cytosine base editor constructed above; when the cell density was 90%, the cells were separated into 12-well plates (JET Biotechnology).

[0182] 6.3 After culturing plated cells for 16-18 hours, transfect them when the cell confluence reaches 80%. The transfection system is as follows: 1.5 μg of mitochondrial genome cytosine base editor plasmid 1 (pGL3-MTS-Talen-up-TALE inserted-Talen-down-riDddAtox-N-UGI-Rta) and 1.5 μg of mitochondrial genome cytosine base editor plasmid 2 (pGL3-PGL3-MTS-Talen-up-TALE inserted-Talen-down-riDddA... tox -C-UGI-Rta) 1.5 μg, EZTrans transfection reagent (Liji Biotechnology) 7.5 μL.

[0183] 6.4 The specific transfection steps are as follows:

[0184] 6.4.1 Preparation of Reagent A: For each well of cells, dilute 3 μg of plasmid into 60 μL of serum-free, antibiotic-free, high-glucose DMEM medium and mix well.

[0185] 6.4.2 Preparation of reagent B: For each well of cells, dilute 7.5 μL of EZ Trans transfection reagent (EZ Trans: plasmid DNA = 2.5:1) into 60 μL of serum-free, antibiotic-free, high-glucose DMEM medium and vortex to mix.

[0186] 6.4.3 Let reagents A and B stand for 5 minutes, then add reagent B to reagent A and shake to mix.

[0187] 6.4.4 Let stand at room temperature for 15 min to form the EZ Trans-DNA complex. Evenly drop the prepared EZ Trans-DNA transfection complex into the corresponding wells of a 12-well plate, and gently shake the culture dish to disperse the EZ Trans-DNA complex evenly.

[0188] 6.4.5 Incubate at 37℃ in a 5% CO2 incubator for 6 hours, remove the culture medium containing the EZ Trans-DNA complex, replace with new culture medium, and incubate for 3 days.

[0189] 6.5 After culturing the transfected cells for 3 days, the cells were digested with trypsin to obtain cells. Further flow cytometry sorting was used to obtain GFP-positive and mCherry-positive cells (FITC fluorescence intensity, PE-Texas Red-A top 10%). On day 3, a portion of the collected cells underwent phenol-chloroform extraction for genomic DNA. The remaining portion was seeded back into 96-well plates, and genomic DNA was extracted on day 6 using the phenol-chloroform method.

[0190] 6.6 Design primers (Table 1, mus-mt-site1-for1; mus-mt-site1-rev1; mus-mt-site1-for2; mus-mt-site1-rev2; mus-mt-site2 / 5 / 6-for1; mus-mt-site2 / 5 / 6-rev1; mus-mt-site2 / 5 / 6-for2; mus-mt-site2 / 5 / 6-rev2; mus-mt-site2 / 5 / 6-for3; mus-mt-site2 / 5 / 6-rev3; mus-mt-si The target sequence was amplified using the following methods: te2 / 5 / 6-for4; mus-mt-site3 / 4-for1; mus-mt-site3 / 4-rev1; mus-mt-site3 / 4-for2; mus-mt-site3 / 4-rev2; mus-mt-site7 / 8-for1; mus-mt-site7 / 8-rev1; mus-mt-site7 / 8-for2; mus-mt-site7 / 8-rev2; mus-mt-site7 / 8-for3; mus-mt-site7 / 8-rev3) and then sent for analysis. Editing efficiency statistics are as follows: Figure 12 As shown.

[0191] In summary, this invention provides a novel double-stranded DNA cytosine deaminase and its derived nuclear and mitochondrial gene cytosine base editors, characterized by high-efficiency editing and the absence of sequence bias. The two editing systems are suitable for both nuclear and mitochondrial genome editing, solving the problems of small editing windows, low editing efficiency, and 5'-TC sequence bias.

[0192] The nucleotide or amino acid sequences involved in the examples are shown below (where the "*" at the end of the sequence indicates a stop codon):

[0193] SEQ ID NO:1: riDddA tox amino acid sequence from position 1 to 144

[0194] QYPCKEEMSAGAGESGRKTISLPEYDGTTTHGVLVLDDGTQIGFTSGNGDPRYTNYRNNGHVEQKSALYMRENNISNATVYHNNTNGTCGYCNTMTATFLPEGATLTVVPPENAVANNSRAIDYVKTYTGTSNDPKISPRYKGN

[0195] SEQ ID NO:2: riDddA tox DNA coding sequence corresponding to amino acids 1 to 144

[0196] CAGTACCCTTGCAAGGAAGAGATGTCTGCCGGAGCTGGCGAGAGCGGCAGAAAGACCATCAGCCTGCCTGAGTACGACGGCACCACAACACACGGCGTGCTGGTGCTGGATGATGGCACCCAGATCGGCTTCACCAGCGGAAATGGCGACCCTAGATACACCAACTACCGGAACAACGGCCACGTGGAACAGAAAAGCGCCCTGTACATGCGGGAA AACAACATCTCTAATGCCACAGTGTACCACAACAATACCAATGGAACCTGTGGCTACTGCAACACCATGACCGCCACCTTCCTGCCAGAGGGCGCTACACTGACCGTGGTCCCCCCGAGAACGCCGTGGCCAACAACAGCAGAGCCATCGACTACGTGAAGACCTACACAGGCACAAGCAACGACCCTAAGATCTCCCCTAGATATAAGGGCAAC

[0197] SEQ ID NO:3: Nuclear gene cytosine base editor plasmid 1 (pCMV-bpNLS-riDddA) tox -N-nSpCas9-2*UGI-bpNLS-Rta) amino acid sequence (where bpNLS is lowercase italic, riDddA) tox -N is uppercase, linker is lowercase, nSpCas9 is uppercase underscore, UGI is uppercase italic, and Rta is uppercase italic underscore.

[0198]

[0199]

[0200] SEQ ID NO:4: Nuclear gene cytosine base editor plasmid 1 (pCMV-bpNLS-riDddA) tox -N-nSpCas9-2*UGI-bpNLS-Rta) DNA sequence (where bpNLS is in lowercase italics, riDddA) tox -N is uppercase, linker is lowercase, nSpCas9 is uppercase underscore, UGI is uppercase italic, and Rta is uppercase italic underscore.

[0201]

[0202]

[0203]

[0204]

[0205]

[0206] SEQ ID NO:5: Nuclear gene cytosine base editor plasmid 2 (pCMV-bpNLS-riDddA) tox -C-SaKKH-nCas9-UGI-bpNLS-Rta) amino acid sequence (where bpNLS is lowercase italic, riDddA) tox -C is uppercase, linker is lowercase, SaKKH-nCas9 is uppercase underscore, UGI is uppercase italic, and Rta is uppercase italic underscore.

[0207]

[0208]

[0209] SEQ ID NO:6: Nuclear gene cytosine base editor plasmid 2 (pCMV-bpNLS-riDddA) tox -C-SaKKH-nCas9-UGI-bpNLS-Rta) DNA sequence (where bpNLS is in lowercase italics, riDddA) tox -C for uppercase, linker for lowercase, SaKKH-nCas9 for uppercase underscore, UGI for uppercase italic, Rta for uppercase italic underscore:

[0210]

[0211]

[0212]

[0213] SEQ ID NO:7: DNA sequence of nuclear gene cytosine base editor sgRNA expression plasmid 1 (pGL3-U6-SpCas9 sgRNAinsert site-scaffold-hPGK-GFP) (where scaffold is uppercase):

[0214] gagggcctatttcccatgattccttcatatttgcatatacgatacaaggctgttagagagataattggaattaatttgactgtaaacacaaagatattagtacaaaatacgtgacgtagaaagtaataatttcttgggtagtttgcagttttaaaattatgttttaaaatggactatcatatgcttaccgtaacttgaaagtatttcgatttcttggctttatatatcttgtggaaaggacgaaacaccgtgagaccgagagagggtctcaGTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGCtttttttaaagaattctcgacctcgagacaaatggcagtattcatccacaattttaaaagaaaaggggggattggggggtacagtgcaggggaaagaatagtagacataatagcaacagacatacaaactaaagaattacaaaaacaaattacaaaaattcaaaattttcgggtttattacagggacagcagagatccactttggccgcggctcgagggggttggggttgcgccttttccaaggcagccctgggtttgcgcagggacgcggctgctctgggcgtggttccgggaaacgcagcggcgccgaccctgggactcgcacattcttcacgtccgttcgcagcgtcacccggatcttcgccgctacccttgtgggccccccggcgacgcttcctgctccgcccctaagtcgggaaggttccttgcggttcgcggcgtgccggacgtgacaaacggaagccgcacgtctcactagtaccctcgcagacggacagcgccagggagcaatggcagcgcgccgaccgcgatgggctgtggccaatagcggctgctcagcagggcgcgccgagagcagcggccgggaaggggcggtgcgggaggcggggtgtggggcggtagtgtgggccctgttcctgcccgcgcggtgttccgcattctgcaagcctccggagcgcacgtcggcagtcggctccctcgttgaccgaatcaccgacctctctccccagggggatccatggtgagcttaccatggtgagcaagggcgaggagctgttcaccggggtggtgcccatcctggtcgagctggacggcgacgtaaacggccacaagttcagcgtgtccggcgagggcgagggcgatgccacctacggcaagctgaccctgaagttcatctgcaccaccggcaagctgcccgtgccctggcccaccctcgtgaccaccctgacctacggcgtgcagtgcttcagccgctaccccgaccacatgaagcagcacgacttcttcaagtccgccatgcccgaaggctacgtccaggagcgcaccatcttcttcaaggacgacggcaactacaagacccgcgccgaggtgaagttcgagggcgacaccctggtgaaccgcatcgagctgaagggcatcgacttcaaggaggacggcaacatcctggggcacaagctggagtacaactacaacagccacaacgtctatatcatggccgacaagcagaagaacggcatcaaggtgaacttcaagatccgccacaacatcgaggacggcagcgtgcagctcgccgaccactaccagcagaacacccccatcggcgacggccccgtgctgctgcccgacaaccactacctgagcacccagtccgccctgagcaaagaccccaacgagaagcgcgatcacatggtcctgctggagttcgtgaccgccgccgggatcactctcggcatggacgagctgtacaagtaa

[0215] SEQ ID NO:8: DNA sequence of sgRNA expression plasmid 2 (pGL3-U6-SaCas9 sgRNAinsert site-scaffold-hPGK-mCherry) (where scaffold is uppercase) for nuclear gene cytosine base editor

[0216]

[0217] SEQ ID NO:9: Mitochondrial gene cytosine base editor 1 (pGL3-MTS-Talen-up-TALE-insertsite-ccdB-Talen-down-riDddA) tox -N-UGI-Rta) amino acid sequence (where MTS is lowercase italic, Taylor-up and Taylor-down are lowercase underscores, ccdB is lowercase italic underscore, TALE insert site DNA sequence is "-", riDddAtox-N is uppercase, linker is lowercase, UGI is uppercase italic, and Rta is uppercase italic underscore)

[0218]

[0219] SEQ ID NO:10: Mitochondrial gene cytosine base editor 1 (pGL3-MTS-Talen-up-TALE-insert site-ccdB-Talen-down-riDddA) tox -N-UGI-Rta) DNA sequence (where MTS is lowercase italic, Taylor-up and Taylor-down are lowercase underlined, ccdB is lowercase italic underlined, TALE insert site is lowercase bold, riDddAtox-N is uppercase, linker is lowercase, UGI is uppercase italic, and Rta is uppercase italic underlined):

[0220]

[0221]

[0222] SEQ ID NO:11: Mitochondrial gene cytosine base editor 2 (pGL3-MTS-Talen-up-TALEinsert site-ccdB-Talen-down-riDddA) tox -C-UGI-Rta) amino acid sequence (where MTS is lowercase italic, Taylor-up and Taylor-down are lowercase underscores, ccdB is lowercase italic underscore, TALE insert site DNA sequence is "-", riDddAtox-C is uppercase, linker is lowercase, UGI is uppercase italic, and Rta is uppercase italic underscore)

[0223]

[0224] SEQ ID NO:12: Mitochondrial cytosine editor plasmid 2 (PGL3-MTS-Talen-up-TALE insertsite-ccdB-Talen-down-riDddA) tox -C-UGI-Rta) DNA sequence (where MTS is lowercase italic, Taylor-up and Taylor-down are lowercase underlined, ccdB is lowercase italic underlined, TALE insert site is lowercase bold, riDddAtox-N is uppercase, linker is lowercase, UGI is uppercase italic, and Rta is uppercase italic underlined)

[0225]

[0226]

[0227] SEQ ID NO:13: Amino acid sequence of the nuclear localization signal peptide

[0228] KRTADGSEFEPKKKRKV

[0229] SEQ ID NO:14: Amino acid sequence of mitochondrial localization signal peptide

[0230] MLGFVGRVAAAPASGALRRLTPSASLPPAQLLLRAAPTAVHPVRDYAAQTSESGGGGSPG

[0231] SEQ ID NO:15: Rta polypeptide amino acid sequence

[0232] RDSREGMFLPKPEAGSAISDVFEGREVCQPKRIRPFHPPGSPWANRPLPASLAPTPTGPVHEPVGSLTPAPVPQPLDPAPAVTPEASHLLEDPDEETSQAVKALREMADTVIPQKEEAAICGQMDLSHPPPRGHLDELTTTLESMTEDLNLDSPLTPELNEILDTFLNDECLLHAMHISTGLSIFDTSLF

[0233] References

[0234] 1.Urnov, FD, et al., Genome editing with engineered zinc fingernucleases. Nat Rev Genet, 2010.11(9):p.636-46.

[0235] 2.Christian,M.,et al.,Targeting DNA double-strand breaks with TALeffector nucleases.Genetics,2010.186(2):p.757-61.

[0236] 3.Lander,E.S.,The Heroes of CRISPR.Cell,2016.164(1-2):p.18-28.

[0237] 4.Garneau,J.E.,et al.,The CRISPR / Cas bacterial immune system cleavesbacteriophage and plasmid DNA.Nature,2010.468(7320):p.67-71.

[0238] 5.Komor,A.C.,et al.,Programmable editing of a target base in genomicDNA without double-stranded DNA cleavage.Nature,2016.533(7603):p.420-4.

[0239] 6.Gaudelli,N.M.,et al.,Programmable base editing of A*T to G*C ingenomic DNA without DNA cleavage.Nature,2017.551(7681):p.464-471.

[0240] 7.Andrews,R.M.,et al.,Reanalysis and revision of the Cambridgereference sequence for human mitochondrial DNA.Nat Genet,1999.23(2):p.147.

[0241] 8.Larsen,N.B.,M.Rasmussen,and L.J.Rasmussen,Nuclear and mitochondrialDNA repair:similar pathways?Mitochondrion,2005.5(2):p.89-108.

[0242] 9.Fontana,G.A.and H.L.Gahlon,Mechanisms of replication and repair inmitochondrial DNA deletion formation.Nucleic Acids Res,2020.48(20):p.11244-11258.

[0243] 10.Pfanner,N.,B.Warscheid,and N.Wiedemann,Mitochondrial proteins:frombiogenesis to functional networks.Nat Rev Mol Cell Biol,2019.20(5):p.267-284.

[0244] 11.Srivastava,S.and C.T.Moraes,Manipulating mitochondrial DNAheteroplasmy by a mitochondrially targeted restriction endonuclease.Hum MolGenet,2001.10(26):p.3093-9.

[0245] 12.Reddy,P.,et al.,Selective elimination of mitochondrial mutationsin the germline by genome editing.Cell,2015.161(3):p.459-469.

[0246] 13.Tanaka,M.,et al.,Gene therapy for mitochondrial disease bydelivering restriction endonuclease SmaI into mitochondria.J Biomed Sci,2002.9(6Pt 1):p.534-41.

[0247] 14.Alexeyev,M.F.,et al.,Selective elimination of mutant mitochondrialgenomes as therapeutic strategy for the treatment of NARP and MILSsyndromes.Gene Ther,2008.15(7):p.516-23.

[0248] 15.Urnov,F.D.,et al.,Highly efficient endogenous human genecorrection using designed zinc-finger nucleases.Nature,2005.435(7042):p.646-51.

[0249] 16.Peeva,V.,et al.,Linear mitochondrial DNA is rapidly degraded bycomponents of the replication machinery.Nat Commun,2018.9(1):p.1727.

[0250] 17.Gammage,P.A.,C.T.Moraes,and M.Minczuk,Mitochondrial GenomeEngineering:The Revolution May Not Be CRISPR-Ized.Trends Genet,2018.34(2):p.101-110.

[0251] 18.Mok,B.Y.,et al.,A bacterial cytidine deaminase toxin enablesCRISPR-free mitochondrial base editing.Nature,2020.583(7817):p.631-637.

[0252] 19.Mok,B.Y.,et al.,CRISPR-free base editors with enhanced activityand expanded targeting scope in mitochondrial and nuclear DNA.Nat Biotechnol,2022.

Claims

1. A cytosine base editing system, characterized in that, The cytosine base editing system includes a first fusion protein and a second fusion protein; The first fusion protein, from its N-terminus to its C-terminus, comprises, in sequence: a nuclear localization signal sequence fragment, a first polypeptide, a first nuclease, a uracil DNA glycosylase inhibitor fragment, a nuclear localization signal sequence fragment, and a transcription activation domain fragment; the second fusion protein, from its N-terminus to its C-terminus, comprises, in sequence: a nuclear localization signal sequence fragment, a second polypeptide, a second nuclease, a uracil DNA glycosylase inhibitor fragment, a nuclear localization signal sequence fragment, and a transcription activation domain fragment; the nuclear localization signal sequence fragment, the uracil DNA glycosylase inhibitor fragment, and the transcription activation domain fragment are linked to the first polypeptide, the second polypeptide, the first nuclease, and / or the second nuclease via linkers; Specifically, the first nuclease and the second nuclease break on the DNA strand; the first nuclease and the second nuclease have different recognition sites; and the first polypeptide and the second polypeptide, in their bound state, edit cytosine into thymine. The amino acid sequence of the first polypeptide is shown in SEQ ID NO: 17; the amino acid sequence of the second polypeptide is shown in SEQ ID NO:

30.

2. The cytosine base editing system as described in claim 1, characterized in that, The amino acid sequence of the first nuclease is shown in SEQ ID NO: 21; the amino acid sequence of the second nuclease is shown in SEQ ID NO: 32; or, the amino acid sequence of the first nuclease is shown in SEQ ID NO: 32; the amino acid sequence of the second nuclease is shown in SEQ ID NO:

21.

3. The cytosine base editing system as described in claim 2, characterized in that, The amino acid sequence of the nuclear localization signal sequence fragment is shown in SEQ ID NO: 13; and / or, the amino acid sequence of the uracil DNA glycosylation inhibitor fragment is shown in SEQ ID NO: 25; and / or, the amino acid sequence of the transcription activation domain fragment is shown in SEQ ID NO:

15.

4. The cytosine base editing system according to any one of claims 1-3, characterized in that, The cytosine base editing system also includes a guide RNA structure for guiding the cytosine base editing system to a target site.

5. The cytosine base editing system as described in claim 4, characterized in that, The backbone of the guide RNA structure is shown in SEQ ID NO:

34.

6. A cytosine base editing system, characterized in that, The cytosine base editing system includes a first fusion protein and a second fusion protein; The first fusion protein, from N-terminus to C-terminus, comprises, in sequence: a mitochondrial targeting sequence fragment, a nuclease, a first polypeptide, a uracil DNA glycosylase inhibitor fragment, and a transcription activation domain fragment; the second fusion protein, from N-terminus to C-terminus, comprises, in sequence: a mitochondrial targeting sequence fragment, a nuclease, a second polypeptide, a uracil DNA glycosylase inhibitor fragment, and a transcription activation domain fragment; the mitochondrial targeting sequence fragment, the uracil DNA glycosylase inhibitor fragment, and the transcription activation domain fragment are linked to the first polypeptide, the second polypeptide, and / or the nuclease via linkers; The nuclease causes breaks in the DNA strand; the first polypeptide and the second polypeptide, in a bound state, edit cytosine into thymine; The amino acid sequence of the first polypeptide is shown in SEQ ID NO: 17; the amino acid sequence of the second polypeptide is shown in SEQ ID NO:

30. The nuclease is a transcription activator-like effector nuclease; the cytosine base editing system is a cytosine base editing system targeting mitochondrial genes.

7. The cytosine base editing system as described in claim 6, characterized in that, The amino acid sequence of the left arm of the nuclease is shown in SEQ ID NO: 36, and the amino acid sequence of the right arm of the nuclease is shown in SEQ ID NO:

38.

8. The cytosine base editing system as described in claim 7, characterized in that, The amino acid sequence of the mitochondrial targeting sequence fragment is shown in SEQ ID NO: 14; and / or, the amino acid sequence of the uracil DNA glycosylation inhibitor fragment is shown in SEQ ID NO: 25; and / or, the amino acid sequence of the transcription activation domain fragment is shown in SEQ ID NO:

15.

9. The cytosine base editing system according to any one of claims 6-8, characterized in that, The cytosine base editing system also includes a guide RNA structure for guiding the cytosine base editing system to a target site.

10. The cytosine base editing system as described in claim 9, characterized in that, The backbone of the guide RNA structure is shown in SEQ ID NO:

34.

11. A polynucleotide, characterized in that, The polynucleotide encodes the cytosine base editing system as described in any one of claims 1-10.

12. A construct, characterized in that, The construct comprises the polynucleotide as described in claim 11.

13. The construct as claimed in claim 12, characterized in that, The plasmid backbone of the construct is a eukaryotic expression vector plasmid backbone; and / or, the construct contains the nucleotide sequence shown in SEQ ID NO: 4 and the nucleotide sequence shown in SEQ ID NO: 6; or, the construct contains the nucleotide sequence shown in SEQ ID NO: 10 and the nucleotide sequence shown in SEQ ID NO:

12.

14. The construct as claimed in claim 13, characterized in that, The plasmid backbone is selected from one or more of pCMV, pSV2 and pGL3; and / or, the construct further comprises the nucleotide sequences shown in SEQ ID NO: 7 and SEQ ID NO:

8.

15. An expression system, characterized in that, The expression system comprises a construct as described in any one of claims 12-14; the host cell of the expression system is selected from eukaryotic cells and prokaryotic cells.

16. The expression system as described in claim 15, characterized in that, The eukaryotic cells were selected from mouse cells and human cells.

17. The expression system as described in claim 16, characterized in that, The eukaryotic cells were selected from mouse neuroma cells and human embryonic kidney cancer cells.

18. The expression system as described in claim 17, characterized in that, The eukaryotic cells were selected from N2a cells and HEK293T cells.

19. A composition for nuclear genome base editing or mitochondrial gene base editing, characterized in that, The composition comprises the cytosine base editing system as described in any one of claims 1-10, the polynucleotide as described in claim 11, the construct as described in any one of claims 12-14, or the expression system as described in any one of claims 15-18.

20. The use of a cytosine base editing system as described in any one of claims 1-10, a polynucleotide as described in claim 11, a construct as described in any one of claims 12-14, or an expression system as described in any one of claims 15-18 in the construction of animal models or crop breeding.

21. The use of a cytosine base editing system as described in any one of claims 1-10, a polynucleotide as described in claim 11, a construct as described in any one of claims 12-14, or an expression system as described in any one of claims 15-18 in nuclear genome base editing or mitochondrial gene base editing; The application is for non-therapeutic purposes.

22. A base editing method, characterized in that, The base editing method comprises administering to target cells the cytosine base editing system as described in any one of claims 1-10, the polynucleotide as described in claim 11, or the construct as described in any one of claims 12-14; or comprises contacting the target cells with the expression system as described in any one of claims 15-18. The base editing method described is not for therapeutic purposes.

23. The base editing method as described in claim 22, characterized in that, The base editing method is performed in vitro; The target cells are eukaryotic cells.

Citation Information

Patent Citations

  • Fusion protein with cytosine deamination function and application thereof

    CN118599013A