A cytosine base editor based on a modified apoBEC3A from cynomolgus monkey
By modifying the APOBEC3A cytosine deaminase of cynomolgus monkeys, a BE4-mA3A-delSVR single-base editor was constructed, which solved the problems of insufficient editing efficiency and specificity in the CBE system, and achieved a highly efficient gene editing effect, which is particularly suitable for the treatment of genetic diseases in eukaryotic cells.
Patent Information
- Application Number
- CN202111217333.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-19
- Publication Date
- 2026-02-24
- Estimated Expiration
- 2041-10-19
AI Technical Summary
The existing CBE system has insufficient editing efficiency and specificity of cytosine deaminase, making it difficult to effectively treat single-gene genetic diseases.
Using the APOBEC3A cytosine deaminase mutant from cynomolgus monkeys, amino acids 27-29 were deleted. Combined with nCas9, a uracil DNA glycosylase inhibitor, and sgRNA, a BE4-mA3A-delSVR single-base editor was constructed, enabling efficient C→T editing via the R-loop region.
It improves editing efficiency and product purity, reduces off-target efficiency, and provides a more efficient and specific gene editing tool suitable for gene therapy in eukaryotic cells.
Smart Images

Figure CN115992123B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a cytosine base editor based on APOBEC3A modified from cynomolgus monkeys, belonging to the field of genetic engineering technology. Background Technology
[0002] Gene editing technology is rapidly developing, enabling the study of gene function, pathogenesis of genetic diseases, development of new drugs, and applications in gene therapy and crop improvement by modifying specific genes. In recent years, various derivative technologies based on the CRISPR-Cas system have been widely applied in life sciences and medicine, with single-base editors (BEs) becoming a crucial component of gene editing technology. BEs are designed based on the CRISPR-Cas9 system. Wild-type Cas9 proteins are modified and linked to cytosine deaminase or adenine deaminase, then guided by sgRNA (small guide RNA), allowing direct editing of single bases without causing DNA double-strand breaks (DSBs). BEs are mainly divided into two categories: cytosine base editors (CBEs) and adenine base editors (ABEs). In 2016, the CBE system developed by David R. Liu's team used cytosine deaminase to deaminate the corresponding cytosine (C) in the non-complementary strand into uracil (U). During DNA replication or repair, U is recognized as thymine (T), while the corresponding guanine (G) on the complementary strand is converted into adenine (A), ultimately achieving the conversion from C>T on the non-complementary strand and G>A on the complementary strand. In 2017, David R. Liu's team developed the ABE system, which operates on a similar principle to the CBE system, except that the cytosine deaminase is replaced with an adenine deaminase. This system can perform the editing of A>G on the non-complementary strand and T>C on the complementary strand, further expanding the types of single-base editors.
[0003] The Cas9 protein in the BE system can be further optimized. One type is dCas9 (catalytically dead Cas9), which binds to the target gene but does not cleave it, and the other is nCas9 (Cas9 nickase), which has single-stranded DNA nicking activity. Neither type of Cas9 protein produces DBS, thus avoiding mismatches in non-homologous end-joining (NHEJ) and eliminating the inefficiency of homology-directed repair (HDR). Subsequently, a uracil glycosylase inhibitor (UGI) was added to BE to inhibit the excision of the intermediate product U, improving the efficiency of C>T editing on the DNA strand. Currently, the BE system is developing rapidly. The CBE system has been optimized to the fourth generation BE4, and the ABE system has been optimized to version ABE7.10. Some dual-base editors that combine the functions of CBE and ABE systems have even been developed, enabling simultaneous C>T and A>G conversion at the same target site.
[0004] The traditional CBE system uses rat Apobec1 (rA1) as the cytosine deaminase. In 2018, Jason M. Gehrke et al. attempted to replace rA1 in the third-generation CBE system (BE3) with a modified human APOBEC3A (eA3A). They then used BE3 carrying the hA3A-N57G mutation in a study on the treatment of β-thalassemia, achieving nearly 40 times higher editing precision compared to BE3. However, in HBB... -28 The editing efficiency of single-base sites is only around 30%. Furthermore, since most pathogenic mutations in single-gene genetic diseases are point mutations, and the BE system can edit individual bases in the DNA sequence, it can correct the pathogenic mutations in genetic patients, achieving "molecular surgery" and eradicating the genetic disease at its root cause. It is also known that approximately 75,000 human genome mutation sites are associated with genetic diseases, of which an estimated 50% are potential therapeutic targets for the CBE system. Therefore, developing and optimizing more efficient and accurate single-base editors is of great significance for the treatment of single-gene genetic diseases. Summary of the Invention
[0005] The cytosine deaminases used in existing CBE systems are mostly rat Apobec1 (rA1), with a few being human APOBEC3A (hA3A) and APOBEC3G (hA3G). There is currently no application of cytosine deaminases from non-human primates such as cynomolgus monkeys in CBE systems.
[0006] To address the aforementioned problems, the present invention aims to provide a gene editing tool with high editing efficiency, high product purity, and strong editing window specificity.
[0007] The first objective of this invention is to provide a cytosine deaminase mutant derived from cynomolgus monkeys, wherein the mutant is based on the parent enzyme with the amino acid sequence shown in SEQ ID NO:1, by deleting three consecutive amino acids from position 27 to 29. Specifically, serine at position 27, valine at position 28, and arginine at position 29 are deleted.
[0008] In one embodiment, the amino acid sequence of the mutant is shown in SEQ ID NO:2.
[0009] A second objective of this invention is to provide a gene encoding the aforementioned cytosine deaminase mutant.
[0010] In one embodiment, the nucleotide sequence of the gene is shown in SEQ ID NO:3.
[0011] A third object of the present invention is to provide an expression cassette containing the gene encoding the above-described cytosine deaminase mutant.
[0012] In one embodiment, the expression cassette further comprises a promoter, nCas9 (D10A), a uracil DNA glycosylation inhibitor (UGI), a nuclear localization sequence (NLS), and a termination sequence.
[0013] In one embodiment, the promoter initiates the expression of a gene encoding the cytosine deaminase mutant, which, in the following order of connection, consists of the promoter, the cytosine deaminase mutant, nCas9 (D10A), a uracil DNA glycosylation inhibitor (UGI), a nuclear localization sequence (NLS), and a termination sequence.
[0014] In one embodiment, the promoter includes, but is not limited to, the CMV promoter.
[0015] In one embodiment, the nucleotide sequence of nCas9(D10A) is shown in SEQ ID NO:4, the nucleotide sequence of UGI is shown in SEQ ID NO:5, the nucleotide sequence of NLS is shown in SEQ ID NO:6, and the nucleotide sequence of the termination sequence is shown in SEQ ID NO:7.
[0016] The fourth objective of this invention is to provide a CBE single-base editor, BE4-mA3A-delSVR, based on a modified cytosine deaminase (APOBEC3A) from cynomolgus monkeys. This CBE single-base editing system comprises four parts: a first part is a transfection efficiency indicator containing red fluorescent protein (dTomato) and a promoter; a second part is an sgRNA transcription unit containing a carrying frame for inserting the sgRNA sequence and its promoter; a third part contains the aforementioned expression cassette; and a fourth part contains the high-copy origin of replication (ori) from the coliform factor (colE1) and the ampicillin resistance selection gene AmpR.
[0017] In one embodiment, the promoter includes a CMV promoter and a U6 promoter.
[0018] In one embodiment, the nucleotide sequence of the red fluorescent protein is shown in SEQ ID NO:8.
[0019] In one embodiment, the nucleotide sequence of the sgRNA transcription unit is shown in SEQ ID NO:9.
[0020] The fifth objective of this invention is to provide the application of the aforementioned CBE single-base editor in the field of gene editing.
[0021] In one embodiment of the present invention, an sgRNA sequence is designed according to the target gene and inserted into the sgRNA transcription unit of the above-mentioned CBE single base editing system to obtain a CBE single base editing system with specific target gene. The CBE single base editing system is introduced into the recipient cell to achieve the mutation of the target base C to T.
[0022] In one embodiment of the present invention, the cell is a eukaryotic cell.
[0023] In one embodiment of the present invention, the eukaryotic cell is a mammalian cell.
[0024] In one embodiment of the present invention, the mammalian cells include HEK293T cells.
[0025] The sixth objective of this invention is to provide a method for constructing the CBE single-base editing system BE4-mA3A-delSVR, the specific construction steps of which are as follows:
[0026] (1) The red fluorescent protein dTomato gene and its CMV enhancer and promoter were amplified using pSpCas9(BB)-2A-dTomato(PX458) plasmid as a template, and the vector backbone Part 1 was obtained.
[0027] (2) Using BE3-rA1 plasmid as a template, the carrying frame with inserted sgRNA sequence and its U6 promoter were amplified to obtain vector backbone Part 2;
[0028] (3) The nucleotide sequence of APOBEC3A (mA3A-B5) was amplified using cynomolgus monkey cDNA as a template to obtain the vector backbone Part 3;
[0029] (4) Using plasmid BE4-rA1 as a template, amplify to obtain a vector backbone including nCas9(D10A), UGI, NLS, termination sequence, CMV enhancer and promoter gene sequence Part 4;
[0030] (5) Using plasmid BE4-rA1 as a template, the vector backbone of the high copy origin of replication of colE1 and the ampicillin resistance selection gene and its ampicillin promoter sequence was amplified.
[0031] (6) Connect the five fragments in the order of Part1, Part2, Part3, Part4 and Part5 to obtain the vector plasmid BE4-mA3A-B5.
[0032] (7) Using plasmid pcDNA3.1-mA3A-GFP as a template, PCR amplification was performed using specific primers to obtain the mA3A sequence with the SVR sequence deleted (i.e., mA3A-delSVR, see...). Figure 5 );
[0033] (8) The mA3A sequence of the SVR sequence deleted in step (7) is ligated with the BE4-mA3A-B5 vector of the cynomolgus monkey whose APOBEC3A (i.e. mA3A) has been removed by double digestion with BamHI and SmaI to obtain the vector plasmid BE4-mA3A-delSVR.
[0034] (9) Design an sgRNA sequence that binds to the target gene and insert it into the carrying frame in step (2) to obtain a single base editor that can target the corresponding sites of different target genes.
[0035] In one embodiment, the pSpCas9(BB)-2A-dTomato(PX458) plasmid is constructed by replacing GFP with dTomato in the pSpCas9(BB)-2A-GFP(PX458) plasmid, wherein the nucleotide sequence of dTomato is shown in SEQ ID NO:8.
[0036] In one embodiment, the plasmid pcDNA3.1-mA3A-GFP is constructed by ligating the mA3A fragment and the GFP fragment to a vector that has been double-digested with EcoRI and XbaI using pcDNA3.1-2xFLAG-SREBP-2 plasmid template.
[0037] In one embodiment, the nucleotide sequence of the mA3A fragment (NCBI reference sequence) is shown in SEQ ID NO:12.
[0038] In one embodiment, the nucleotide sequence of the mA3A-delSVR fragment is shown in SEQ ID NO:3.
[0039] In one embodiment, the nucleotide sequence of the GFP fragment is shown in SEQ ID NO:10.
[0040] This invention also provides the application of the above-mentioned mutants in the field of gene editing.
[0041] This invention also provides the application of the above expression cassette in the field of gene editing.
[0042] The beneficial effects of this invention are:
[0043] 1. The single-base editor BE4-mA3A-delSVR constructed in this invention contains an sgRNA transcription unit. When using it for gene editing, the BE4-mA3A-delSVR plasmid is transfected into cells. The sgRNA carrying the target gene on the plasmid guides the fusion protein to bind to the target gene DNA region through the base complementary pairing principle. Then, the mA3A of the SVR sequence is deleted and binds to the single-stranded DNA (ssDNA) in the R-loop region formed by the nCas9 protein, sgRNA and genomic DNA. Cytosine (C) in a certain range (20bp protospacer sequence) of this ssDNA is deaminated to uracil (U). Then, U is converted to thymine (T) through DNA replication or repair, and finally the direct replacement of CG base pairs to TA base pairs is achieved. Furthermore, since nCas9 only has single-stranded DNA cleavage enzyme activity, it will not cleave the double-stranded DNA of the target gene between 2-3 bases upstream of the PAM (Protospacer Adjacent Motif) sequence, thus preventing the formation of DSB and greatly reducing off-target efficiency.
[0044] 2. This invention knocked out amino acids 27-29 of the cytosine deaminase mA3A from cynomolgus monkeys, successfully constructing a cytosine base editor, BE4-mA3A-delSVR, which exhibits superior editing efficiency, product purity, and editing specificity compared to the original BE4-rA1. This editor was successfully used for C>T editing of the target sites in the site3, RNF2, EMX1, HEK2, and HEK4 genes in HEK293T cells. The single-base editor constructed in this invention specifically performs C>T conversion at C5 and C6 positions, demonstrating high editing efficiency and product purity. This fully demonstrates that the modified mA3A-delSVR is a superior cytosine deaminase for optimizing CBE gene editing systems compared to rA1 and even the original mA3A. This provides a more flexible and controllable prototype, as well as new ideas and directions for further optimization and modification of CBE systems, enriching the gene editing toolkit and offering a potential tool for gene therapy of genetic diseases. (Note: site3, HEK2, and HEK4 are names used in published articles. The gene name for site3 is LINC01509, the gene name for HEK2 is AC114971.1, and the gene name for HEK4 is DNMT3B.) Attached Figure Description
[0045] Figure 1 : Plasmid map of BE4-mA3A-B5, a cytosine base editor;
[0046] Figure 2 : Plasmid map of the overexpression plasmid pcDNA3.1-mA3A-GFP.
[0047] Figure 3 Plasmid map of BE4-mA3A-delSVR, a cytosine base editor;
[0048] Figure 4 Electrophoresis image of BE4-mA3A-B5 (~11kb) plasmid digested with BamHI and SmaI;
[0049] Figure 5 Electrophoresis images of the reference sequence (mA3A) of cynomolgus monkey APOBEC3A and its mutant with the SVR sequence deleted (mA3A-delSVR) from NCBI.
[0050] Figure 6 Sanger sequencing was used to verify the editing efficiency of different pyrimidine single-base editors (BE3-rA1, BE4-rA1, BE4-mA3A and BE4-mA3A-delSVR) at the site3 gene target site (n=3).
[0051] Figure 7Sanger sequencing was used to verify the editing efficiency of different pyrimidine single-base editors (BE3-rA1, BE4-rA1, BE4-mA3A and BE4-mA3A-delSVR) at the RNF2 gene target site (n=3).
[0052] Figure 8 Sanger sequencing was used to verify the editing efficiency of different pyrimidine single-base editors (BE3-rA1, BE4-rA1, BE4-mA3A and BE4-mA3A-delSVR) at the EMX1 gene target site (n=3).
[0053] Figure 9 Sanger sequencing was used to verify the editing efficiency of different pyrimidine single-base editors (BE3-rA1, BE4-rA1, BE4-mA3A and BE4-mA3A-delSVR) at the HEK2 gene target site (n=3).
[0054] Figure 10 Sanger sequencing was used to verify the editing efficiency of different pyrimidine single-base editors (BE3-rA1, BE4-rA1, BE4-mA3A and BE4-mA3A-delSVR) at the HEK4 gene target site (n=3).
[0055] Figure 11 Comparison of C>T editing efficiency of BE3-rA1, BE4-rA1, BE4-mA3A and BE4-mA3A-delSVR editors at the target sites of the site3, RNF2, EMX1, HEK2 and HEK4 genes (n=3);
[0056] Figure 12 Comparison of the distribution of products (T, G, A) at site3-C5, RNF2-C6, EMX1-C5, EMX1-C6, HEK2-C6, and HEK4-C5 sites using the BE3-rA1, BE4-rA1, BE4-mA3A, and BE4-mA3A-delSVR editor (n=3);
[0057] Figure 13 The average editing efficiency of Cs at 5 gene loci was analyzed based on the position of cytosine C in the protospacer (PAM is located at positions 21-23). Different shapes and gray levels represent different editors (n=3,6,9).
[0058] Figure 14 Comparison of average editing efficiency of BE4-mA3A and BE4-mA3A-delSVR editors for Cs at 5 gene loci (n=3,6,9). Detailed Implementation
[0059] The plasmids involved in the following examples:
[0060] BE4 plasmid: Addgene Plasmid#100802.
[0061] pSpCas9(BB)-2A-dTomato(PX458) plasmid: Based on the pSpCas9(BB)-2A-GFP(PX458)(Addgene Plasmid#48138) plasmid, our laboratory chemically synthesized dTomato (SEQ ID NO:8) to replace GFP and constructed it.
[0062] BE3-rA1 plasmid: provided by Professor Yang Hui, which has been reported in the article Off-target RNA mutation induced by DNAbase editing and its elimination by mutagenesis. In this invention, BE3 in the article is BE3-rA1.
[0063] BE4-rA1 plasmid: Constructed in our laboratory, based on BE4 plasmid (Addgene Plasmid#100802), with the chemically synthesized dTomato gene sequence (SEQ ID NO:8) and the U6promoter+sgRNA scaffold+CMV enhancer+CMVpromoter fragment of BE3-rA1 plasmid sequentially inserted into the NotⅠ restriction site.
[0064] pcDNA3.1-mA3A-GFP plasmid: Constructed in our laboratory, using pcDNA3.1-2xFLAG-SREBP-2 (Addgene Plasmid#26807) plasmid as the backbone, the vector was double-digested with EcoRI and XbaI, and then chemically synthesized mA3A fragments (reference sequence NP_001306288.1 on NCBI, SEQ ID NO:12) and GFP fragments (SEQ ID NO:10) were ligated to the digested vector. See the plasmid map below. Figure 2 .
[0065] BE4-mA3A-B5 plasmid: Constructed in our laboratory; specific construction method is described in Example 1. Plasmid map can be found here. Figure 1 .
[0066] The cynomolgus monkeys involved in this invention were purchased from Guangzhou Xiangguan Biotechnology Co., Ltd. (Production License Number: SCXK(Yue)2018 - 0043). Before the experiment, their health conditions were confirmed to be good through records and veterinary examinations, and the animal facilities met the national experimental animal standards (GB14925 - 2010). Subsequently, cynomolgus monkeys that had developed hypercholesterolemia after 19 months of high - sugar and high - fat diet treatment and could spontaneously recover to normal blood lipid levels were selected for the construction of the cytosine base editor (BE4 - mA3A - B5) of this invention.
[0067] Example 1: Construction of cytosine base editor (BE4 - mA3A - B5)
[0068] (1) Obtaining the target gene (mA3A - B5)
[0069] In this invention, blood samples of cynomolgus monkeys were obtained using Paxgene tubes, and the RNA of the cynomolgus monkeys was extracted using TIANGEN's blood RNA extraction kit. The cDNA sequence of the cynomolgus monkeys was obtained using Takara's reverse transcription kit. Then, using its cDNA sequence as a template, the mA3A - B5 gene fragment (i.e., Part3, SEQ ID NO: 11) was amplified by PCR using the primer pair 5’ - TATAGGGAGAGCCGCCACCATGGAAGCCAGCCCAG - 3’ and 5’ - ACCAGAAGAACCACCAGAGTTTCCCTGATTCTGG - 3’. The PCR product was purified using TIANGEN's agarose gel DNA extraction kit. The PCR reaction system is shown in Table 1.
[0070] Table 1 PCR reaction system
[0071]
[0072] The reaction procedure is as follows: pre - denaturation at 95°C for 3 min; 95°C for 15 s, 60°C for 15 s, 72°C for 25 s, for 35 cycles; extension at 72°C for 5 min, cooling to 4°C, finally obtaining mA3A - B5 (Part3).
[0073] (2) Preparation of linearized plasmid vector (BE4)
[0074] Using the BE4 - rA1 plasmid as a template, PCR amplification was carried out using the primer pair 5’ - TCTGGTGGTTCTTCTGGTGGTTCTAGCGGC - 3’ and 5’ - GGTGGCGGCTCTCCCTATAGTGAGTCGTAT - 3’ to obtain the PCR product Part4, including nCas9(D10A), UGI, NLS, termination sequence, and CMV enhancer and promoter gene sequences;
[0075] Using the BE4-rA1 plasmid as a template, PCR amplification was performed using primer pairs 5'-CGGTGGCTTCGATAGCCCTACAGTTGCCT-3' and 5'-CTACTAGGACAGAATAGGCAACTGTAGGGC-3' to obtain PCR product Part5, which includes the colE1 high copy origin of replication and the ampicillin resistance selection gene and its ampicillin promoter sequence.
[0076] PCR products Part 4 and Part 5 were purified using TIANGEN's agarose gel DNA extraction kit.
[0077] (3) Acquisition of the red fluorescent protein (dTomato) gene
[0078] Using the pSpCas9(BB)-2A-dTomato(PX458) plasmid as a template, PCR amplification was performed using primers 5'-TAGAGATCCGCGCCACCATGGTGAGC-3' and 5'-GAAGGCACAGTTACTTGTACAGCTCG-3' to obtain the red fluorescent protein dTomato gene, along with its CMV enhancer and promoter, forming the vector backbone Part 1. Part 1 of the PCR product was then purified using a TIANGEN agarose gel DNA extraction kit.
[0079] (4) Acquisition of the sgRNA carrying framework and the U6 promoter gene
[0080] Using the BE3-rA1 plasmid as a template, PCR amplification was performed using the following primers: 5'-GCTCACATGTGAGGGCCTATTTCCC-3' and 5'-ATAGGCCCTCACATGTGAGCAAAAG-3', to obtain the carrier framework and its U6 promoter for inserting the sgRNA sequence, i.e., vector backbone Part 2. The PCR product Part 2 was then purified using TIANGEN's agarose gel DNA extraction kit.
[0081] (5) Construction of the vector plasmid BE4-mA3A-B5 (homological recombination method)
[0082] The five PCR fragments Part1, Part2, Part3, Part4 and Part5 purified in steps (1) to (4) were ligated using the MultiS One Step Cloning Kit from Novizan to obtain ligation products. These products were then transformed into Escherichia coli DH5α and plated on LB plates containing 0.05% Amp (100 μg / mL ampicillin) resistance and incubated overnight at 37°C with the plates inverted.
[0083] Three colonies were selected from each cloning plate and inoculated into liquid LB medium containing 0.1% ampicillin (100 μg / mL) resistance. The culture was then incubated for at least 8 hours, and the bacterial culture was sent to Genewiz for Sanger sequencing verification. The successfully inserted cloning vector was amplified, and the vector plasmid was obtained using an endotoxin-free plasmid extraction kit from Kangwei Century and named BE4-mA3A-B5. It was stored at -20℃ for later use. The BE4-mA3A-B5 plasmid map is shown below. Figure 1 As shown.
[0084] Example 2: Construction of a cytosine single-base editor based on cynomolgus monkey APOBEC3A (mA3A) and its mutants
[0085] (1) Preparation of linearized plasmid vector (BE4-mA3A-B5)
[0086] The BE4-mA3A-B5 vector obtained in Example 1 was double-digested with the rapid digestive enzymes BamHI and SmaI from Takara to remove the cynomolgus macaque APOBEC3A (i.e., the mA3A-B5 gene fragment with the nucleotide sequence shown in SEQ ID NO:11) from the vector, so that the cytosine deaminase could be replaced with the mA3A gene fragment with the nucleotide sequence shown in SEQ ID NO:12. The enzyme digestion system and reaction conditions are shown in Table 2. Then, the digested product was purified using a standard DNA product purification kit from TIANGEN to obtain a linearized vector (see Table 2). Figure 4 ).
[0087] Table 2. BamHI and SmaI double digestion system
[0088]
[0089] (2) Obtaining the target gene (mA3A and mA3A-delSVR)
[0090] Using the plasmid pcDNA3.1-mA3A-GFP overexpressing mA3A as a template, the mA3A gene sequence was obtained by PCR amplification using primers 5'-CCCAGACACTTGATGGATCCAAACACGTTCACTTTC-3' and 5'-GGACTCTGAGGTCCCGGGAGTCTCGCTGCC-3'. The mA3A-delSVR gene fragment was obtained by PCR amplification using the plasmid pcDNA3.1-mA3A-GFP overexpressing mA3A as a template (see [link to original text]). Figure 5 The PCR products were then purified using TIANGEN's agarose gel DNA extraction kit. During PCR, homologous arms complementary to the vector cut were added to both ends of the primers to facilitate cloning into the vector. The PCR reaction system is shown in Table 3.
[0091] Table 3 PCR reaction system
[0092]
[0093] The reaction procedure was as follows: pre-denaturation at 95℃ for 3 min; 95℃ for 15 s, 60℃ for 15 s, 72℃ for 25 s, for 35 cycles; extension at 72℃ for 5 min, followed by cooling to 4℃, to finally obtain mA3A and mA3A-delSVR gene fragments.
[0094] (3) Construction of the vector plasmid BE4-mA3A-delSVR (homological recombination method)
[0095] The mA3A or mA3A-delSVR gene fragments purified in step (2) were ligated to the linearized vector BE4-mA3A-B5 purified in step (1) using the ClonExpress II One Step Cloning Kit from Novizan. The ligation products were transformed into Escherichia coli DH5α and plated on LB plates containing 0.05% Amp (100 μg / mL ampicillin) resistance and incubated overnight at 37°C with the plates inverted.
[0096] Three colonies from each cloning plate were inoculated into liquid LB medium containing 0.1% ampicillin (100 μg / mL) and cultured for at least 8 hours. The bacterial culture was then sent to Genewiz for Sanger sequencing verification. The successfully inserted cloning vectors were amplified, and the endotoxin-free plasmid extraction kit from Kangwei Century was used to obtain the vector plasmids, which were named BE4-mA3A and BE4-mA3A-delSVR and stored at -20℃ for later use. The BE4-mA3A-delSVR plasmid map is shown below. Figure 3 As shown.
[0097] Example 3: Application method of cytosine base editor BE4-mA3A-delSVR
[0098] (1) Insertion of sgRNA into a specific target gene
[0099] Since the single-base editor is based on the CRISPR / Cas9 system, its targeting specificity consists of two parts: one is the complementary base pairing between the sgRNA and the target DNA sequence, and the other is determined by the Cas9 protein and the short DNA sequence located at the 3' end of the target DNA sequence. The target DNA sequence is called the protospacer (in bacterial innate immune defense, foreign DNA fragments are called protospacers), and the short DNA sequence located at the 3' end of the target DNA sequence is called the PAM (protospacer adjacent motif).
[0100] To better compare with existing studies, the sgRNA used is a sequence that has been reported in the literature, as shown in Table 4.
[0101] Table 4: sgRNA and its PAM sequence listing
[0102]
[0103] 1) Linearized vectors were obtained by digesting the original BE3-rA1 plasmid, BE4-rA1 plasmid and the cytosine base editor (BE4-mA3A and BE4-mA3A-delSVR) constructed in Example 2 with BbsI. The digestion system is shown in Table 5.
[0104] 2) Synthesize cloning primers for sgRNAs targeting site3, RNF2, EMX1, HEK2, and HEK4 as shown in Table 6, respectively, and self-ligate the cloning primers into double-stranded oligonucleotide fragments by heat shock annealing (sgRNA self-ligation system is shown in Table 7). Ligate the double-stranded oligonucleotide fragments to the linearized vector from step 1) (ligation system is shown in Table 8). Transform the ligated vector into competent E. coli cells (DH5α), extract plasmids and sequence them for identification, and finally obtain a series of cytosine single-base editors targeting site3, RNF2, EMX1, HEK2, and HEK4 genes.
[0105] Table 5. BBSI digestion system
[0106]
[0107] Table 6 sgRNA Cloning Primer Table
[0108]
[0109] Table 7 sgRNA self-ligation system
[0110]
[0111] Table 8 Connection System
[0112]
[0113] (2) Cell transfection
[0114] Preparation of high-glucose DMEM complete medium: The high-glucose DMEM medium contains 10% fetal bovine serum (FBS) and 1% triple antibiotics (penicillin-streptomycin-gentamicin).
[0115] 1) Cell culture: Take the frozen HEK293T cells out of the liquid nitrogen tank and thaw them quickly in a 37°C water bath. Add the thawed cell suspension to 10 mL of high-glucose DMEM complete medium, centrifuge to collect the cell pellet, resuspend the cells in high-glucose DMEM complete medium and add them to a cell culture dish containing 10 mL of high-glucose DMEM complete medium, and place it in a 37°C constant temperature cell culture incubator containing 5% CO2 for culture.
[0116] 2) Cell plating: Once HEK293T cells have confluently grown into the culture dish, digest the cells with 0.25% trypsin for 1 min, terminate the digestion with high-glucose DMEM complete medium, centrifuge at 1500 rpm for 2 min, discard the supernatant, resuspend the cells in 1 mL of high-glucose DMEM complete medium, and plate them at 1×10⁻⁶ cells / mL. 5 Cells / well density: Cells were seeded in 6-well plates (each well containing 2 mL of high-glucose DMEM complete medium);
[0117] 3) Cell transfection: In a 1.5 mL EP tube, dilute 3.75 μL of transfection reagent Lipofectamine 3000 with 125 μL of Opti-MEM medium and mix thoroughly to obtain mixture A; In another 1.5 mL EP tube, dilute 2.5 μg of the cytosine monobase editor targeting site3, RNF2, EMX1, HEK2 or HEK4 genes constructed in step (1) with 125 μL of Opti-MEM medium and add 5 μL of LP3000 auxiliary transfection reagent and mix thoroughly to obtain mixture B; Mix mixture A and B in a 1:1 ratio and incubate at room temperature for 15 min to obtain a mixture; Add the mixture to the 6-well plate in which cells were just seeded in step 2), mix gently to avoid damaging the cells; Then place the 6-well plate back into an incubator at 37°C and 5% CO2 and incubate the cells for 72 h for subsequent experiments. During this period, observe and photograph the cells every 24 h using an inverted microscope.
[0118] (3) Flow sorting
[0119] Cells transfected for 72 hours were first digested with 0.25% trypsin and digestion was terminated with high-glucose DMEM complete medium. Cell pellet was collected by centrifugation at 1500 rpm for 2 min. Cells were then resuspended and washed with 2 mL of PBS containing 2% fetal bovine serum (FBS). After centrifugation at 1500 rpm for 2 min, the washing was repeated twice. Cells were then resuspended with 1 mL of PBS containing 2% FBS. The cell suspension was then collected through a flow cytometry tube with a filter membrane to remove larger cell debris and prevent clogging of the instrument during sorting.
[0120] The prepared cell suspension and corresponding collection tubes containing 2 mL of high-glucose DMEM complete medium were placed on ice and then subjected to flow cytometry sorting. Because the cytosine base editor constructed in this invention carries red fluorescent protein (dTomato), dTomato-positive cells were directly collected and sorted based on the excitation wavelength (554 nm) and emission wavelength (581 nm). The BE3-rA1 plasmid carries green fluorescent protein (GFP), and GFP-positive cells were directly collected and sorted based on its excitation wavelength (488 nm) and emission wavelength (507 nm). At least 100,000 cells were obtained from each sample.
[0121] (4) Sanger sequencing validation
[0122] The dTomato-positive or GFP-positive cells collected in step (3) were centrifuged at 1500 rpm for 2 min, and the precipitate was collected. Cell DNA was then extracted using the Novizan Cell / Tissue DNA Extraction Kit. Using the DNA as a template, a specific sequence (~400 bp) of the target gene was amplified by high-fidelity enzyme PCR, specifically the amplicon containing the 20 bp fragment that targets sgRNA. The sequencing primers are shown in Table 9. 5 μL of the PCR product was subjected to agarose gel electrophoresis. PCR products with band sizes matching the expectations were sent to Genewiz for Sanger sequencing. The peak diagram of the Sanger sequencing was viewed using Chromas software, and the proportion of each base was analyzed using SnapGene Viewer software.
[0123] Table 9. Sanger sequencing primers for sgRNA
[0124]
[0125]
[0126] Example 4: Comparison of editing efficiency of cytosine base editors
[0127] HEK293T cells were transfected with a series of editors targeting site3, RNF2, EMX1, HEK2 and HEK4 genes (such as BE3-rA1, BE4-rA1, BE4-mA3A and BE4-mA3A-delSVR). Successfully transfected cells were sorted out and the editing status of each editor at the same gene site was analyzed. The specific implementation method is the same as steps (2) to (4) of Example 3.
[0128] The comparison results of editing efficiency of different editors at the site3, RNF2, EMX1, HEK2, and HEK4 gene target sites are shown in the figure below. Figure 6 , Figure 7 , Figure 8 , Figure 9 and Figure 10The results showed that, compared with the original BE3-rA1 and BE4-rA1 plasmids, the cytosine base editors carrying mA3A or mA3A-delSVR had a certain advantage in C>T editing efficiency, with an overall improvement of at least 1-fold, especially at the EMX1 gene target site, where the improvement was nearly 5-fold. Furthermore, compared with the unmodified BE4-mA3A, the BE4-mA3A-delSVR with the SVR sequence deletion showed significantly enhanced editing specificity. Specifically, the editing site on site3 changed from C4, C5, and C14 to primarily editing the C5 site, while on RNF2 and HEK2 genes, it primarily edited the C6 site. The editing efficiency at C5 and C6 sites remained relatively stable. Except for a decrease in editing efficiency at the HEK4 gene target site, the editing efficiency at other gene target sites remained comparable to or even higher than that of BE4-mA3A (at the site3-C5 site). (See...) Figure 11 .
[0129] Example 5: Comparison of product purity of cytosine base editor
[0130] HEK293T cells were transfected with a series of editors targeting site3, RNF2, EMX1, HEK2 and HEK4 genes (such as BE3-rA1, BE4-rA1, BE4-mA3A and BE4-mA3A-delSVR). Successfully transfected cells were sorted out and the editing status of each editor at the same gene site was analyzed. The specific implementation method is the same as steps (2) to (4) of Example 3.
[0131] The results of the comparison of product purity at the site3, RNF2, EMX1, HEK2, and HEK4 gene targeting sites using different editors are shown below. Figure 12 It is known that BE4-mA3A-delSVR, which deletes the SVR sequence, primarily edits the C5 and C6 sites compared to BE4-mA3A. Therefore, this invention analyzed the product purity of different editors at site3-C5, RNF2-C6, EMX1-C5, EMX1-C6, HEK2-C6, and HEK4-C5 sites. The results showed that compared to the original BE3-rA1 and BE4-rA1 plasmids, the cytosine base editors carrying mA3A or mA3A-delSVR improved product purity at both C5 and C6 sites. The proportion of C edited to T was greater than 70%, significantly higher than that of BE3-rA1 and BE4-rA1 (a higher proportion of C edited to T indicates higher product purity, meaning a significantly reduced proportion of non-T products). Especially at the EMX1-C6 site, the purity was nearly 10 times higher than the original editor (see [link to EMX1-C6]). Figure 12Furthermore, compared to BE4-mA3A, the product purity of BE4-mA3A-delSVR with the SVR sequence deleted remained at a comparable level at the C5 and C6 sites, and the product purity at the EMX1-C6 sites reached approximately 96%.
[0132] Example 6: Comparison of editing windows of the cytosine base editor
[0133] Editing of C bases in all tested genes (site3, RNF2, EMX1, HEK2, and HEK4) was integrated, and the average editing efficiency of C bases at these five gene sites was analyzed based on their position in the original spacer (PAM at positions 21-23). BE3 and BE4 had relatively small editing windows, effectively editing only between C4 and C6, with low editing efficiency (maximum not exceeding 30%). Compared to BE3-rA1 and BE4-rA1, BE4-mA3A showed a larger editing window (C3-C14 sites) and higher overall editing efficiency. BE4-mA3A-delSVR had the smallest editing window, editing only C5 and C6 sites, exhibiting higher specificity, and maintaining stable peak editing efficiency and product purity. Figure 13 ).
[0134] Comparing BE4-mA3A and BE4-mA3A-delSVR separately, it is clear that the editing window of BE4-mA3A-delSVR is more specific, showing a clear preference for C5 and C6 sites, and the C>T editing efficiency at these two sites is similar to that of BE4-mA3A. Figure 14 Therefore, the SVR sequence unique to cynomolgus monkeys has a significant impact on the activity and selectivity of RNA editing enzymes. It is further speculated that the SVR sequence may influence the size of the editing pocket, leading to more precise editing. Thus, this invention provides a good idea and direction for future editor optimization. The BE4-mA3A-delSVR editor constructed in this invention enriches the gene editing toolkit and has great application potential in gene therapy for single-gene genetic diseases.
[0135] Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Anyone skilled in the art can make various modifications and alterations without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention should be determined by the claims. SEQUENCE LISTING <110> Jiangnan University <120> A cytosine base editor based on APOBEC3A modified from cynomolgus monkeys <130> BAA211146A <160> 12 <170> PatentIn version 3.3 <210> 1 <211> 202 <212> PRT <213> Artificial Sequence <400> 1 Met Glu Ala Ser Pro Ala Ser Arg Pro Arg His Leu Met Asp Pro Asn 1 5 10 15 Thr Phe Thr Phe Asn Phe Asn Asn Asp Leu Ser Val Arg Gly Arg His 20 25 30 Gln Thr Tyr Leu Cys Tyr Glu Val Glu Arg Leu Asp Asn Gly Thr Trp 35 40 45 Val Pro Met Asp Glu Arg Arg Gly Phe Leu Cys Asn Lys Ala Lys Asn 50 55 60 Val Pro Cys Gly Asp Tyr Gly Cys His Ala Glu Leu Cys Phe Leu Gly 65 70 75 80 Glu Val Pro Ser Trp Gln Leu Asp Pro Ala Gln Thr Tyr Arg Val Thr 85 90 95 Trp Phe Ile Ser Trp Ser Pro Cys Phe Arg Arg Gly Cys Ala Glu Gln 100 105 110 Val Arg Ala Phe Leu Gln Glu Asn Thr His Met Arg Leu Arg Ile Phe 115 120 125 Ala Ala Arg Ile Tyr Asp Tyr Asp Leu Leu Tyr Gln Glu Ala Leu Arg 130 135 140 Thr Leu Arg Asp Ala Gly Ala Gln Val Ser Ile Met Thr Tyr Glu Glu 145 150 155 160 Phe Lys His Cys Trp Asp Thr Phe Val Asp Arg Gln Gly Arg Pro Phe 165 170 175 Gln Pro Trp Asp Gly Leu Asp Glu His Ser Gln Ala Leu Ser Gly Arg 180 185 190 Leu Arg Asp Ile Leu Gln Asn Gln Gly Asn 195 200 <210> 2 <211> 199 <212> PRT <213> Artificial Sequence <400> 2 Met Glu Ala Ser Pro Ala Ser Arg Pro Arg His Leu Met Asp Pro Asn 1 5 10 15 Thr Phe Thr Phe Asn Phe Asn Asn Asp Leu Gly Arg His Gln Thr Tyr 20 25 30 Leu Cys Tyr Glu Val Glu Arg Leu Asp Asn Gly Thr Trp Val Pro Met 35 40 45 Asp Glu Arg Arg Gly Phe Leu Cys Asn Lys Ala Lys Asn Val Pro Cys 50 55 60 Gly Asp Tyr Gly Cys His Ala Glu Leu Cys Phe Leu Gly Glu Val Pro 65 70 75 80 Ser Trp Gln Leu Asp Pro Ala Gln Thr Tyr Arg Val Thr Trp Phe Ile 85 90 95 Ser Trp Ser Pro Cys Phe Arg Arg Gly Cys Ala Glu Gln Val Arg Ala 100 105 110 Phe Leu Gln Glu Asn Thr His Met Arg Leu Arg Ile Phe Ala Ala Arg 115 120 125 Ile Tyr Asp Tyr Asp Leu Leu Tyr Gln Glu Ala Leu Arg Thr Leu Arg 130 135 140 Asp Ala Gly Ala Gln Val Ser Ile Met Thr Tyr Glu Glu Phe Lys His 145 150 155 160 Cys Trp Asp Thr Phe Val Asp Arg Gln Gly Arg Pro Phe Gln Pro Trp 165 170 175 Asp Gly Leu Asp Glu His Ser Gln Ala Leu Ser Gly Arg Leu Arg Asp 180 185 190 Ile Leu Gln Asn Gln Gly Asn 195 <210> 3 <211> 600 <212> DNA <213> Artificial Sequence <400> 3 atggaagcca gcccagcatc caggcccaga cacttgatgg atccaaacac gttcactttc 60 aactttaaca atgaccttgg acggcaccag acctacttgt gctacgaggt ggagcgcctg 120 gacaatggca cctgggtccc gatggacgag cgcaggggct ttctatgcaa caaggctaag 180 aatgttccct gtggtgatta tggctgccac gcggagctgt gcttcctggg cgaggttcct 240 tcttggcagt tggacccggc ccagacgtac agggtcactt ggttcatctc ctggagcccc 300 tgcttcagga ggggctgtgc cgagcaagtg cgtgcgttcc ttcaggagaa cacacacatg 360 agactgcgca tctttgctgc ccgcatctat gattacgatc tcctgtatca ggaggcactg 420 cgaacgctgc gggatgctgg ggcccaagtc tccatcatga cctacgagga atttaagcac 480 tgctgggaca cctttgtgga ccgccaggga cgtcccttcc agccctggga tggactagat 540 gagcacagcc aagccctgag tgggaggctt cgggacattc tccagaatca gggaaactga 600 <210> 4 <211> 4101 <212> DNA <213> Artificial sequence <400> 4 gataaaaagt attctattgg tttagccatc ggcactaatt ccgttggatg ggctgtcata 60 accgatgaat acaagtacc ttcaagaaa tttaggtgt tgggacac agaccgtcat 120 tcgattaaaa agaatcttat cggtgccctc ctattcgata gtggcgaac ggcagaggcg 180 actcgcctga aacgaaccgc tcggagagg tatacacgtc gcaagaccg atatgttac 240 ttacaagaaa ttttagcaa tgagatggcc aaagttgacg attctctt tcaccgtttg 300 gaagagtcct tccttgtcga agaggacaag aaacatgaac ggcacccat ctttggaac 360 atagtagatg aggtggcata tcatgaaaag tacccaacga tttacacct cagaaaaaag 420 ctagttgact siactgataa agcggacctg agttaatct acttggctct tgcccatatg 480 aaagttcc gtgggcactt tctcattgag ggtgatctaa atccggacaa ctcggatgtc 540 gandaactgt tcatccagtt agtacaacc tataatcagt tgtttgaaga gaaccctata 600 aatgcaagtg gcgtggatgc gaaggctatt cttagcgccc gccctctaa atcccgacgg 660 ctagaaaacc tgatcgcaca attacccgga gagagaaaa atggttgtt cggtaacctt 720 atagcgctct cactaggcct gataccaat tttaagtcga acttcgactt agctgaagat 780 gccaaattgc agcttagtaa ggacacgtac gatgacgatc tcgacaatct actggcacaa 840 attggagatc agtatgcgga cttattttg gctgccaaaa accttagcga tgcaatcctc 900 ctatctgaca tactgagagt tatactgag attaccaagg cgccgttatc cgcttcaatg 960 atcaaaaggt acgatgaaca tcaccaagac ttgacacttc tcaaggccct agtccgtcag 1020 caactgcctg agaaatataa ggaaatattc tttgatcagt cgaaaaacgg gtacgcaggt 1080 tatattgacg gcggagcgag tcaagaggaa ttctacaagt tttcaaacc catattagag 1140 aagatggatg ggacggaaga gttgcttgta aaactcaatc gcgaagatct actgcgaaag 1200 cagcggactt tcgacaacgg tagcattcca catcaaatcc acttaggcga attgcatgct 1260 atacttagaa ggcaggagga tttttatccg ttcctcaaag acaatcgtga aaagattgag 1320 aaaatcctaa cctttcgcat accttactat gtgggacccc tggcccgagg gaactctcgg 1380 ttcgcatgga tgacaagaaa gtccgaagaa acgattactc catggaattt tgaggaagtt 1440 gtcgataaag gtgcgtcagc tcaatcgttc atcgagagga tgaccaactt tgacaagaat 1500 ttaccgaacg aaaaagtatt gcctaagcac agtttacttt acgagtattt cacagtgtac 1560 aatgaactca cgaaagttaa gtatgtcact gagggcatgc gtaaacccgc ctttctaagc 1620 1680 aagcaattga aagagacta ctttaagaaa attgaatgct tcgattctgt cgagatctcc 1740 ggggtagaag atcgatttaa tgcgtcactt ggtacgtatc atgacctcct aaagataatt 1800 aaagataagg acttcctgga taacgaag aatgaagata tctttagaaga tatagtgttg 1860 actcttaccc tctttgaaga tcgggaaatg attgagaaa gactaaaaac atacgctcac 1920 ctgttcgacg ataaggttat gaaacagtta aagaggcgtc gctatacggg ctggggacga 1980 ttgtcgcgga aacttatcaa cgggataaga gacaagcaaa gtggtaaaac tattctcgat 2040 tttctaaaga gcgacggctt cgccaatagg aactttatgc agctgatcca tgatgactct 2100 ttaaccttca aagaggattat acaaaaggca caggtttccg gacaagggga ctcattgcac 2160 gaacatattg cgaatcttgc tggttcgcca gccatcaaaa agggcatact ccagacagtc 2220 aaagtagtgg atgagctagt taaggtcatg ggacgtcaca aaccggaaaa cattgtaatc 2280 gagatggcac gcgaaaatca aacgactcag aaggggcaaa aaaacagtcg agagcggatg 2340 aagagaatag aagagggtat taaagaactg ggcagccaga tcttaaagga gcatcctgtg 2400 gaaaataccc aattgcagaa cgagaaactt tacctctatt acctacaaaa tggaagggac 2460 atgtatgttg atcaggaact ggacataaac cgtttatctg attacgacgt cgatcacatt 2520 gtaccccaat cctttttgaa ggacgattca atcgacaata aagtgcttac acgctcggat 2580 aagaaccgag ggaaaagtga caatgttcca agcgaggaag tcgtaaagaa aatgaagaac 2640 tattggcggc agctcctaaa tgcgaaactg ataacgcaaa gaaagttcga taacttaact 2700 aaagctgaga ggggtggctt gtctgaactt gacaaggccg gattttata acgtcagctc 2760 gtggaaaccc gccaaatcac aaagcatgtt gcacagatac tagattcccg aatgaatacg 2820 aaatacgacg agaacgataa gctgattcgg gaagtcaaag taatcacttt aaagtcaaaa 2880 ttggtgtcgg acttcagaaa ggattttcaa ttctataaag ttagggagat aaataactac 2940 caccatgcgc acgacgctta tcttaatgcc gtcgtaggga ccgcactcat windowaatac 3000 ccgaagctag aaagtgagtt tgtgtatggt gattacaaag ttatgacgt ccgtaagatg 3060 atcgcgaaaa gcgaacagga gataggcaag gctacagcca atactctt ttattctaac 3120 attatgaatt tctttagac ggaatcact ctggcaacg gagagatacg caacgacct 3180 ttaattgaaa ccaatgggga gandaggtgaa atcgtatggg ataagggccg ggactcgcg 3240 acggtgagaa aagttttgtc catgccccaa gtcacatag taaagaaac tgaggtgcag 3300 accggagggt ttcaagga atcgattctt ccaaaagga atagtgataa gctcatcgct 3360 cgtaaaaagg actgggaccc gaaaaagtac ggtggctcg atagccctac agttgcctat 3420 tctgtcctag tagtggcaaa agttgagaag ggaaatcca agaaactgaa gtcagtcaa 3480 gatttattgg ggataacgat ttggagcgc tcgtctttg aaagaaccc catcgacttc 3540 cttgaggcga aaggttacaa ggaagtaaaaaggatctca taatttaac accaaagtat 3600 agtctgtttg agttagaaaa tggccgaaaa cggatgttgg ctagcgccgg agagcttcaa 3660 aaggggaacg aactcgcact accgtctaaa tacgtgaatt tcctgtattt agcgtcccat 3720 tacgagaagt tgaaaggttc acctgaagat aacgaacaga agcaactttt tgttgagcag 3780 cacaaacatt atctcgacga aatcatagag caaatttcgg aattcagtaa gagagtcatc 3840 ctagctgatg ccaatctgga caaagtatta agcgcataca acaagcacag ggataaaccc 3900 atacgtgagc aggcggaaaa tattatccat ttgtttactc ttaccaacct cggcgctcca 3960 gccgcattca agtattttga cacaacgata gatcgcaaac gatacacttc taccaaggag 4020 gtgctagacg cgacactgat tcaccaatcc atcacgggat tatatgaaac tcggatagat 4080 ttgtcacagc ttgggggtga c 4101 <210> 5 <211> 249 <212> DNA <213> Artificial Sequence <400> 5 actaatctgt cagatattat tgaaaaggag accggtaagc aactggttat ccaggaatcc 60 atcctcatgc tcccagagga ggtggaagaa gtcattggga acaagccgga aagcgatata 120 ctcgtgcaca ccgcctacga cgagagcacc gacgagaatg tcatgcttct gactagcgac 180 gcccctgaat acaagccttg ggctctggtc atacaggata gcaacggtga gaacaagatt 240 aagatgctc 249 <210> 6 <211> 21 <212> DNA <213> Artificial sequence <400> 6 cccaagaaga agaggaaagt c 21 <210> 7 <211> 225 <212> DNA <213> Artificial sequence <400> 7 ctgtgccttc tagttgccag ccatctgttg tttgcccctc ccccgtgcct tccttgaccc 60 tggaaggtgc cactcccact gtcctttcct aataaaatga ggaaattgca tcgcattgtc 120 tgagtaggtg tcattctatt ctggggggtg gggtggggca ggacagcaag ggggaggatt 180 gggttgacaa tagcaggcat gctggggatg cggtgggctc tatgg 225 <210> 8 <211> 705 <212> DNA <213> Artificial sequence <400> 8 atggtgagca agggcgagga ggtcatcaaa gagttcatgc gcttcaaggt gcgcatggag 60 ggctccatga acggccacga gttcgagatc gagggcgagg gcgagggccg cccctacgag 120 ggcacccaga ccgccaagct gaaggtgacc aagggcggcc ccctgccctt cgcctgggac 180 atcctgtccc cccagttcat gtacggctcc aaggcgtacg tgaagcaccc cgccgacatc 240 cccgattaca agaagctgtc cttccccgag ggcttcaagt gggagcgcgt gatgaacttc 300 gaggacggcg gtctggtgac cgtgacccag gactcctccc tgcaggacgg cacgctgatc 360 tacaaggtga agatgcgcgg caccaacttc ccccccgacg gccccgtaat gcagaagaaa 420 accatgggct gggaggcctc caccgagcgc ctgtaccccc gcgacggcgt gctgaagggc 480 gagatccacc aggccctgaa gctgaaggac ggcggccact acctggtgga gttcaagacc 540 atctacatgg ccaagaagcc cgtgcaactg cccggctact actacgtgga caccaagctg 600 gacatcacct cccacaacga ggactacacc atcgtggaac agtacgagcg ctccgagggc 660 cgccaccacc tgttcctgta cggcatggac gagctgtaca agtaa 705 <210> 9 <211> 343 <212> DNA <213> Artificial sequence <400> 9 gagggcctat ttcccatgat tccttcatat ttgcatatac gatacaaggc tgttagagag 60 ataattggaa ttaatttgac tgtaaacaca aagatattag tacaaaatac gtgacgtaga 120 aagtaataat ttcttgggta gtttgcagtt ttaaaattat gttttaaaat ggactatcat 180 atgcttaccg taacttgaaa gtatttcgat ttcttggctt tatatatctt gtggaaagga 240 cgaaacaccg ggtcttcgag aagacctgtt ttagagctag aaatagcaag ttaaaataag 300 gctagtccgt tatcaacttg aaaaagtggc accgagtcgg tgc 343 <210> 10 <211> 720 <212> DNA <213> Artificial Sequence <400> 10 atggtgagca agggcgagga gctgttcacc ggggtggtgc ccatcctggt cgagctggac 60 ggcgacgtaa acggccacaa gttcagcgtg tccggcgagg gcgagggcga tgccacctac 120 ggcaagctga ccctgaagtt catctgcacc accggcaagc tgcccgtgcc ctggcccacc 180 ctcgtgacca ccctgaccta cggcgtgcag tgcttcagcc gctaccccga ccacatgaag 240 cagcacgact tcttcaagtc cgccatgccc gaaggctacg tccaggagcg caccatcttc 300 ttcaaggacg acggcaacta caagacccgc gccgaggtga agttcgaggg cgacaccctg 360 gtgaaccgca tcgagctgaa gggcatcgac ttcaaggagg acggcaacat cctggggcac 420 aagctggagt acaactacaa cagccacaac gtctatatca tggccgacaa gcagaagaac 480 ggcatcaagg tgaacttcaa gatccgccac aacatcgagg acggcagcgt gcagctcgcc 540 gaccactacc agcagaacac ccccatcggc gacggccccg tgctgctgcc cgacaaccac 600 tacctgagca cccagtccgc cctgagcaaa gaccccaacg agaagcgcga tcacatggtc 660 ctgctggagt tcgtgaccgc cgccgggatc actctcggca tggacgagct gtacaagtaa 720 <210> 11 <211> 609 <212> DNA <213> Artificial Sequence <400> 11 atggaagcca gcccagcatc caggcccaga cacttgatgg atccaaacac gttcactttc 60 aactttaaca atgacctttc ggtccgtgga cggcaccaga cctacttgtg ctacgaggtg 120 gagcgcctgg acaatggcac ctgggtcccg atggacgagc gcaggggctt tctatgcaac 180 aaggctaaga atgttccctg tggtgattat ggctgccacg cggagctgtg cttcctgggc 240 gaggttcctt cttggcagtt ggacccggcc cagacgtaca gggtcacttg gttcatctcc 300 tggagcccct gcttcaggag gggctgtgcc ggcaagtgc gtgcgttcct tcaggagaac 360 acacacatga gactgcgcat ctttgctgcc cgcatctatg attacgattt cctgtatcag gaggcactgc gacgctgcg ggatgctggg gcccaagtct ccatcatgac ctacgagga 540. tttaagcact gctgggacac ctttgtggac cgccagggac gtcccttcca gccctgggat ggactagatg agcacagcca agccctgagt gggaggcttc gggacattct ccagaatcag ggaaactga 609 <210> 12 <211> 609 <212> DNA <213> The snowstorm <400> 12 60. atggaagcca gcccagcatc caggcccaga cacttgatgg atccaaacac gttcactttc 120. aactttaaca atgacctttc ggtccgtgga cggcaccaga cctacttgtg ctacgaggtg 180. gcgcctgg acaatggcac ctgggtcccg atggacgagc gcaggggctt tctatgcaac aaggctaaga atgttccctg tggtgattat ggctgccacg cggagctgtg cttcctgggc 240 gaggttcctt cttggcagtt ggacccggcc cagacgtaca gggtcacttg gttcatctcc 300 tggagcccct gcttcaggag gggctgtgcc ggcaagtgc gtgcgttcct tcaggagaac 360 acacacatga gactgcgcat ctttgctgcc cgcatctatg attack cctgtatcag gaggcactgc gacgctgcg ggatgctggg gcccaagtct ccatcatgac ctacgagga 540. tttaagcact gctgggacac ctttgtggac cgccagggac gtcccttcca gccctgggat ggactagatg agcacagcca agccctgagt gggaggcttc gggacattct ccagaatcag ggaaactga 609
Claims
1. A cytosine deaminase mutant derived from cynomolgus monkeys, characterized in that, The mutant is based on the parent enzyme with the amino acid sequence shown in SEQ ID NO:1, by deleting three consecutive amino acids from position 27 to 29.
2. The gene encoding the cytosine deaminase mutant of claim 1.
3. An expression box, characterized in that, It contains the gene described in claim 2.
4. The expression box as described in claim 3, characterized in that, The expression cassette further comprises a promoter, an nCas9-encoded nucleic acid, a uracil DNA glycosylase inhibitor-encoded nucleic acid, a nuclear localization sequence, and a termination sequence; the promoter initiates the expression of the gene of claim 2, and the sequence in the following order is: promoter, gene of claim 2, nCas9-encoded nucleic acid, uracil DNA glycosylase inhibitor-encoded nucleic acid, nuclear localization sequence, and termination sequence.
5. The expression box as described in claim 3 or 4, characterized in that, The nucleotide sequence encoding nCas9 is shown in SEQ ID NO:4, the nucleotide sequence encoding the uracil DNA glycosylation inhibitor is shown in SEQ ID NO:5, the nucleotide sequence of the nuclear localization sequence is shown in SEQ ID NO:6, and the nucleotide sequence of the termination sequence is shown in SEQ ID NO:
7.
6. A CBE single-base editor based on cytosine deaminase modified from cynomolgus monkeys, characterized in that, The CBE single-base editing system comprises four parts: the first part is the transfection efficiency indicator, which includes a red fluorescent protein and a promoter; the second part is the sgRNA transcription unit, which includes a carrying frame for inserting the sgRNA sequence and its promoter. The third part comprises the expression cassette as described in any one of claims 3 to 5, and the fourth part comprises the high copy origin of coliform factors ori and the ampicillin resistance selection gene AmpR.
7. The CBE single-base editor as described in claim 6, characterized in that, The encoding nucleotide sequence of the red fluorescent protein is shown in SEQ ID NO:8, and the nucleotide sequence of the sgRNA transcription unit is shown in SEQ ID NO:
9.
8. The application of the CBE single-base editor according to claim 6 or 7 in the preparation of gene editing tools.
9. The application as described in claim 8, characterized in that, The sgRNA sequence is designed according to the target gene and inserted into the sgRNA transcription unit of the CBE single base editing system described in claim 6 or 7 to obtain a CBE single base editing system with specific target gene. Then, the CBE single base editor described in claim 6 or 7 is introduced into the recipient cell to achieve the mutation of the target base cytosine to thymine.
10. The use of the cytosine deaminase mutant of claim 1, the gene of claim 2, or the expression cassette of any one of claims 3 to 5 in the preparation of gene editing tools.