A cytosine base editor based on human modified APOBEC3A

By inserting specific amino acid residues into APOBEC3A, the efficient CBE single-base editor BE4-hA3A-SVR+ was constructed, which solved the shortcomings of existing CBE in editing efficiency and product purity, and achieved more efficient and purer gene editing effects.

CN115992122BActive Publication Date: 2025-05-16JIANGNAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111215861.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-19
Publication Date
2025-05-16
Estimated Expiration
2041-10-19

AI Technical Summary

Technical Problem

Existing cytosine base editors (CBEs) have shortcomings in editing efficiency and product purity, especially the modified APOBEC3A is less efficient in catalyzing C>T mutations.

Method used

A new cytosine deaminase mutant was constructed by inserting three amino acid residues of serine, valine and arginine into the amino acid sequence of human APOBEC3A to construct the efficient CBE single-base editor BE4-hA3A-SVR+.

Benefits of technology

The efficiency of targeted site C>T editing in site3, RNF2, EMX1, HEK2 and HEK4 genes in HEK293T cells was significantly improved. The editing window expanded to the C3-C14 site, and the editing efficiency was nearly 5 times higher than the original version, and the product purity can reach 97%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115992122B_ABST
    Figure CN115992122B_ABST
Patent Text Reader

Abstract

The present invention discloses a cytosine base editor based on human-modified APOBEC3A, belonging to the technical field of genetic engineering. The editing window of the CBE single-base editor BE4-hA3A-SVR<supgt;+< / supgt> constructed in the present invention is the C3-C14 site, and the editing efficiency is increased by nearly 5 times at most compared with BE4-rA1, increased by nearly 1 time compared with BE4-hA3A, and the product purity can reach up to 97%. Therefore, the hA3A-SVR<supgt;+< / supgt> constructed in the present invention is a cytosine deaminase superior to rA1 for optimizing the CBE gene editing system, and a gene editing tool with high editing efficiency, high product purity and wide editing window is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a cytosine base editor based on human modified APOBEC3A, belonging to the technical field of genetic engineering. Background Art

[0002] At present, gene editing technology is developing rapidly. By modifying specific genes, it can study gene functions, the pathogenesis of genetic diseases, develop new drugs, and be used for gene therapy and crop improvement. In recent years, various derivative technologies based on the CRISPR-Cas system have been widely used in the fields of life sciences and medicine. Among them, base editors (BE) have become an important part of gene editing technology. BE is designed based on the CRISPR-Cas9 system. The wild-type Cas9 protein is modified and connected to cytosine deaminase or adenine deaminase, and then guided by sgRNA (small guide RNA), a single base is directly edited without generating double stranded breaks (DSB) of DNA. BE is mainly divided into two categories: cytosine base editor (CBE) and adenine base editor (ABE). In 2016, the cytosine deaminase used in the CBE system developed by Dvid RLiu.'s team can convert the corresponding cytosine (C) in the non-complementary chain into uracil (U) through deamination. During DNA replication or repair, U is recognized as thymine (T), and the corresponding guanine (G) on the complementary chain will become adenine (A), ultimately achieving the conversion of C>T on the non-complementary chain and G>A on the complementary chain. In 2017, Dvid R Liu.'s team developed the ABE system, which has a similar principle to the CBE system, except that the cytosine deaminase is replaced by adenine deaminase, which can complete the editing of A>G on the non-complementary chain and T>C on the complementary chain, further supplementing the types of single-base editors.

[0003] The Cas9 protein in the BE system can be further optimized. One is dCas9 (Catalytically dead Cas9) without endonuclease activity that can bind to the target gene but does not cut the target gene, and the other is nCas9 (Cas9 nickase) with single-strand DNA nickase activity. Both Cas9 proteins will not produce DBS, thus avoiding the mismatch of non-homologous end-joining (NHEJ) and the low efficiency of homologous recombination repair (HDR). Subsequently, uracil glycosylase inhibitor (UGI) was added to BE to inhibit the excision of the intermediate product U, thereby improving the editing efficiency of C>T on the DNA chain. At present, the BE system has developed rapidly. The CBE system has been optimized to the fourth generation BE4, and the ABE system has also been optimized to the ABE7.10 version. Even some dual-base editors that integrate the functions of the CBE and ABE systems have been developed, which can simultaneously achieve the conversion of C>T and A>G at the same target site.

[0004] The cytosine deaminase used in the traditional CBE system is rat Apobec1 (rA1). In 2018, Jason M Gehrke et al. tried to use human modified APOBEC3A (eA3A) to replace rA1 in the third-generation CBE system (BE3). They performed site-directed mutagenesis on human APOBEC3A (hA3A) and constructed a series of BE3 mutant editors carrying single amino acid mutations in hA3A, among which the hA3A-N57G targeted (HBB -28 site) with high editing efficiency and adjacent (HBB -25 The editing efficiency of the hA3A-N57G site was the lowest among all mutants. They then used BE3 carrying hA3A-N57G for the treatment of β-thalassemia, and improved its editing accuracy by nearly 40 times compared with BE3, but in HBB -28 The editing efficiency of the site is only about 30%. This suggests that the modified hA3A is better than rA1, but the overall editing efficiency is low, so further research is needed on the ability of hA3A to catalyze C>T mutations.

[0005] In addition, since most of the pathogenic mutations of single-gene genetic diseases are point mutations, the BE system can edit single bases of DNA sequences, thereby correcting the pathogenic site mutations of genetic patients, achieving "molecular surgery" and curing genetic diseases on the basis of pathogenic substances. It is known that there are about 75,000 human genome site mutations related to genetic diseases, of which about 50% are estimated to be potential therapeutic targets for the CBE system. Therefore, the development and optimization of more efficient and accurate single-base editors are of great significance for the treatment of single-gene genetic diseases. Summary of the invention

[0006] The editing accuracy of the BE3 mutant editor using the modified human APOBEC3A (eA3A) is high, but the overall editing efficiency is low and the purity of the editing product is unknown. It is now urgent to construct a cytosine base editor with high product purity, especially improved editing efficiency.

[0007] In order to solve the above problems, the present invention provides a gene editing tool with high editing efficiency and high product purity.

[0008] The first object of the present invention is to provide a mutant of human cytosine deaminase, wherein the mutant is based on the parent enzyme with an amino acid sequence such as SEQ ID NO: 1, and three amino acid residues, serine (S), valine (V) and arginine (R), are inserted in sequence between the 26th and 27th amino acids.

[0009] In one embodiment, the amino acid sequence of the mutant is shown in SEQ ID NO:2.

[0010] The second object of the present invention is to provide a gene encoding the above cytosine deaminase mutant.

[0011] In one embodiment, the nucleotide sequence of the mutant gene is shown in SEQ ID NO:3.

[0012] The third object of the present invention is to provide an expression cassette, which comprises the gene encoding the cytosine deaminase mutant.

[0013] In one embodiment, the expression cassette further contains a promoter, nCas9 (D10A), a uracil DNA glycosylase inhibitor (UGI), a nuclear localization sequence NLS, and a termination sequence.

[0014] In one embodiment, the promoter drives the expression of the gene encoding the above-mentioned cytosine deaminase mutant, which is connected in the order of promoter, the gene described in the second purpose, nCas9 (D10A), uracil DNA glycosylase inhibitor, NLS and termination sequence.

[0015] In one embodiment, the promoter includes a CMV promoter and a U6 promoter.

[0016] In one embodiment, the nucleotide sequence of the nCas9 (D10A) is as shown in SEQ ID NO:4, the nucleotide sequence of UGI is as shown in SEQ ID NO:5, the nucleotide sequence of the NLS is as shown in SEQ ID NO:6, and the nucleotide sequence of the termination sequence is as shown in SEQ ID NO:7.

[0017] The fourth object of the present invention is to provide a CBE single base editor BE4-hA3A-SVR based on human cytosine deaminase (APOBEC3A) + The CBE single-base editing system consists of four parts. The first part is the transfection efficiency indicator part, which includes a red fluorescent protein (dTomato) and a promoter; the second part is the sgRNA transcription unit, which includes a carrying frame for inserting the sgRNA sequence and its promoter; the third part includes the above-mentioned expression cassette, and the fourth part includes the high-copy replication origin ori from the colicin factor (colE1) and the ampicillin resistance screening gene AmpR.

[0018] In one embodiment, the promoter includes a CMV promoter and a U6 promoter.

[0019] In one embodiment, the nucleotide sequence of the red fluorescent protein is as shown in SEQ ID NO:8.

[0020] In one embodiment, the nucleotide sequence of the sgRNA transcription unit is shown in SEQ ID NO:9.

[0021] The fifth object of the present invention is to provide the application of the above-mentioned CBE single-base editor in the field of gene editing.

[0022] In one embodiment, an sgRNA sequence is designed according to the target gene and inserted into the sgRNA transcription unit of the above-mentioned CBE single-base editing system to obtain a CBE single-base editing system with a specific targeted gene. The CBE single-base editing system is introduced into a recipient cell to achieve a mutation of the target base C to T, thereby obtaining cells containing a single base mutation.

[0023] In one embodiment, the cell is a eukaryotic cell.

[0024] In one embodiment, the eukaryotic cell is a mammalian cell.

[0025] In one embodiment, the mammalian cell comprises a human embryonic kidney epithelial cell.

[0026] The sixth object of the present invention is to provide the CBE single base editing system BE4-hA3A-SVR + The construction method, the specific construction steps are as follows:

[0027] (1) Using the pSpCas9(BB)-2A-dTomato(PX458) plasmid as a template, the red fluorescent protein dTomato gene and its CMV enhancer and promoter were amplified to obtain the vector skeleton Part 1;

[0028] (2) Using the BE3-rA1 plasmid as a template, amplify the carrier frame for inserting the sgRNA sequence and its U6 promoter to obtain the vector skeleton Part 2;

[0029] (3) Using cynomolgus macaque cDNA as a template, the nucleotide sequence of APOBEC3A (mA3A-B5) was amplified to obtain the vector backbone Part 3;

[0030] (4) Using plasmid BE4-rA1 as a template, amplification was performed to obtain the vector backbone Part 4 including nCas9 (D10A), UGI, NLS, bGH poly (A) signal, and CMV enhancer and promoter gene sequences;

[0031] (5) Using plasmid BE4-rA1 as a template, amplify the vector backbone Part 5 containing the colE1 high-copy replication origin, the ampicillin resistance selection gene, and its ampicillin promoter sequence;

[0032] (6) Connect the five fragments in the order of Part 1, Part 2, Part 3, Part 4, and Part 5 to obtain the vector plasmid BE4-mA3A-B5.

[0033] (7) Using the pcDNA3.1+hA3A+MYC DDK plasmid as a template, PCR amplification was performed to obtain the hA3A sequence with the SVR sequence inserted, i.e., hA3A-SVR + ;

[0034] (8) Connect the hA3A sequence inserted with the SVR sequence in step (7) with the BE4-mA3A-B5 vector after double restriction digestion to remove the mA3A-B5 fragment, to obtain the vector plasmid BE4-hA3A-SVR + .

[0035] (9) Designing a sgRNA sequence that binds to the target gene and inserting it into the carrying framework in step (2) to obtain a single base editor that can target the corresponding sites of different target genes.

[0036] In one embodiment, the pSpCas9(BB)-2A-dTomato(PX458) plasmid is constructed by replacing GFP with dTomato based on the plasmid pSpCas9(BB)-2A-GFP(PX458), and the nucleotide sequence of dTomato is shown in SEQ ID NO:8.

[0037] In one embodiment, the pcDNA3.1+hA3A+MYC DDK plasmid is constructed by connecting the hA3A fragment and a chemically synthesized MYCDDK tag fragment to the vector after double digestion with EcoRⅠ and XbaⅠ based on the pcDNA3.1-2xFLAG-SREBP-2 plasmid.

[0038] In one embodiment, the nucleotide sequence of the hA3A fragment is shown in SEQ ID NO:10.

[0039] In one embodiment, the nucleotide sequence of mA3A-B5 is as shown in SEQ ID NO:11.

[0040] The present invention also provides application of the mutant in the field of gene editing.

[0041] The present invention also provides application of the above expression cassette in the field of gene editing.

[0042] Beneficial effects of the present invention:

[0043] 1. Single-base editor BE4-hA3A-SVR constructed by the present invention + The sgRNA transcription unit is contained on the sgRNA. When using it for gene editing, only BE4-hA3A-SVR + Plasmids are transfected into cells. The sgRNA targeting the target gene carried on the plasmid guides the fusion protein to bind to the targeted genomic DNA region through the principle of complementary base pairing. Then, hA3A inserted with the SVR sequence binds to the single-stranded DNA (ssDNA) in the R-loop region formed by the nCas9 protein, sgRNA and target gene DNA, deaminates cytosine (C) within a certain range of the ssDNA (20bp protospacer sequence) into uracil (U), and then converts U into thymine (T) through DNA replication or repair, ultimately achieving a direct replacement of the CG base pair to the TA base pair. In addition, since nCas9 only has single-stranded DNA nickase activity, it will not cut the double-stranded DNA of the target gene between 2-3 bases upstream of the PAM (Protospacer Adjacent Motif) sequence, thereby not forming DSB, which greatly reduces the off-target efficiency.

[0044] 2. The present invention is based on inserting three amino acid residues of SVR in sequence between the 26th and 27th positions of the amino acid sequence of human cytosine deaminase hA3A, and successfully constructs a cytosine base editor BE4-hA3A-SVR that is superior to the original BE4-rA1 in terms of editing efficiency, product purity and editing window width. + The single-base editor constructed by the present invention has an editing window of C3-C14, and its editing efficiency is nearly 5 times higher than that of BE4-rA1, and nearly 1 times higher than that of BE4-hA3A, and the product purity can reach up to 97%. It fully proves that the modified hA3A-SVR + It is a cytosine deaminase that is superior to rA1 for optimizing the CBE gene editing system, thus providing a new idea and direction for further optimizing and transforming the CBE system, enriching the toolkit for gene editing, and also providing a potential tool for gene therapy of genetic diseases. (Note: site3, HEK2, and HEK4 are named in published articles. The gene name of site3 is LINC01509, the gene name of HEK2 is AC114971.1, and the gene name of HEK4 is DNMT3B.) BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 :Plasmid map of cytosine base editor BE4-mA3A-B5;

[0046] Figure 2 : Plasmid map of overexpression plasmid pcDNA3.1+hA3A+MYCDDK;

[0047] Figure 3 :Cytosine base editor BE4-hA3A-SVR + Plasmid map of

[0048] Figure 4 : Electrophoresis of BE4-mA3A-B5 (~11 kb) plasmid double digested with BamHI and SmaI;

[0049] Figure 5 :Human APOBEC3A gene sequence (hA3A) and its mutant with SVR sequence inserted (hA3A-SVR + ) electropherogram;

[0050] Figure 6 :Sanger sequencing validates different cytosine single base editors (BE3-rA1, BE4-rA1, BE4-hA3A and BE4-hA3A-SVR +) Editing efficiency at site3 gene targeting site (n=3);

[0051] Figure 7 :Sanger sequencing validates different cytosine single base editors (BE3-rA1, BE4-rA1, BE4-hA3A and BE4-hA3A-SVR + ) Editing efficiency at the RNF2 gene targeting site (n=3);

[0052] Figure 8 :Sanger sequencing validates different cytosine single base editors (BE3-rA1, BE4-rA1, BE4-hA3A and BE4-hA3A-SVR + ) Editing efficiency at the EMX1 gene targeting site (n=3);

[0053] Fig. 9 :Sanger sequencing validates different cytosine single base editors (BE3-rA1, BE4-rA1, BE4-hA3A and BE4-hA3A-SVR + ) Editing efficiency at the HEK2 gene targeting site (n=3);

[0054] Fig.10 :Sanger sequencing validates different cytosine single base editors (BE3-rA1, BE4-rA1, BE4-hA3A and BE4-hA3A-SVR + ) Editing efficiency at the HEK4 gene targeting site (n=3);

[0055] Fig.11 : BE3-rA1, BE4-rA1, BE4-hA3A and BE4-hA3A-SVR + Comparison of the distribution of products (T, G, A) of all targeted sites of the editor in the site3 gene (n=3);

[0056] Fig.12 : BE3-rA1, BE4-rA1, BE4-hA3A and BE4-hA3A-SVR + Comparison of the distribution of editor products (T, G, A) at all targeted sites of the RNF2 gene (n=3);

[0057] Fig.13 : BE3-rA1, BE4-rA1, BE4-hA3A and BE4-hA3A-SVR + Comparison of the distribution of editor products (T, G, A) at all targeted sites of the EMX1 gene (n=3);

[0058] Fig.14: BE3-rA1, BE4-rA1, BE4-hA3A and BE4-hA3A-SVR + Comparison of the distribution of editor products (T, G, A) at all targeted sites of the HEK2 gene (n=3);

[0059] Fig.15 : BE3-rA1, BE4-rA1, BE4-hA3A and BE4-hA3A-SVR + Comparison of the distribution of editor products (T, G, A) at all targeted sites of the HEK4 gene (n=3);

[0060] Fig.16 : BE4-hA3A and BE4-hA3A-SVR + Comparison of the average editing efficiency of Cs editors in 5 gene loci (n=3, 6, 9).

[0061] Fig.17 : BE3-rA1, BE4-rA1, BE4-hA3A and BE4-hA3A-SVR + Comparison of C>T editing efficiency of editors at site3, RNF2, EMX1, HEK2, and HEK4 gene targeting sites (n=3); DETAILED DESCRIPTION

[0062] The plasmids involved in the following examples are:

[0063] BE4 plasmid: Addgene Plasmid #100802.

[0064] pSpCas9(BB)-2A-dTomato(PX458) plasmid: This was constructed by chemically synthesizing dTomato (SEQ ID NO: 8) to replace GFP based on the pSpCas9(BB)-2A-GFP(PX458) (Addgene Plasmid #48138) plasmid.

[0065] BE3-rA1 plasmid: donated by Professor Yang Hui, reported in the article Off-target RNA mutation induced by DNA base editing and its elimination by mutagenesis. The BE3 in the article is BE3-rA1 in the present invention.

[0066] BE4-rA1 plasmid: constructed in our laboratory, based on the BE4 plasmid (Addgene Plasmid #100802), the chemically synthesized dTomato gene sequence (SEQ ID NO: 8) and the U6 promoter + sgRNA scaffold + CMV enhancer + CMV promoter fragment of the BE3-rA1 plasmid were inserted into the NotⅠ restriction site.

[0067] pcDNA3.1+hA3A+MYCDDK plasmid: constructed in our laboratory, using pcDNA3.1-2xFLAG-SREBP-2 (Addgene Plasmid #26807) plasmid as the backbone, the vector was double-digested with EcoRⅠ and XbaⅠ, and the hA3A fragment (SEQ ID NO: 10) and a chemically synthesized MYCDDK tag fragment were ligated to the digested vector ( Figure 2 ).

[0068] BE4-mA3A-B5 plasmid: constructed in this laboratory. The specific construction method is shown in Example 1.

[0069] The cynomolgus monkeys involved in the present invention were purchased from Guangzhou Xiangguan Biotechnology Co., Ltd. (production license number: SCXK (Guangdong)

[0070] 2018-0043), and the animals were in good health as confirmed by records and veterinary examinations before the experiment, and the animal facilities met the national standards for experimental animals (GB14925-2010). Subsequently, cynomolgus monkeys that had been treated with a high-sugar and high-fat diet for 19 months and had developed hypercholesterolemia and could recover to normal blood lipids on their own were selected for the construction of the cytosine single-base editor (BE4-mA3A-B5) of the present invention.

[0071] Example 1: Construction of a cytosine single-base editor (BE4-mA3A-B5)

[0072] (1) Acquisition of target gene (mA3A-B5)

[0073] The present invention uses a Paxgene tube to obtain a blood sample of a cynomolgus monkey, and extracts RNA from the cynomolgus monkey using a blood RNA extraction kit of TIANGEN, obtains a cDNA sequence of the cynomolgus monkey using a reverse kit of Takara, uses the cDNA sequence as a template, performs PCR amplification using primer pairs 5'-TATAGGGAGAGCCGCCACCATGGAAGCCAGCCCAG-3' and 5'-ACCAGAAGAACCACCAGAGTTTCCCTGATTCTGG-3' to obtain a mA3A-B5 gene fragment (i.e., Part 3), and then purifies the PCR product using an agarose gel DNA extraction kit of TIANGEN. The PCR reaction system is shown in Table 1.

[0074] Table 1 PCR reaction system

[0075]

[0076]

[0077] The reaction procedure was as follows: pre-denaturation at 95°C for 3 min; 35 cycles of 95°C for 15 s, 60°C for 15 s, and 72°C for 25 s; extension at 72°C for 5 min, and cooling to 4°C to finally obtain mA3A-B5 (Part 3).

[0078] (2) Preparation of linearized plasmid vector (BE4)

[0079] Using BE4-rA1 plasmid as a template, PCR amplification was performed with primers 5'-TCTGGTGGTTCTTCTGGTGGTTCTAGCGGC-3' and 5'-GGTGGCGGCTCTCCCTATAGTGAGTCGTAT-3' to obtain PCR product Part4, including nCas9 (D10A), UGI, NLS, and CMV enhancer and promoter gene sequences;

[0080] Using BE4-rA1 plasmid as template, PCR amplification was performed with primer pair 5'-CGGTGGCTTCGATAGCCCTACAGTTGCCT-3' and 5'-CTACTAGGACAGAATAGGCAACTGTAGGGC-3' to obtain PCR product Part5, including colE1 high copy replication origin and ampicillin resistance selection gene and its ampicillin promoter sequence.

[0081] PCR products Part 4 and Part 5 were purified using TIANGEN's agarose gel DNA extraction kit.

[0082] (3) Obtaining the gene for red fluorescent protein (dTomato)

[0083] Using the pSpCas9(BB)-2A-dTomato(PX458) plasmid as a template, PCR amplification was performed using primers 5'-TAGAGATCCGCGCCACCATGGTGAGC-3' and 5'-GAAGGCACAGTTACTTGTACAGCTCG-3' to obtain the red fluorescent protein dTomato gene and its CMV enhancer and promoter, i.e., vector backbone Part 1. The PCR product Part 1 was then purified using TIANGEN's agarose gel DNA extraction kit.

[0084] (4) Obtaining the gene carrying frame of sgRNA and U6 promoter

[0085] Using the BE3-rA1 plasmid as a template, the following primers 5'-GCTCACATGTGAGGGCCTATTTCCC-3' and 5'-ATAGGCCCTCACATGTGAGCAAAAG-3' were used for PCR amplification to obtain the carrier frame for inserting the sgRNA sequence and its U6 promoter, i.e., the vector backbone Part 2. The PCR product Part 2 was then purified using the TIANGEN agarose gel DNA extraction kit.

[0086] (5) Construction of vector plasmid BE4-mA3A-B5 (homologous recombination method)

[0087] The five PCR fragments Part 1, Part 2, Part 3, Part 4 and Part 5 purified in steps (1) to (4) were ligated using the MultiS One Step Cloning Kit from Novozymes to obtain ligation products, which were transformed into Escherichia coli DH5α and spread on LB plates containing 0.05% Amp (ampicillin at a concentration of 100 μg / mL) resistance and cultured inverted at 37°C overnight.

[0088] Select 3 colonies on each cloning plate and inoculate them into liquid LB medium containing 0.1% Amp (100 μg / mL ampicillin) for more than 8 hours, and then send the bacterial solution to Genewise for Sanger sequencing verification. The clone vector with the fragment successfully inserted was expanded and cultured, and then the vector plasmid was obtained using the endotoxin-free plasmid extraction kit of Kangwei Century Company and named BE4-mA3A-B5, and stored at -20℃ for future use. The BE4-mA3A-B5 plasmid map is shown below: Figure 1 shown.

[0089] Example 2: Construction of cytosine single-base editors based on human APOBEC3A (hA3A) and its mutants

[0090] (1) Preparation of linearized plasmid vector (BE4-mA3A-B5)

[0091] The BE4-mA3A-B5 vector obtained in Example 1 was double-digested with Takara's fast-cutting enzymes BamHI and SmaI to remove the APOBEC3A of the cynomolgus monkey in the vector, so that the cytosine deaminase can be replaced next. The enzyme digestion system and reaction conditions are shown in Table 2. Then, the enzyme digestion product was purified using the TIANGEN common DNA product purification kit to obtain a linearized vector (see Figure 4 ).

[0092] Table 2 BamHI and SmaI double restriction enzyme digestion system

[0093]

[0094] () Acquisition of target genes (hA3A and hA3A-SVR + )

[0095] Using the hA3A overexpression plasmid pcDNA3.1+hA3A+MYC DDK constructed in our laboratory as a template, the gene sequence of hA3A was obtained by PCR amplification using the primer pair 5'-CCAGACACTTGATGGATCCACACATATTCACTTCCAAC-3' and 5'-GGACTCTGAGGTCCCGGGAGTCTCGCTGCC-3'; similarly, using the hA3A overexpression plasmid pcDNA3.1+hA3A+MYC DDK as a template, the gene sequence of hA3A was obtained by PCR amplification using the primer pair 5'-CCAGACACTTGATGGATCCACACATATTCACTTCCAACTTTAACAATGGCATTTCGGTCCGTGGAAGGCATAAG-3' and 5'-GGACTCTGAGGTCCCGGGAGTCTCGCTGCC-3' + Gene fragments (see Figure 5 ). Then, the PCR product was purified using TIANGEN's agarose gel DNA extraction kit. When performing PCR, homology arms complementary to the vector cutout were added to both ends of the primers for cloning into the vector. The PCR reaction system is shown in Table 3.

[0096] Table 3 PCR reaction system

[0097]

[0098] The reaction procedure was as follows: pre-denaturation at 95°C for 3 min; 35 cycles of 95°C for 15 s, 60°C for 15 s, and 72°C for 25 s; extension at 72°C for 5 min, cooling to 4°C, and finally obtaining the hA3A gene fragment.

[0099] (3) Construction of vector plasmid BE4-hA3A-SVR + (Homologous recombination method)

[0100] hA3A or hA3A-SVR purified in step (2) was cloned into ClonExpress II One Step Cloning Kit from Novozymes + The gene fragments were ligated with the linearized vector BE4-mA3A-B5 purified in step (1) to obtain ligation products, which were transformed into Escherichia coli DH5α and spread on LB plates containing 0.05% Amp (ampicillin at a concentration of 100 μg / mL) resistance and cultured inverted at 37°C overnight.

[0101] Three colonies were selected from each cloning plate and inoculated into liquid LB medium containing 0.1% Amp (100 μg / mL ampicillin) for more than 8 hours, and then the bacterial solution was sent to Genewise for Sanger sequencing verification. The clone vector with the fragment successfully inserted was expanded and cultured, and then the vector plasmid was obtained using the endotoxin-free plasmid extraction kit of Kangwei Century Company and named BE4-hA3A and BE4-hA3A-SVR. + , stored at -20℃ for future use. BE4-hA3A-SVR + The plasmid map of Figure 3 shown.

[0102] Example 3: Cytosine base editor BE4-hA3A-SVR + Application Methods

[0103] (1) Inserting sgRNA that specifically targets a gene

[0104] Since the single-base editor is based on the CRISPR / Cas9 system, its targeting specificity is composed of two parts, one is the base complementary pairing between the sgRNA and the target DNA sequence, and the other is determined by the Cas9 protein and the short DNA sequence at the 3' end of the target DNA sequence. The target DNA sequence is called the protospacer (the foreign DNA fragment is called the protospacer during the natural immune defense of bacteria), and the short DNA sequence at the 3' end of the target DNA sequence is called PAM (protospacer adjacent motif).

[0105] In order to better compare with existing studies, the sgRNA used is a sequence that has been reported in the literature, see Table 4 for details.

[0106] Table 4: sgRNA and its PAM sequence list

[0107]

[0108] 1) The original BE3-rA1 plasmid, BE4-rA1 plasmid, and the cytosine base editors (BE4-hA3A and BE4-hA3A-SVR) constructed in Example 2 were digested with BbsI. + ) to obtain the linearized vector, and the enzyme digestion system is shown in Table 5;

[0109] 2) Chemically synthesize cloning primers of sgRNAs of target sites 3, RNF2, EMX1, HEK2 and HEK4 in Table 6 respectively, and self-link the cloning primers into double-stranded oligonucleotide fragments by heat shock annealing (see Table 7 for sgRNA self-linking system), connect the double-stranded oligonucleotide fragments with the linearized vector in step 1) (see Table 8 for ligation system), transform the linked vector into competent Escherichia coli cells (DH5α), sequence and identify, and extract plasmids, and finally obtain a series of cytosine single base editors targeting site 3, RNF2, EMX1, HEK2 and HEK4 genes.

[0110] Table 5 BbsI restriction enzyme system

[0111]

[0112] Table 6 sgRNA cloning primers

[0113]

[0114] Table 7 sgRNA self-ligation system

[0115]

[0116] Table 8 Connection system

[0117]

[0118] (2) Cell transfection

[0119] Preparation of high-glucose DMEM complete medium: High-glucose DMEM medium contains 10% fetal bovine serum (FBS) and 1% triple antibody (penicillin-streptomycin-gentamicin).

[0120] 1) Cell culture: cryopreserved human embryonic kidney epithelial cells (HEK293T cells) were taken out from a liquid nitrogen tank, rapidly thawed in a 37°C water bath, and the thawed cell suspension was added to 10 mL of high-glucose DMEM complete medium. The cell pellet was collected by centrifugation, and the cells were resuspended in high-glucose DMEM complete medium and added to a cell culture dish containing 10 mL of high-glucose DMEM complete medium, and cultured in a 37°C constant temperature cell culture incubator containing 5% CO2.

[0121] 2) Cell plating: When HEK293T cells have grown all over the culture dish, digest the cells with 0.25% trypsin for 1 min, terminate the digestion with high-glucose DMEM complete medium, centrifuge at 1500 rpm for 2 min, discard the supernatant, resuspend the cells with 1 mL high-glucose DMEM complete medium, and plate at 1×10 5 The cells were seeded in 6-well plates (each well contained 2 mL of high-glucose DMEM complete medium);

[0122] 3) Cell transfection: In a 1.5 mL EP tube, dilute 3.75 μL of transfection reagent Lipofectamine 3000 with 125 μL Opti-MEM medium, mix thoroughly, and obtain mixed solution A; in another 1.5 mL EP tube, dilute 2.5 μg of the cytosine single base editor targeting site3, RNF2, EMX1, HEK2 or HEK4 gene constructed in step (1) with 125 μL Opti-MEM medium, and add 5 μL P3000 auxiliary transfection reagent, mix thoroughly, and obtain mixed solution B; mix mixed solutions A and B in a ratio of 1:1, incubate at room temperature for 15 min, and obtain a mixture; add the mixture to the 6-well plate where the cells have just been plated in step 2), mix gently to avoid damaging the cells; then put the 6-well plate back into a 37°C, 5% CO2 incubator to incubate the cells for 72 h, and then use them for subsequent experiments. During this period, observe and photograph them with an inverted microscope every 24 h.

[0123] (3) Flow sorting

[0124] The cells were first digested with 0.25% trypsin after transfection for 72 hours, and the digestion was terminated with high-glucose DMEM complete medium. The cell pellet was collected by balanced centrifugation at 1500 rpm for 2 minutes, and then 2 mL of PBS containing 2% fetal bovine serum (FBS) was added to resuspend and wash the cells. The cells were balanced centrifuged at 1500 rpm for 2 minutes. After repeated washing twice, the cells were resuspended with 1 mL of PBS containing 2% FBS, and then the cell suspension was collected through a flow sorting tube with a filter membrane to remove larger cell debris aggregates to avoid clogging of the instrument during sorting.

[0125] The prepared cell suspension and the corresponding collection tube containing 2mL of high-glucose DMEM complete medium were placed on ice, and then flow sorting was performed. Because the cytosine base editor constructed by the present invention carries red fluorescent protein (dTomato), dTomato positive cells were directly sorted and collected according to the excitation wavelength (554nm) and emission wavelength (581nm). BE3-rA1 plasmid carries green fluorescent protein (GFP), and GFP positive cells were directly sorted and collected according to its excitation wavelength (488nm) and emission wavelength (507nm). At least 100,000 cells were obtained for each sample.

[0126] (4) Sanger sequencing verification

[0127] The dTomato-positive cells or GFP-positive cells collected in step (3) were balanced centrifuged at 1500rpm for 2 minutes, and the precipitate was collected. Then, the cell DNA was extracted using the cell / tissue DNA extraction kit of Novogene. Using DNA as a template, PCR amplification was performed using a high-fidelity enzyme to obtain PCR products. The sequencing primers are shown in Table 9. Take 5 μL of PCR product for agarose gel electrophoresis, and send the PCR product with the expected band size to Jinweizhi Company for Sanger sequencing. The peak diagram of Sanger sequencing was viewed by using Chromas software, and the results of Sanger sequencing were analyzed by SnapGene Viewer software to analyze the proportion of each base.

[0128] Table 9 Sanger sequencing primers for sgRNA

[0129]

[0130] Example 4: Comparison of editing efficiency of cytosine base editors

[0131] A series of editors targeting site3, RNF2, EMX1, HEK2, and HEK4 genes (such as BE3-rA1, BE4-rA1, BE4-hA3A, and BE4-hA3A-SVR) + ) were transfected into HEK293T cells respectively, the successfully transfected cells were sorted out, and the editing status of each editor at the same gene site was analyzed. The specific implementation method was the same as steps (2) to (4) of Example 3.

[0132] The comparison results of editing efficiency of different editors at site3, RNF2, EMX1, HEK2 and HEK4 gene targeting sites are shown in the table below. Figure 6 , Figure 7 , Figure 8 , Fig. 9 and Fig.10The results showed that compared with the original BE3-rA1 and BE4-rA1 plasmids, the plasmids carrying hA3A or hA3A-SVR + The cytosine base editors of BE4-hA3A-SVR have obvious advantages in the editing efficiency of C>T, especially BE4-hA3A-SVR + The editing efficiency of the editor at the EMX1 gene target site was increased by nearly 5 times. In addition, BE4-hA3A-SVR with SVR sequence inserted + Compared with the unmodified BE4-hA3A, it is nearly doubled. Fig.16 .

[0133] Example 5: Comparison of product purity of cytosine base editors

[0134] A series of editors targeting site3, RNF2, EMX1, HEK2, and HEK4 genes (such as BE3-rA1, BE4-rA1, BE4-hA3A, and BE4-hA3A-SVR) + ) were transfected into HEK293T cells respectively, the successfully transfected cells were sorted out, and the editing status of each editor at the same gene site was analyzed. The specific implementation method was the same as steps (2) to (4) of Example 3.

[0135] The comparison results of product purity of different editors at site3, RNF2, EMX1, HEK2 and HEK4 gene targeting sites are shown in the figure. Fig.11 , Fig.12 , Fig.13 , Fig.14 and Fig.15 The results showed that BE4-hA3A-SVR with SVR sequence inserted + The product purity of BE4-hA3A and the unmodified BE4-hA3A was improved, and the proportion of C edited to T was greater than 80%, which was much higher than that of BE3-rA1 and BE4-rA1 (the higher the proportion of C edited to T, the higher the product purity, that is, the proportion of non-T products produced was greatly reduced). + The purity of the products was nearly 10 times higher than that of the original editor (see Fig.11 , Fig.12 , Fig.13 , Fig.14 and Fig.15 ), especially at the EMX1-C6 site, BE4-hA3A-SVR + The purity of the product is nearly 15 times higher than that of the original editor BE4-rA1 (see Fig.13). In addition, BE4-hA3A-SVR with SVR sequence inserted + Compared with BE4-hA3A, the overall product purity was further improved, with the product purity of more than half of the sites reaching over 90%, especially at the EMX1-C6 site, where the product purity could reach about 97%.

[0136] Example 6: Comparison of editing windows of cytosine base editors

[0137] The editing of C bases in all tested genes (site3, RNF2, EMX1, HEK2, and HEK4) was integrated, and the average editing efficiency of C in these five gene sites was sorted out according to the position of C base in the original spacer (PAM is located at positions 21-23). ​​Among them, the editing window of BE3-rA1 and BE4-rA1 is small, and only between C4-C6 can be effectively edited, and the editing efficiency is low (no more than 30%). Compared with BE3-rA1 and BE4-rA1, it can be seen that BE4-hA3A and BE4-hA3A-SVR + The editing window is significantly expanded (C3-C14 sites) and the overall editing efficiency is higher (up to 70%) ( Fig.17 ).

[0138] Comparison of BE4-hA3A and BE4-hA3A-SVR alone + It can be clearly seen that BE4-hA3A-SVR + The editing window of BE4-hA3A is not significantly different from that of BE4-hA3A, but it has a certain preference. The editing efficiency of the middle positions in the window (such as C9 and C10 sites) is relatively low, and the C>T editing efficiency at these two sites is similar to that of BE4-hA3A ( Fig.17 ).

[0139] Although the present invention has been disclosed as above in the form of a preferred embodiment, it is not intended to limit the present invention. Anyone familiar with this technology can make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention should be based on the definition of the claims. SEQUENCE LISTING <110> Jiangnan University <120> A cytosine base editor based on human modified APOBEC3A <130> BAA211147A <160> 11 <170> PatentIn version 3.3 <210> 1 <211> 199 <212> PRT <213> Artificial Sequence <400> 1 Met Glu Ala Ser Pro Ala Ser Gly Pro Arg His Leu Met Asp Pro His 1 5 10 15 Ile Phe Thr Ser Asn Phe Asn Asn Gly Ile Gly Arg His Lys Thr Tyr 20 25 30 Leu Cys Tyr Glu Val Glu Arg Leu Asp Asn Gly Thr Ser Val Lys Met 35 40 45 Asp Gln His Arg Gly Phe Leu His Asn Gln Ala Lys Asn Leu Leu Cys 50 55 60 Gly Phe Tyr Gly Arg His Ala Glu Leu Arg Phe Leu Asp Leu Val Pro 65 70 75 80 Ser Leu Gln Leu Asp Pro Ala Gln Ile Tyr Arg Val Thr Trp Phe Ile 85 90 95 Ser Trp Ser Pro Cys Phe Ser Trp Gly Cys Ala Gly Glu Val Arg Ala 100 105 110 Phe Leu Gln Glu Asn Thr His Val Arg Leu Arg Ile Phe Ala Ala Arg 115 120 125 Ile Tyr Asp Tyr Asp Pro Leu Tyr Lys Glu Ala Leu Gln Met Leu Arg 130 135 140 Asp Ala Gly Ala Gln Val Ser Ile Met Thr Tyr Asp Glu Phe Lys His 145 150 155 160 Cys Trp Asp Thr Phe Val Asp His Gln Gly Cys Pro Phe Gln Pro Trp 165 170 175 Asp Gly Leu Asp Glu His Ser Gln Ala Leu Ser Gly Arg Leu Arg Ala 180 185 190 Ile Leu Gln Asn Gln Gly Asn 195 <210> 2 <211> 202 <212> PRT <213> Artificial Sequence <400> 2 Met Glu Ala Ser Pro Ala Ser Gly Pro Arg His Leu Met Asp Pro His 1 5 10 15 Ile Phe Thr Ser Asn Phe Asn Asn Gly Ile Ser Val Arg Gly Arg His 20 25 30 Lys Thr Tyr Leu Cys Tyr Glu Val Glu Arg Leu Asp Asn Gly Thr Ser 35 40 45 Val Lys Met Asp Gln His Arg Gly Phe Leu His Asn Gln Ala Lys Asn 50 55 60 Leu Leu Cys Gly Phe Tyr Gly Arg His Ala Glu Leu Arg Phe Leu Asp 65 70 75 80 Leu Val Pro Ser Leu Gln Leu Asp Pro Ala Gln Ile Tyr Arg Val Thr 85 90 95 Trp Phe Ile Ser Trp Ser Pro Cys Phe Ser Trp Gly Cys Ala Gly Glu 100 105 110 Val Arg Ala Phe Leu Gln Glu Asn Thr His Val Arg Leu Arg Ile Phe 115 120 125 Ala Ala Arg Ile Tyr Asp Tyr Asp Pro Leu Tyr Lys Glu Ala Leu Gln 130 135 140 Met Leu Arg Asp Ala Gly Ala Gln Val Ser Ile Met Thr Tyr Asp Glu 145 150 155 160 Phe Lys His Cys Trp Asp Thr Phe Val Asp His Gln Gly Cys Pro Phe 165 170 175 Gln Pro Trp Asp Gly Leu Asp Glu His Ser Gln Ala Leu Ser Gly Arg 180 185 190 Leu Arg Ala Ile Leu Gln Asn Gln Gly Asn 195 200 <210> 3 <211> 609 <212> DNA <213> Artificial Sequence <400> 3 atggaagcca gcccagcatc cgggcccaga cacttgatgg atccacacat attcacttcc 60 aactttaaca atggcatttc ggtccgtgga aggcataaga cctacctgtg ctacgaagtg 180. gagcgcctgg acaatggcac ctcggtcaag atggaccagc acaggggctt tctacacaac caggctaaga atcttctctg tggcttttac ggccgccatg cggagctgcg cttcttggac 240 ctggttcctt ctttgcagtt ggacccggcc cagatctaca gggtcacttg gttcatctcc 300 tggagcccct gcttctcctg gggctgtgcc gggagtgc gtgcgttcct tcaggagaac 360 acacacgtga gactgcgtat cttcgctgcc cgcatctatg attacgaccc cctatataag gaggcactgc aaatgctgcg ggatgctggg gcccaagtct ccatcatgac ctacgatgaa 540. tttaagcact gctgggacac ctttgtggac caccaggat gtcccttcca gccctgggat ggactagatg agcacagcca agccctgagt gggaggctgc gggccattct ccagaatcag ggaaactga 609 <210> 4 <211> 4101 <212> DNA <213> The snowstorm <400> 4 gataaaaagt attctattgg tttagccatc ggcactaatt ccgttggatg ggctgtcata accgatgaat acaagtacc ttcaagaaa tttaggtgt tgggacac agaccgtcat 120 tcgattaaaa agaatcttat cggtgccctc ctattcgata gtggcgaac ggcagaggcg 180 actcgcctga aacgaaccgc tcggagagg tatacacgtc gcaagaccg atatgttac 240 ttacaagaaa ttttagcaa tgagatggcc aaagttgacg attctctt tcaccgtttg 300 gaagagtcct tccttgtcga agaggacaag aaacatgaac ggcacccat ctttggaac 360 atagtagatg aggtggcata tcatgaaaag tacccaacga tttacacct cagaaaaaag 420 ctagttgact siactgataa agcggacctg agttaatct acttggctct tgcccatatg 480 aaagttcc gtgggcactt tctcattgag ggtgatctaa atccggacaa ctcggatgtc 540 gandaactgt tcatccagtt agtacaacc tataatcagt tgtttgaaga gaaccctata 600 aatgcaagtg gcgtggatgc gaaggctatt cttagcgccc gccctctaa atcccgacgg 660 ctagaaaacc tgatcgcaca attacccgga gagagaaaa atggttgtt cggtaacctt 720 atagcgctct cactaggcct gataccaat tttaagtcga acttcgactt agctgaagat 780 gccaaattgc agcttagtaa ggacacgtac gatgacgatc tcgacaatct actggcacaa 840 attggagatc agtatgcgga cttattttg gctgccaaaa accttagcga tgcaatcctc 900 ctatctgaca tactgagagt tatactgag attaccaagg cgccgttatc cgcttcaatg 960 atcaaaaggt acgatgaaca tcaccaagac ttgacacttc tcaaggccct agtccgtcag 1020 caactgcctg agaaatataa ggaaatattc tttgatcagt cgaaaaacgg gtacgcaggt 1080 tatattgacg gcggagcgag tcaagaggaa ttctacaagt tttcaaacc catattagag 1140 aagatggatg ggacggaaga gttgcttgta aaactcaatc gcgaagatct actgcgaaag 1200 cagcggactt tcgacaacgg tagcattcca catcaaatcc acttaggcga attgcatgct 1260 atacttagaa ggcaggagga tttttatccg ttcctcaaag acaatcgtga aaagattgag 1320 aaaatcctaa cctttcgcat accttactat gtgggacccc tggcccgagg gaactctcgg 1380 ttcgcatgga tgacaagaaa gtccgaagaa acgattactc catggaattt tgaggaagtt 1440 gtcgataaag gtgcgtcagc tcaatcgttc atcgagagga tgaccaactt tgacaagaat 1500 ttaccgaacg aaaaagtatt gcctaagcac agtttacttt acgagtattt cacagtgtac 1560 aatgaactca cgaaagttaa gtatgtcact gagggcatgc gtaaacccgc ctttctaagc 1620 1680 aagcaattga aagagacta ctttaagaaa attgaatgct tcgattctgt cgagatctcc 1740 ggggtagaag atcgatttaa tgcgtcactt ggtacgtatc atgacctcct aaagataatt 1800 aaagataagg acttcctgga taacgaag aatgaagata tctttagaaga tatagtgttg 1860 actcttaccc tctttgaaga tcgggaaatg attgagaaa gactaaaaac atacgctcac 1920 ctgttcgacg ataaggttat gaaacagtta aagaggcgtc gctatacggg ctggggacga 1980 ttgtcgcgga aacttatcaa cgggataaga gacaagcaaa gtggtaaaac tattctcgat 2040 tttctaaaga gcgacggctt cgccaatagg aactttatgc agctgatcca tgatgactct 2100 ttaaccttca aagaggattat acaaaaggca caggtttccg gacaagggga ctcattgcac 2160 gaacatattg cgaatcttgc tggttcgcca gccatcaaaa agggcatact ccagacagtc 2220 aaagtagtgg atgagctagt taaggtcatg ggacgtcaca aaccggaaaa cattgtaatc 2280 gagatggcac gcgaaaatca aacgactcag aaggggcaaa aaaacagtcg agagcggatg 2340 aagagaatag aagagggtat taaagaactg ggcagccaga tcttaaagga gcatcctgtg 2400 gaaaataccc aattgcagaa cgagaaactt tacctctatt acctacaaaa tggaagggac 2460 atgtatgttg atcaggaact ggacataaac cgtttatctg attacgacgt cgatcacatt 2520 gtaccccaat cctttttgaa ggacgattca atcgacaata aagtgcttac acgctcggat 2580 aagaaccgag ggaaaagtga caatgttcca agcgaggaag tcgtaaagaa aatgaagaac 2640 tattggcggc agctcctaaa tgcgaaactg ataacgcaaa gaaagttcga taacttaact 2700 aaagctgaga ggggtggctt gtctgaactt gacaaggccg gattttata acgtcagctc 2760 gtggaaaccc gccaaatcac aaagcatgtt gcacagatac tagattcccg aatgaatacg 2820 aaatacgacg agaacgataa gctgattcgg gaagtcaaag taatcacttt aaagtcaaaa 2880 ttggtgtcgg acttcagaaa ggattttcaa ttctataaag ttagggagat aaataactac 2940 caccatgcgc acgacgctta tcttaatgcc gtcgtaggga ccgcactcat windowaatac 3000 ccgaagctag aaagtgagtt tgtgtatggt gattacaaag ttatgacgt ccgtaagatg 3060 atcgcgaaaa gcgaacagga gataggcaag gctacagcca atactctt ttattctaac 3120 attatgaatt tctttagac ggaatcact ctggcaacg gagagatacg caacgacct 3180 ttaattgaaa ccaatgggga gandaggtgaa atcgtatggg ataagggccg ggactcgcg 3240 acggtgagaa aagttttgtc catgccccaa gtcacatag taaagaaac tgaggtgcag 3300 accggagggt ttcaagga atcgattctt ccaaaagga atagtgataa gctcatcgct 3360 cgtaaaaagg actgggaccc gaaaaagtac ggtggctcg atagccctac agttgcctat 3420 tctgtcctag tagtggcaaa agttgagaag ggaaatcca agaaactgaa gtcagtcaa 3480 gatttattgg ggataacgat ttggagcgc tcgtctttg aaagaaccc catcgacttc 3540 cttgaggcga aaggttacaa ggaagtaaaaaggatctca taatttaac accaaagtat 3600 agtctgtttg agttagaaaa tggccgaaaa cggatgttgg ctagcgccgg agagcttcaa 3660 aaggggaacg aactcgcact accgtctaaa tacgtgaatt tcctgtattt agcgtcccat 3720 tacgagaagt tgaaaggttc acctgaagat aacgaacaga agcaactttt tgttgagcag 3780 cacaaacatt atctcgacga aatcatagag caaatttcgg aattcagtaa gagagtcatc 3840 ctagctgatg ccaatctgga caaagtatta agcgcataca acaagcacag ggataaaccc 3900 atacgtgagc aggcggaaaa tattatccat ttgtttactc ttaccaacct cggcgctcca 3960 gccgcattca agtattttga cacaacgata gatcgcaaac gatacacttc taccaaggag 4020 gtgctagacg cgacactgat tcaccaatcc atcacgggat tatatgaaac tcggatagat 4080 ttgtcacagc ttgggggtga c 4101 <210> 5 <211> 249 <212> DNA <213> Artificial Sequence <400> 5 actaatctgt cagatattat tgaaaaggag accggtaagc aactggttat ccaggaatcc 60 atcctcatgc tcccagagga ggtggaagaa gtcattggga acaagccgga aagcgatata 120 ctcgtgcaca ccgcctacga cgagagcacc gacgagaatg tcatgcttct gactagcgac 180 gcccctgaat acaagccttg ggctctggtc atacaggata gcaacggtga gaacaagatt 240 aagatgctc 249 <210> 6 <211> 21 <212> DNA <213> Artificial sequence <400> 6 cccaagaaga agaggaaagt c 21 <210> 7 <211> 225 <212> DNA <213> Artificial sequence <400> 7 ctgtgccttc tagttgccag ccatctgttg tttgcccctc ccccgtgcct tccttgaccc 60 tggaaggtgc cactcccact gtcctttcct aataaaatga ggaaattgca tcgcattgtc 120 tgagtaggtg tcattctatt ctggggggtg gggtggggca ggacagcaag ggggaggatt 180 gggttgacaa tagcaggcat gctggggatg cggtgggctc tatgg 225 <210> 8 <211> 705 <212> DNA <213> Artificial sequence <400> 8 atggtgagca agggcgagga ggtcatcaaa gagttcatgc gcttcaaggt gcgcatggag 60 ggctccatga acggccacga gttcgagatc gagggcgagg gcgagggccg cccctacgag 120 ggcacccaga ccgccaagct gaaggtgacc aagggcggcc ccctgccctt cgcctgggac 180 atcctgtccc cccagttcat gtacggctcc aaggcgtacg tgaagcaccc cgccgacatc 240 cccgattaca agaagctgtc cttccccgag ggcttcaagt gggagcgcgt gatgaacttc 300 gaggacggcg gtctggtgac cgtgacccag gactcctccc tgcaggacgg cacgctgatc 360 tacaaggtga agatgcgcgg caccaacttc ccccccgacg gccccgtaat gcagaagaaa 420 accatgggct gggaggcctc caccgagcgc ctgtaccccc gcgacggcgt gctgaagggc 480 gagatccacc aggccctgaa gctgaaggac ggcggccact acctggtgga gttcaagacc 540 atctacatgg ccaagaagcc cgtgcaactg cccggctact actacgtgga caccaagctg 600 gacatcacct cccacaacga ggactacacc atcgtggaac agtacgagcg ctccgagggc 660 cgccaccacc tgttcctgta cggcatggac gagctgtaca agtaa 705 <210> 9 <211> 343 <212> DNA <213> Artificial Sequence <400> 9 gagggcctat ttcccatgat tccttcatat ttgcatatac gatacaaggc tgttagagag 60 ataattggaa ttaatttgac tgtaaacaca aagatattag tacaaaatac gtgacgtaga 120 aagtaataat ttcttgggta gtttgcagtt ttaaaattat gttttaaaat ggactatcat 180 atgcttaccg taacttgaaa gtatttcgat ttcttggctt tatatatctt gtggaaagga 240 cgaaacaccg ggtcttcgag aagacctgtt ttagagctag aaatagcaag ttaaaataag 300 gctagtccgt tatcaacttg aaaaagtggc accgagtcgg tgc 343 <210> 10 <211> 600 <212> DNA <213> Artificial Sequence <400> 10 atggaagcca gcccagcatc cgggcccaga cacttgatgg atccacacat attcacttcc 60 aactttaaca atggcattgg aaggcataag acctacctgt gctacgaagt ggagcgcctg 120 gacaatggca cctcggtcaa gatggaccag cacaggggct ttctacacaa ccaggctaag 180 aatcttctct gtggctttta cggccgccat gcggagctgc gcttcttgga cctggttcct 240 tctttgcagt tggacccggc ccagatctac agggtcactt ggttcatctc ctggagcccc 300 tgcttctcct ggggctgtgc cggggaagtg cgtgcgttcc ttcaggagaa cacacacgtg 360 agactgcgta tcttcgctgc ccgcatctat gattacgacc ccctatataa ggaggcactg 420 caaatgctgc gggatgctgg ggcccaagtc tccatcatga cctacgatga atttaagcac 480 tgctgggaca cctttgtgga ccaccaggga tgtcccttcc agccctggga tggactagat 540 gagcacagcc aagccctgag tgggaggctg cgggccattc tccagaatca gggaaactga 600 <210> 11 <211> 609 <212> DNA <213> Artificial sequence <400> 11 atggaagcca gcccagcatc caggcccaga cacttgatgg atccaaacac gttcactttc 60 aactttaaca atgacctttc ggtccgtgga cggcaccaga cctacttgtg ctacgaggtg 120 gagcgcctgg acaatggcac ctgggtcccg atggacgagc gcaggggctt tctatgcaac 180 aaggctaaga atgttccctg tggtgattat ggctgccacg cggagctgtg cttcctgggc 240 gaggttcctt cttggcagtt ggacccggcc cagacgtaca gggtcacttg gttcatctcc 300 tggagcccct gcttcaggag gggctgtgcc gagcaagtgc gtgcgttcct tcaggagaac 360 acacacatga gactgcgcat ctttgctgcc cgcatctatg attacgattt cctgtatcag 420 gaggcactgc gacgctgcg ggatgctggg gcccaagtct ccatcatgac ctacgagga 540. tttaagcact gctgggacac ctttgtggac cgccagggac gtcccttcca gccctgggat ggactagatg agcacagcca agccctgagt gggaggcttc gggacattct ccagaatcag ggaaactga 609

Claims

1. A cytosine deaminase mutant derived from human, characterized in that: The mutant is based on the parent enzyme with an amino acid sequence such as SEQ ID NO: 1, and three amino acid residues, serine, valine and arginine, are sequentially inserted between the 26th and 27th amino acids.

2. A gene encoding the cytosine deaminase mutant according to claim 1.

3. An expression cassette, characterized in that Comprising the gene according to claim 2.

4. The expression cassette according to claim 3, wherein The expression cassette also contains a promoter, nCas9-D10A, a uracil DNA glycosylase inhibitor, a nuclear localization sequence and a termination sequence; the promoter starts the expression of the gene according to claim 2, and the sequence is, in order of connection, a promoter, the gene according to claim 2, nCas9-D10A, a uracil DNA glycosylase inhibitor, a nuclear localization sequence and a termination sequence; the nucleotide sequence of the nCas9-D10A is shown in SEQ ID NO:4, the nucleotide sequence of the uracil DNA glycosylase inhibitor is shown in SEQ ID NO:5, the nucleotide sequence of the nuclear localization sequence is shown in SEQ ID NO:6, and the nucleotide sequence of the termination sequence is shown in SEQ ID NO:

7.

5. A CBE single-base editor based on human cytosine deaminase, characterized in that: The CBE single-base editing system consists of four parts. The first part is the transfection efficiency indicator part, which includes a red fluorescent protein and a promoter. The second part is the sgRNA transcription unit, which includes a carrier frame for inserting the sgRNA sequence and its promoter. The third part comprises the expression cassette according to any one of claims 3 to 4, and the fourth part comprises the high copy replication origin ori from the colicin factor and the ampicillin resistance selection gene AmpR.

6. The CBE single-base editor according to claim 5, wherein The nucleotide sequence of the red fluorescent protein is shown in SEQ ID NO:8, and the nucleotide sequence of the sgRNA transcription unit is shown in SEQ ID NO:9.

Citation Information

Patent Citations

  • Construction method of novel base conversion editing system and application thereof

    CN110835629A

  • Novel base conversion editing system and application thereof

    CN110835634A