A cytosine deaminase editor and an editing system
By designing a combined optimization of cytosine deaminoglycemia editor, using C-to-T editing of targeted chains to achieve G-to-A editing of non-targeted chains, the problem that existing tools cannot efficiently implement G-to-A editing, improve editing efficiency and product purity, and expand the scope of application.
Patent Information
- Application Number
- CN202510042739.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-01-10
AI Technical Summary
Existing base editing tools cannot efficiently implement G-to-A editing on non-targeted chains.
A cytosine deamination editor was designed to realize C-to-T editing of the targeted chain by combining the optimized enUn1Cas12f1 protein and evoCDA1 deaminase, thereby indirectly realizing G-to-A editing of the non-targeted chain.
This editor not only improves the editing efficiency and product purity of the targeted chain, but also expands the application scope of existing base editing tools and realizes efficient editing of the non-targeted chain G-to-A.
Smart Images

Figure CN119776330B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of genome editing, and particularly to a cytosine deaminase editor and an editing system. Background Art
[0002] The emergence of base editing tools has significantly broadened the application of the CRISPR / Cas system, expanding from the initial gene knockout and gene regulation to more precise base-level editing. Currently, base editing is mainly achieved through tools such as deaminases and glycosylases, supporting base conversions or transversions. Since these enzymes generally prefer to act on single-stranded DNA and bind to the R-loop structure formed by Cas9 during targeting, existing base editors mainly edit bases on the non-target strand (NTS), including the classical deaminase-based editing tools BE4max and ABE8e, which can achieve C-to-T and A-to-G conversions respectively. It is worth noting that the DddA deaminase, as a double-stranded deaminase, has a different editing mode.
[0003] In addition, glycosylase-based base editors such as gGBE, AYBE, CGBE, and gTBE achieve base conversions or transversions, including C-to-G, T-to-S, A-to-Y, and G-to-Y editing, by cleaving the target base in the non-target strand (NTS) to generate an apurinic (AP) site and relying on the DNA repair mechanism. However, there is still a lack of a tool in existing base editing tools that can efficiently achieve G-to-A editing on the non-target strand.
[0004] Previous studies have found that base editors constructed based on inactivated SpCas9 (dSpCas9) occasionally exhibit rare off-target activities, enabling them to achieve C-to-T editing on the template strand (TS), thereby indirectly causing G-to-A changes on the non-target strand. However, conventional base editing systems usually use Cas9Nickase mutants, under which DNA repair is more inclined to use the non-target strand as a template, making these off-target edits difficult to detect after repair. Summary of the Invention
[0005] Aiming at the limitation that existing cytosine base editing tools cannot achieve G-to-A editing on the non-target strand, the present invention provides a cytosine deaminase editor and an editing system.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] In a first aspect, the present invention provides a cytosine deaminase editor, which is sequentially composed of a deaminase evoCDA1, an XTEN sequence, and an enUn1Cas12f1 protein connected in sequence from the N-terminus to the C-terminus. The amino acid sequences of the deaminase evoCDA1, the XTEN sequence, and the enUn1Cas12f1 protein are as shown in SEQ ID NO.4 and SEQ ID NO.5 in sequence. The enUn1Cas12f1 has 8 amino acid mutations compared to the existing Un1Cas12f1 protein (NCBI accession number: A0A482D308.1), which are 143R, 244R, 393R, 437R, 447R, 331E, 147R, and 203R in sequence.
[0008] Through the combination of amino acid mutations, the present invention can effectively increase the stability of the protein-nucleic acid complex and thereby enhance its editing efficiency. The optimized enUn1Cas12f1 has the ability to efficiently cleave double-stranded DNA, and when this optimized protein is combined with the evoCDA1 deaminase with high cytosine deamination ability, it can achieve preferential C-to-T editing of the TS strand.
[0009] As a preference of the cytosine deaminase editor of the present invention, the 473rd amino acid of the enUn1Cas12f1 protein also has an alanine mutation. Further performing an alanine mutation on the 473rd amino acid of enUn1Cas12f1 can achieve more efficient C-to-T editing of the target strand at the target site and efficient G-to-A editing of the non-target strand, thereby reducing the amino acid mutation of the C-to-T editing efficiency of the non-target strand to improve the purity of the editing product.
[0010] As a further preference of the cytosine deaminase editor of the present invention, a UGI is further connected to the C-terminus of the cytosine deaminase editor, and its amino acid sequence is as shown in SEQ ID NO.3. This cytosine deaminase editor has the ability to edit C-to-T of the target strand at the target site and an editing window far from the PAM end, effectively expanding the application range of existing base editing tools.
[0011] As a further preference of the cytosine deaminase editor of the present invention, the cytosine deaminase editor further includes a DNA-binding protein, the DNA-binding protein is an HMG-D protein, and its coding gene sequence is as shown in SEQ ID NO.7. The DNA-binding protein is located between the enUn1Cas12f1 protein and the deaminase evoCDA1.
[0012] Higher targeted strand editing efficiency and editing product purity can be obtained by adding HMG-D protein to the cytosine deamination editor, and the editing window range remains unchanged. By means of amino acid mutations of enUn1Cas12f1 near the PAM region, the PAM recognition sequence can be extended from TTTR to NTTR.
[0013] In a second aspect, the present invention provides a cytosine deamination editing system comprising the cytosine deamination editor, and the cytosine deamination editor is guided by sgRNA to a target editing site.
[0014] In a third aspect, the present invention provides a polynucleotide encoding the cytosine deamination editor or the cytosine deamination editing system.
[0015] In a fourth aspect, the present invention provides a vector comprising the polynucleotide.
[0016] In a fifth aspect, the present invention provides a cell comprising the vector.
[0017] In a sixth aspect, the present invention provides a gene editing pharmaceutical composition comprising the cytosine deamination editor, the cytosine deamination editing system, the polynucleotide, the vector or the cell.
[0018] Based on the fact that the cytosine deamination editor performs C-to-T editing on the targeted strand during deamination catalysis and indirectly causes G-to-A changes in the non-targeted strand, the pharmaceutical composition provided by the present invention is expected to further expand the application scenarios in this field and promote the wide application of precise gene editing in the fields of disease treatment, basic research, etc.
[0019] The present invention has the following beneficial effects:
[0020] Compared with the existing cytosine base tools, the cytosine deamination editor proposed by the present invention has a change in the targeted object, which changes from editing the non-targeted strand to editing the targeted strand. The original base editing tool cannot achieve G-to-A editing of the non-targeted strand, and the cytosine deamination editor of the present invention can indirectly achieve it through C-to-T editing of the targeted strand. The editing window of the cytosine deamination editor of the present invention is located at 11-25 nt away from the PAM end, forming a complement to the editing window of the existing base editing tool. At the same time, the improved enUn1Cas12f1 used in the present invention itself has higher DNA double-strand cleavage activity compared with the existing Cas12f tools, expanding the application potential of Cas12f. And due to some amino acid mutations being near the PAM sequence, the PAM recognition range of enUn1Cas12f1 is extended from TTTR to NTTR, having a broader targeting range. In summary, the development of this tool expands the application scope of the existing cytosine base editing tools and has broad application prospects. Brief Description of the Drawings
[0021] Figure 1 It is a schematic diagram for the TSminiCBE to achieve targeted strand C-to-T editing.
[0022] Figure 2 It is a graph showing the editing efficiency of different combinations of deaminases and Un1Cas12f1 protein at two gene loci.
[0023] Figure 3 It is the comparison result of the editing efficiency between the improved enUn1Cas12f and EnAsCas12f. Among them, A is the comparison of the editing efficiency at each site in HEK293T cells; B is the statistical analysis of the editing efficiency in HEK 293T cells; C is the comparison of the editing efficiency at each site in HeLa cells; D is the statistical analysis of the editing efficiency in HeLa cells.
[0024] Figure 4 It is the evaluation of the average editing efficiency of TSminiCBE on 12 targets in different cell lines. A is HeLa cells, and B is AGS cells.
[0025] Figure 5 It is the comparison of the editing efficiency between HMG-TSminiCBE and TSminiCBE after adding DNA binding protein. Among them, A is the schematic diagram of the editing tool after adding different DNA binding proteins; B is the comparison of the editing efficiency at each site in HEK293T cells; C is the statistical analysis of the editing efficiency of different editors in HEK 293T cells; D is the purity of the editing products of different editors in HEK 293T cells; E is the statistical analysis of the Indel frequency of different editors in HEK 293T cells.
[0026] Figure 6 It is the exploration of the target adjacent motif (PAM) sequence between enUn1Cas12f1 and HMG-TSminiCBE. A is the comparison of the editing efficiency of the reporting system of enUn1Cas12f1 and Un1Cas12f1 under the PAM of NTTR in HEK293T cells; B is the statistical analysis of the editing efficiency of enUn1Cas12f1 and Un1Cas12f1 under the NTTR PAM of the reporting system; C is the comparison of the editing efficiency of the endogenous gene locus of enUn1Cas12f1, Un1Cas12f1 and enAsCas12f under the PAM of NTTR in HEK 293T cells; D is the statistical analysis of the editing efficiency of enUn1Cas12f1, Un1Cas12f1 and enAsCas12f under the PAM of NTTR; E is the editing efficiency of HMG-TSminiCBE under the NTTR PAM in HEK 293T cells. Detailed Implementation Manner
[0027] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments, but it should not be construed as a limitation of the present invention. Unless otherwise specified, the technical means used in the following embodiments are conventional means well known to those skilled in the art. The materials, reagents, etc. used in the following embodiments can be obtained from commercial sources unless otherwise specified.
[0028] The sources of the vectors involved in the following examples are as follows:
[0029] pCMV-Un1Cas12f1 vector: Addgene, catalog number: 174820.
[0030] pGL3-U6-gRNA4.1 vector: Addgene, catalog number: 107721.
[0031] pCMV-enAsCas12f vector: Addgene, catalog number: 204637, which contains the EnAsCas12f (or denoted as enAsCas12f) protein. For comparison, its effect is compared with that of the cytosine deaminase editor of the present invention.
[0032] U6-AsCas12f-v4.1 vector: Addgene, catalog number: 204636.
[0033] Example 1: Combining different deaminases with Un1Cas12f1 protein to achieve C-to-T editing of the target strand
[0034] Taking two sites of the RIT1 gene in the HEK 293T cell genome as the target editing sites, different deaminases were combined with Un1Cas12f1 protein for C-to-T editing. Based on the characteristic that the Un1Cas12f1 single RuvC domain completes DNA double-strand cleavage, the possible working mode of Un1Cas12f1-deaminase for C-to-T editing of the TS strand was speculated, as Figure 1 shown.
[0035] I. Preparation of combinations of different deaminases and Un1Cas12f1 protein and sgRNA expression vectors
[0036] 1. Construction of expression vectors for combinations of different deaminases and Un1Cas12f1 protein
[0037] The nucleotide sequence based on the Un1Cas12f1 protein (amino acid sequence as shown in SEQ ID NO.1) and the nucleotide sequences of deaminases (evoCDA1, hA3A, rAPOBEC, Anc689) were ligated at the N-terminus of the Un1Cas12f1 protein using XTEN (amino acid sequence as shown in SEQ ID NO.2). Recombinant vectors were constructed using homologous recombination methods, inserting XTEN and deaminases into pCMV-Un1Cas12f1 to obtain vectors Un1Cas12f1-evoCDA1, Un1Cas12f1-hA3A, Un1Cas12f1-rAPOBEC1, Un1Cas12f1-Anc689, in which Un1Cas12f1 was introduced with D326A and D510A mutations to become a nuclease-inactivated form. Except for Un1Cas12f1-Anc689 which was ligated with 2 UGI proteins (amino acid sequence as shown in SEQ ID NO.3) at the C-terminus, the remaining vectors were each ligated with only 1 UGI protein at the C-terminus. After all vectors were constructed, plasmid sequencing was required to confirm the absence of mutants before they could be used in downstream experiments.
[0038] The sequence of the evoCDA1 is as shown in SEQ ID NO.4, the NCBI accession number of Anc689 is WYV95014.1, the NCBI accession number of hA3A is NP_001180218.1, and the NCBI accession number of rAPOBEC1 is WYV95015.1.
[0039] SEQ ID NO.1: MAKNTITKTLKLRIVRPYNSAEVEKIVADEKNNREKIALEKNKDKVKEACSKHLKVAAYCTTQVERNACLFCKARKLDDKFYQKLRGQFPDAVFWQEISEIFRQLQKQAAEIYNQSLIELYYEIFIKGKGIANASSVEHYLSDVCYTRAAELFKNAAIASGLRSKIKSNFRLKELKNMKSGLPTTKSDNFPIPLVKQKGGQYTGFEISNHNSDFIIKIPFGRWQVKKEIDKYRPWEKFDFEQVQKSPKPISLLLSTQRRKRNKGWSKDEGTEAEIKKVMNGDYQTSYIEVKRGSKIGEKSAWMLNLSIDVPKIDKGVDPSIIGGIDVGVKSPLVCAINNAFSRYSISDNDLFHFNKKMFARRRILLKKNRHKRAGHGAKNKLKPITILTEKSERFRKKLIERWACEIADFFIKNKVGTVQMENLESMKRKEDSYFNIRLRGFWPYAEMQNKIEFKLKQYGIEIRKVAPNNTSKTCSKCGHLNNYFNFEYRKKNKFPHFKCEKCNFKENADYNAALNISNPKLKSTKEEP。
[0040] SEQ ID NO.2: SGSETPGTSESATPES。
[0041] SEQ ID NO.3: SGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKMLSGGS。
[0042] SEQ ID NO.4: STDAEYVRIHEKLDIYTFKKQFSNNKKSVSHRCYVLFELKRRGERRACFWGYAVNKPQSGTERGIHAEIFSIRKVEEYLRDNPGQFTINWYSSWSPCADCAEKILEWYNQELRGNGHTLKIWVCKLYYEKNARNQIGLWNLRDNGVGLNVMVSEHYQCCRKIFIQSSHNQLNENRWLEKTLKRAEKRRSELSIMFQVKILHTTKSPAV。
[0043] 2. Construction of sgRNA expression vector targeting the target site
[0044] (1) Design and synthesize spacer oligonucleotides targeting 2 target sites (Table 1). The lowercase letters in the sequence are the linker parts.
[0045] Table 1 Spacer oligonucleotide sequences at different sites of the RIT1 gene
[0046]
[0047]
[0048] (2) Prepare the Buffer buffer for annealing. The formula is shown in Table 2:
[0049] Table 2 Composition of Buffer buffer
[0050] Reagent Dosage NaCl 0.08766g 10 mM Tris-HCl buffer (pH = 8.5) 0.2 mL <![CDATA[ddH 2 O]]> 30 mL
[0051] (3) Anneal the spacer oligonucleotide chains to obtain the spacer annealing product. The annealing system and program are shown in Tables 3 and 4.
[0052] Table 3 Annealing system
[0053] Reagent Dosage Oligonucleotide sequence forward (10 μM) 1 μL Oligonucleotide sequence reverse (10 μM) 1 μL Annealing Buffer 2 μL <![CDATA[ddH 2 O]]> 6 μL Total system 10 μL
[0054] Table 4 Annealing program
[0055] Temperature Time Number of cycles 95℃ 5 min 1 95℃ 30s 1 From 85 °C to 25 °C -1 °C / cycle 60 cycles From 25 °C to 5 °C -2 °C / cycle 10 cycles 4℃ Forever 1
[0056] (5) Using the pGL3-U6-gRNA4.1 plasmid as a template, obtain the expression plasmid with the target spacer sequence by enzymatic digestion and ligation. The specific operations are as follows:
[0057] 1) Use the F primer: 5'-agctaggtctccttttatttttttaaagaattctcgacctcgagacaaatg-3', R primer: 5'-tctctcggtctcagttgcattcctttctttgtttcgagggttactttc-3' to amplify pGL3-U6-gRNA4.1 to linearize it and add a Bsa I restriction site at the end. The amplification system and program are shown in Tables 5 and 6.
[0058] Table 5 PCR amplification system of pGL3-U6-gRNA4.1 plasmid
[0059] Reagent Dosage 2×Phanta Flash Master Mix (Dye Plus) 12.5 μL Template plasmid 1 μL F (10 μM) 0.5 μL R (10 μM) 0.5 μL <![CDATA[ddH 2 O]]> 10.5 μL Total system 25 μL
[0060] 2) Use Bsa I enzyme to perform a single digestion reaction on the linearized vector to generate sticky ends.
[0061] Table 6 pGL3-U6-gRNA4.1 plasmid PCR amplification program
[0062]
[0063] 3) Purify and recover the backbone of the sgRNA expression plasmid after digestion. Use T4 DNA Ligase to ligate the annealed product of the spacer targeting the target site with the backbone of the sgRNA expression plasmid overnight at 16°C. The ligation system is shown in Table 7 below.
[0064] Table 7 Ligation system
[0065]
[0066]
[0067] (6) Transform the ligation product into DH5α competent cells. Pick monoclonal colonies the next day for Sanger sequencing. Expand the culture of the bacterial solution with correct sequencing results and extract the plasmid. The obtained plasmid is the sgRNA expression plasmid targeting the target site.
[0068] II. Verification of editing efficiency by transfection of HEK 293T cells
[0069] Perform transient transfection on the plasmid components targeting the same site and measure the editing efficiency.
[0070] 1. Seed cells
[0071] Resuscitate the cryopreserved HEK 293T cells and culture them in a 10 cm culture dish. Add 7 mL of complete medium (90% high-glucose DMEM + 10% fetal bovine serum + working concentration of penicillin-streptomycin) and culture at 37°C under 5% CO 2 conditions. When the cell density reaches 90%, seed the cells into a 24-well plate and continue culturing.
[0072] 2. Cell transfection and sorting
[0073] (1) When the cell density in the 24-well plate grows to 75% - 85%, use the EZ trans transfection reagent to co-transfect each component expression plasmid into the cells according to the instructions. The transfection mixture system is shown in Table 8 below:
[0074] Table 8 Transfection mixture system
[0075] Composition Dosage pCMV-Un1Cas12f1-deaminase 900 ng sgRNA expression plasmid 300 ng EZ trans transfection reagent 2.5 μL DMEM culture medium 100 μL
[0076] (2) Let the mixed system stand at room temperature for 10 minutes.
[0077] (3) Add the above-prepared and statically settled transfection solution to each well of the cells.
[0078] (4) After 8 hours of transfection, remove the culture medium containing the transfection reagent and add 700 μL of complete medium.
[0079] (5) After 72 hours of transfection, remove the culture medium, wash each well of the cells with 200 μL of PBS solution, then digest and collect the cells into a 1.5 mL centrifuge tube, and resuspend the cell pellet with 260 μL of PBS solution.
[0080] (6) Filter the resuspended cells through a cell strainer to form a single-cell suspension, add it to a flow tube, and perform FACS sorting using a flow cytometer to collect 10,000 - 20,000 cells with the top 20% GFP fluorescence intensity.
[0081] 3. Detection of C-to-T editing efficiency
[0082] Centrifuge the above-collected cells and add cell lysis buffer for lysis. Use Phanta Super-Fidelity DNA Polymerase to amplify the DNA sequence containing the target site. The PCR reaction system and conditions are shown in Tables 9 and 10 below.
[0083] Table 9 PCR amplification system for DNA sequence containing the target site
[0084] Composition Dosage Buffer 12.5 μL dNTP Mix 0.5 μL Primer F 0.5 μL Primer R 0.5 μL Phanta Super-Fidelity DNA Polymerase 0.5 μL Cell lysate 1 - 2 μL <![CDATA[ddH 2 O]]> Make up to 25 μL
[0085] Table 10 PCR amplification program for DNA sequence containing the target site
[0086]
[0087] After the PCR products are detected by agarose gel electrophoresis and the bands are specific, they are purified and recovered, and then Sanger sequencing and targeted deep sequencing are performed respectively for editing efficiency analysis. The results are as Figure 2 shown. The red-marked ones are TS, and the black ones are NTS. The results show that the combination of Un1Cas12f1-evoCDA1 can effectively achieve G-to-A editing of the NTS strand (C-to-T editing of the TS strand).
[0088] Example 2: Comparison of editing efficiency between enUn1Cas12f1 and enAsCas12f
[0089] Based on the combination of computational simulation and saturation mutagenesis strategies, this study found that combining 8 point mutations can significantly improve the cleavage efficiency of Un1Cas12f1 at the eukaryotic cell level. The present invention names this mutant protein enUn1Cas12f1 (amino acid sequence as SEQ ID NO.5). enAsCas12f is currently the Cas12f protein with the best editing efficiency. To illustrate that enUn1Cas12f1 developed in the present invention has better editing performance, in this example, the editing efficiencies of enUn1Cas12f1 and enAsCas12f at the cell level were compared.
[0090] SEQ ID NO.5: MAKNTITKTLKLRIVRPYNSAEVEKIVADEKNNREKIAL EKNKDKVKEACSKHLKVAAYCTTQVERNACLFCKARKLDDKFYQKLRGQFPDAVFWQEISEIFRQLQKQAAEIYNQSLIELYYEIFIKGKGIANASSVEHYLSRVCYRRAAELFKNAAIASGLRSKIKSNFRLKELKNMKSGLPTTKSDNFPIPLVKQKGGQYRGFEISNHNSDFIIKIPFGRWQVKKEIDKYRPWEKFDFEQVRKSPKPISLLLSTQRRKRNKGWSKDEGTEAEIKKVMNGDYQTSYIEVKRGSKIGEKSAWMLNLSIDVPKIDKGVDPSIIGGIDVGVKEPLVCAINNAFSRYSISDNDLFHFNKKMFARRRILLKKNRHKRAGHGAKNKLKPITILTEKSRRFRKKLIERWACEIADFFIKNKVGTVQMENLESMKRKEDSYFNRRLRGFWPYARMQNKIEFKLKQYGIEIRKVAPNNTSKTCSKCGHLNNYFNFEYRKKNKFPHFKCEKCNFKENADYNAALNISNPKLKSTKEEP.
[0091] I. Development of enUn1Cas12f1
[0092] 1. Obtain mutant candidate sites using computational simulation methods
[0093] (1) Preparation steps before mutation simulation with Discovery Studio 2019 software. After entering the software interface, first load the protein structure file (7l49), select the Prepare Protein option under the Macromolecules column, and select the Clean Protein operation from the dropdown. Enter the Change Forcefield option under the Simulation column, and select the CHARMm force field in Forcefield from the dropdown to assign to the protein structure.
[0094] (2) Mutation process with Discovery Studio 2019 software. Select all atoms of the entire protein, enter the Design Protein option under the Macromolecules column, select Calculation Mutation Energy (Stability) in Mutation Energy from the dropdown, select all amino acid sites, and select all amino acid mutation types, then the calculation can start. The result is output as an Excel file, and the smaller the value, the more stable the structure after mutation.
[0095] 2. Screening for saturation mutations and forward single-point mutations
[0096] (1) Construct a fluorescence reporter system for rapid characterization of editing efficiency at the cellular level. The reporter system architecture is CMV-BFP-P2A-Target-GFP, where the length of Target before GFP is 31bp, causing a frameshift mutation in GFP and not emitting light under normal circumstances. After using Un1Cas12f1 to target and cut Target, the Indel generated by DNA repair has a probability of restoring the downstream GFP to the normal coding frame, and it can emit light normally. The higher the cleavage activity, the higher the GFP fluorescence ratio, which can be quickly detected using a flow analyzer.
[0097] (2) Steps for screening saturation mutations. Use two degenerate bases, KNB and MNB, to cover 20 amino acids. For the candidate mutation site, use KNB and MNB to construct a small amino acid mutation library. After extracting the plasmid, perform cell transfection with the above reporter system. The specific transfection steps are the same as in Example 1. 48 hours after transfection, digest the cells into a flow tube and use a flow analyzer to detect the GFP fluorescence ratio, then it can be known which mutation library in which amino acid site can improve the editing efficiency. Then split all the amino acid types included in this library to construct single-point mutation plasmids, and perform transfection and fluorescence ratio analysis here to obtain the forward single-point mutations that can improve the editing efficiency.
[0098] (3) Construction of enUn1Cas12f1 by single-point mutation combination. Based on the single-point mutations obtained through saturation mutagenesis screening, using the Discovery Studio 2019 software, following the same process as above, perform energy calculations for combined mutations, select the top-ranked mutation combinations for experimental analysis, select the triple-mutation combination with the highest efficiency improvement, randomly design Un1Cas12f1 variants containing more mutation sites for fluorescence reporter system testing, and select the variant with the optimal efficiency as enUn1Cas12f1.
[0099] II. Construction of gRNA4.1 and As-sgRNA vectors
[0100] Design and synthesize spacer oligonucleotides targeting 8 genes (Table 11). The lowercase letters in the sequence are the linker parts (for As-sgRNA, just change the linker sequence Forward to GAAC, use U6-AsCas12f-v4.1 as the backbone after digestion, and keep the rest of the information unchanged). The subsequent steps of ligation, transformation, monoclonal sequencing, and plasmid acquisition are the same as in Example 1.
[0101] Table 11 Spacer oligonucleotide sequences of eight target genes
[0102]
[0103]
[0104] III. Cell transfection and sorting
[0105] Use a 24-well plate for cell transfection. The specific steps of culturing and transfecting HEK 293T cells are the same as in Example 1. The culture conditions and transfection steps of HeLa cells are the same as those of HEK 293T cells.
[0106] IV. Detection of editing efficiency
[0107] Lyse the sorted positive cells to obtain genomic DNA. The steps of PCR amplification of the target fragment and library construction for sequencing are the same as in Example 1. The statistical results of editing efficiency are as Figure 3 shown. enUn1Cas12f1 shows higher editing efficiency than EnAsCas12f in both HEK 293T and HeLa cell lines.
[0108] Example 3: Development of TSminiCBE and determination of editing efficiency
[0109] The enUn1Cas12f1 developed based on the above embodiments was combined with the evoCDA1 deaminase to develop a TS strand base editor with higher editing activity. Alanine scanning mutations were performed on the domain that cleaves the TS strand to construct a TSminiCBE editor with higher editing activity, and the editing efficiency of TSminiCBE was tested in multiple cell lines.
[0110] I. Development of TSminiCBE
[0111] 1. enUn1Cas12f1 further improves editing efficiency
[0112] Eight point mutations contained in enUn1Cas12f1 were introduced into the Un1Cas12f1-evoCDA1 vector constructed in Example 1 based on homologous recombination (while restoring the nuclease activity of Un1Cas12f1, A510D and A326D), and the editing efficiency was tested at three gene loci in HEK293T cells. The gene locus information is shown in the following table (Table 12).
[0113] Table 12 Nucleotide sequences targeting CBE genes
[0114]
[0115]
[0116] 2. Construction of TSminiCBE by alanine scanning mutation
[0117] Alanine scanning mutations were performed on the domain that interacts with the TS strand (RuvC+TNB) to construct mutant proteins, and the editing efficiency was detected at two gene loci in HEK 293T cells (the gene locus information is shown in Table 13 below). It was determined that the 473A mutation could effectively improve the editing efficiency of the TS strand. Based on this, evoCDA1-enUn1Cas12f1-473A was determined to be TSminiCBE, and its amino acid sequence is shown in SEQ ID NO.6, which contains a signal peptide sequence and a linker sequence.
[0118] SEQ ID NO.6: MPKKKRKVSTDAEYVRIHEKLDIYTFKKQFSNNKKSVS HRCYVLFELKRRGERRACFWGYAVNKPQSGTERGIHAEIFSIRKVEEYLRDNPGQFTINWYSSWSPCADCAEKILEWYNQELRGNGHTLKIWVCKLYYEKNARNQIGLWNLRDNGVGLNVMVSEHYQCCRKIFIQSSHNQLNENRWLEKTLKRAEKRRSELSIMFQVKILHTTKSPAVSGSETPGTSESATPESPKKKRKVGIHGVPAAMAKNTITKTLKLRIVRPYNSAEVEKIVADEKNNREKIALEKNKDKVKEACSKHLKVAAYCTTQVERNACLFCKARKLDDKFYQKLRGQFPDAVFWQEISEIFRQLQKQAAEIYNQSLIELYYEIFIKGKGIANASSVEHYLSRVCYRRAAELFKNAAIASGLRSKIKSNFRLKELKNMKSGLPTTKSDNFPIPLVKQKGGQYRGFEISNHNSDFIIKIPFGRWQVKKEIDKYRPWEKFDFEQVRKSPKPISLLLSTQRRKRNKGWSKDEGTEAEIKKVMNGDYQTSYIEVKRGSKIGEKSAWMLNLSIDVPKIDKGVDPSIIGGIDVGVKEPLVCAINNAFSRYSISDNDLFHFNKKMFARRRILLKKNRHKRAGHGAKNKLKPITILTEKSRRFRKKLIERWACEIADFFIKNKVGTVQMENLESMKRKEDSYFNRRLRGFWPYARMQNKIEFKLKQYGIEIRKVAPNNTSATCSKCGHLNNYFNFEYRKKNKFPHFKCEKCNFKENADYNAALNISNPKLKSTKEEPKRPAATKKAGQAKKKKSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKMLSGGSPKKKRKV*。
[0119] Table 13 Nucleotide sequences targeting the DNMT1 gene and the HNRNPK gene
[0120]
[0121] II. Editing efficiency test of multiple cell lines
[0122] The gRNA4.1 vector used for constructing HeLa and AGS cell lines follows the same design principle as in Example 1. The specific information is shown in Table 14 below. The lowercase letters in the sequence are the linker part. The subsequent construction process and plasmid acquisition are the same as in Example 1.
[0123] Table 14 Spacer oligonucleotide sequences at different sites in HeLa and AGS cell lines
[0124]
[0125]
[0126] AGS cells were cultured using RPM 1640 medium, and the rest was the same as the culture methods of HeLa and HEK 293T cells. The cell transfection and sorting processes were the same as in Example 1.
[0127] The steps of library construction, sequencing, and editing efficiency detection were the same as those in Example 1. The results are as Figure 4 shown. TSminiCBE could achieve an average editing efficiency higher than 30% in both cell lines, and the editing window was located after 11 nt.
[0128] Example 4: Comparison of editing efficiency between HMG-TSminiCBE and TSminiCBE
[0129] Based on the DNA-binding protein HMG-D (gene sequence as SEQ ID NO.7) and Ssod7, they were ligated into TSminiCBE to construct an enhanced version of the efficiency. Three sgRNAs were designed for the target sites of IRF6 and PCSK9 genes in the HEK 293T cell genome to measure the editing efficiency of HMG-TSminiCBE and TSminiCBE.
[0130] SEQ ID NO.7: ATGAGCGACAAGCCAAAGAGACCTCTGAGCGCCTAC ATGCTGTGGCTGAACAGCGCTAGAGAATCCATCAAGCGCGAAAACCCCGGCATCAAGGTGACCGAGGTGGCCAAGCGGGGAGGCGAGCTGTGGCGGGCCATGAAGGACAAGTCTGAATGGGAGGCCAAGGCCGCTAAAGCCAAAGACGACTACGACAGAGCCGTGAAGGAATTCGAGGCAAATGGCGGCAGCAGCGCCGCCAACGGCGGAGGCGCCAAGAAAAGAGCCAAGCCTGCTAAGAAGGTCGCCAAGAAGTCCAAGAAAGAGGAATCTGATGAGGACGATGATGACGAGAGCGAG。
[0131] I. Construction of TSminiCBE vector incorporated with DNA binding protein
[0132] 1. Preparation of vector incorporated with DNA binding protein
[0133] The schematic diagram of TSminiCBE enhancer protein constructed based on HMG-D and Ssod7 proteins is shown in A of Figure 5 as shown, and the above vector was constructed using the homologous recombination method.
[0134] 2. Preparation of sgRNA vector targeting three sites and determination of editing efficiency
[0135] Design and synthesize spacer oligonucleotide sequences targeting three sites (Table 15), where the lowercase letters in the sequences are linker parts. The annealing ligation of the gRNA4.1 vector and the subsequent plasmid extraction procedures are the same as those in Example 1.
[0136] Table 15 Spacer oligonucleotide sequences targeting three target sites
[0137]
[0138] 3. Detection of editing efficiency in HEK 293T cells
[0139] The transfection and sorting steps of HEK 293T cells are the same as those in Example 1. The sorted cells were lysed to obtain the genome and the DNA sequences near the editing sites were amplified. The PCR steps and procedures are the same as those in Example 1. After amplification, the products were library constructed and deep sequenced. The related steps are the same as those in Example 1. The results are shown in Figure 5As shown, it indicates that the addition of HMG-D can further improve the G-to-A editing efficiency of TSminiCBE on the NTS strand (C-to-T editing of the TS strand), and the editing efficiency of HMG-D-M (amino acid sequence as SEQ ID NO.8, gene sequence as SEQ ID NO.9) located between evoCDA1 and enUn1Cas12f1-473A is the highest.
[0140] SEQ ID NO.8: MPKKKRKVSTDAEYVRIHEKLDIYTFKKQFSNNKKSVSHRCYVLFELKRRGERRACFWGYAVNKPQSGTERGIHAEIFSIRKVEEYLRDNPGQFTINWYSSWSPCADCAEKILEWYNQELRGNGHTLKIWVCKLYYEKNARNQIGLWNLRDNGVGLNVMVSEHYQCCRKIFIQSSHNQLNENRWLEKTLKRAEKRRSELSIMFQVKILHTTKSPAVSGSETPGTSESATPESMSDKPKRPLSAYMLWLNSARESIKRENPGIKVTEVAKRGGELWRAMKDKSEWEAKAAKAKDDYDRAVKEFEANGGSSAANGGGAKKRAKPAKKVAKKSKKEESDEDDDDESESGSETPGTSESATPESPKKKRKVGIHGVPAAMAKNTITKTLKLRIVRPYNSAEVEKIVADEKNNREKIALEKNKDKVKEACSKHLKVAAYCTTQVERNACLFCKARKLDDKFYQKLRGQFPDAVFWQEISEIFRQLQKQAAEIYNQSLIELYYEIFIKGKGIANASSVEHYLSRVCYRRAAELFKNAAIASGLRSKIKSNFRLKELKNMKSGLPTTKSDNFPIPLVKQKGGQYRGFEISNHNSDFIIKIPFGRWQVKKEIDKYRPWEKFDFEQVRKSPKPISLLLSTQRRKRNKGWSKDEGTEAEIKKVMNGDYQTSYIEVKRGSKIGEKSAWMLNLSIDVPKIDKGVDPSIIGGIDVGVKEPLVCAINNAFSRYSISDNDLFHFNKKMFARRRILLKKNRHKRAGHGAKNKLKPITILTEKSRRFRKKLIERWACEIADFFIKNKVGTVQMENLESMKRKEDSYFNRRLRGFWPYARMQNKIEFKLKQYGIEIRKVAPNNTSATCSKCGHLNNYFNFEYRKKNKFPHFKCEKCNFKENADYNAALNISNPKLKSTKEEPKRPAATKKAGQAKKKKSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKMLSGGSPKKKRKV。
[0141] SEQ ID NO.9: ATGCCCAAGAAGAAGAGGAAAGTC AGTACCGACGCCGAGTACGTGCGGATCCACGAGAAGCTGGATATCTATACATTCAAGAAGCAGTTTAGCAACAATAAGAAGTCCGTGTCTCACAGATGCTACGTGCTGTTCGAGCTGAAGCGGAGAGGAGAGAGGCGCGCCTGTTTTTGGGGCTATGCCGTGAACAAGCCACAGTCTGGAACCGAGAGGGGAATCCACGCAGAGATCTTCAGCATCAGGAAGGTGGAGGAGTACCTGCGCGACAACCCCGGCCAGTTTACAATCAATTGGTATAGCTCCTGGAGCCCTTGCGCCGATTGTGCCGAGAAGATCCTGGAGTGGTACAACCAGGAGCTGAGGGGCAATGGCCACACCCTGAAGATCTGGGTGTGCAAGCTGTACTATGAGAAGAACGCCAGGAATCAGATCGGCCTGTGGAACCTGCGCGACAATGGCGTGGGCCTGAACGTGATGGTGTCCGAGCACTATCAGTGCTGTCGCAAGATCTTTATCCAGTCTAGCCACAATCAGCTGAACGAGAATCGGTGGCTGGAGAAGACACTGAAGAGAGCCGAGAAGCGGAGAAGCGAGCTGTCCATCATGTTTCAGGTGAAGATCCTGCACACCACAAAGTCTCCCGCCGTG AGCGGCAGCGAGACTCCCGGGACC TCAGAGTCCGCCACACCCGAAAGTATGAGCGACAAGCCAAAGAGACCTCTGAGCGCCTACATGCTGTGGCTGAACAGCGCTAGAGAATCCATCAAGCGCGAAAACCCCGGCATCAAGGTGACCGAGGTGGCCAAGCGGGGAGGCGAGCTGTGGCGGGCCATGAAGGACAAGTCTGAATGGGAGGCCAAGGCCGCTAAAGCCAAAGACGACTACGACAGAGCCGTGAAGGAATTCGAGGCAAATGGCGGCAGCAGCGCCGCCAACGGCGGAGGCGCCAAGAAAAGAGCCAAGCCTGCTAAGAAGGTCGCCAAGAAGTCCAAGAAAGAGGAATCTGATGAGGACGATGATGACGAGAGCGAG TCCGGCAGCGAGACACCCGGCACCA GCGAAAGCGCCACCCCTGAGTCTCCTAAGAAGAAACGGAAGGTGGGCATACACGGCGTGCCAGCCGCCATGGCCAAGAACACAATCACAAAAACCCTGAAGCTGAGAATCGTGCGGCCCTACAATAGCGCCGAAGTGGAAAAAATCGTGGCCGACGAGAAGAATAACCGGGAAAAAATCGCTCTGGAAAAGAACAAGGACAAGGTGAAGGAAGCCTGTAGCAAGCACCTGAAGGTGGCCGCCTACTGCACCACCCAGGTGGAAAGAAACGCCTGCCTGTTCTGCAAGGCCAGAAAGTTGGATGATAAGTTCTACCAGAAGCTGAGAGGCCAGTTCCCCGACGCCGTGTTCTGGCAGGAGATCAGCGAGATTTTCAGACAGCTGCAGAAGCAGGCCGCTGAGATCTACAACCAGAGCCTGATCGAGCTGTACTACGAGATCTTTATCAAGGGCAAGGGCATCGCTAATGCTTCTTCAGTGGAACACTACCTGAGCAGAGTGTGCTACAGAAGAGCCGCCGAGCTGTTTAAAAATGCCGCTATCGCCAGCGGCCTGAGAAGCAAGATCAAGTCTAATTTTCGGCTTAAGGAACTGAAGAACATGAAGTCCGGTCTTCCAACAACCAAGTCCGACAATTTCCCTATCCCTCTGGTGAAGCAGAAAGGAGGCCAGTATAGAGGCTTTGAGATCAGCAATCACAACTCCGACTTCATCATCAAGATCCCGTTCGGCAGATGGCAAGTGAAGAAAGAGATCGACAAGTACAGACCTTGGGAGAAGTTCGACTTCGAACAGGTTAGAAAGTCTCCTAAGCCTATCAGCCTGCTGCTGTCTACCCAGCGGAGAAAGCGGAACAAGGGATGGTCCAAAGATGAGGGCACAGAGGCGGAGATCAAGAAGGTGATGAACGGCGACTACCAGACCAGCTATATC
[0142] GAGGTGAAGCGCGGCAGCAAGATCGGCGAGAAGTCTGCCTGGATGCTG
[0143] AATCTGAGCATTGATGTGCCTAAGATCGACAAGGGCGTGGACCCTAGCA
[0144] TCATCGGCGGAATCGATGTTGGCGTCAAGGAGCCCCTGGTGTGTGCCAT
[0145] CAACAACGCCTTCAGCAGATACAGCATCTCTGATAACGACCTGTTCCACT
[0146] TCAACAAAAAGATGTTCGCCAGAAGACGCATCCTGCTGAAAAAAAACAG
[0147] ACACAAGCGGGCCGGACACGGCGCCAAGAACAAACTCAAACCTATTAC
[0148] AATCCTGACCGAGAAGTCTAGAAGATTCCGGAAGAAGCTGATCGAGAGA
[0149] TGGGCCTGTGAAATCGCCGACTTCTTCATCAAGAACAAGGTGGGAACCG
[0150] TGCAGATGGAAAACCTGGAAAGCATGAAGAGAAAGGAGGACAGCTACT
[0151] TCAACAGAAGACTGCGGGGCTTCTGGCCCTATGCCAGAATGCAGAACAA
[0152] GATCGAGTTCAAACTGAAACAGTACGGCATCGAAATCAGAAAAGTCGCC
[0153] CCTAACAACACCTCTGCTACCTGCAGCAAATGCGGCCATCTGAACAACT
[0154] ACTTTAACTTCGAGTACCGGAAAAAGAACAAGTTCCCCCACTTTAAGTG
[0155] CGAGAAGTGCAACTTCAAGGAGAACGCCGACTACAACGCCGCCCTGAAC
[0156] ATTAGCAACCCCAAGCTGAAGAGCACCAAGGAGGAACCT AAGCGGCCT
[0157] GCTGCTACAAAGAAGGCAGGCCAAGCCAAGAAGAAGAAG TCTGGTGGT
[0158] TCTACTAATCTGTCAGATATTATTGAAAAGGAGACCGGTAAGCAACTGG
[0159] TTATCCAGGAATCCATCCTCATGCTCCCAGAGGAGGTGGAAGAAGTCAT
[0160] TGGGAACAAGCCGGAAAGCGATATACTCGTGCACACCGCCTACGACGAG
[0161] AGCACCGACGAGAATGTCATGCTTCTGACTAGCGACGCCCCTGAATACA
[0162] AGCCTTGGGCTCTGGTCATACAGGATAGCAACGGTGAGAACAAGATTAA
[0163] GATGCTCTCTGGTGGTTCT CCCAAGAAGAAGAGGAAAGTCTAA 。
[0164] The underlined part is the signal peptide NLS sequence, and the double-underlined part is the linker region.
[0165] Example 5: PAM sequence analysis of HMG-TSminiCBE and enUn1Cas12f1
[0166] 1. Construct the reporter system plasmid
[0167] Based on the PAM sequence of NTTN, 16 reporter system plasmids were constructed and co-transfected with the enUn1Cas12f1 plasmid into HEK293T cells only. The transfection procedure was the same as in Example 1.
[0168] 48 hours after transfection, the cells were collected for flow cytometry analysis. The analysis process is as follows:
[0169] (1) Digest the cells with 50 μL of trypsin for 4 minutes, stop the digestion with 350 μL of culture medium, and transfer to a flow tube;
[0170] (2) When performing flow cytometry analysis, first obtain the live cell gate based on FSC-A and SSC-A, and then select the transfected positive cells based on PE-Texas Red-A;
[0171] (3) Analyze the fluorescence ratio under FITC-A for transfection-positive cells. The higher the fluorescence ratio, the higher the editing efficiency.
[0172] 2. Verify the editing efficiency at the endogenous gene locus
[0173] The sequence information of the endogenous gene locus is shown in Table 16 below. The steps of cell transfection, sorting, and subsequent determination of editing efficiency are the same as those in Example 1. The analysis results are as Figure 6 shown. Both HMG-TSminiCBE and enUn1Cas12f1 can achieve the recognition of the NTTR PAM sequence.
[0174] Table 16 Nucleotide sequences of the targeted gene locus
[0175]
[0176]
[0177]
[0178]
[0179] The above examples demonstrate an innovative tool, TSminiCBE, with efficient targeted strand C-to-T editing ability. This tool not only changes the editing object of the traditional BE system but also achieves efficient and high-purity editing of the target site. The development of the TSminiCBE technology fills a piece of the puzzle in the targeting range of existing cytosine base editing technologies and has broad application prospects in the field of cytosine bases.
[0180] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present invention.
[0181] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention also intends to include these modifications and variations.
Claims
1. A cytosine deamination editor, characterized in that The cytosine deamination editor is sequentially connected from the N-terminus to the C-terminus by the deaminase evoCDA1, the XTEN sequence and the enUn1Cas12f1 protein, and the amino acid sequences of the deaminase evoCDA1 and the enUn1Cas12f1 protein are shown in SEQ ID NO.4 and SEQ ID NO.5, respectively.
2. The cytosine deamination editor according to claim 1, characterized in that The amino acid at position 473 of the enUn1Cas12f1 protein also has an alanine mutation.
3. The cytosine deamination editor according to claim 2, characterized in that The C-terminus of the cytosine deamination editor is also connected to a UGI, whose amino acid sequence is shown in SEQ ID NO.
3.
4. The cytosine deamination editor according to claim 3, characterized in that The cytosine deamination editor also comprises a DNA binding protein, which is located between the enUn1Cas12f1 protein and the deaminase evoCDA1.
5. The cytosine deamination editor according to claim 4, characterized in that The DNA binding protein is HMG-D protein, and its encoding gene sequence is shown in SEQ ID NO.
7.
6. A cytosine deamination editing system comprising the cytosine deamination editor according to any one of claims 1 to 5, characterized in that: The cytosine deamination editor is guided by sgRNA to the target editing site.
7. A polynucleotide encoding the cytosine deamination editor according to any one of claims 1 to 5 or the cytosine deamination editing system according to claim 6.
8. A vector comprising the polynucleotide of claim 7.
9. A cell comprising the vector of claim 8.
10. A gene editing pharmaceutical composition, characterized in that: It comprises the cytosine deamination editor according to any one of claims 1 to 5, the cytosine deamination editing system according to claim 6, the polynucleotide according to claim 7, the vector according to claim 8 or the cell according to claim 9.
Citation Information
Patent Citations
Casz compositions and methods of use
CN111886336A
UN1Cas12f1 mutant and application thereof
CN118792282A