A base editor and construction method and application thereof
By fusing cytidine deaminase or deoxyadenosine deaminase to dCas12f, miniature base editors (miniBEs) were constructed, solving the problems of large size and insufficient activity of existing base editors and achieving more efficient gene editing results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-28
- Publication Date
- 2026-05-15
AI Technical Summary
The existing CRISPR/Cas9 and CRISPR/Cas12a systems have large base editors, which limits their application prospects. Furthermore, miniaturized base editors such as miniABE and miniCBE have insufficient activity, making it difficult to meet the needs of scientific research and clinical practice.
By fusing cytidine deaminase or deoxyadenosine deaminase with mutant mini nucleases, represented by dCas12f, at different sites, miniature base editors, including miniABEs and miniCBEs, are constructed to achieve CT and AG single base substitutions.
It enables base editing with smaller size, higher activity and diverse activity windows, broadens the application range of single base editing tools, and provides a wider range and more precise gene editing capabilities.
Smart Images

Figure CN116355100B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of gene editing technology, and in particular to a base editor, its construction method, and its application. Background Technology
[0002] Novel gene editing technologies, represented by CRISPR / Cas9, offer advantages such as high editing efficiency, greatly advancing gene editing technology. Since 2016, researchers have developed various DNA base editing tools based on CRISPR / Cas9 and CRISPR / Cas12a (Cpf1), enabling highly efficient and precise point mutations mediated at the DNA level without causing DNA double-strand breaks. Currently, two main types of base editors have been reported: cytosine base editors (CBE, which mediates C·G--T·A mutations) and adenine base editors (ABE, which mediates A·T--G·C mutations). CBE involves fusing a specific cytidine deaminase with a mutant nuclease (such as spCas9 carrying D10A or a mutation, called nspCas9; or a functionally inactive Cas12a, called dCas12a) and uracil glycosylation inhibitor (UGI). The resulting fusion protein can mediate C·G--T·A mutations under the guidance of sgRNA. ABE involves fusing a directionally evolved E. coli-derived adenosine deaminase (ecTadA*) dimer or monomer with a mutant nuclease (such as spCas9 carrying D10A or a mutation, called nspCas9; or a functionally inactive Cas12a, called dCas12a). The resulting fusion protein can mediate A·T--G·C mutations under the guidance of sgRNA. In addition, researchers fused a specific cytidine deaminase, a directed evolution of E. coli-derived adenosine deaminase (ecTadA*), mutant nucleases (such as spCas9 carrying D10A or mutation, referred to as nCas9), and uracil glycosylation inhibitor (UGI) to construct ACBE. The resulting fusion protein can mediate the simultaneous occurrence of C·G--T·A mutation and A·T--G·C mutation at specific sites under the guidance of sgRNA (Nat Biotechnol, 2020. 38(7): 861-864; BMC Biol, 2020. 18(1): 131; Nat Biotechnol, 2020. 38(7): 856-860; Nat Biotechnol, 2020. 38(7): 865-869). Existing ABE, CBE, and ACBE tools mainly rely on the CRISPR / Cas9 and CRISPR / Cas12a (Cpf1) systems. These two systems are too large, and the proteins of Cas9 and Cas12a are close to or exceed 1,000 amino acids, which limits their application prospects.
[0003] In recent years, researchers have discovered several new CRISPR systems, such as Casd12f, Cas12j, and TnpB, which are only 400–700 amino acids in size. These new CRISPR systems exhibit significant DNA cleavage activity in mammalian cells. [1-6] Therefore, it holds promise for miniaturization research of base editors. Some researchers have constructed functional miniaturized ABEs (miniABEs) using the inactivated Un1Casd12f1 (dCasd12f), but their base editing activity is less than 10%, which is insufficient to meet the needs of research and clinical applications. Currently, miniaturized CBEs (miniCBEs) and ACBE tools still lack research.
[0004] 1.Karvelis, T. et al. PAM recognition by miniature CRISPR-Casd12fnucleases triggers programmable double-stranded DNA targetcleavage. Nucleic Acids Research 48, 5016-5023 (2020).
[0005] 2. Kim, DY et al. Efficient CRISPR editing with a hypercompactCasd12f1and engineered guide RNAs delivered by adeno-associated virus. Nature Biotechnology (2021).
[0006] 3.Xu,XSet al.Engineered miniature CRISPR-Cas system for mammaliangenome regulation and editing.Mol Cel 81,4333-+(2021).
[0007] 4.Karvelis, T. et al. Transposon-associated TnpB is a programmable RNA-guided DNA endonuclease. Nature (2021).
[0008] 5. Altae-Tran, H. et al. The widespread IS200 / IS605 transposon family encodes diverse programmable RNA-guided endonucleases. Science 374, 57-65 (2021).
[0009] 6. Wu, Z. et al. Programmed genome editing by a miniature CRISPR-Casd12fnuclease. Nat Chem Bio l17, 1132-1138 (2021). Summary of the Invention
[0010] To develop miniaturized base editors ABE, CBE, and ACBE (collectively referred to as miniBEs), this invention provides a base editor, its construction method, and its applications. The base editors in this invention are primarily fusion proteins that generate point mutations in genes.
[0011] This invention fused cytidine deaminase / deoxyadenosine deaminase and its mutants to different sites (N-terminus, C-terminus, and internal insertion site) of mutant small nucleases, such as dCas12f, to obtain novel fusion proteins. This invention also developed a miniaturized base editor capable of effectively performing CT and AG single-base substitutions. Depending on the type of deaminase and its fusion site within the mutant small nuclease represented by dCas12f, the newly constructed base editor also exhibits different activity windows. This invention effectively broadens the application of single-base editing tools.
[0012] The objective of this invention can be achieved through the following technical solutions:
[0013] The first objective of this invention is to provide a base editor for fusion proteins miniBEs formed by placing cytidine deaminase or deoxyadenosine deaminase at the N-terminus, different internal sites, or the C-terminal fusion site of a mutant nuclease dCasd12f, wherein the amino acid sequence of dCasd12f is shown in SEQ ID NO.1.
[0014] In one embodiment of the present invention, an enzyme cleavage site SpeI-BamHI-Xba is introduced into the N-terminus, different internal sites, or the C-terminus of dCasd12f, and different deaminases are fused to obtain the fusion protein miniBEs.
[0015] In one embodiment of the present invention, the different insertion site information within the dCasd12f is as follows:
[0016]
[0017] In one embodiment of the present invention, the deoxyadenosine deaminase is selected as TadA-8e or a TadA-8e mutant, or a substance obtained by adding an NLS sequence and a linker sequence to the 5' and 3' of a TadA-8e or TadA-8e mutant, respectively.
[0018] In one embodiment of the present invention, a base editor is provided as a fusion protein miniABEs-8e. An NLS sequence and a linker sequence are added to the 5' and 3' of TadA-8e, respectively, to obtain an NLS-TadA-8e-linker. The NLS-TadA-8e-linker is fused to a fusion site at the N-terminus, C-terminus, or internal fusion site of dCasd12f to form a fusion protein. The amino acid sequence of the NLS-TadA-8e-linker is shown in SEQ ID NO.6.
[0019] In one embodiment of the present invention, miniABEs with significant AG editing activity can be obtained by fusing NLS-TadA-8e-linker at the N-terminal site, the two C-terminal sites of dCasd12f, and the 128 and 130 sites inside d12f.
[0020] In one embodiment of the present invention, the base editor is a fusion protein miniABEs-2-8e. The fusion protein miniABEs-2-8e refers to the fusion protein obtained by introducing the restriction enzyme site SpeI-BamHI-Xba at the N-terminus, different internal sites, or the C-terminus of dCasd12f and fusing it with TadA-8e-TadA-8e, wherein the amino acid sequence of TadA-8e-TadA-8e is shown in SEQ ID NO.7.
[0021] In one embodiment of the present invention, miniABEs with significant AG editing activity can be obtained by fusing TadA-8e-TadA-8e at the N-terminal site, the two C-terminal sites of dCasd12f, and the 128 and 130 sites inside dCasd12f.
[0022] In one embodiment of the present invention, a new fusion protein is obtained by fusing a TadA-8e mutant with dCasd12f instead of TadA-8e. The TadA-8e mutant refers to the V106W, F84M-N108Y, V28G-N46C, V28G-A48G, and V28G-A48G-I49A-V82T-N108Y mutations of TadA-8e.
[0023] F84M-N108Y indicates that amino acid F mutates to M at amino acid position 84 and amino acid N mutates to Y at amino acid position 108, meaning that both positions mutate simultaneously. Other positions can be interpreted similarly.
[0024] In one embodiment of the present invention, the cytidine deaminase is selected as rAPOBEC1-YE1, rAPOBEC1-YE1 mutant, APOBEC3A mutant, or a substance obtained by adding NLS sequence and linker sequence to the 5' and 3' of rAPOBEC1-YE1, rAPOBEC1-YE1 mutant, and APOBEC3A mutant, respectively.
[0025] In one embodiment of the present invention, the base editor is a fusion protein miniCBEs-YE1, which refers to a fusion protein obtained by introducing the restriction enzyme site SpeI-BamHI-Xba at different sites inside or at the C-terminus of dCasd12f and fusing it with rAPOBEC1-YE1.
[0026] In one embodiment of the present invention, the base editor is a fusion protein miniCBEs-YE1. The fusion protein miniCBEs-YE1 refers to the fusion protein obtained by introducing the restriction enzyme site SpeI-BamHI-Xba at different sites inside or at the C-terminus of dCasd12f and fusing it with NLS-YE1-linker. The amino acid sequence of NLS-YE1-linker is shown in SEQ ID NO.8.
[0027] In one embodiment of the present invention, miniCBEs-YE1 with significant CT editing activity can be obtained by fusing the NLS-YE1-linker at the N-terminal site and the C-terminal CL site of dCasd12f.
[0028] In one embodiment of the present invention, the base editor is a fusion protein miniCBEs-3A130. The fusion protein miniCBEs-3A130 refers to the fusion protein obtained by introducing the restriction enzyme site SpeI-BamHI-Xba at the N-terminus, different internal sites, or the C-terminus of dCasd12f and fusing it with the APOBEC3A mutant. The amino acid sequence of the APOBEC3A mutant is shown in SEQ ID NO.9.
[0029] In one embodiment of the present invention, miniCBEs-3A130 with significant CT editing activity can be obtained by fusing the APOBEC3A mutant at the N-terminal site and the C-terminal CL site of dCasd12f.
[0030] In one embodiment of the present invention, given all the above-mentioned fusion proteins, the base dCasd12f can be mutated to obtain a fusion protein based on a dCasd12f point mutation. The dCasd12f point mutation refers to a combination of D143R-T147R-E151A point mutations in dCasd12f.
[0031] D143R-T147R-E151A indicates that amino acid D mutates to R at amino acid position 143, amino acid T mutates to R at amino acid position 147, and amino acid E mutates to A at amino acid position 151, meaning that all three positions mutate simultaneously.
[0032] This invention also provides a method for constructing the base editor, comprising the following steps:
[0033] A base editor was constructed by fusing cytidine deaminase or deoxyadenosine deaminase with different sites of dCasd12f.
[0034] In one embodiment of the present invention, NLS sequences and linker sequences may be added to cytidine deaminase or deoxyadenosine deaminase 5' and 3', respectively.
[0035] The present invention also provides applications of the base editor, wherein the base editor is used as a cytosine base editor (CBE, which can mediate C·G--T·A mutation) or an adenine base editor (ABE, which can mediate A·T--G·C mutation) for mutation of specific bases at specific sites in DNA.
[0036] The present invention also provides a polynucleotide encoding the fusion protein in the base editor.
[0037] The present invention also provides a carrier containing the aforementioned polynucleotide.
[0038] The present invention also provides a host cell containing the base editor, or containing the vector.
[0039] The present invention also provides a kit comprising reagents for constructing the base editor.
[0040] This invention fuses cytidine deaminase and its mutants to different sites of mutant nucleases, such as Cas12f1, to obtain novel fusion proteins capable of effectively performing CT base mutations on cytosine at different positions in the prespacer sequence. Similarly, by fusing deoxyadenosine deaminase and its variants to different sites of mutant nucleases, such as Cas12f1, the resulting novel fusion proteins can effectively perform AG and / or CT base mutations on adenine and / or cytosine at different positions in the prespacer sequence. The fusion proteins obtained by this invention, based on different insertion sites, exhibit varying mutagenic ranges. Compared to existing technologies, the base editor of this invention is smaller and has a more diverse active window, enabling wider, more precise, and safer CT and AG single-base substitutions. This effectively broadens the application of single-base editing tools and has high application value. Compared to existing technologies, the beneficial effects of this invention are mainly reflected in the following aspects:
[0041] 1. This invention provides the first miniaturized base editors (miniCBEs) based on dCasd12f. Furthermore, compared to previously reported methods, the miniaturized ABEs (miniABEs) based on dCasd12f in this invention exhibit higher base editing activity and a more diverse activity window.
[0042] 2. This invention identifies insertion sites within mutant miniaturized nucleases, represented by dCasd12f, for deaminase fusion. By combining these with deoxyadenosine deaminase and its variants / cytidine deaminase and its variants, the resulting novel fusion proteins can achieve effective AG base mutations in adenine located at different positions in the prespacer sequence (the TTTR sequence of the prespacer adjacent motif (PAM) is defined as positions -3 to 0, and the prespacer sequence after the TTTR is defined as positions 1 to 20), or effective CT base mutations in cytosine located at different positions in the prespacer sequence (the TTTR sequence of the prespacer adjacent motif (PAM) is defined as positions -3 to 0, and the prespacer sequence after the TTTR is defined as positions 1 to 20). Furthermore, the mutation range of fusion proteins based on different deaminases and different insertion sites varies. Based on this, a new gene editing composition is provided, enabling a wider range and more precise AG and CT single-base substitutions. Attached Figure Description
[0043] Figure 1Schematic diagrams of various fusion protein constructions. The N- and C-termini of dCasd12f are fused with nuclear insertion signals (NLS). In the construction of miniBEs, N-terminal fusion involves fusing a deaminase to the N-terminus of dCasd12f; C-terminal fusion involves fusing a deaminase to the C-terminus of dCasd12f. Internal insertion involves fusing a deaminase to different sites within dCasd12f.
[0044] Figure 2 Base editing characteristics of miniABEs. Base editing characteristics of miniABEs constructed using TadA-8e (abbreviated as 8e)(a) monomers and TadA-8e-TadA-8e (abbreviated as 2-8e)(b) dimers as deaminases. The sgRNA used to detect base editing characteristics was d12f-sg89 (sequence CACACACACAGTGGGCTACC, PAM sequence TTTG). Each group had n = 2 replicates.
[0045] Figure 3 Base editing characteristics of MiniCBEs-YE1. Base editing characteristics of miniCBEs constructed using the rAPOBEC1-YE1 (YE1) mutant deaminase. The sgRNA used to detect the base editing characteristics was d12f-sg89 (sequence CACACACACAGTGGGCTACC, PAM sequence TTTG). Each group had n = 2 replicates.
[0046] Figure 4 The effect of the V106W point mutation in TadA-8e on the base editing characteristics of miniABEs. The V106W point mutation was introduced into TadA-8e in miniABEs-8e and miniABEs-2-8e, and base editing characteristics were analyzed using Sanger sequencing. The sgRNA used to detect base editing characteristics was d12f-sg89 (sequence CACACACACAGTGGGCTACC, PAM sequence TTTG). Each group had n = 2 replicates.
[0047] Figure 5 The effect of different point mutation combinations of TadA-8e on the base editing characteristics of miniBEs. Point mutation combinations F84M-N108Y (a), V28G-N46C (b), V28G-A48G (c), and V28G-A48G-I49A-V82T-N108Y(GGATY) (d) were introduced into TadA-8e of miniABEs-8e, and base editing characteristics were analyzed using Sanger sequencing. The sgRNA used to detect base editing characteristics was d12f-sg89 (sequence CACACACACAGTGGGCTACC, PAM sequence TTTG). Each group had n = 2 replicates.
[0048] Figure 6 Base editing signature of MiniCBEs-3A130. Base editing signature of miniCBEs constructed using the human APOBEC3A-Y130F (abbreviated as 3A130) mutant deaminase. The sgRNA used to detect the base editing signature was d12f-sg89 (sequence CACACACACAGTGGGCTACC, PAM sequence TTTG). Each group had n = 2 replicates.
[0049] Figure 7 The effect of specific point mutation combinations in .dCasd12f on the base editing characteristics of miniABEs-8e was investigated. The D143R-T147R-E151A point mutation combination (RRA) was introduced into .dCasd12f, and Sanger sequencing was used to detect its effect on the base editing characteristics of miniABEs-8e. The sgRNA used for detecting the base editing characteristics was d12f-sg89 (sequence CACACACACAGTGGGCTACC, PAM sequence TTTG). Each group had n = 2 replicates.
[0050] Figure 8 The effect of specific point mutation combinations in .dCasd12f on the base editing characteristics of the miniABEs-8e mutant. The D143R-T147R-E151A point mutation combination (RRA) was introduced into dCasd12f in miniABEs-2-8e(V106W)(a) and miniABEs-8e(V106W)(b), and Sanger sequencing was used to detect its effect on the base editing characteristics of miniABEs-8e. The sgRNA used to detect the base editing characteristics was d12f-sg89 (sequence CACACACACAGTGGGCTACC, PAM sequence TTTG). Each group had n = 2 replicates.
[0051] Figure 9 The effect of specific point mutation combinations in .dCasd12f on the base editing characteristics of miniCBE mutants. The D143R-T147R-E151A point mutation combination (RRA) was introduced into dCasd12f in miniCBEs-YE1(a), miniCBEs-3A130(b), and miniBEs-8e(GGATY)(c), and Sanger sequencing was used to detect its effect on the base editing characteristics of miniCBEs. The sgRNA used to detect the base editing characteristics was d12f-sg89 (sequence CACACACACAGTGGGCTACC, PAM sequence TTTG). Each group had n = 2 replicates.
[0052] Figure 10The base editing characteristics of preferred variants of .RRA miniBEs-8e(GGATY) at three sites: d12f-sg5, d12f-sg9, and d12f-sg21. The effects of Sanger sequencing on the base editing characteristics of miniCBEs were analyzed. Each group had n = 3 replicates. Detailed Implementation
[0053] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.
[0054] Example 1: Construction of Fusion Protein
[0055] The open reading frame (ORF) of the SV-40-NLS-dCasd12f-Nucleoplasmin NLS-linker optimized with human codons was synthesized, wherein the amino acid sequence of dCasd12f is shown in SEQ ID NO.1, the amino acid sequence of SV-40-NLS is shown in SEQ ID NO.2, the amino acid sequence of Nucleoplasmin NLS is shown in SEQ ID NO.3, and the amino acid sequence of NucleoplasminNLS-linker is shown in SEQ ID NO.4.
[0056] In this embodiment, dCasd12f was used as a mutant of un1Casd12f1 (double point mutations of D326A and D510A, resulting in loss of DNA cleavage activity). Subsequently, restriction enzyme sites (SpeI-BamHI-XbaI) were introduced at the N-terminus, different internal sites, and the C-terminus of dCasd12f through point mutations to fuse different deaminases (site information is detailed in Table 1), obtaining the fusion protein miniBEs for subsequent experiments. Figure 1 ).
[0057] In this article, expressions like "D326A" indicate that 326 represents the 28th amino acid position, D represents the amino acid before the mutation at the 326th amino acid position, and A represents the amino acid after the mutation at the 326th amino acid position. Both D and A are abbreviations for amino acids. That is, D326A means that the 326th amino acid position is changed from aspartic acid to alanine. Other expressions like "D326A" in the above or below will be explained in a similar way.
[0058] In this embodiment, the fusion protein and green fluorescent protein EGFP are co-expressed using the 2A peptide to indicate the expression of the fusion protein and to be used for subsequent flow cytometry sorting.
[0059] In this embodiment, an sgRNA expression vector (sgRNA backbone sequence as shown in SEQ ID NO.5, NNNN is the target sequence of sgRNA) was also constructed to express UGI-2A-mCherry simultaneously with the expression of a specific sgRNA. UGI can inhibit the activity of uracil glycosylation enzyme and improve CT mutation efficiency. The red fluorescent protein mCherry is used to indicate the expression status of the vector and can be used for subsequent flow cytometry sorting.
[0060] Table 1. Different insertion sites within dCasd12f
[0061]
[0062] Example 2: Detection of point mutation frequency and characteristics in HEK293T cells transfected with different miniABEs and sgRNA expression vectors.
[0063] Different miniABEs-8e were obtained by fusing the NLS-TadA-8e-linker (SEQ ID NO.6) to the N-terminus, different internal sites, and the C-terminus of dCasd12f (abbreviated as d12f) (Table 1). To evaluate their editing activity, different miniABEs-8e were co-transfected with single-guide RNA (sgRNA) (sgRNA simultaneously expressing uracil glycosylation inhibitor UGI and 2A-mCherry) into cultured 293T cells. After 72 hours of cell culture, double-positive cells simultaneously expressing EGFP and mCherry were collected by FACS.
[0064] The sgRNA used in this embodiment is d12f-sg89 (sequence CACACACACAGTGGGCTACC, PAM sequence TTTG).
[0065] Genomic DNA was extracted from collected double-positive cells, and directional PCR was performed using primers that specifically amplify the d12f-sg89 site. The PCR products were then sequenced by Sanger sequencing to verify the editing efficiency of different miniABEs at the d12f-sg89 site.
[0066] The results showed that fusing the NLS-TadA-8e-linker with the N-terminal site, the two C-terminal sites, and sites 128 and 130 within d12f yielded miniABEs with significant AG editing activity. Figure 2a) The results showed that d12f-N-TadA-8e obtained by N-terminal fusion had the broadest activity window (A2-A18, with the TTTR sequence of the adjacent motif (PAM) in the pre-interstitial region defined as position -3 to 0, and the pre-interstitial space sequence after the TTTR defined as positions 1-20), but the highest activity site was A4. The activity windows of d12f-C-TadA-8e and d12f-CL-TadA-8e obtained by C-terminal fusion were slightly narrower (A4-A18), with the highest activity site being A18. The activity sites of d12f-128-TadA-8e and d12f-130-TadA-8e obtained by internal fusion were similar (A4-A6), with the highest activity site being A4. In summary, by changing the fusion site, miniABEs with different activity windows can be obtained.
[0067] Traditional ABEs are constructed by tandemly combining the expression sequences of heterologous / homologous TadA deaminase wild-type / mutant (TadA-TadA) and then fusing them to different sites on the Cas expression sequence. Therefore, in this invention, TadA-8e is constructed in a similar manner as TadA-8e-TadA-8e (abbreviated as 2-8e, SEQ ID NO.7), fused to the N-terminus, different internal sites, and the C-terminus of dCasd12f (Table 1) to obtain different miniABEs-2-8e. To evaluate their editing activity, different miniABEs-2-8e were co-transfected with single-guide RNA (sgRNA) (sgRNA simultaneously expressing uracil glycosylation inhibitor UGI and 2A-mCherry) into cultured 293T cells. After 72 hours of cell culture, double-positive cells simultaneously expressing EGFP and mCherry were collected by FACS. The sgRNA used in this embodiment is d12f-sg89 (sequence CACACACACAGTGGGCTACC, PAM sequence TTTG).
[0068] Genomic DNA was extracted from collected double-positive cells, and directional PCR was performed using primers that specifically amplify the d12f-sg89 site. The PCR products were then sequenced by Sanger sequencing to verify the editing efficiency of different miniABEs at the d12f-sg89 site.
[0069] The results showed that, similar to the TadA-8e monomer, fusion of 2-8e at the N-terminal site, the two C-terminal sites, and the 128 and 130 sites within d12f could yield miniABEs with significant AG editing activity. Figure 2(b) The results showed that d12f-N-2-8e obtained by N-terminal fusion had the broadest activity window (A2-A18, with the TTTR sequence of the pre-interstitial motif (PAM) defined as positions -3 to 0, and the pre-interstitial spacer sequence after the TTTR defined as positions 1-20), but the highest activity site was A4. The activity windows of d12f-C-2-8e and d12f-CL-2-8e obtained by C-terminal fusion were slightly narrower (A4-A18), with the highest activity sites being A4 and A18, respectively. The activity sites of d12f-128-2-8e and d12f-130-2-8e obtained by internal fusion were similar (A4-A8), with similar activities at each site. Compared with the TadA-8e monomer (miniABEs-8e), the activity window of the deaminase dimer form 2-8e (miniABEs-2-8e) was somewhat altered, but overall similar, and the base editing activities were also very close. However, the activity windows of d12f-128-2-8e and d12f-130-2-8e obtained from the internal fusion site have changed significantly (from A4-A6 to A4-A8), which further increases the diversity of the activity windows of Casd12f-based miniABEs.
[0070] Example 3: Detection of point mutation frequency and characteristics in HEK293T cells transfected with different miniCBEs and sgRNA expression vectors.
[0071] rAPOBEC1-YE1 (YE1 for short) is one of the most commonly used and safest cytidine deaminase mutants. In this embodiment, YE1 was first selected for the construction of miniCBEs. Different miniCBEs-YE1 were obtained by fusing the NLS-YE1-linker (SEQ ID NO.8) to the N-terminus, different internal sites, and the C-terminus of dCasd12f (Table 1). To evaluate its editing activity, different miniCBEs-YE1 were co-transfected with single-guide RNA (sgRNA) (sgRNA simultaneously expressing uracil glycosylation inhibitor UGI and 2A-mCherry) into cultured 293T cells. After 72 hours of cell culture, double-positive cells simultaneously expressing EGFP and mCherry were collected by FACS. The sgRNA used in this embodiment was d12f-sg89 (sequence CACACACACAGTGGGCTACC, PAM sequence TTTG).
[0072] Genomic DNA was extracted from collected double-positive cells, and directional PCR was performed using primers that specifically amplify the d12f-sg89 site. The PCR products were then sequenced by Sanger sequencing to verify the editing efficiency of different miniCBEs-YE1 at the d12f-sg89 site.
[0073] The results showed that fusing the NLS-YE1-linker with the N-terminal and C-terminal CL sites yielded miniCBEs-YE1 with significant CT editing activity. Figure 3 However, the editing activity was below 20%, and functional miniCBEs-YE1 could not be obtained from internal fusion sites. The results showed that the activity window of d12f-N-YE1 obtained by N-terminal fusion was C3-C5. The activity window of d12f-CL-YE1 obtained by C-terminal fusion was C3.
[0074] Example 4: Effects of TadA-8e mutant on the activity window and base editing specificity of miniABEs
[0075] Different mutants of TadA-8e significantly affect its off-target effects and the specificity of base editing. Among them, TadA-8e (V106W) can reduce the off-target effects of TadA-8e at the RNA level. Therefore, in this embodiment, the V106W point mutation was introduced into functional miniABEs-8e and miniABEs-2-8e to construct miniABEs-8e (V106W) and miniABEs-2-8e (V106W). To evaluate their editing activity, different miniABEs-8e (V106W) and miniABEs-2-8e (V106W) were co-transfected with single-guide RNA (sgRNA) (sgRNA simultaneously expressing uracil glycosylation inhibitor UGI and 2A-mCherry) into cultured 293T cells. After 72 hours of cell culture, double-positive cells simultaneously expressing EGFP and mCherry were collected by FACS. The sgRNA used in this embodiment is d12f-sg89 (sequence CACACACACAGTGGGCTACC, PAM sequence TTTG).
[0076] Genomic DNA was extracted from collected double-positive cells, and directional PCR was performed using primers that specifically amplify the d12f-sg89 site. The PCR products were then sequenced by Sanger sequencing to verify the editing efficiency of different miniABEs at the d12f-sg89 site.
[0077] The results showed that the functional miniABEs-8e, fused from the N-terminal site, the two C-terminal sites, and sites 128 and 130 within d12f, still exhibited significant AG editing activity after the introduction of V106W. Figure 4a) The activity window is similar to that of miniABEs-8e. However, the activity window of d12f-N-8e (V106W) obtained by N-terminal fusion is significantly narrower (A2-A6, with the TTTR sequence of the adjacent motif (PAM) in the pre-interstitial region defined as position -3 to 0, and the pre-interstitial spacer sequence after the TTTR defined as positions 1-20), with the highest active site being A4. The activity window and highest active site of d12f-CL-8e (V106W) obtained by C-terminal fusion are not significantly changed. The active sites and highest active sites of d12f-128-8e (V106W) and d12f-130-8e (V106W) obtained by internal fusion are also not significantly changed.
[0078] The functional miniABEs-2-8e, which incorporates N-terminal sites, two C-terminal sites, and sites 128 and 130 within d12f, still exhibits significant AG editing activity after the introduction of V106W. Figure 4 (b) The activity window is similar to that of miniABEs-2-8e. However, the activity window of d12f-N-2-8e (V106W) obtained by C-terminal fusion is slightly increased (A2-A18, with the TTTR sequence of the pre-interstitial motif (PAM) defined as position -3 to 0, and the pre-interstitial spacer sequence after TTTR defined as positions 1-20), with the highest active site being A4. The activity window and highest active site of d12f-CL-2-8e (V106W) obtained by N-terminal fusion are not significantly changed. The active sites and highest active sites of d12f-128-2-8e (V106W) and dCasd12f-130-2-8e (V106W) obtained by internal fusion are also not significantly changed, but the base editing activity is significantly altered.
[0079] Example 5: Effects of TadA-8e mutant on the activity window and base editing specificity of miniABEs
[0080] Different mutants of TadA-8e significantly affect its off-target effects and the specificity of base editing. Previous studies have found that specific amino acid mutations and combinations of mutations can alter the base editing properties of TadA-8e, resulting in base editors with CT editing activity.
[0081] Therefore, in this embodiment, some point mutations and combinations of point mutations were introduced into the functional miniABEs-8e, constructing miniABEs-8e(F84M-N108Y), miniABEs-8e(V28G-N46C), miniABEs-8e(V28G-A48G), and miniABEs-8e(V28G-A48G-I49A-V82T-N108Y, abbreviated as GGATY). To evaluate its editing activity, different miniABEs-8e(F84M-N108Y), miniABEs-8e(V28G-N46C), miniABEs-8e(V28G-A48G), and miniBEs-8e(GGATY) were co-transfected with single-guide RNA (sgRNA) (sgRNA simultaneously expressing uracil glycosylation inhibitor UGI and 2A-mCherry) into cultured 293T cells. After 72 hours of cell culture, double-positive cells expressing both EGFP and mCherry were collected using FACS. The sgRNA used in this example was d12f-sg89 (sequence CACACACACAGTGGGCTACC, PAM sequence TTTG).
[0082] Genomic DNA was extracted from collected double-positive cells, and directional PCR was performed using primers that specifically amplify the d12f-sg89 site. The PCR products were then sequenced by Sanger sequencing to verify the editing efficiency of different miniBEs at the d12f-sg89 site.
[0083] The results showed that d12f-N-8e (F84M-N108Y) obtained by N-terminal fusion could simultaneously mediate AG and CT editing (A3-C4, the TTTR sequence of the adjacent motif (PAM) in the anterior septum was defined as position -3 to 0, and the anterior septum sequence after TTTR was defined as position 1-20). Figure 5 a) The miniABEs-8e(F84M-N108Y) fused at positions 128, 130 and C-terminus did not show significant editing activity (<10%).
[0084] The fusion of the N-terminal and C-terminal CL sites with 8e (V28G-N46C) yields miniCBEs with significant CT editing activity (editing site C3). Figure 5 (b) and the CT editing activity of d12f-N-8e(V28G-N46C) is >20%. Internal fusion sites cannot obtain functional miniBEs.
[0085] Fusion of 8e (V28G-A48G) at the C-terminal CL site yields a miniCBE with significant CT editing activity (editing site C3). Figure 5c). Functional miniBEs cannot be obtained at N-terminal and internal fusion sites.
[0086] After fusing NLS-8e (GGATY) with two sites at the N-terminus, two sites at the C-terminus, and the 130 site inside d12f, miniCBEs with significant CT editing activity can be obtained (all with an activity window of C3). Figure 5 d). Among them, the CT editing activity of d12f-N-8e(GGATY) and d12f-CL-8e(GGATY) obtained by fusing N-terminus and C-terminus is >30%, with the highest reaching 34%.
[0087] In summary, introducing point mutations into TadA-8e can not only change the activity window of miniABEs-8e, but also change its base editing specificity, resulting in miniCBEs with more precise windows and higher base editing activity, such as d12f-N-8e(GGATY) and d12f-CL-8e(GGATY).
[0088] Example 6: MiniCBEs constructed from the APOBEC3A mutant and their sgRNA expression vector were transfected into HEK293T cells to detect point mutation frequency and characteristics.
[0089] The human APOBEC3A (Y130F) mutant (abbreviated as 3A130) is one of the commonly used cytidine deaminases. In SpCas9-based CBEs, its activity is higher than that of YE1, but it can lead to non-Cas9-dependent DNA off-target effects.
[0090] In this embodiment, 3A130 was selected for the construction of miniCBEs. Different miniCBEs-3A130 were obtained by fusing 3A130 (SEQ ID NO. 9) to the N-terminus, internal 128, 130, and C-terminus of dCasd12f. To evaluate their editing activity, different miniCBEs-3A130 were co-transfected with single-guide RNA (sgRNA) (sgRNA simultaneously expressing uracil glycosylation inhibitor UGI and 2A-mCherry) into cultured 293T cells. After 72 hours of cell culture, double-positive cells simultaneously expressing EGFP and mCherry were collected by FACS. The sgRNA used in this embodiment was d12f-sg89 (sequence CACACACACAGTGGGCTACC, PAM sequence TTTG).
[0091] Genomic DNA was extracted from collected double-positive cells, and directional PCR was performed using primers that specifically amplify the d12f-sg89 site. The PCR products were then sequenced by Sanger sequencing to verify the editing efficiency of different miniCBEs at the d12f-sg89 site.
[0092] The results showed that fusing 3A130 with N-terminal and C-terminal CL sites yielded miniCBEs-3A130 with significant CT editing activity. Figure 6 Functional miniCBEs-3A130 could not be obtained from internal fusion sites 128 and 130. The results showed that the activity window of d12f-N-3A130 obtained by N-terminal fusion was C3-C9, while the activity window of d12f-CL-3A130 obtained by C-terminal fusion was C3-C20.
[0093] Example 7: Effect of dCasd12f point mutation on miniBEs base editing activity
[0094] Studies have found that specific point mutations and combinations of Casd12f can enhance its activity, with the D143R-T147R-E151A point mutation combination (referred to as RRA) showing the most significant effect. To evaluate the impact of RRA mutations on the editing activity of miniBEs, RRA mutations were introduced into different miniBEs, and then different miniBEs-RRA were co-transfected with single-guide RNA (sgRNA) (sgRNA simultaneously expressing uracil glycosylation inhibitor UGI and 2A-mCherry) into cultured 293T cells. After 72 hours of cell culture, double-positive cells simultaneously expressing EGFP and mCherry were collected by FACS. The sgRNA used in this example was d12f-sg89 (sequence CACACACACAGTGGGCTACC, PAM sequence TTTG).
[0095] Genomic DNA was extracted from collected double-positive cells, and directional PCR was performed using primers that specifically amplify the d12f-sg89 site. The PCR products were then sequenced by Sanger sequencing to verify the editing efficiency of different miniBEs at the d12f-sg89 site.
[0096] The results of introducing RRA mutations into MiniABEs-8e showed that fusion of the NLS-TadA-8e-linker at the N-terminal, C-terminal, and sites 128, 130, 211, 226, and 229 within d12f yielded miniABEs with significant AG editing activity. Figure 7Among them, the RRA d12f-N-TadA-8e obtained by N-terminal fusion has the broadest activity window (A2-A18, with the TTTR sequence of the pre-interstitial motif (PAM) defined as position -3 to 0, and the pre-interstitial spacer sequence after TTTR defined as positions 1-20), but the highest activity site is A4. The RRA d12f-C-TadA-8e and d12f-CL-TadA-8e obtained by C-terminal fusion have slightly narrower activity windows (A4-A18), with the highest activity site being A18. The RRA miniABEs obtained through internal fusion, including RRAAd12f-128-TadA-8e, RRAAd12f-130-TadA-8e, and RRA d12f-229-8e, have similar active sites (A4-A8), with the highest activity site being A4. The activity window of RRA d12f-226-8e is A4-A6, with the highest activity site being A4. The activity window of RRA d12f-211-8e is A18. Furthermore, compared to the corresponding miniABEs-8e without RRA mutations, RRA miniABEs-8e exhibits stronger base editing activity and a somewhat altered activity window. Additionally, RRA miniABEs-8e constructed from the fusion sites 211, 226, and 229, which originally showed no significant base editing activity (<10%), also exhibited significant AG editing activity. In summary, RRA point mutations can significantly enhance the AG editing activity of miniABEs and also have a certain influence on the activity window.
[0097] Example 8: Effect of dCasd12f point mutation on the base editing activity of other miniBEs
[0098] To assess the impact of RRA mutations on the editing activity of miniBEs, RRA mutations were introduced into different miniBEs (including miniABEs-2-8e(V106W), miniABEs-8e(V106W), miniCBEs-YE1, miniCBEs-3A130, and miniBEs-8e(GGATY)). Then, the different miniBEs-RRA mutations were co-transfected with single-guide RNA (sgRNA) (sgRNA simultaneously expressing uracil glycosylation inhibitor UGI and 2A-mCherry) into cultured 293T cells. After 72 hours of cell culture, double-positive cells simultaneously expressing EGFP and mCherry were collected by FACS. The sgRNA used in this example was d12f-sg89 (sequence CACACACACAGTGGGCTACC, PAM sequence TTTG).
[0099] Genomic DNA was extracted from collected double-positive cells, and directional PCR was performed using primers that specifically amplify the d12f-sg89 site. The PCR products were then sequenced by Sanger sequencing to verify the editing efficiency of different miniBEs at the d12f-sg89 site.
[0100] The results of introducing RRA mutations into miniABEs-8e(V106W) showed that miniABEs-8e(V106W) obtained by fusing the N-terminal, C-terminal, and 128 sites within d12f sites still exhibited significant AG editing activity. Figure 8 a) and the editing activity is higher than that of the corresponding RRA-free miniABEs-8e (V106W), but the activity window is similar.
[0101] The results of introducing RRA mutations into miniABEs-2-8e(V106W) showed that miniABEs-8e(V106W) obtained by fusing the N-terminal site, C-terminal site, and sites 128 and 130 inside d12f still exhibited significant AG editing activity. Figure 8 b) and the editing activity is higher than that of the corresponding RRA-free miniABEs-8e (V106W), but the activity window is similar.
[0102] The base results of introducing RRA mutations into miniCBEs-YE1 showed that miniCBEs-YE1 obtained by fusing the N-terminal site, C-terminal site, and the 130 site inside d12f exhibited significant AG editing activity. Figure 9 a) and the editing activity is higher than that of the corresponding miniCBEs-YE1 without RRA mutation, but the absolute value of the base editing activity is not high (≤20%).
[0103] The base results of introducing RRA mutations into miniCBEs-3A130 showed that miniCBEs-3A130 obtained by fusing the N-terminal site, the C-terminal site, and sites 128 and 130 inside d12f exhibited significant AG editing activity. Figure 9 b) and the editing activity is higher than that of the corresponding miniCBEs-3A130 without RRA mutation.
[0104] The results of introducing RRA mutations into miniBEs-8e(GGATY) showed that miniCBEs-3A130, obtained by fusing the N-terminal site, C-terminal site, and sites 128 and 130 within d12f, exhibited significant AG editing activity. Figure 9c) The editing activity is higher than that of the corresponding RRA-free miniBEs-8e (GGATY). Furthermore, the results indicate that RRA miniBEs-8e (GGATY) has a very specific activity window (C3), and the highest editing activity is close to 50% (up to 44%).
[0105] Example 9: Analysis of the base editing characteristics of preferred RRA miniBEs-8e (GGATY) at multiple sites
[0106] To comprehensively evaluate the base editing properties of RRA miniBEs-8e (GGATY), RRA miniBEs-8e (GGATY) obtained by fusing the N-terminal site, C-terminal site, and site 130 within d12f was preferred. This was co-transfected into cultured 293T cells with three single-guide RNAs (sgRNAs) (the sgRNAs simultaneously expressed uracil glycosylation inhibitors UGI and 2A-mCherry). After 72 hours of cell culture, double-positive cells simultaneously expressing EGFP and mCherry were collected by FACS. The sgRNAs used in this example were d12f-sg5 (sequence GGCAAGGGTCTTGATGCATC, PAM sequence TTTA), d12f-sg9 (sequence TACTTTGTCCTCCGGTTCTG, PAM sequence TTTG), and d12f-sg21 (sequence CCCCCACAGGATTGTAATAA, PAM sequence TTTA).
[0107] Genomic DNA was extracted from collected double-positive cells, and directional PCR was performed using primers that specifically amplified information at the d12f-sg89 site. The PCR products were then sequenced by Sanger to verify the editing efficiency of different miniBEs at the d12f-sg5, d12f-sg9, and d12f-sg21 sites.
[0108] The results showed that RRA d12f N-8e(GGATY) exhibited significant base editing activity at d12f-sg5, d12f-sg9, and d12f-sg21 sites (CT editing efficiency of 20%–40%), and the editing window was highly specific (C3). RRA d12f130-8e(GGATY) showed significantly lower base editing activity at all sites than RRA d12f N-8e(GGATY), with the highest base editing activity at d12f-sg21 (CT editing efficiency of ~25%), and a slightly wider editing window (C3–C4). RRA d12f CL-8e (GGATY) exhibits significant base editing activity at d12f-sg5, d12f-sg9, and d12f-sg21 sites (CT editing efficiency of 10%–44%), with the editing window significantly affected by the site. At d12f-sg5, it shows significant editing ability for cytosine at sites C3–C17 (CT editing efficiency >10%), while at d12f-sg9 and d12f-sg21, it mainly shows significant editing ability for cytosine at site C3 (CT editing efficiency >10%).
[0109] The above description of the embodiments is provided to enable those skilled in the art to understand and use the invention. It will be apparent to those skilled in the art that various modifications can be made to these embodiments, and the general principles described herein can be applied to other embodiments without inventive effort. Therefore, the present invention is not limited to the above embodiments, and any improvements and modifications made by those skilled in the art based on the disclosure of the present invention without departing from the scope of the invention should be within the protection scope of the present invention. sequence list <110> Fudan University <120> A base editor, its construction method and application <160> 9 <170> SIPOSequenceListing 1.0 <210> 1 <211> 529 <212> PRT <213> Artificial Sequence <400> 1 Met Ala Lys Asn Thr Ile Thr Lys Thr Leu Lys Leu Arg Ile Val Arg 1 5 10 15 Pro Tyr Asn Ser Only Glu Val Glu Lys Ile Val Only Asp Glu Lys Asn 20 25 30 Asn Arg Glu Lys Ile Ala Leu Glu Lys Asn Lys Asp Lys Val Lys Glu 35 40 45 Only Cys Served Lys Leu Lys Val Only Tyr Cys Thr Thr Gln Val 50 55 60 Glu Arg Asn Ala Cys Leu Phe Cys Lys Ala Arg Lys Leu Asp Lys 65 70 75 80 Phe Tyr Gln Lys Leu Arg Gly Gln Phe Pro Asp Ala Val Phe Trp Gln 85 90 95 Glu Ile Is Glu Ile Phe Arg Gln Leu Gln Lys Gln Ala Ala Glu Ile 100 105 110 Tyr Asn Gln Ser Leu To Glu Leu Tyr Tyr Glu To Phe To Lys Gly 115 120 125 Lys Gly Ile Ala Asn Ala Ser Ser Val Glu His Tyr Leu Ser Asp Val 130 135 140 Cys Tyr Thr Arg Ala Ala Glu Leu Phe Lys Asn Ala Ala Ile Ala Ser 145 150 155 160 Gly Leu Arg Ser Lys Ile Lys Asn Phe Arg Leu Lys Glu Leu Lys 165 170 175 Asn Met Lys Ser Gly Leu Pro Thr Thr Lys Ser Asp Asn Phe Pro Ile 180 185 190 Pro Leu Val Lys Gln Lys Gly Gly Gln Tyr Thr Gly Phe Glu Ile Ser 195 200 205 Asn His Asn Ser Asp Phe Ile Ile Lys Ile Pro Phe Gly Arg Trp Gln 210 215 220 Val Lys Lys Glu Ile Asp Lys Tyr Arg Pro Trp Glu Lys Phe Asp Phe 225 230 235 240 Glu Gln Val Gln Lys Ser Pro Lys Pro Ile Ser Leu Leu Leu Ser Thr 245 250 255 Gln Arg Arg Lys Arg Asn Lys Gly Trp Ser Lys Asp Glu Gly Thr Glu 260 265 270 Ala Glu Ile Lys Lys Val Met Asn Gly Asp Tyr Gln Thr Ser Tyr Ile 275 280 285 Glu Val Lys Arg Gly Ser Lys Ile Cys Glu Lys Ser Ala Trp Met Leu 290 295 300 Asn Leu Ser Ile Asp Val Pro Lys Ile Asp Lys Gly Val Asp Pro Ser 305 310 315 320 Ile Ile Gly Gly Ile Ala Val Gly Val Lys Ser Pro Leu Val Cys Ala 325 330 335 Ile Asn Asn Ala Phe Ser Arg Tyr Ser Ile Ser Asp Asn Asp Leu Phe 340 345 350 His Phe Asn Lys Lys Met Phe Ala Arg Arg Arg Ile Leu Leu Lys Lys 355 360 365 Asn Arg His Lys Arg Ala Gly His Gly Ala Lys Asn Lys Leu Lys Pro 370 375 380 Thr Ile Leu Thr Glu Lys Ser Glu Arg Phe Arg Lys Lys Leu Ile 385 390 395 400 Glu Arg Trp Ala Cys Glu Ile Ala Asp Phe Phe Ile Lys Asn Lys Val 405 410 415 Gly Thr Val Gln Met Glu Asn Leu Glu Ser Met Lys Arg Lys Glu Asp 420 425 430 Ser Tyr Phe Asn Ile Arg Leu Arg Gly Phe Trp Pro Tyr Ala Glu Met 435 440 445 Gln Asn Lys with Glu Phe Lys Leu Lys Gln Tyr Gly with Glu with Arg 450 455 460 Lys Val Ala Pro Asn Asn Thr Ser Lys Thr Cys Ser Lys Cys Gly His 465 470 475 480 Leu Asn Asn Tyr Phe Asn Phe Glu Tyr Arg Lys Lys Asn Lys Phe Pro 485 490 495 His Phe Lys Cys Glu Lys Cys Asn Phe Lys Glu Asn Ala Ala Tyr Asn 500 505 510 Ala Ala Leu Asn Ile Ser Asn Pro Lys Leu Lys Ser Thr Lys Glu Glu 515 520 525 Pro <210> 2 <211> 15 <212> PRT <213> Artificial Sequence <400> 2 Pro Lys Lys Lys Arg Lys Val Gly Ile His Gly Val Pro Ala Ala 1 5 10 15 <210> 3 <211> 16 <212> PRT <213> Artificial Sequence <400> 3 Lys Arg Pro Ala Ala Thr Lys Lys Ala Gly Gln Ala Lys Lys Lys Lys 1 5 10 15 <210> 4 <211> 23 <212> PRT <213> Artificial Sequence <400> 4 Lys Arg Pro Ala Ala Thr Lys Lys Ala Gly Gln Ala Lys Lys Lys Lys 1 5 10 15 Glu Phe Thr Ser Gly Ser Gly 20 <210> 5 <211> 116 <212> DNA <213> Artificial Sequence <400> 5 accgcttcac cgagtgaagg tgggctgctt gcatcagcct aatgtcgaga agtgctttct 60 tcggaaagta accctcgaaa caaagaaagg aatgcaacnn nnnnntttta tttttt 116 <210> 6 <211> 211 <212> PRT <213> Artificial Sequence <400> 6 Thr Ser Gly Ser Pro Lys Lys Lys Arg Lys Val Ser Glu Val Glu Phe 1 5 10 15 Ser His Glu Tyr Trp Met Arg His Ala Leu Thr Leu Ala Lys Arg Ala 20 25 30 Arg Asp Glu Arg Glu Val Pro Val Gly Ala Val Leu Val Leu Asn Asn 35 40 45 Arg Val Ile Gly Glu Gly Trp Asn Arg Ala Ile Gly Leu His Asp Pro 50 55 60 Thr Ala His Ala Glu Ile Met Ala Leu Arg Gln Gly Gly Leu Val Met 65 70 75 80 Gln Asn Tyr Arg Leu Ile Asp Ala Thr Leu Tyr Val Thr Phe Glu Pro 85 90 95 Cys Val Met Cys Ala Gly Ala Met Ile His Ser Arg Ile Gly Arg Val 100 105 110 Val Phe Gly Val Arg Asn Ser Lys Arg Gly Ala Ala Gly Ser Leu Met 115 120 125 Asn Val Leu Asn Tyr Pro Gly Met Asn His Arg Val Glu Ile Thr Glu 130 135 140 Gly Ile Leu Ala Asp Glu Cys Ala Ala Leu Leu Cys Asp Phe Tyr Arg 145 150 155 160 Met Pro Arg Gln Val Phe Asn Ala Gln Lys Lys Ala Gln Ser Ser Ile 165 170 175 Asn Ser Gly Gly Ser Ser Gly Gly Ser Ser Gly Ser Glu Thr Pro Gly 180 185 190 Thr Ser Glu Ser Ala Thr Pro Glu Ser Ser Gly Gly Ser Ser Gly Gly 195 200 205 Ser Ser Arg 210 <210> 7 <211> 405 <212> PRT <213> Artificial Sequence <400> 7 Pro Lys Lys Lys Arg Lys Val Ser Glu Val Glu Phe Ser His Glu Tyr 1 5 10 15 Trp Met Arg His Ala Leu Thr Leu Ala Lys Arg Ala Arg Asp Glu Arg 20 25 30 Glu Val Pro Val Gly Ala Val Leu Val Leu Asn Asn Arg Val Ile Gly 35 40 45 Glu Gly Trp Asn Arg Ala Ile Gly Leu His Asp Pro Thr Ala His Ala 50 55 60 Glu Ile Met Ala Leu Arg Gln Gly Gly Leu Val Met Gln Asn Tyr Arg 65 70 75 80 Leu Ile Asp Ala Thr Leu Tyr Val Thr Phe Glu Pro Cys Val Met Cys 85 90 95 Ala Gly Ala Met Ile His Ser Arg Ile Gly Arg Val Val Phe Gly Val 100 105 110 Arg Asn Ser Lys Arg Gly Ala Ala Gly Ser Leu Met Asn Val Leu Asn 115 120 125 Tyr Pro Gly Met Asn His Arg Val Glu Ile Thr Glu Gly Ile Leu Ala 130 135 140 Asp Glu Cys Ala Ala Leu Leu Cys Asp Phe Tyr Arg Met Pro Arg Gln 145 150 155 160 Val Phe Asn Ala Gln Lys Lys Ala Gln Ser Ser Ile Asn Ser Gly Gly 165 170 175 Ser Ser Gly Gly Ser Ser Gly Ser Glu Thr Pro Gly Thr Ser Glu Ser 180 185 190 Ala Thr Pro Glu Ser Ser Gly Gly Ser Ser Gly Gly Ser Ser Ser Ser 195 200 205 Glu Val Glu Phe Ser His Glu Tyr Trp Met Arg His Ala Leu Thr Leu 210 215 220 Ala Lys Arg Ala Arg Asp Glu Arg Glu Val Pro Val Gly Ala Val Leu 225 230 235 240 Val Leu Asn Asn Arg Val Ile Gly Glu Gly Trp Asn Arg Ala Ile Gly 245 250 255 Leu His Asp Pro Thr Ala His Ala Glu Ile Met Ala Leu Arg Gln Gly 260 265 270 Gly Leu Val Met Gln Asn Tyr Arg Leu Ile Asp Ala Thr Leu Tyr Val 275 280 285 Thr Phe Glu Pro Cys Val Met Cys Ala Gly Ala Met Ile His Ser Arg 290 295 300 Ile Gly Arg Val Val Phe Gly Val Arg Asn Ser Lys Arg Gly Ala Ala 305 310 315 320 Gly Ser Leu Met Asn Val Leu Asn Tyr Pro Gly Met Asn His Arg Val 325 330 335 Glu Ile Thr Glu Gly Ile Leu Ala Asp Glu Cys Ala Ala Leu Leu Cys 340 345 350 Asp Phe Tyr Arg Met Pro Arg Gln Val Phe Asn Ala Gln Lys Lys Ala 355 360 365 Gln Ser Ser Ile Asn Ser Gly Gly Ser Ser Gly Gly Ser Ser Gly Ser 370 375 380 Glu Thr Pro Gly Thr Ser Glu Ser Ala Thr Pro Glu Ser Ser Gly Gly 385 390 395 400 Ser Ser Gly Gly Ser 405 <210> 8 <211> 286 <212> PRT <213> Artificial Sequence <400> 8 Pro Lys Lys Lys Arg Lys Val Gly Ser Ser Ser Glu Thr Gly Pro Val 1 5 10 15 Ala Val Asp Pro Thr Leu Arg Arg Arg Ile Glu Pro His Glu Phe Glu 20 25 30 Val Phe Phe Asp Pro Arg Glu Leu Arg Lys Glu Thr Cys Leu Leu Tyr 35 40 45 Glu Ile Asn Trp Gly Gly Arg His Ser Ile Trp Arg His Thr Ser Gln 50 55 60 Asn Thr Asn Lys His Val Glu Val Asn Phe Ile Glu Lys Phe Thr Thr 65 70 75 80 Glu Arg Tyr Phe Cys Pro Asn Thr Arg Cys Ser Ile Thr Trp Phe Leu 85 90 95 Ser Tyr Ser Pro Cys Gly Glu Cys Ser Arg Ala Ile Thr Glu Phe Leu 100 105 110 Ser Arg Tyr Pro His Val Thr Leu Phe Ile Tyr Ile Ala Arg Leu Tyr 115 120 125 His His Ala Asp Pro Glu Asn Arg Gln Gly Leu Arg Asp Leu Ile Ser 130 135 140 Ser Gly Val Thr Ile Gln Ile Met Thr Glu Gln Glu Ser Gly Tyr Cys 145 150 155 160 Trp Arg Asn Phe Val Asn Tyr Ser Pro Ser Asn Glu Ala His Trp Pro 165 170 175 Arg Tyr Pro His Leu Trp Val Arg Leu Tyr Val Leu Glu Leu Tyr Cys 180 185 190 Ile Ile Leu Gly Leu Pro Pro Cys Leu Asn Ile Leu Arg Arg Lys Gln 195 200 205 Pro Gln Leu Thr Phe Phe Thr Ile Ala Leu Gln Ser Cys His Tyr Gln 210 215 220 Arg Leu Pro Pro His Ile Leu Trp Ala Thr Gly Leu Lys Ser Gly Ser 225 230 235 240 Glu Thr Pro Gly Thr Ser Glu Ser Ala Thr Pro Glu Ser Ser Ser Gly 245 250 255 Gly Ser Ser Gly Gly Ser Ser Gly Ser Glu Thr Pro Gly Thr Ser Glu 260 265 270 Ser Ala Thr Pro Glu Ser Ser Gly Gly Ser Ser Gly Gly Ser 275 280 285 <210> 9 <211> 198 <212> PRT <213> Artificial Sequence <400> 9 Glu Ala Ser Pro Ala Ser Gly Pro Arg His Leu Met Asp Pro His Ile 1 5 10 15 Phe Thr Ser Asn Phe Asn Asn Gly Ile Gly Arg His Lys Thr Tyr Leu 20 25 30 Cys Tyr Glu Val Glu Arg Leu Asp Asn Gly Thr Ser Val Lys Met Asp 35 40 45 Gln His Arg Gly Phe Leu His Asn Gln Ala Lys Asn Leu Leu Cys Gly 50 55 60 Phe Tyr Gly Arg His Ala Glu Leu Arg Phe Leu Asp Leu Val Pro Ser 65 70 75 80 Leu Gln Leu Asp Pro Ala Gln Ile Tyr Arg Val Thr Trp Phe Ile Ser 85 90 95 Trp Ser Pro Cys Phe Ser Trp Gly Cys Ala Gly Glu Val Arg Ala Phe 100 105 110 Leu Gln Glu Asn Thr His Val Arg Leu Arg Ile Phe Ala Ala Arg Ile 115 120 125 Phe Asp Tyr Asp Pro Leu Tyr Lys Glu Ala Leu Gln Met Leu Arg Asp 130 135 140 Ala Gly Ala Gln Val Ser Ile Met Thr Tyr Asp Glu Phe Lys His Cys 145 150 155 160 Trp Asp Thr Phe Val Asp His Gln Gly Cys Pro Phe Gln Pro Trp Asp 165 170 175 Gly Leu Asp Glu His Ser Gln Ala Leu Ser Gly Arg Leu Arg Ala Ile 180 185 190 Leu Gln Asn Gln Gly Asn 195
Claims
1. A base editor, characterized in that, Choose from one of the following: (1) d12f N-YE1: The fusion protein obtained by fusing the NLS-YE1-linker as shown in SEQ ID NO.8 to the N-terminus of dCasd12f; (2) d12f CL-YE1: The fusion protein obtained by fusing the NLS-YE1-linker as shown in SEQ ID NO.8 to the C-terminal CL site of dCasd12f; (3) d12f N-3A130: The fusion protein obtained by fusing 3A130, as shown in SEQ ID NO.9, to the N-terminus of dCasd12f; (4) d12f CL-3A130: The fusion protein obtained by fusing 3A130, as shown in SEQ ID NO.9, to the C-terminal CL site of dCasd12f; (5) RRA d12f N-YE1: The fusion protein is obtained by introducing the RRA mutation into d12f N-YE1, where d12f N-YE1 is the fusion protein obtained by fusing the NLS-YE1-linker as shown in SEQ ID NO.8 to the N-terminus of dCasd12f. (6) RRA d12f 130-YE1: The fusion protein is d12f 130-YE1, which is obtained by introducing the RRA mutation into d12f 130-YE1. d12f 130-YE1 is a fusion protein obtained by fusing the NLS-YE1-linker as shown in SEQ ID NO. 8 into position 130 inside dCasd12f. (7) RRA d12f CL-YE1: The fusion protein is d12f CL-YE1 with the introduction of the RRA mutation, wherein d12f CL-YE1 is a fusion protein obtained by fusing the NLS-YE1-linker as shown in SEQ ID NO.8 to the CL site at the C-terminus of dCasd12f. (8) RRA d12f N-3A130: The fusion protein is obtained by introducing the RRA mutation into d12f N-3A130, where d12f N-3A130 is the fusion protein obtained by fusing 3A130 as shown in SEQ ID NO. 9 to the N-terminus of dCasd12f. (9) RRA d12f 128-3A130: The fusion protein is obtained by introducing the RRA mutation into d12f 128-3A130, where d12f 128-3A130 is the fusion protein obtained by fusing 3A130 as shown in SEQ ID NO. 9 into position 128 inside dCasd12f. (10) RRA d12f 130-3A130: The fusion protein is obtained by introducing the RRA mutation into d12f 128-3A130, wherein d12f 130-3A130 is a fusion protein obtained by fusing 3A130 as shown in SEQ ID NO. 9 into position 130 inside dCasd12f. (11) RRA d12f CL-3A130: The fusion protein is d12f N-3A130 with the introduction of the RRA mutation, wherein d12f CL-3A130 is a fusion protein obtained by fusing 3A130 as shown in SEQ ID NO. 9 to the CL site at the C-terminus of dCasd12f. Among them, RRA mutation refers to the introduction of the D143R-T147R-E151A point mutation combination in dCasd12f; The amino acid sequence of dCasd12f is shown in SEQ ID NO.
1.
2. A polynucleotide, characterized in that, Encodes the fusion protein in the base editor of claim 1.
3. A carrier, characterized in that, It contains the polynucleotide described in claim 2.
4. A host cell, characterized in that, It contains the base editor of claim 1, or the carrier of claim 3.
5. A reagent kit, characterized in that, It contains reagents for constructing the base editor of claim 1.