apoBEC3F-be4max editor

By modifying the APOBEC3F protein and fusing it with the nCas9 protein, the APOBEC3F-BE4max editor was designed, which solved the problems of insufficient editing efficiency and accuracy in the existing technology and achieved efficient and accurate cytosine base editing.

CN121294478BActive Publication Date: 2026-05-01TIANJIN TUMOR HOSPITAL
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TIANJIN TUMOR HOSPITAL
Filing Date
2025-12-12
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

There is still room for improvement in the editing efficiency and precision of existing APOBEC3 family deaminases in gene editing, especially since the application potential of APOBEC3F has not been fully explored, making it difficult to meet the needs for higher efficiency and more precise regulation.

Method used

By modifying the C-terminal domain of the APOBEC3F protein, an APOBEC3F-BE4max editor was designed, including specific amino acid substitutions such as W310D/S216T/H198D and N240G/Y359A. Combined with the nCas9 protein and the uracil glycosylation inhibitor UGI, a highly efficient and high-fidelity base editing system was formed.

Benefits of technology

It achieves efficient conversion of cytosine to thymine, significantly improves editing efficiency, optimizes the editing window, reduces off-target effects, and enhances editing accuracy, significantly surpassing existing tools.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121294478B_ABST
    Figure CN121294478B_ABST
Patent Text Reader

Abstract

This invention discloses an APOBEC3F-BE4max editor. This invention relates to the field of biotechnology and provides a base editor containing a cytosine deaminase variant and an nCas9 protein; the cytosine deaminase variant is obtained by modifying the C-terminal domain of the wild-type APOBEC3F protein; the modification includes mutating the C-terminal domain (SEQ ID NO:1) of the wild-type APOBEC3F protein by either W310D / S216T / H198D or W310D / S216T / H198D / N240G / Y359A mutation. The former achieves an average editing efficiency of 68.2% and expands the editing window to positions 3-9; the latter precisely restricts the editing window to positions 5-6 while maintaining an editing efficiency of 42.5%, and both also have low off-target rates, demonstrating their superior specificity.
Need to check novelty before this filing date? Find Prior Art

Description

APOBEC3F-BE4max Editor Technical Field

[0001] This invention relates to the field of biotechnology, and more specifically to the APOBEC3F-BE4max editor. Background Technology

[0002] APOBEC3 family deaminases (including A3A, A3B, and A3G) have shown broad application potential in cytosine base editors (CBEs) [Yang, L., Huo, Y., Wang, M., Zhang, D., Zhang, T., Wu, H., Rao, X., Meng, H., Yin, S., Mei, J., et al. (2024). Engineering APOBEC3Adeaminase for highly accurate and efficient base editing. Nat Chem Biol 20, 1176-1187., Jin, S., Fei, H., Zhu, Z., Luo, Y., Liu, J., Gao, S., Zhang, F., Chen, YH, Wang, Y., and Gao, C. (2020). Rationally Designed APOBEC3BCytosine Base Editors with Improved Specificity. Mol Cell 79, 728-740.e726., Lee, S., Ding, N., Sun, Y., Yuan, T., Li, J., Yuan, Q., Liu, L., Yang, J., Wang, Q., Kolomeisky, AB, et al. (2020). Single C-to-T substitution using engineered APOBEC3G-nCas9 base editors with minimum genome- and transcriptome-wide off-target effects. Sci Adv 6, eaba1773.]. In the molecular mechanisms of gene editing, these enzymes can efficiently and precisely convert cytosine (C) on the genomic DNA strand into uracil (U) through specific biochemical reactions. This allows for precise and targeted editing of the genome sequence through the cell's own DNA repair mechanisms, providing crucial technical support for gene function research, disease model construction, and gene therapy exploration.

[0003] However, although deaminases such as A3A, A3B, and A3G have been widely used in CBEs and achieved certain results, there is still ample room for improvement in their editing efficiency and precision, given the development needs of gene editing technology for higher efficiency and more precise regulation. Especially for APOBEC3F (A3F), although previous studies have shown that it possesses strong deaminase activity and specificity in vitro [Haché, G., Liddament, MT, and Harris, RS (2005). Theretroviral hypermutation specificity of APOBEC3F and APOBEC3G is governed by the C-terminal DNA cytosine deaminase domain. J Biol Chem 280, 10920-10924.], its application potential in base editing has not yet been fully explored.

[0004] It is worth noting that in related studies conducted within the BE3 editing system framework, comparative analysis of the editing efficiency of different APOBEC3 family deaminases revealed that A3F's editing efficiency is second only to A3A and A3B, and it exhibits the least bias between methylated and unmethylated GC sequences [Yang, L., Huo, Y., Wang, M., Zhang, D., Zhang,T., Wu, H., Rao, X., Meng, H., Yin, S., Mei, J., et al. (2024). Engineering APOBEC3A deaminase for highly accurate and efficient base editing. Nat ChemBiol 20, 1176-1187.]. This suggests that it is more likely to maintain relatively uniform and stable editing behavior under complex genomic backgrounds and diverse gene sequence characteristics. This unique characteristic, from both a molecular basis and application potential perspective, indicates that A3F has the potential to become an ideal molecular framework for developing novel and superior cytosine base editors (CBEs). It is expected to provide a new molecular tool basis for gene editing technology to break through the current efficiency and accuracy bottlenecks and expand into a wider range of application scenarios. Therefore, targeted modification and in-depth research on A3F also have important scientific value and application prospects. Summary of the Invention

[0005] The purpose of this invention is to provide an APOBEC3F-BE4max editor.

[0006] Firstly, this invention claims protection for a base editor.

[0007] The base editor claimed in this invention contains a cytosine deaminase variant and the nCas9 protein.

[0008] The cytosine deaminase variant is obtained by modifying the C-terminal domain of the wild-type APOBEC3F protein. The amino acid sequence of the C-terminal domain of the wild-type APOBEC3F protein is shown in SEQ ID NO:1.

[0009] The modification includes replacing the 127th amino acid of SEQ ID NO:1 with W instead of D (corresponding to W310D in the embodiment, the specific position in the embodiment is the position in the full-length APOBEC3F protein, the same below), replacing the 33rd amino acid with S instead of T (corresponding to S216T in the embodiment), and replacing the 15th amino acid with H instead of D (corresponding to H198D in the embodiment).

[0010] In some embodiments of the present invention, the cytosine deaminase variant is heA3F; the amino acid sequence of heA3F is obtained by replacing the 127th amino acid of SEQ ID NO:1 with D instead of W, the 33rd amino acid with T instead of S instead of S, and the 15th amino acid with D instead of H instead of H (corresponding to W310D / S216T / H198D in the embodiments).

[0011] Furthermore, based on the above modifications, the modifications also include replacing the amino acid at position 57 of SEQ ID NO:1 with G (corresponding to N240G in the embodiment) and replacing the amino acid at position 176 with A (corresponding to Y359A in the embodiment).

[0012] In some embodiments of the present invention, the cytosine deaminase variant is haA3F-GA; the amino acid sequence of haA3F-GA is obtained by replacing the amino acid at position 127 of SEQ ID NO:1 with W instead of D, the amino acid at position 33 with S instead of T, the amino acid at position 15 with H instead of D, the amino acid at position 57 with N instead of G, and the amino acid at position 176 with Y instead of A (corresponding to W310D / S216T / H198D / N240G / Y359A in the embodiments).

[0013] The nCas9 protein may be the Cas9(D10A) protein.

[0014] As needed, the base editor may also contain all or part of the following: nuclear localization signal, linker peptide, and uracil glycosylation inhibitor UGI.

[0015] Furthermore, the base editor may be a fusion protein formed from the N-terminus to the C-terminus by sequentially fusing the nuclear localization signal NLS (denoted as nuclear localization signal 1), the cytosine deaminase variant described in the first aspect above, the linker peptide (denoted as linker peptide 1), the Cas9 (D10A) protein, the linker peptide (denoted as linker peptide 2), the uracil glycosylase inhibitor UGI, the linker peptide (denoted as linker peptide 3), the uracil glycosylase inhibitor UGI, and the nuclear localization signal NLS (denoted as nuclear localization signal 2).

[0016] Wherein, the amino acid sequence of the nuclear localization signal 1 may be as shown in positions 1-19 of SEQ ID NO:3 (or positions 1-19 of SEQ ID NO:5); the amino acid sequence of the linker peptide 1 may be as shown in positions 209-240 of SEQ ID NO:3 (or positions 209-240 of SEQ ID NO:5); the amino acid sequence of the Cas9 (D10A) protein may be as shown in positions 241-1607 of SEQ ID NO:3 (or positions 241-1607 of SEQ ID NO:5); the amino acid sequence of the linker peptide 2 may be as shown in positions 1608-1617 of SEQ ID NO:3 (or positions 1608-1617 of SEQ ID NO:5); the amino acid sequence of the uracil glycosylation inhibitor UGI may be as shown in positions 1618-1700 or 1711-1793 of SEQ ID NO:3 (or positions 1-19 of SEQ ID NO:5); The amino acid sequence of the linker peptide 3 may be as shown in positions 1701-1710 of SEQ ID NO:3 (or positions 1701-1710 of SEQ ID NO:5); the amino acid sequence of the nuclear localization signal 2 may be as shown in positions 1794-1810 of SEQ ID NO:3 (or positions 1794-1810 of SEQ ID NO:5).

[0017] In some embodiments of the present invention, the base editor is heA3F-BE4max; the amino acid sequence of heA3F-BE4max is shown in SEQ ID NO:3.

[0018] In some embodiments of the present invention, the base editor is haA3F-GA-BE4max; the amino acid sequence of haA3F-GA-BE4max is shown in SEQ ID NO:5.

[0019] Secondly, this invention claims protection for a base editing system.

[0020] The base editing system claimed in this invention contains the base editor and guide RNA described in the first aspect above.

[0021] Thirdly, the present invention claims protection for a cytosine deaminase variant, which is the cytosine deaminase variant described in the first aspect above.

[0022] Fourthly, this invention claims protection for any of the following biological materials:

[0023] (A1) Nucleic acid molecules encoding the cytosine deaminase variant described in the third aspect above;

[0024] (A2) An expression cassette, recombinant vector, recombinant microorganism, or recombinant cell containing the nucleic acid molecule described in (A1);

[0025] (A3) Encodes the nucleic acid molecule of the base editor described in the preceding section;

[0026] (A4) An expression cassette, recombinant vector, recombinant microorganism, or recombinant cell containing the nucleic acid molecule described in (A3);

[0027] (A5) Nucleic acid molecules that encode the base editing system described in the second aspect above;

[0028] (A6) An expression cassette or recombinant vector or recombinant microorganism or recombinant cell containing the nucleic acid molecule described in (A5).

[0029] In this invention, the expression cassette refers to DNA capable of expressing the cytosine deaminase variant described in the third aspect above, or the base editor described in the first aspect above, or the base editing system described in the second aspect above, in a host cell. The expression cassette may include all regulatory sequences necessary to initiate the expression of the coding sequence (i.e., the nucleic acid molecule described in (A1) or (A3) or (A5) above) of the cytosine deaminase variant described in the third aspect above, or the base editor described in the first aspect above, or the base editing system described in the second aspect above, under compatible conditions. The regulatory sequences, under these conditions, guide the expression of the coding sequence in a suitable host cell of the cytosine deaminase variant described in the third aspect above, or the base editor described in the first aspect above, or the base editing system described in the second aspect above. The regulatory sequences include, but are not limited to, leader sequences, polyadenylated sequences, propeptide sequences, promoters, signal sequences, and transcription terminators. At a minimum, the regulatory sequences shall include a promoter and termination signals for transcription and translation.

[0030] In this invention, the recombinant vector can be a cloning vector or an expression vector. As used herein, the term "vector" refers to a nucleic acid delivery vehicle into which polynucleotides can be inserted. When a vector enables the expression of a protein encoded by the inserted polynucleotide, it is called an expression vector. Vectors can be introduced into host cells through transformation, transduction, or transfection, allowing the genetic material elements they carry to be expressed in the host cells. Vectors are well known to those skilled in the art and include, but are not limited to, plasmids, phage particles, cosmids, and artificial chromosomes.

[0031] In this invention, the recombinant microorganism may specifically be bacteria, yeast, algae, or fungi carrying the nucleic acid molecules. The bacteria may be any of the following: prokaryotic microorganisms, Gram-negative bacteria, Escherichia coli, or Escherichia coli.

[0032] In this invention, the recombinant cell may specifically be an animal cell or a plant cell carrying the nucleic acid molecule.

[0033] Fifthly, the present invention claims a kit for achieving the conversion of cytosine to thymine in a target DNA sequence.

[0034] The kit claimed in this invention contains the cytosine deaminase variant described in the third aspect above, or the base editor described in the first aspect above, or the base editing system described in the second aspect above, or the biological material described in the fourth aspect above.

[0035] Sixthly, the present invention claims protection for any of the following applications:

[0036] (B1) The application of the cytosine deaminase variant described in the third aspect above, or the base editor described in the first aspect above, or the base editing system described in the second aspect above, or the biomaterial described in the fourth aspect above, or the kit described in the fifth aspect above, in achieving the conversion of cytosine to thymine in the target DNA sequence; the application is a non-disease diagnostic and therapeutic application.

[0037] (B2) The use of the cytosine deaminase variant described in the third aspect above, or the base editor described in the first aspect above, or the base editing system described in the second aspect above, or the biomaterial described in the fourth aspect above, in the preparation of products for achieving the conversion of cytosine to thymine in a target DNA sequence. The product may be a reagent or a kit.

[0038] (B3) The use of the cytosine deaminase variant described in the third aspect above, or the base editor described in the first aspect above, or the base editing system described in the second aspect above, or the biomaterial described in the fourth aspect above, in the preparation of a medicament for treating a disease; wherein the disease is caused by a pathogenic point mutation that can be corrected by cytosine-to-thymine (CT) base editing.

[0039] In a seventh aspect, the present invention claims a method for achieving the conversion of cytosine to thymine in a target DNA sequence.

[0040] The method for achieving the conversion of cytosine to thymine in a target DNA sequence, as claimed in this invention, may include the step of contacting the target DNA sequence with the base editing system described in the second aspect above. The method may be a non-disease diagnosis or treatment method.

[0041] This invention, through systematic verification, confirms that the A3F-CBEs (i.e., the APOBEC3F-BE4max editor) based on ESM and structure-guided modification significantly outperforms existing tools in editing performance: the high-efficiency variant heA3F (W310D / S216T / H198D) achieves an average editing efficiency of 68.2% at 10 genomic sites, which is up to 1.9 times and 3.3 times higher than A3A-BE4max and Anc689-BE4max, respectively, and the editing window is extended to positions 3-9; the high-precision variant haA3F (N240G / Y359A) precisely restricts the editing window to positions 5-6 while maintaining an editing efficiency of 42.5%, which is 3.0 times higher than haA3A-G at most sites. Off-target analysis showed that the heA3F and haA3F variants induced significantly less off-target RNA editing compared to A3A-BE4max. Although heA3Fs showed higher off-target activity than haA3A-A, comparable to A3A, haA3Fs maintained high fidelity on both DNA and RNA substrates, confirming their superior specificity. Attached Figure Description

[0042] Figure 1 shows the structural and rational engineering modifications of APOBEC3F to enhance cytosine base editing efficiency. a) is a heatmap showing the editing frequency of CT at the FANCF / SPRK3 / CTLA locus of 19 potential CBE mutants for enhancing A3F activity, data from n=3 independent biological replicates. b) is the average editing frequency of the three CBE variants at each prototypical spacer position across nine endogenous sites (VEGFA site 1, PPP1R12C site 1, TET2 site, PPP1R12C site 2, PPP1R12C site 3, AFF1 site, MAGEA1 site 1, DNMT3B site, and HEK4OT2 site).

[0043] Figure 2 shows computer-guided enhanced CT editing of the APOBEC3F engineered variant. a) shows the CT editing frequency of the evolved CBE variant at the HIRA / FANCF-site2 site, data from n=3 independent biological replicates (signatures >10%). b) shows the mean CT editing efficiency of the combinatorial A3F mutant at the HIRA site. Data are expressed as mean ± standard deviation of n=3 independent biological replicates. c) shows the mean editing frequency of each prototypical spacer sequence position at eight endogenous sites (DNMT3B, AFF1, PPP1R12C3, PPP1R12C2, PSMB2, VEGFA1, SPPK32, and TEF2) in HEK293T cells for both CBE variants.

[0044] Figure 3 illustrates the rational design of A3F-guided CBEs with higher fidelity. a) is a heatmap showing the editing frequencies of eight potential CBE mutants for improving heA3F specificity at the EMX1-site1 / ABL1 site CT, data from n=3 independent biological replicates. b) shows the mean editing frequencies of the prototype spacer sequence at each of the eight endogenous sites (EMX1 site 2, RP11 site, TP53 site, PPP1R12C site 4, EMX1 site 3, PPP1R12C site 5, PDCD1 site, and PPP1R12C site 2) of the four high-fidelity CBE variants in HEK293T cells. c) shows the insertion / deletion frequencies of the four high-fidelity CBE variants at the eight endogenous sites in (b). Statistical significance was assessed using a two-tailed Student's t-test.

[0045] Figure 4 compares the performance of the A3F-guided base editor with that of a state-of-the-art CBE. a) is a violin plot showing the CT editing efficiency of five high-efficiency CBE variants at nine endogenous sites (EMX1 site 1, ABL1 site, EMX2 site 2, PPP1R12C site 4, PPP1R12C site 5, EMX1 site 3, PDCD1 site, TP53 site, and AFF1 site) in HEK293T cells at positions 1-20. The horizontal red line represents the median. b) shows the specificity-efficiency correlation of high-fidelity CBE variants. X-axis: specificity ratio (maximum CT efficiency at the target site / maximum CT efficiency in the flanking region); Y-axis: mean editing efficiency (positions 1-20, n=8 sites, from e). Flanking positions are defined as the region with the highest editing efficiency ±1 nucleotide.

[0046] Figure 5 shows the off-target effects of heA3F and haA3F in HEK293T cells. a represents the target and off-target editing efficiencies of A3A, Anc689, haA3A-A, heA3F, and heA3F-359 (color gradient from light to dark) at EMX1 and HEK2 genomic loci in HEK293T cells. Data are expressed as mean ± standard deviation (n=3). b represents the target and off-target editing efficiencies of haA3A-G, eA3A, haA3F, and haA3F-GA (color gradient from light to dark) at HEK2 and HEK3 genomic loci in HEK293T cells. Data are expressed as mean ± standard deviation (n=3).

[0047] Figure 6 shows the non-gRNA-dependent off-target effects of heA3F and haA3F in HEK293T cells. a) shows the cumulative CT editing frequencies of haA3A-A, haA3F-GA, heA3F, and A3A at the cytosine sites of the R-loop prototype spacer sequence. b) shows the off-target CU conversion analysis of the four CBE variants across the cellular transcriptome. Data are presented as mean ± standard deviation (n=3). Statistical significance was assessed using the two-tailed Student's test.

[0048] In each figure, all mutation sites marked refer to mutations based on A3F-CTD, and all mutation positions are counted with the first amino acid of the full length of APOBEC3F as the first position. Detailed Implementation

[0049] The construction and optimization scheme of the APOBEC3F-BE4max editor of this invention is as follows:

[0050] 1. ESM combined with structurally guided mutation strategy

[0051] To achieve precise optimization of the APOBEC3F-BE4max editor, a two-stage screening process was adopted. The first stage used three ESM protein language models (esm-1v, esm-msa-1b, and esm-if1) to score residue-level mutations in the C-terminal domain (A3F-CTD) sequence of APOBEC3F, predicting functional enhancement sites from UniProt / UniRef50 data and screening the top 10% of candidate sites. The second stage assessed evolutionary conservation using a position-specific scoring matrix (PSSM), further refining 10 core point mutations (such as H198D and S216T). Structure-guided design, combined with homologous protein crystal structures, compared and aligned substrate-binding interface residues (such as Y359) to design enhancing mutations that regulate DNA affinity (such as Y359A) and high-fidelity mutations that introduce steric hindrance (such as N240G), thus systematically balancing editing efficiency and specificity.

[0052] 2. APOBEC3F Remodeling Direction and Optimization Path

[0053] The APOBEC3F-BE4max editor is built with the APOBEC3F C-end structure domain (CTD) as the core functional module, and integrates the BE4max system (nCas9-UGI) to form the basic editor skeleton. The initial mutant library was designed with 15 rational mutants based on key sites in the literature (such as W310). Subsequent optimization focused on two main directions: the efficiency enhancement pathway widened the editing window through the W310D mutation and combined it with ESM-guided S216T / H198D / W310D (heA3F), which improved efficiency by 1.5–2 times compared to single mutations; the high-fidelity pathway (haA3F, which is a further modification of heA3F) used N240G / Y307G to compress the editing window to position 5–6 (using the base far from PAM in the 20bp of the target sequence as the first base, calculated towards PAM, the same below) (width ≤2nt). The N240G+Y359A combination maintained efficient editing while reducing CA / G conversion rate by 40%, and significantly reduced bypass editing and Indel frequency (reduction of up to 60%).

[0054] 3. High-performance CBE parameter standards and verification

[0055] In terms of editing efficiency, heA3F outperformed A3A-BE4max by 1.9 times (e.g., PDL1-C7) and Anc689-BE4max by 3.3 times (e.g., MAGEA1-C4) within the standard window (positions 4–9). haA3F achieved an average efficiency >45% across 8 gene loci, representing a maximum improvement of 3.0 times compared to haA3A-G (EMX1-C6). Regarding specificity control, haA3F exhibited a CA / G conversion rate ≤0.9% and an indel frequency ≤0.2%, significantly superior to the traditional BE4max system. Off-target effect assessment showed that the editing rate at gRNA-dependent off-target sites (e.g., HEK3) was <0.1%, deaminase-dependent background editing was reduced by 50% compared to Anc689, and RNA off-target levels were lower than A3A-BE4max as verified by RNA-seq.

[0056] 4. Experimental verification system and technological breakthroughs

[0057] Validation employed a stepwise strategy: initial screening was conducted on HEK293T cells to test for FANCF / SPRK3 / CTLA sites, with further validation extending to HeLa. Editing efficiency was analyzed using CRISPResso2 (parameters: -wc10-w20), and off-target effects were assessed using whole-transcriptome RNA-seq combined with GATK variant detection. Regarding core mutational functions, W310D expands the editing window, N240G compresses the window to positions 5–6, and Y359A inhibits the editing of unmethylated cytosine, collectively laying the foundation for efficient and safe editing.

[0058] The present invention will now be described in further detail with reference to specific embodiments. The given embodiments are merely illustrative of the invention and not intended to limit its scope. The embodiments provided below can serve as a guide for further improvements by those skilled in the art and do not constitute a limitation on the invention in any way.

[0059] Unless otherwise specified, the experimental methods used in the following examples are conventional methods, performed according to the techniques or conditions described in the literature in this field or according to the product instructions. Unless otherwise specified, the materials and reagents used in the following examples are commercially available.

[0060] The amino acid sequences and corresponding coding nucleotide sequences of the various cytosine deaminases involved in the following examples are as follows:

[0061] The amino acid sequence of the wild-type full-length A3F (FL) is shown in SEQ ID NO:7. In this sequence, positions 1-184 are the N-terminal domain; positions 185-373 are the C-terminal domain (CTD, i.e., SEQ ID NO:1).

[0062] The nucleotide sequence encoding the wild-type full-length A3F (FL) is shown in SEQ ID NO:8. In this sequence, positions 1-552 encode the N-terminal domain; positions 553-1119 encode the C-terminal domain (CTD) (i.e., SEQ ID NO:2).

[0063] The amino acid sequence of A3F-CTD_N214H is obtained by replacing the 214th amino acid N with H in the above wild-type full-length A3F and deleting amino acids 1-184; the corresponding coding nucleotide sequence is obtained by replacing AAT with CAC in positions 643-645 of the coding nucleotide sequence of the above wild-type full-length A3F and deleting nucleotides 1-552.

[0064] The amino acid sequence of A3F-CTD_Y196D is obtained by replacing the 196th amino acid Y of the above wild-type full-length A3F with D and deleting amino acids 1-184; the corresponding coding nucleotide sequence is obtained by replacing TAT at positions 589-591 of the coding nucleotide sequence of the above wild-type full-length A3F with GAC and deleting nucleotides 1-552.

[0065] The amino acid sequence of A3F-CTD_C259A is obtained by replacing amino acid C at position 259 of the above-mentioned wild-type full-length A3F with A and deleting amino acids 1-184; the corresponding coding nucleotide sequence is obtained by replacing TGT at positions 778-780 of the coding nucleotide sequence of the above-mentioned wild-type full-length A3F with GCG and deleting nucleotides 1-552.

[0066] The amino acid sequence of A3F-CTD_F302K is obtained by replacing amino acid F at position 302 of the above-mentioned wild-type full-length A3F with K and deleting amino acids 1-184; the corresponding coding nucleotide sequence is obtained by replacing TTC at positions 907-909 of the coding nucleotide sequence of the above-mentioned wild-type full-length A3F with AAG and deleting nucleotides 1-552.

[0067] The amino acid sequence of A3F-CTD_W310K is obtained by replacing amino acid W at position 310 of the above wild-type full-length A3F with K and deleting amino acids 1-184; the corresponding coding nucleotide sequence is obtained by replacing TGG at positions 931-933 of the coding nucleotide sequence of the above wild-type full-length A3F with AAG and deleting nucleotides 1-552.

[0068] The amino acid sequence of A3F-CTD_W310D is obtained by replacing amino acid W at position 310 of the above wild-type full-length A3F with D and deleting amino acids 1-184; the corresponding coding nucleotide sequence is obtained by replacing TGG at positions 931-933 of the coding nucleotide sequence of the above wild-type full-length A3F with GAC and deleting nucleotides 1-552.

[0069] The amino acid sequence of A3F-CTD_Y314A is obtained by replacing the 314th amino acid Y of the above wild-type full-length A3F with A and deleting amino acids 1-184; the corresponding coding nucleotide sequence is obtained by replacing TAC at positions 943-945 of the coding nucleotide sequence of the above wild-type full-length A3F with GCC and deleting nucleotides 1-552.

[0070] The amino acid sequence of A3F-CTD_Q315A is obtained by replacing the 315th amino acid Q of the above-mentioned wild-type full-length A3F with A, and deleting amino acids 1-184; the corresponding coding nucleotide sequence is obtained by replacing CAG at positions 946-948 of the coding nucleotide sequence of the above-mentioned wild-type full-length A3F with GCC, and deleting nucleotides 1-552.

[0071] The amino acid sequence of A3F-CTD_F363D is obtained by replacing amino acid F at position 363 of the above wild-type full-length A3F with D and deleting amino acids 1-184; the corresponding coding nucleotide sequence is obtained by replacing TTT at positions 1090-1092 of the coding nucleotide sequence of the above wild-type full-length A3F with GAC and deleting nucleotides 1-552.

[0072] The amino acid sequence of A3F-CTD_7x is obtained by replacing amino acid Y at position 196 with D, amino acid C at position 259 with A, amino acid F at position 302 with K, amino acid W at position 310 with K, amino acid Y at position 314 with A, amino acid Q at position 315 with A, amino acid F at position 363 with D, and deleting amino acids 1-184. The corresponding coding nucleotide sequence is obtained by replacing amino acid Y at position 314 with A, amino acid Q at position 315 with A, amino acid F at position 363 with D, and deleting amino acids 1-184. The sequence obtained by replacing TAT at positions 589-591 with GAC, TGT at positions 778-780 with GCG, TTC at positions 907-909 with AAG, TGG at positions 931-933 with AAG, TAC at positions 943-945 with GCC, CAG at positions 946-948 with GCC, and TTT at positions 1090-1092 with GAC, and deleting nucleotides 1-552.

[0073] The amino acid sequence of A3F-CTD_S216A is obtained by replacing the 216th amino acid S with A in the wild-type full-length A3F and deleting amino acids 1-184; the corresponding coding nucleotide sequence is obtained by replacing TCC with GCC in positions 649-651 of the coding nucleotide sequence of the wild-type full-length A3F and deleting nucleotides 1-552.

[0074] The amino acid sequence of A3F-CTD_H247R is obtained by replacing the 247th amino acid H of the above-mentioned wild-type full-length A3F with R and deleting amino acids 1-184; the corresponding coding nucleotide sequence is obtained by replacing CAC with AGG at positions 742-744 of the coding nucleotide sequence of the above-mentioned wild-type full-length A3F and deleting nucleotides 1-552.

[0075] The amino acid sequence of A3F-CTD_W310A is obtained by replacing amino acid W at position 310 of the above wild-type full-length A3F with A and deleting amino acids 1-184; the corresponding coding nucleotide sequence is obtained by replacing TGG at positions 931-933 of the coding nucleotide sequence of the above wild-type full-length A3F with GCG and deleting nucleotides 1-552.

[0076] The amino acid sequence of A3F-CTD_D311K is obtained by replacing amino acid D at position 311 of SEQ ID NO:1 with K and deleting amino acids 1-184; the corresponding coding nucleotide sequence is obtained by replacing GAC at positions 934-936 of the coding nucleotide sequence of the wild-type full-length A3F with AAG and deleting nucleotides 1-552.

[0077] The amino acid sequence of A3F-CTD_T312P is obtained by replacing amino acid T at position 312 of SEQ ID NO:1 with P and deleting amino acids 1-184; the corresponding coding nucleotide sequence is obtained by replacing ACC at positions 937-939 of the coding nucleotide sequence of the wild-type full-length A3F with CCT and deleting nucleotides 1-552.

[0078] The specific gRNA sequences (target sequences) and primer sequences targeting each gene locus involved in the following examples are shown in Tables 1 and 2.

[0079]

[0080]

[0081]

[0082] Example 1: Development of the APOBEC3F-BE4max Editor

[0083] I. Experimental Materials and Methods

[0084] 1. Cell Culture and Transfection

[0085] All cell lines used in this invention were ATCC products. HEK293T and HeLa cells were cultured in DMEM medium supplemented with 10% fetal bovine serum (FBS) under a constant temperature and humidity incubator at 37°C and 5% CO2. For transfection experiments, cells were seeded at an appropriate density in 24-well plates (Corning, USA) and transfected using polyethyleneimine (Polysciences, USA) according to the manufacturer's instructions. A simplified procedure was as follows: 600 ng of base editor (BE) plasmid and 300 ng of sgRNA expression plasmid were mixed with 50 μl of PEI-containing Opti-MEM (Gibco, USA) and transfected for 24 hours. After transfection, cells were cultured for 3 days in fresh medium containing 5 μg / mL puromycin (Merck, USA) for selection. Finally, genomic DNA was extracted using QuickExtract DNA extraction solution (Epicentre, USA). Target genomic regions of interest (200-300 bp) were amplified by PCR for high-throughput DNA sequencing analysis.

[0086] 2. Plasmid construction

[0087] The compact A3F system was synthesized by AZENTA. PCR products were purified by agarose gel electrophoresis, digested with DpnI restriction endonuclease (NEB, USA), and then assembled according to the manufacturer's instructions using the Gibson or Golden Gate method. All gRNA expression plasmids were assembled using the RNF2 sgRNA expression plasmid as a template via the Golden Gate method, with the protospacer sequence embedded in the primers. The specific construction methods for the relevant plasmids are as follows:

[0088] (1) Construction of sgRNA plasmid

[0089] The method for constructing sgRNA according to this invention is as follows: The corresponding gene sequence of human hg38 is obtained from the NCBI genome database. The desired editing location is located using SnapGene software, and the editing site is determined. PAM motifs (NGG) conforming to the requirements of the spCas9 system are screened. Modification primers are then designed. If the 5' end of the target sequence is not G-started, a G base is added after the 5'-CACC-3' of the forward primer, and a C base is added to the reverse primer simultaneously. BsmBI restriction sites (sgRNA primer sequences synthesized by Suzhou Genewiz Biotechnology Co., Ltd.) are introduced at both ends of all primers. The primers are cloned into the backbone of the BsmBI-digested lentiGuide-Puro vector (addgene, #52963) using the Golden Gate assembly method. Subsequently, the forward and reverse primers, T4 DNA ligase buffer, and ddH2O are mixed in proportion, and the mixture is annealed at 95°C for 2 minutes and then naturally cooled to 25°C to form a double-stranded DNA adapter. The annealed product is then combined with the linearized vector, BsmBI restriction enzyme, and T4... DNA ligase, buffer, and BSA were mixed and subjected to 25 cycles of 37°C for 3 minutes and 25°C for 4 minutes to complete a single-step enzyme digestion-ligation reaction. The reaction was terminated by heat treatment at 80°C for 5 minutes. The ligation product was transformed into Trans-TI chemocompetent cells, incubated on ice for 30 minutes, then heat-shocked at 42°C for 45 seconds. After thawing on ice, the cells were plated on LB agar plates containing ampicillin and incubated at 37°C for 12-16 hours. Single colonies were picked, amplified, and sent for Sanger sequencing verification. Correct clones were then extracted using plasmids and preserved. The specific gRNA sequences (target sequences) targeting each gene locus involved in this invention are shown in Tables 1 and 2.

[0090] (2) Construction of editor plasmid

[0091] The following uses the expression plasmid of the base editor heA3F-BE4max as an example to illustrate the construction method of the editor plasmid of the present invention.

[0092] The editor plasmid of this invention is constructed using seamless cloning technology. The lentiCas9-Blast vector (addgene, #52962) is used as a template to replace the Cas9-nuclear localization sequence-flag-P2A-BSD sequence on the original template with the encoding gene (SEQ ID NO:4) of the exogenous fragment base editor heA3F-BE4max (fusion protein), thus obtaining the expression plasmid of the base editor heA3F-BE4max. The specific method is as follows: Specific primers for the target fragment were designed using Diva software, and the primer annealing temperature and fragment length were recorded. Using the lentiCas9-Blast vector (addgene, #52962) and the encoding gene (SEQ ID NO:4) of the exogenous fragment base editor heA3F-BE4max (fusion protein) as templates, a PCR reaction system containing the corresponding primers and DMSO was prepared using PrimeSTAR Max DNA Polymerase for amplification. This yielded the lentiCas9-Blast vector backbone sequence without the Cas9-nuclear localization sequence-flag-P2A-BSD sequence and the encoding gene (SEQ ID NO:4) of the exogenous fragment base editor heA3F-BE4max (fusion protein). The reaction procedure included pre-denaturation at 95°C for 3 minutes, followed by 35 cycles of 95°C for 30 seconds, 65°C for 30 seconds, and 72°C for 60 seconds, with a final extension at 72°C for 5 minutes. The PCR products were separated by 1% TAE agarose gel electrophoresis, and the target band was excised under UV light and purified using a SanPrep column DNA gel extraction kit. DpnI restriction enzyme was added to the gel-extracted product, and the mixture was digested at 37°C for 60 minutes to remove the original template. The product concentration was then measured. Based on the number of inserted fragments, the required amount was calculated according to the principle that the optimal amount per fragment is 0.02 × fragment base pairs. The lentiCas9-Blast vector backbone sequence (excluding the Cas9-nuclear localization sequence-flag-P2A-BSD sequence) and the encoding gene (SEQ ID NO:4) of the exogenous fragment base editor heA3F-BE4max (fusion protein) were mixed with SuperFusion Cloning Mix and reacted at 50℃ for 60 minutes to complete seamless cloning and ligation. Finally, the ligation product was transformed with bacteria, plated on antibiotic plates, single colonies were picked and cultured, and verified by Sanger sequencing to obtain the correct editor plasmid.

[0093] In the final base editor heA3F-BE4max expression plasmid, the structure of the base editor heA3F-BE4max encoding gene from the 5' end to the 3' end is as follows: NLS coding sequence - coding sequence of the CTD variant (W310 / S216T / H198D) of cytosine deaminase APOBEC3F - linker peptide coding sequence - Cas9 (D10A) protein coding sequence - linker peptide coding sequence - uracil glycosylase inhibitor UGI coding sequence - linker peptide coding sequence - uracil glycosylase inhibitor UGI coding sequence - NLS coding sequence.

[0094] The corresponding base editor plasmid was obtained by replacing the coding sequence of the CTD variant (W310 / S216T / H198D) of cytosine deaminase APOBEC3F with other CTD variants of APOBEC3F, wild-type CTD, or the full-length coding gene of APOBEC3F (see above for specific sequences).

[0095] 3. Strains and culture conditions

[0096] *Escherichia coli* Trans5α was used as the cloning host and cultured in lysozyme (LB; containing 1% (w / v) tryptone, 0.5% (w / v) yeast extract and 1% (w / v) NaCl) at 37°C. To screen plasmids, ampicillin (Sigma-Aldrich, USA) was added to the medium to a final concentration of 100 mg / L to identify positive clones.

[0097] 4. High-throughput sequencing and data analysis of genomic DNA samples

[0098] The construction and analysis of the next-generation sequencing library followed the previously described methodology [Yang, L., Huo, Y., Wang, M., Zhang, D., Zhang, T., Wu, H., Rao, X., Meng, H., Yin, S., Mei, J., et al. (2024). Engineering APOBEC3A deaminase for highly accurate and efficient baseediting. Nat Chem Biol 20, 1176-1187.]. A simplified workflow is as follows: purified PCR fragments underwent end repair, 5' phosphorylation, and dA tailing in a single reaction using End PrepEnzyme Mix. Subsequently, a TA ligation reaction was performed to attach adapters to both ends of the fragments. The resulting PCR products were purified, quantified, and sequenced on the Illumina HiSeq platform according to the manufacturer's protocol.

[0099] Amplicon sequencing data were analyzed using batch processing mode of CRISPResso2 (v2.0.45) with window parameters set to -wc 10 -w 20. Base transition frequencies were extracted from the output file “Nucleotide_percentage_summary.txt”, while indel frequencies were obtained from “CRISPRessoBatch_quantification_of_editing_frequency.txt” for base editor experiments. Information on all genomic loci and deep sequencing primers used for sgRNA is provided in Table 1.

[0100] 5. RNA editing analysis

[0101] Transcriptome-wide CU RNA editing analysis was performed using the previously described method [Su, J., Han, C., Zhou, Y., Shan, J., Zhou, X., and Yuan, F. (2024). SaProt: Protein Language Modeling with Structure-aware Vocabulary. bioRxiv, 2023.2010.2001.560349.] with slight modifications. HEK293T cells were transfected with either the base editor construct or the nCas9 (D10A) control editor (with deaminase removed) and cultured for 48 hours. Total RNA was extracted and RNA was sequenced. Raw FASTQ files were processed using Fastp to remove adapter sequences, low-quality bases (Q<20), and unpaired reads were discarded. Cleaned reads were aligned to the human reference genome (hg38, UCSC) using STARv2.7.11b with default parameters. SAM files were sorted and indexed using SAMtools.

[0102] To identify and quantify RNA editing events, we analyzed aligned reads using REDItools2, with the parameter -s 2 specifying strand specificity. Only cytosine on the sense strand and guanine on the antisense strand were considered potential CU editing sites. Sites with coverage <20 or a mapping / reads quality score <30 were excluded to reduce false positives. For each sample, the percentage of editing was calculated by dividing the number of cytosine (or guanine) converted to uracil (or adenine) by the total number of cytosine (or guanine) remaining after filtration. Data are presented as mean ± standard deviation of three independent biological replicates.

[0103] 6. Evolutionary screening of protein language models using APOBEC3F

[0104] The evolutionary screening process consisted of two rounds of analysis. In the first round, we evaluated protein structure and adaptive landscape to screen for potentially efficient mutation sites. Using the overall structure of the AlphaFold2 enzyme, we identified the enzyme-substrate binding sites and amino acid sites in disordered regions. Subsequently, we evaluated the functional scores of a given amino acid sequence under saturation mutations using three models (ESM-1v, ESM-MSA-1b, and transformer-inverse fold (ESM-if1)). Based on the combined results of the first round, we selected the top 10% of mutation sites. In the second round, we filtered using the PSSM score matrix, considering higher scores compared to existing amino acids to obtain potentially efficient mutation sites for A3F. The top variants of A3F were further rationally selected based on their functional domains.

[0105] 7. Statistical Analysis and Repeatability

[0106] Unless otherwise specified, all experimental data are presented as mean ± standard deviation of three biological replicates. Statistical comparisons between the control and experimental groups were performed using Student's t-test or log-rank test (GraphPad Prism 8). A p-value < 0.05 was considered statistically significant.

[0107] II. Results and Analysis

[0108] 1. Structural and rational engineering modifications of APOBEC3F to achieve enhanced cytosine base editing.

[0109] This invention systematically utilizes the wild-type full-length A3F (FL) and its C-terminal structural domain (CTD) to construct a series of A3F-based editors. Based on existing structural data, rational design principles, and reported beneficial mutations [Bohn, MF, Shandilia, SM, Albin, JS, Kouno, T., Anderson, BD, McDougle, RM, Carpenter, MA, Rathore, A., Evans, L., Davis, AN, et al. (2013). Crystal structure of the DNA cytosine deaminase APOBEC3F: the catalytically active and HIV-1 Vif-binding domain. Structure 21, 1042-1050. Chen, Q., Xiao, X., Wolfe, A., and Chen, XS (2016). The in vitro Biochemical Characterization of an HIV-1 Restriction Factor APOBEC3F: Importance of Loop 7 on Both CD1 and CD2 for DNA Binding and Deamination. J Mol Biol 428, 2661-2670. Wan, L., Nagata, T., and Katahira, M. (2018). Influence of the DNA sequence / length and pH on deaminase activity, as well as the roles of the amino acid residuesaround the catalytic center of APOBEC3F. Phys Chem Chem Phys 20, 3109-3117. Wan, L., Kamba, K., Nagata, T., and Katahira, M. (2020).[An insight into the dependence of the deamination rate of human APOBEC3F on the length of single-stranded DNA, which is affected by the concentrations of APOBEC3F and single-stranded DNA. Biochim Biophys Acta Gen Subj 1864, 129346.] We generated 15 candidate A3F variants using molecular cloning technology. These variants were co-transfected with gRNA expression vectors targeting three genomic loci (SPRK3, CTLA, and FANCF loci) in HEK293T cells, and the editing efficiency and indel frequency were quantified by high-throughput sequencing and CRISPResso2 analysis [Clement, K., Rees, H., Canver, MC, Gehrke, JM, Farouni, R., Hsu, JY, Cole, MA, Liu, DR, Joung, JK, Bauer, DE, and Pinello, L. (2019). CRISPResso2 provides accurate and rapid genome editing sequence analysis. NatBiotechnol 37, 224-226.]. Among the candidate variants, several variants (especially those carrying the W310 mutation) showed higher editing efficiency compared to wild-type A3F and its CTD (Figure 1a). Among them, the variant A3F-CTD_W310D-BE4max (hereinafter referred to as A3F-W310D; for the sake of brevity, BE4max will be omitted in the following description, and the same applies to other editors provided in this invention) has editing efficiency comparable to or higher than the current gold standard BE4max and A3A-BE4max, and the editing window is slightly smaller.

[0110] To comprehensively evaluate its broad applicability, we tested A3F-W310D at nine additional genomic sites with different sequence characteristics (VEGFA site 1, PPP1R12C site 1, TET2 site, PPP1R12C site 2, PPP1R12C site 3, AFF1 site, MAGEA1 site 1, DNMT3B site, and HEK4 OT2 site). We found that its editing efficiency was significantly superior to BE4max and A3F-CTD at all targets, with no obvious sequence background bias (Figure 1b). These combined data indicate that A3F-W310D is a highly efficient and broadly applicable CBE, strongly supporting A3F as a potential candidate molecule for further development of high-performance base editors.

[0111] 2. Enhanced CT Editing Based on Computer Simulation of APOBEC3F Engineering Retrofit

[0112] To further improve the editing performance of the A3F-W310D, we employed an innovative computer simulation method to guide its functional optimization. Specifically, we utilized evolutionary scale modeling (ESM), an advanced protein language model (PLM) trained on comprehensive sequence databases such as UniProt and UniRef50, which can accurately predict the relationship between protein sequences and functions. Learning inverse folding from millions of predicted structures. In C.Kamalika, J. Stefanie, S. Le, S. Csaba, N. Gang, and S. Sivan, eds.Proceedings of the 39th International Conference on Machine Learning.PMLR.Rives, A., Meier, J., Sercu, T., Goyal, S., Lin, Z., Liu, J., Guo, D., Ott, M., Zitnick, CL, Ma, J., and Fergus, R. (2021). Biological structure and function emerge from scaling unsupervised learning to 250 million Based on the synergistic analysis of proteinsequences. Proc Natl Acad Sci US A118. Mansoor, S., Baek, M., Juergens, D., Watson, JL, and Baker, D. (2023). Zero-shot mutation effect prediction on protein stability and function using RoseTTAFold. Protein Sci 32, e4780., we performed a comprehensive residue-level mutation scoring of the A3F-CTD sequence (Figure 3a).After obtaining the initial results, the top 10% of high-scoring substitution mutations were further rigorously screened using a position-specific scoring matrix (PSSM), ultimately selecting 10 candidate point mutations that were consistently predicted to significantly enhance function (Figure 2a). These mutations were precisely introduced into the A3F-W310D backbone using site-directed mutagenesis, and their editing activity was evaluated in HEK293T cells at two representative genomic sites (EMX1 and HEK3).

[0113] Experimental results showed that multiple mutation sites (especially H198D, S216T, and G285R) significantly improved CT editing efficiency compared to the W310D variant alone (Figure 2a). To comprehensively evaluate the potential synergistic effects of these mutations, we constructed several combined mutants. After systematic functional comparison, we found that the three-mutant W310D+S216T+H198D (hereinafter referred to as heA3F, i.e., high-efficiency A3F) performed best under all test conditions (Figure 2b).

[0114] The complete amino acid sequence of the heA3F editor is shown in SEQ ID NO:3, where positions 1-19 are the nuclear localization signal NLS (denoted as nuclear localization signal 1); positions 20-208 are heA3F; positions 209-240 are the linker (denoted as linker peptide 1); positions 241-1607 are the Cas9(D10A) protein; positions 1608-1617 are the linker (denoted as linker peptide 2); positions 1618-1700 are the amino acid sequence of the uracil glycosylase inhibitor UGI; positions 1701-1710 are the linker (denoted as linker peptide 3); positions 1711-1793 are the amino acid sequence of the uracil glycosylase inhibitor UGI; and positions 1794-1810 are the nuclear localization signal NLS (denoted as nuclear localization signal 2).

[0115] The encoding nucleotide sequence of the heA3F editor is shown in SEQ ID NO:4, wherein positions 1-57 are the coding sequence for nuclear localization signal 1; positions 58-624 are the coding sequence for heA3F; positions 625-720 are the coding sequence for linker peptide 1; positions 721-4821 are the coding sequence for Cas9(D10A) protein; positions 4822-4851 are the coding sequence for linker peptide 2; positions 4852-5100 are the coding sequence for uracil glycosylase inhibitor UGI; positions 5101-5130 are the coding sequence for linker peptide 3; positions 5131-5379 are the coding sequence for uracil glycosylase inhibitor UGI; and positions 5380-5430 are the coding sequence for nuclear localization signal 2.

[0116] Subsequently, we comprehensively tested heA3F at multiple endogenous sites (DNMT3B, AFF1, PPP1R12C3, PPP1R12C2, PSMB2, VEGFA1, SPPK32, and TEF2), finding that its editing efficiency was significantly higher than A3F-W310D, with a slightly larger editing window (approximately 1-2 nucleotides more), while maintaining no obvious sequence background bias (Figure 2c). These combined data strongly demonstrate that the ESM-based APOBEC3F engineering strategy can successfully develop a more efficient and widely applicable cytosine base editor, providing a new tool option for the field of genome editing.

[0117] 3. Rationally design high-fidelity A3F-oriented CBEs

[0118] High-fidelity base editors with narrow editing windows and low off-target activity are crucial for safe and precise therapeutic genome editing. Previous studies have shown that by carefully modifying the substrate-binding interface of cytidine deaminases, their nucleotide interaction kinetics can be finely regulated, thereby significantly enhancing editing specificity. Based on this scientific principle, we set out to optimize the heA3F system and design a high-fidelity APOBEC3F variant.By integrating reported crystal structure data and detailed homology comparisons of APOBEC3F with other cytidine deaminases (especially APOBEC3A), we accurately identified key residues that may be involved in substrate coordination [Grünewald, J., Zhou, R., Garcia, SP, Iyer, S., Lareau, CA, Aryee, MJ, and Joung, JK (2019). Transcriptome-wide off-target RNAediting induced by CRISPR-guided DNA base editors. Nature 569, 433-437. Yang,L., Huo, Y., Wang, M., Zhang, D., Zhang, T., Wu, H., Rao, X., Meng, H., Yin,S., Mei, J., et al. (2024). Engineering APOBEC3A deaminase for highly accurate and efficient base editing. Nat Chem Biol 20, 1176-1187.]. Gehrke, JM, Cervantes, O., Clement, MK, Wu, Y., Zeng, J., Bauer, DE, Pinello, L., and Joung, JK (2018). An APOBEC3A-Cas9 base editor with minimizedbystander and off-target activities. Nat Biotechnol 36, 977-982. Fang, Y., Xiao, X., Li, SX, Wolfe, A., and Chen, XS (2018). Molecular Interactions of a DNA Modifying Enzyme APOBEC3F Catalytic Domain with a Single-Stranded DNA. J Mol Biol 430, 87-101. Based on these structural biology insights, we rationally designed eight heA3F point mutations, which were predicted to effectively limit the editing window or finely regulate DNA binding affinity (Figure 3a).

[0119] Among these mutations, those obtained through homology inference (such as N240G corresponding to N57G in A3A and Y307G corresponding to Y130G) showed a significant improvement in editing accuracy in experiments (Figure 3a). Interestingly, the Y359A mutation, predicted to disrupt nucleotide binding, actually slightly enhanced editing activity at both test sites (Figure 3a). To comprehensively evaluate the combined effects of these mutations, we constructed several combined mutants. After systematic functional testing, the constructs heA3F-N240G, heA3F-N240G+Y359A, and heA3F-Y307G (all referred to as haA3F, i.e., high-fidelity A3F) exhibited the best editing properties. To validate the broad applicability of these variants, we performed targeted sequencing analysis on eight additional genomic sites (EMX1 site 2, RP11 site, TP53 site, PPP1R12C site 4, EMX1 site 3, PPP1R12C site 5, PDCD1 site, and PPP1R12C site 2). The results showed that these haA3F (high-fidelity A3F) variants, while maintaining robust CT editing activity, can strictly limit the editing window to 5-6 sites (Figure 3b), a characteristic of significant value for therapeutic applications. Furthermore, analysis of editing byproducts indicated that the indel frequencies of heA3F-N240G (haA3F-G) and heA3F-N240G+Y359A (haA3F-GA) were not increased compared to heA3F (Figure 3c). These important findings demonstrate that structure-guided rational design and transdeaminase homology alignment strategies can successfully develop A3F base editors with excellent fidelity, providing a new technological option for precise genome editing.

[0120] 4. Performance comparison between A3F-guided base editors and existing CBEs

[0121] To comprehensively evaluate the performance of our designed A3F-guided base editor, we conducted a series of rigorous control experiments, comparing the developed high-efficiency (heA3F) and high-fidelity (haA3F) variants with a variety of widely used CBEs, including A3A, Anc689, and the advanced haA3A variant. [Yang, L., Huo, Y., Wang, M., Zhang, D., Zhang, T., Wu, H., Rao, X., Meng, H., Yin, S., Mei, J., et al. (2024). Engineering APOBEC3A deaminase for highly accurate and efficient base editing. Nat ChemBiol 20, 1176-1187. Wang, X., Li, J., Wang, Y., Yang, B., Wei, J., Wu, J., Wang, R., Huang, X., Chen, J., and Yang, L. (2018). Efficient base editing inmethylated regions with a human A systematic comparison was performed using APOBEC3A-Cas9 fusion. Nat Biotechnol 36, 946-949. Koblan, LW, Doman, JL, Wilson, C., Levy, JM, Tay, T., Newby, GA, Maianti, JP, Raguram, A., and Liu, DR (2018). Improving cytidine and adenine base editors by expression optimization and ancestral reconstruction. Nat Biotechnol 36, 843-846. In the editing efficiency analysis, HEK293T cells were co-transfected with specified gRNAs and expression vectors encoding different editors, testing multiple genomic sites. The results showed that heA3F and heA3F+Y359A exhibited strong editing efficiency at 9 test sites (EMX1 site 1, ABL1 site, EMX2 site 2, PPP1R12C site 4, PPP1R12C site 5, EMX1 site 3, PDCD1 site, TP53 site, and AFF1 site), with average efficiency superior to the benchmark editor (Figure 4a).

[0122] To further explore the editing specificity, we compared the performance of haA3F-G, haA3F-GA, haA3A-G, and eA3A [Gehrke, JM, Cervantes, O., Clement, MK, Wu, Y., Zeng, J., Bauer, DE, Pinello, L., and Joung, JK (2018). An APOBEC3A-Cas9base editor with minimized bystander and off-target activities. NatBiotechnol 36, 977-982.] at multiple endogenous sites (EMX1 site 2, RP11 site, PP1R12C site 4, PDL1 site 3, PDL1 site 1, PDL1 site 2, TP53 site, and VEGFA site 2). These comparisons revealed site-dependent differences in editing activity: all four constructs consistently exhibited narrower editing windows and enhanced accuracy. Notably, haA3F-GA achieves an optimal balance between editing efficiency and specificity (Figure 4b), highlighting the superior performance of engineered haA3F variants. These important findings establish the A3F editing system as a highly versatile and powerful genome editing tool for both basic research and potential clinical translation.

[0123] The complete amino acid sequence of the haA3F-GA editor is shown in SEQ ID NO:5, where positions 1-19 are the nuclear localization signal NLS (denoted as nuclear localization signal 1); positions 20-208 are haA3F-GA; positions 209-240 are the linker (denoted as linker peptide 1); positions 241-1607 are the Cas9(D10A) protein; positions 1608-1617 are the linker (denoted as linker peptide 2); positions 1618-1700 are the amino acid sequence of the uracil glycosylase inhibitor UGI; positions 1701-1710 are the linker (denoted as linker peptide 3); positions 1711-1793 are the amino acid sequence of the uracil glycosylase inhibitor UGI; and positions 1794-1810 are the nuclear localization signal NLS (denoted as nuclear localization signal 2).

[0124] The nucleotide sequence encoding the haA3F-GA editor is shown in SEQ ID NO:6, wherein positions 1-57 are the coding sequence for nuclear localization signal 1; positions 58-624 are the coding sequence for haA3F-GA; positions 625-720 are the coding sequence for linker peptide 1; positions 721-4821 are the coding sequence for Cas9(D10A) protein; positions 4822-4851 are the coding sequence for linker peptide 2; positions 4852-5100 are the coding sequence for uracil glycosylase inhibitor UGI; positions 5101-5130 are the coding sequence for linker peptide 3; positions 5131-5379 are the coding sequence for uracil glycosylase inhibitor UGI; and positions 5380-5430 are the coding sequence for nuclear localization signal 2.

[0125] 5. Off-target analysis of heA3F and haA3F in HEK293T cells

[0126] Given that CBEs can induce both gRNA-dependent and deaminase-dependent off-target effects, we comprehensively evaluated the off-target effects of heA3F and haA3F variants in HEK293T cells using multiple methods. For gRNA-dependent DNA off-target analysis, we quantified the cumulative CT switching at each known off-target site corresponding to each gRNA. The results showed that the off-target editing level of heA3F was comparable to that of Anc689 and A3A (Figure 5a). Notably, the heA3F variant exhibited comparable or better specificity at these sites (Figure 5a). In the analysis of potential off-target sites at HEK3, both haA3F and haA3A-G variants induced some degree of background editing, but haA3F showed higher specificity than haA3A-G at most sites (Figure 5b). To more comprehensively assess deaminase-dependent off-target activity, we employed an orthogonal R-loop assay at two genomic loci [Doman, JL, Raguram, A., Newby, GA, and Liu, DR (2020). Evaluation and minimization of Cas9-independent off-target DNA editing by cytosine base editors. Nat Biotechnol 38, 620-628.] to test highly efficient variants. The results showed that heA3F exhibited high DNA off-target activity comparable to A3A, while haA3F-GA showed significantly reduced off-target levels (Figure 6a). To further investigate potential RNA off-target effects, we performed whole-transcriptome RNA sequencing on cells expressing base editors. Quantitative analysis of accidental CU switching indicated that heA3F and haA3F variants induced significantly less RNA off-target editing than A3A-BE4max (Figure 6b). Taken together, these results show that although heA3Fs exhibit increased off-target activity relative to haA3A-A (comparable to A3A), haA3Fs maintain high fidelity on both DNA and RNA substrates, supporting their potential as precise and safe genome engineering tools.

[0127] The present invention has been described in detail above. Those skilled in the art will recognize that the invention can be practiced in a wide range of ways with equivalent parameters, concentrations, and conditions without departing from its spirit and scope, and without requiring unnecessary experiments. While specific embodiments have been provided, it should be understood that further modifications can be made to the invention. In summary, according to the principles of the invention, this application is intended to include any changes, uses, or improvements to the invention, including changes made using conventional techniques known in the art that depart from the scope disclosed herein.

Claims

1. A cytosine deaminase variant, characterized by: The cytosine deaminase variant is obtained by modifying the C-terminal domain of the wild-type APOBEC3F protein; the amino acid sequence of the C-terminal domain of the wild-type APOBEC3F protein is shown in SEQ ID NO:1; the cytosine deaminase variant is heA3F; the amino acid sequence of heA3F is obtained by replacing the 127th amino acid with W and the 33rd amino acid with T and the 15th amino acid with H and D in SEQ ID NO:

1.

2. A cytosine deaminase variant, characterized by: The cytosine deaminase variant is obtained by modifying the C-terminal domain of the wild-type APOBEC3F protein; the amino acid sequence of the C-terminal domain of the wild-type APOBEC3F protein is shown in SEQ ID NO:1; the cytosine deaminase variant is haA3F-GA; the amino acid sequence of haA3F-GA is obtained by replacing the amino acid at position 127 (W with D), position 33 (S with T), position 15 (H with D), position 57 (N with G), and position 176 (Y with A) of SEQ ID NO:

1.

3. A base editor, characterized in that: The base editor contains the cytosine deaminase variant of claim 1 or 2 and the nCas9 protein.

4. A base editing system, characterized in that: The base editing system comprises the base editor and guide RNA as described in claim 3.

5. A biomaterial, characterized in that: The biological material is any one of (A1)-(A9) below: (A1) a nucleic acid molecule encoding a cytosine deaminase variant as described in claim 1 or 2; (A2) an expression cassette, recombinant vector, or recombinant cell containing the nucleic acid molecule described in (A1); (A3) a recombinant microorganism containing the nucleic acid molecule described in (A1); (A4) a nucleic acid molecule encoding a base editor as described in claim 3; (A5) an expression cassette, recombinant vector, or recombinant cell containing the nucleic acid molecule described in (A4); (A6) a recombinant microorganism containing the nucleic acid molecule described in (A4); (A7) a nucleic acid molecule encoding a base editing system as described in claim 4; (A8) an expression cassette, recombinant vector, or recombinant cell containing the nucleic acid molecule described in (A7); and (A9) a recombinant microorganism containing the nucleic acid molecule described in (A7).

6. A kit for achieving the conversion of cytosine to thymine in target DNA, characterized in that: The kit contains the cytosine deaminase variant of claim 1 or 2, the base editor of claim 3, the base editing system of claim 4, or the biological material of claim 5.

7. The non-disease therapeutic application of the cytosine deaminase variant of claim 1 or 2, the base editor of claim 3, the base editing system of claim 4, the biomaterial of claim 5, or the kit of claim 6 in achieving the conversion of cytosine to thymine in target DNA.

8. A method for achieving the conversion of cytosine to thymine in target DNA, characterized in that: The method includes the step of contacting the target DNA with the base editing system of claim 4; the method is a non-disease therapeutic method.

Citation Information

Patent Citations

  • Using split deaminases to limit unwanted off-target base editor deamination

    CN111093714A

  • Human modified APOBEC3A-based cytosine base editor

    CN115992122A