Fusion protein, and base editing system for realizing protein multimerization and use thereof

WO2026194220A1PCT designated stage Publication Date: 2026-09-24ZHUHAI SHU TONG MEDICAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/130299
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-20
Filing Date
2025-10-27
Publication Date
2026-09-24

Smart Images

  • Figure CN2025130299_24092026_PF_FP_ABST
    Figure CN2025130299_24092026_PF_FP_ABST
Patent Text Reader

Abstract

Provided are a fusion protein, and a base editing system for realizing protein multimerization and the use thereof. The protein is a Cas9 protein into which an antigen epitope is inserted or in which an antigen epitope replaced. The antigen epitope is an antigen epitope in a protein recruitment system, and can realize the multimerization of the fusion protein. The base editing system comprises the fusion protein and an editing enzyme fusion protein. The editing enzyme fusion protein comprises an antibody containing an antigen epitope and an editing enzyme, and can be used in fields such as basic research, gene therapy, in-vivo animal editing, and agricultural breeding.
Need to check novelty before this filing date? Find Prior Art

Description

A fusion protein and a base editing system for achieving protein polymerization and their applications Technical Field

[0001] This application relates to the field of gene editing technology, specifically to a fusion protein and a base editing system for achieving protein polymerization, and their applications. Background Technology

[0002] With the development of gene editing technology, our understanding of life has entered the genetic level. The discovery and confirmation of various hereditary diseases and gene disorders have fully demonstrated the decisive role of DNA in biological traits. Alterations in DNA sequence, such as base deletions, substitutions, and insertions, can cause phenotypic changes or trigger diseases. As an emerging research field in life sciences, gene editing technology can directly edit specific sequences in DNA, providing a powerful tool for gene function research, gene detection, and therapy, greatly promoting the development of life sciences.

[0003] Currently, gene editing technologies mainly include zinc finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), and clustered regularly interspaced short palindromic repeats (CRISPR) / CRISPR-associated protein 9 (CRISPR / Cas9). Among these, the CRISPR / Cas9 system occupies an increasingly important position in the field of gene editing due to its advantages such as high efficiency, high specificity, and scalability. The Cas9 protein targets a specific DNA sequence via sgRNA and causes a double-strand break in that DNA. Mutations, including addition, deletion, or substitution, are then introduced using the cell's own repair mechanisms. However, double-strand breaks are extremely prone to causing non-target mutations and chromosomal rearrangements and translocations, and it is difficult to efficiently and accurately mutate specific bases. Furthermore, the CRISPR / Cas9 gene editing technology alone cannot correct single-base mutations, making it difficult to meet the gene therapy requirements for single-base mutation genetic diseases.

[0004] To address this challenge, base editing technology has emerged. Base editing technology based on the CRISPR / Cas9 system can achieve precise base editing without cutting double-stranded DNA. In this technology, SpCas9 undergoes an amino acid mutation (D10A) to become the single-strand-cutting nickase nCas9 (D10A). Furthermore, editing enzymes such as deaminases or glycosylation enzymes are needed to replace bases, thus achieving base editing. Current base editing technologies mainly include: adenine base editors (ABE) mediating A-to-G conversions; cytosine base editors (CBE) achieving C-to-T conversions; glycosylation enzyme base editors (GBE or CGBE) achieving C-to-G conversions; and through protein evolution and the introduction of new protein components, novel DNA base editors can also achieve other mutation types, such as ACBE converting A to C. Base editing technology does not introduce double-strand breaks, thus avoiding non-target mutations and chromosomal rearrangements and translocations, exhibiting extremely high safety and becoming an indispensable tool in basic research and therapeutic applications.

[0005] However, existing base editors still have the following shortcomings: First, they have fixed editing windows, which means that some bases outside the editing window cannot be edited; or when multiple identical bases exist within the window, they will be edited simultaneously, resulting in unnecessary mutations, also known as bystander editing. Second, each base editor has its own sequence preference, which means that target bases in certain sequence backgrounds cannot be edited or are edited inefficiently, or non-target bases are edited efficiently because they are located in preferred sequences. Third, in addition to producing the target editing product, base editors also produce unnecessary byproducts and indels (insertions or deletions), which seriously hinders their clinical application. Fourth, some base editors also suffer from low editing efficiency, which further hinders their widespread application. Fifth, the development of novel base editors mainly relies on protein element replacement, amino acid mutation, or the introduction of new protein elements. Replacing SpCas9 with other Cas proteins such as SaCas9 or iscB can enable the recognition of other PAM sites. Replacing the editing enzyme with an enzyme from another species, such as replacing the rat APOBEC1 of BE3 with the lamprey pmCDA1, yields CBE that does not reject GC sequence background. Adding auxiliary elements such as Rad51 or eUNG to existing base editors can improve editing efficiency or expand the editing window. Amino acid mutations can be performed on existing editing enzymes, such as mutating ABE8e to ABE9. Therefore, the development of existing base editing technologies heavily relies on protein optimization or the introduction of new protein elements. Sixth, in base editing technology, nCas9(D10A), the editing enzyme, and other protein elements are arranged in a specific sequence and fused together to form a single protein. During the optimization of most base editors, when the editing enzyme is fused to the N-terminus, C-terminus, or inserted into a specific position within nCas9(D10A), different editing characteristics will be exhibited, including editing efficiency, editing window, sequence bias, byproducts, and indels. This means that the interaction between the editing enzyme and DNA changes depending on the spatial location of the enzyme, thus altering the overall editing characteristics. However, editing enzymes are large, have complex conformations, and are diverse. Inserting them into nCas9(D10A) requires rigorous engineering design and validation to ensure the stability and efficiency of base editing. This results in fewer selectable insertion sites and further limits the spatial locations where the editing enzyme can exist. Seventh, protein polymerization can increase local protein concentration and promote catalytic reactions. Current base editing tools are all in monomeric form. If base editing tools are polymerized, the protein concentration at the target site can be increased, thereby greatly improving base editing efficiency. The conformation formed after polymerization is an important factor affecting the catalytic reaction; simply polymerizing the base editing tool cannot efficiently improve base editing efficiency.Therefore, it is necessary to systematically explore the editing efficiency of base editors under different polymerization conditions and find the optimal polymerization conditions.

[0006] The Suntag system is a protein recruitment system containing the antigenic epitope GCN4 and its antibody AntiGCN4 (scFv). The editing enzyme APOBEC1 of BE3 is fused to the C-terminus of AntiGCN4, and GCN4 is fused to the N-terminus or C-terminus of nCas9 (D10A), enabling nCas9 (D10A) to recruit APOBEC1 to the N-terminus or C-terminus, thus establishing the base editing technology BE-PLUS, which separates nCas9 (D10A) and APOBEC1. Compared to BE3, BE-PLUS exhibits higher editing efficiency, fewer non-specific products, and a significantly expanded editing window. Similarly, the Suntag system can be combined with CGBE to alter existing editing characteristics, such as reducing non-specific products and shrinking the editing window. Therefore, the Suntag system is highly flexible, can be combined with different types of deaminases, and exhibits different editing characterization patterns. However, placing GCN4 only at the N-terminus or C-terminus of nCas9(D10A) can only generate two module combination methods and two polymerized structures, meaning only two editing characterization modes. This significantly limits the application value of protein recruitment systems in the field of base editing. In addition, there are many other similar protein recruitment systems, such as the Moontag system. Currently, none of these systems have been used in base editing technology.

[0007] In summary, altering the spatial relationship between editing enzymes and DNA, as well as the polymerization structure of editing tools, affects the editing results. In other words, suitable spatial locations and polymerization structures can optimize the characteristics of base editors, such as window size, efficiency, sequence bias, and byproducts. However, in base editors, the spatial location of editing enzymes is fixed, and there is currently no technology to control the polymerization structure of base editing tools, severely limiting their flexibility and reducing their effectiveness in processing specific DNA sequences. Summary of the Invention

[0008] The purpose of this application is to overcome the shortcomings of the prior art and provide a fusion protein and a base editing system for achieving protein polymerization, as well as their applications.

[0009] To achieve the above objectives, the technical solution adopted in this application is as follows:

[0010] In a first aspect, this application provides a fusion protein, wherein the fusion protein is a Cas9 protein with an internally embedded or replaced antigenic epitope; the antigenic epitope is an antigenic epitope in a protein recruitment system; and the antigenic epitope can realize the polymerization of the fusion protein.

[0011] The fusion protein of this application can change the spatial position and polymerization structure of the editing enzyme to achieve functions such as changing the editing window, improving editing efficiency, adjusting sequence preference, reducing byproducts and indels, and achieving specific single-base editing.

[0012] As a preferred embodiment of the fusion protein described in this application, the Cas9 protein is any one of nCas9 (D10A) protein, FrCas9 (E796A) protein, and LbCas12 (dead) protein.

[0013] As a preferred embodiment of the fusion protein described in this application, the protein recruitment system includes any one of SunTag, MoonTag, ALFA, Flag, and HA.

[0014] In a preferred embodiment of the fusion protein described in this application, the antigenic epitope includes any one of GCN4, gp41, and Alfa.

[0015] As a preferred embodiment of the fusion protein described in this application, the amino acid sequence of the antigenic epitope is shown in any one of SEQ ID No. 21, 81, and 85.

[0016] In a further preferred embodiment of the fusion protein described in this application, the antigenic epitopes are embedded in the nCas9(D10A) protein at positions 1, 25, 54, 62, 170, 203, 213, 231, 249, 257, 308, 355, 365, 400, 532, 573, 584, 674, 719, 765, 768, 776, 782, 790, 808, 819, 826, 831, 846, 868, 890, 910, 924, 945, and 975. The antigenic epitope is located after at least one of the following amino acids: 1010, 1020, 1033, 1050, 1051, 1055, 1059, 1068, 1072, 1102, 1110, 1120, 1130, 1227, 1246, 1248, 1252, 1260, 1276, 1290, 1300, 1302, 1327, 1332, 1340, and 1368, or the antigenic epitope replaces amino acids 1048-1063 of the nCas9(D10A).

[0017] As a further preferred embodiment of the fusion protein described in this application, the antigenic epitopes are embedded in the 1st, 40th, 51st, 55th, 108th, 112th, 120th, 125th, 170th, 192nd, 200th, 215th, 234th, 250th, 265th, 270th, 275th, 290th, 295th, 315th, 320th, 325th, 363rd, 379th, 428th, and 539th epitopes of the FrCas9(E796A) protein. The following at least one amino acid from the following amino acids: 566, 610, 618, 690, 702, 707, 727, 755, 761, 800, 830, 882, 980, 1002, 1043, 1055, 1095, 1111, 1117, 1134, 1184, 1188, 1224, 1248, 1321, 1329, 1333, 1347, and 1372.

[0018] As a further preferred embodiment of the fusion protein described in this application, the antigenic epitope is embedded after at least one amino acid selected from the following amino acids: 1, 85, 212, 270, 372, 406, 441, 478, 571, 625, 655, 713, 730, 760, 770, 774, 780, 807, 825, 847, 865, 965, 985, 1010, 1040, 1055, 1074, 1079, 1087, 1109, 1120, 1121, 1142, 1143, 1158, 1171, 1201, 1212, and 1228 of the LbCas12(dead) protein.

[0019] Secondly, this application provides a nucleic acid molecule encoding the fusion protein described above.

[0020] As a preferred embodiment of the nucleic acid molecule described in this application, the nucleotide sequence of the antigenic epitope in the fusion protein is shown in any one of SEQ ID NO: 22, 82, and 86.

[0021] Thirdly, this application provides a plasmid comprising the aforementioned nucleic acid molecule.

[0022] Fourthly, this application provides a base editor system, comprising the aforementioned fusion protein and editing enzyme fusion protein; the editing enzyme fusion protein comprises an antibody against the antigenic epitope and an editing enzyme.

[0023] The base editor system of this application has the following advantages:

[0024] (1) The plug-in base editing technology (Oligomerization Enhances Base Editing, OEBE) of this application overcomes the limitations of editing capabilities.

[0025] Current base editors are limited by the spatial location and polymerization structure of editing enzymes, which affects editing characteristics such as editing window, editing efficiency, sequence bias, and nonspecific products. This application presents a modular, polymerizable, programmable, and scalable base editor system, OEBE, which can dynamically change the spatial location, polymerization structure, and reprogram the editing characteristics of the editing enzyme, overcoming these limitations. By embedding epitopes at different locations within the Cas protein and fusing the editing enzyme with an antibody, the spatial location of the editing enzyme and the polymerization structure of the base editing tool can be manipulated to optimize their interaction with DNA, thereby improving editing efficiency, window limitations, and nonspecific products. The OEBE system is a base editor based on protein spatial location and protein polymerization design strategies. Its editing results are significantly influenced by the spatial location of the editing enzyme and protein polymerization, a feature not present in previous base editors. OEBE unlocks the potential of traditional base editors, improving editing efficiency, accuracy, window limitations, and safety.

[0026] (2) OEBE adjusts the spatial location, polymerization structure and reprograms the editing characterization of editing enzymes.

[0027] This application's OEBE utilizes numerous embedding sites for antigenic epitopes within nCas9(D10A), allowing the editing enzyme to be fixed at different spatial locations within nCas9(D10A) like a plug-in. Simultaneously, the polymerization ability of the antigenic epitopes promotes the polymerization of nCas9(D10A), thereby altering the editing characterization and significantly increasing the number of module combinations, the number of protein polymerization structures, and the diversity of base results. Furthermore, OEBE can strategically adjust the protein recruitment system and the arrangement of internal components to dynamically modify the spatial interaction between the editing enzyme and DNA, the protein polymerization structure, and generate more editing patterns. In HPV18 oncogene editing and zebrafish embryo editing, efficient editing of specific bases can be achieved through appropriate OEBE component combinations. In contrast, traditional base editing technologies use fixed components that cannot polymerize, resulting in some bases being uneditable.

[0028] (3) OEBE has superior editing capabilities

[0029] Under specific module combinations, OEBE can improve efficiency by 1-3 times and significantly reduce byproducts, or it can precisely edit specific bases—features not found in traditional base editors. Therefore, compared to traditional base editors, OEBE significantly optimizes window limitations, editing efficiency, editing precision, sequence bias, byproducts, and indels. The OEBE system has demonstrated editing results in HPV18 oncogene therapy and zebrafish embryo editing that are unattainable by traditional base editing techniques, further demonstrating its superior editing capabilities.

[0030] (4) OEBE can serve as a powerful and extensible platform for generating a series of novel base editors.

[0031] The OE system uses CRISPR / Cas as the insertion platform, multiple embedded antigenic epitopes on the Cas protein as slots, and protein recruitment systems and editing enzymes as plug-ins. Adjusting the position of the antigenic epitopes alters the spatial location of the editing enzyme and changes the final protein polymerization structure, thus generating a series of novel base editors without protein evolution. The OEBE system utilizes the numerous insertion sites within nCas9 (D10A), giving it strong scalability. By being compatible with different protein recruitment systems (such as SunTag, MoonTag, and ALFA) and other protein elements such as editing enzymes (such as APOBEC1, pmCDA1, TadA8eWQ, TadA9, and R33A), the number of modules and combinations can be increased exponentially. Through adjusting the combinations, OEBE generates a large number of novel base editors with diverse editing capabilities. Therefore, OEBE is a modular, programmable, and scalable base editing platform. As the CRISPR platform, slots, and plug-ins are added, it can continuously expand and generate new base editors, bringing more possibilities and choices to the field of gene editing.

[0032] (5) OEBE allows for customized editing of results.

[0033] In the OEBE system, a large number of novel base editing tools with diverse characteristics can be obtained by using different module combinations. For DNA regions that are difficult to edit efficiently by traditional base editors or prone to bystander editing, OEBE can customize the editing window and sequence preferences of the editing results through specific module combinations, thereby achieving efficient or precise editing of that DNA region. OEBE technology can provide a more refined, efficient, and diverse base editing technology system, rather than just a single tool. Appropriate OEBE module combinations can be selected according to specific needs to customize base editing tools that meet specific editing requirements. This is particularly evident in HPV18 oncogene editing and zebrafish embryo editing, where specific base editing requirements were achieved by appropriately adjusting the component combinations of the OEBE system.

[0034] In a preferred embodiment of the base editor system described in this application, the antibody includes any one of AntiGCN4, Antigp41, and AntiALFA; the amino acid sequence of AntiGCN4 is shown in SEQ ID NO: 23; the amino acid sequence of Antigp41 is shown in SEQ ID NO: 83; the amino acid sequence of AntiALFA is shown in SEQ ID NO: 87; and the editing enzyme includes at least one of APOBEC1, pmCDA1, TadA8eWQ, TadA9, and R33A.

[0035] As a preferred embodiment of the base editor system described in this application, the nucleotide sequence of the antibody is shown in any one of SEQ ID NO: 24, 84, 88.

[0036] In a preferred embodiment of the base editor system described in this application, the functional elements of the editing enzyme fusion protein are sequentially connected in any of the following ways (where "sequential" can be, for example, from the N-terminus to the C-terminus; "sequential connection" means that the functional elements are continuous and there is no inserted sequence between them; or "sequential connection" means that the functional elements are partially or completely discontinuous and there is an inserted sequence between them):

[0037] i. Promoter, antibody, editing enzyme, UGI, GB1, NLS, BGH;

[0038] ii. Promoter, antibody, editing enzyme, GB1, NLS, BGH;

[0039] iii. Promoter, the editing enzyme, the antibody, GB1, NLS, BGH;

[0040] iv. Promoter, antibody, editing enzyme, eUNG, GB1, NLS, BGH;

[0041] v, promoter, antibody, eUNG, editing enzyme, GB1, NLS, BGH.

[0042] As a further preferred embodiment of the base editor system described in this application, the functional elements of the editing enzyme fusion protein are sequentially connected in any of the following ways (where "sequential" can be, for example, from the N-terminus to the C-terminus; "sequential connection" means that the functional elements are continuous and there is no inserted sequence between them; or "sequential connection" means that the functional elements are partially or completely discontinuous and there is an inserted sequence between them):

[0043] i. Promoter, AntiGCN4, APOBEC1, UGI, GB1, NLS, BGH;

[0044] ii. Promoter, AntiGCN4, TadA8eWQ, GB1, NLS, BGH;

[0045] iii. Promoter, AntiGCN4, R33A, GB1, NLS, BGH;

[0046] iv. Promoter, Antigp41, APOBEC1, UGI, GB1, NLS, BGH;

[0047] v, promoter, Antigp41, TadA8eWQ, GB1, NLS, BGH;

[0048] vi, promoter, Antigp41, R33A, GB1, NLS, BGH;

[0049] vii, the promoter, the AntiALFA, the APOBEC1, UGI, GB1, NLS, BGH;

[0050] viii, promoter, AntiALFA, TadA8eWQ, GB1, NLS, BGH;

[0051] ix, promoter, AntiALFA, R33A, GB1, NLS, BGH;

[0052] x, the promoter, AntiGCN4, pmCDA1, UGI, GB1, NLS, BGH;

[0053] xi, promoter, Antigp41, pmCDA1, UGI, GB1, NLS, BGH;

[0054] xii, promoter, TadA8eWQ, AntiGCN4, GB1, NLS, BGH;

[0055] xiii, promoter, TadA8Ewq, Antigp41, GB1, NLS, BGH;

[0056] xiv, promoter, AntiGCN4, TadA9, GB1, NLS, BGH;

[0057] xv, promoter, TadA9, AntiGCN4, GB1, NLS, BGH;

[0058] xvi, the promoter, the Antigp41, the TadA9, GB1, NLS, BGH;

[0059] xvii, promoter, TadA9, Antigp41, GB1, NLS, BGH;

[0060] xviii, promoter, AntiGCN4, R33A, eUNG, GB1, NLS, BGH;

[0061] xix, promoter, AntiGCN4, eUNG, R33A, GB1, NLS, BGH;

[0062] xx, promoter, Antigp41, R33A, eUNG, GB1, NLS, BGH;

[0063] xxi, promoter, Antigp41, eUNG, R33A, GB1, NLS, BGH.

[0064] As a preferred embodiment of the base editor system described in this application, it further includes sgRNA; the sgRNA includes a target site recognition sequence.

[0065] Fifthly, this application provides a cell including the aforementioned base editor system.

[0066] Sixthly, this application provides the application of the fusion protein, the nucleic acid molecule, the plasmid, the base editor system, and the cell in gene editing.

[0067] This application also provides a gene editing method, comprising the following steps: introducing at least one sgRNA and the base editor system together into animal cells to achieve gene editing. Preferably, the introduction method includes plasmid transfection. The animal cells include cells from pigs, cattle, sheep, horses, monkeys, fish (preferably zebrafish), mice, birds, or humans. Preferably, the animal cells are zebrafish 1-cell stage embryos, HeLa cells, or human embryonic kidney 293T cells. The sgRNA and the base editor system are introduced into the zebrafish 1-cell stage embryo via microinjection.

[0068] As a preferred embodiment of the application described in this application, the gene editing includes at least one of editing the base C of the target site sequence to the base T, editing the base C of the target site sequence to the base G, and editing the base A of the target site sequence to the base G.

[0069] Seventhly, this application provides the use of the fusion protein, the nucleic acid molecule, the plasmid, the base editor system, and the cell in the preparation of drugs for tumor gene therapy.

[0070] This application further provides a gene therapy drug for treating tumors, comprising the fusion protein, the nucleic acid molecule, the plasmid, the base editor system, or the cell, and a pharmaceutically acceptable vector.

[0071] Eighthly, this application provides the application of the fusion protein, the nucleic acid molecule, the plasmid, the base editor system, and the cell in agricultural breeding.

[0072] Compared with the prior art, the beneficial effects of this application are as follows:

[0073] (1) By changing the spatial position of the editing enzyme and the protein polymerization structure, the editing window can be changed, including enlargement, shrinking, and movement, so as to cover DNA regions that the original base editor cannot cover; the sequence preference can be changed, including improving the editing efficiency of certain sequences or reducing the editing efficiency of certain sequences, thereby reducing unnecessary base editing within the editing window; by reducing byproducts and indels to different degrees, greatly improving the purity of the editing product; and greatly improving the editing efficiency.

[0074] (2) In the OEBE system, the spatial location of the editing enzyme is programmable, thereby enabling the reprogramming of the original base editor editing characterization, overcoming the limitation of the spatial location of the base editor editing enzyme, and better optimizing the editing performance.

[0075] (3) The base editor system of this application has the characteristics of a modular, polymerizable, programmable, and scalable base editing platform. It can continuously expand with the increase of CRISPR platforms, slots, and plug-ins, constantly generating new base editors and bringing more possibilities and choices to the field of gene editing. By changing the spatial position of editing enzymes and the protein polymerization structure, a series of novel modular base editors with different editing characteristics are generated. OEBE does not require protein evolution and is compatible with multiple protein recruitment systems and multiple editing enzymes, greatly expanding the tool pool of base editing technology. It can provide more refined, efficient, and diversified modular combination methods, bringing more choices and possibilities to the field of base editing.

[0076] (4) The base editor system of this application can be applied to gene editing, including but not limited to applications in basic research, gene therapy (such as tumor or cancer gene therapy), animal live editing (such as zebrafish embryo editing), agricultural breeding, and other fields, as well as for different gene loci and editing needs. Attached Figure Description

[0077] Figure 1 shows a comparison between existing base editors and OEBE; in Figure 1, (A) the editing enzyme (e.g., deaminase) of current (conventional) base editors is usually located at the N-terminus or C-terminus of nCas9 (D10A); (B) in the OEBE system, the antigenic epitope GCN4 is embedded inside nCas9 (D10A), while the antibody is fused with the editing enzyme, thereby enabling the editing enzyme to be located at a specific spatial position and adjust the protein dimerization structure; (C) in the OEBE system, the antigenic epitope gp41 is embedded inside nCas9 (D10A), while the antibody is fused with the editing enzyme, thereby enabling the editing enzyme to be located at a specific spatial position and adjust the protein trimerization structure.

[0078] Figure 2 is a schematic diagram of the insertion sites of nCas9(D10A). In Figure 2, (A) shows the 62 insertion sites inside nCas9(D10A), that is, the antigenic epitope is inserted after the corresponding amino acid. β-sheets and α-helices are represented by rectangles of different gray depths, and irregular curls are represented by straight lines. Among them, 1048-1063 indicates that amino acids 1048 to 1063 are removed and replaced with the antigenic epitope. (B) is an example of inserting the antigenic epitope GCN4 into nCas9(D10A), such as D10A-231GCN4, which means that GCN4 is inserted after the 231st amino acid. Other D10A-GCN4 variants are similar, and the D10A-gp41 and D10A-ALFA variants mentioned later also follow the same naming rules.

[0079] Figure 3 shows a schematic diagram of BE3 and SunTag OEBE3.

[0080] Figure 4 shows the editing efficiency of the SunTag OEBE3 system and BE3 at four target points; the Y-axis of the heatmap represents variants, such as 532 representing D10A-532GCN4.

[0081] Figure 5 shows a comparison of the window diagrams of the SunTag OEBE3 system and BE3; the numbers in the legend represent variants, such as 532 representing D10A-532GCN4.

[0082] Figure 6 shows a comparison of the edited products of SunTag OEBE3 and BE3; the numbers on the X-axis represent variants, such as 532 representing D10A-532GCN4.

[0083] Figure 7 shows a schematic diagram of ABE8eWQ and SunTag OEABE8eWQ.

[0084] Figure 8 shows a comparison of the editing characterization of ABE8eWQ and SunTag OEABE8eWQ; in Figure 8, (A) comparison of editing window and editing efficiency; (B) comparison of editing efficiency and by-products; (C) comparison of single-base editing products.

[0085] Figure 9 shows a schematic diagram of miniCGBE1 and SunTag OEminiCGBE1.

[0086] Figure 10 shows a comparison of the editing characteristics of miniCGBE1 and SunTag OEminiCGBE1; in Figure 10, (A) is a comparison of editing window and editing efficiency; (B) is a comparison of editing product and purity.

[0087] Figure 11 is a schematic diagram of MoonTag OEBE.

[0088] Figure 12 shows a comparison of the editing characterization of BE3 and MoonTag OEBE3; in Figure 12, (A) is a comparison of editing window and editing efficiency; (B) is a comparison of editing product and purity.

[0089] Figure 13 shows a comparison of the editing characteristics of ABE8eWQ and MoonTag OEABE8eWQg; in Figure 13, (A) is a comparison of editing window and editing efficiency; (B) is a comparison of editing product and purity.

[0090] Figure 14 shows a comparison of the editing characteristics of miniCGBE1 and MoonTag OEminiCGBE1; in Figure 14, (A) is a comparison of editing window and editing efficiency; (B) is a comparison of editing product and purity.

[0091] Figure 15 shows a schematic diagram of the SunTag OETarget AID system and the MoonTag OETarget AID system.

[0092] Figure 16 shows the editing characteristics of SunTag OETarget AID and MoonTag OETarget AID; in Figure 16, (A) the editing window of SunTag OETarget AID; (B) the editing window of SunTag OETarget AID; (C) the editing output analysis of SunTag OETarget AID and MoonTag OETarget AID.

[0093] Figure 17 is a schematic diagram of ALFA OEBE.

[0094] Figure 18 shows a comparison of the editing characterization of BE3 and ALFA OEBE3; in Figure 18, (A) comparison of editing efficiency at the four sites; (B) comparison of editing window and editing efficiency; (C) comparison of editing product and purity.

[0095] Figure 19 shows a comparison of the editing characterization of ABE8eWQ and ALFA OEABE8eWQ; in Figure 19, (A) the editing efficiency of the four sites is compared; (B) the editing window and editing efficiency are compared; (C) the editing products and purity are compared; and (D) the proportion of single-base editing products is compared.

[0096] Figure 20 shows a comparison of the editing characterization of miniCGBE1 and ALFA OEminiCGBE1; in Figure 20, (A) comparison of editing efficiency at four sites; (B) comparison of editing products and purity.

[0097] Figure 21 shows a comparison of editing characterization after changing the order of antibodies and editing enzymes within the OEABE8eWQ framework; in Figure 21, (A) is a schematic diagram of the OEABE8eWQ framework order adjustment; (B) is a comparison of editing window and editing efficiency; (C) is a comparison of editing products and purity; (D) is a comparison of the proportion of single-base products; 532gp41-ABE8eWQ indicates the use of the D10A-532gp41 variant and Antigp41-ABE8eWQ; ABE8eWQ-532GCN4 indicates the use of the D10A-532GCN4 variant and ABE8eWQ-AntiGCN4; other combinations follow the same naming rules.

[0098] Figure 22 shows a comparison of editing characterization after changing the order of antibodies and editing enzymes within the OEABE9 framework; in Figure 22, (A) is a schematic diagram of OEABE9 framework order adjustment; (B) is a comparison of editing window and editing efficiency; (C) is a comparison of editing products and purity; and (D) is a comparison of the proportion of single-base products. 532gp41-ABE9 indicates the use of the D10A-532gp41 variant and Antigp41-ABE9; ABE9-532GCN4 indicates the use of the D10A-532GCN4 variant and ABE9-AntiGCN4; other combinations follow the same naming rules.

[0099] Figure 23 shows a comparison of editing characterization after changing the order of antibodies and editing enzymes within the OECGBE1 framework; in Figure 23, (A) is a schematic diagram of OEminiCGBE1 framework order adjustment; (B) is a comparison of editing efficiency at four sites; (C) is a comparison of editing windows; (D) is a comparison of editing products and purity, as well as a comparison of the proportion of single-base products; 1246gp41-R33A-eUNG indicates the use of the D10A-1246gp41 variant and Antigp41-R33A-eUNG; other combinations follow the same naming rules.

[0100] Figure 24 shows OEBE systems based on other Cas proteins; in Figure 24, (A) the SunTag OEABE8eWQ system based on FrCas9 (E796A) (top) can achieve a base A mutation to G at the target site (bottom); (B) the SunTag OEBE3 system based on LbCas12 (dead) (top) can achieve a base C mutation to T at the target site (bottom).

[0101] Figure 25 shows the application of the OEBE system in HPV18 tumor gene editing; in Figure 25, (A) the efficiency of three base editing technologies in editing the target base C to T at the E6 and E7 target sites; where 1246gp41-BE3 represents D10A-1246gp41 and AntiGCN4-BE3; (B) the growth rate of HeLa cells treated with the three base editing technologies.

[0102] Figure 26 shows the performance of the OEBE system in zebrafish embryo editing; in Figure 26, (A) is a schematic diagram of the mRNA structure used for zebrafish embryo injection; an additional P2A-fluorescent protein element was added to the C-terminus of the protein to visualize protein expression; (B) is the visualization of fluorescent embryos for subsequent DNA sequencing; (C) the editing results and phenotype of the TWIST2 site in zebrafish; 231gp41-ABE8eWQ represents the combination of D10A-231gp41 and Antigp41-ABE8eWQ. (D) Analysis of the editing products of ABE8eWQ and 231gp41-ABE8eWQ at the TWIST2 site by amplicon sequencing; the samples included the five embryos with the highest efficiency calculated using Sanger sequencing data; (E) Editing results and phenotypes of the TYR site in zebrafish; (F) Analysis of the editing products of BE3 and 1246gp41-BE3 at the TYR site by amplicon sequencing; the samples included the five embryos with the highest efficiency calculated using Sanger sequencing data.

[0103] Figure 27 shows the spectrum of plasmid P11694.

[0104] Figure 28 shows the spectrum of plasmid 113029.

[0105] Figure 29 shows the spectrum of plasmid 113022. Detailed Implementation

[0106] This application's plug-in base editing technology (Oligomerization Enhances Base Editing, OEBE) embeds antigenic epitopes of a protein recruitment system into different positions within nCas9 (D10A), while the editing enzyme fuses with an antibody of the protein recruitment system. This allows the editing enzyme to exist at a specific spatial location through protein recruitment. Because the antigenic epitopes GCN4 and gp41 have polymerizing capabilities, the base editing tool can be polymerized. Most importantly, the spatial localization of the editing enzyme and the polymerized structure of the protein can be adjusted by changing the embedding position of the antigenic epitopes or the fusion order of the antibody and the editing enzyme. To better illustrate the purpose, technical solution, and advantages of this application, a detailed study and demonstration of the OEBE system's editing characterization mode, modularity, polymerization, and scalability are presented. The following will further illustrate this application with specific embodiments. Those skilled in the art should understand that the specific embodiments described herein are only for explaining this application and are not intended to limit this application.

[0107] The technical principle of OEBE is shown in Figure 1. It mainly involves three editing methods: (1) editing the base C of the target site sequence to the base T; (2) editing the base C of the target site sequence to the base G; (3) editing the base A of the target site sequence to the base G.

[0108] Unless otherwise specified, the experimental methods used in the examples are conventional methods; the materials and reagents used are commercially available unless otherwise specified. The primers used in the examples were synthesized by Suzhou Genewiz Biotechnology Co., Ltd.; the PCR reagents used were 2× [reagent name missing] from Beijing TransGen Biotech Co., Ltd. PCR SuperMix, catalog number AS111-02.

[0109] To illustrate the OEBE system of this application in detail, the following embodiments include base editing for genomic sites in Table 1. Those skilled in the art can design base editing schemes for other gene sites based on the following embodiments.

[0110] Table 1 shows the target sites, sequence information, and editing methods involved in the embodiments.

[0111] Example 1

[0112] 1. Construct sgRNA plasmids targeting the target site

[0113] The components of the sgRNA plasmids include: ampicillin resistance gene, ori replicon, U6 promoter, spacer sequence, and scaffold backbone sequence. The template plasmid was purchased from the Miaoling plasmid platform, catalog number P11694 (Figure 27). Spacer sequences from the four sites HEK3, FANCF, ZAP70, and VEGFA were used to construct sgRNA plasmids. DNA fragments were obtained by PCR amplification. Primer 1 contained the 20nt sequence at the 3' end of the U6 promoter, the 20nt sequence of the spacer, and the 20nt sequence at the 5' end of the scaffold. Primer 2 was the reverse complementary sequence of the 20nt sequence at the 3' end of the U6 promoter. Primers were synthesized by Genewiz and then cloned into the plasmid vector using the Gibson assembly method. The constructed plasmids were named HEK3-sgRNA, FANCF-sgRNA, ZAP70-sgRNA, and VEGFA-sgRNA. After construction, the correct sgRNA plasmid sequence was confirmed to be free of mutations by routine sequencing comparison. Single colonies with completely correct sequences were selected for amplification and plasmid extraction.

[0114] The target site Spacer sequence is shown in Table 1.

[0115] Scaffold sequence (SEQ ID NO: 18):

[0116] Primer 1 (SEQ ID NO: 19):

[0117] GTGGAAAGGACGAAACACCG (Spacer 20 nt sequence)GTTTTAGAGCTAGAAATAGC;

[0118] Primer 2 (SEQ ID NO: 20): CGGTGTTTCGTCCTTTCCAC.

[0119] 2. Positive control plasmid

[0120] The BE3 plasmid was purchased from the Miaoling plasmid platform, catalog number P1236. The BE3 plasmid (Figure 3) contains an nCas9 (D10A) sequence with DNA single-strand cutting activity, an APOBEC1 editing enzyme sequence that enables the editing of base C to T, and a UGI sequence to improve editing efficiency.

[0121] 3. Construct the Suntag OEBE3 plasmid

[0122] The AntiGCN4, APOBEC1, and UGI sequences were derived from plasmid pST1374-scFv-APOBEC-UGI-GB1, purchased from Addgene (catalog number 113029) (Figure 28). The GCN4, nCas9(D10A), and nuclear localization signal peptide sequences were derived from plasmid pST1374-GCN4-D10A, purchased from Addgene (catalog number 113022) (Figure 29). The antigenic epitope GCN4 of the Suntag protein recruitment system was embedded at different positions within nCas(D10A) and named D10A-GCN4, as shown in Figure 2. The antibody AntiGCN4 was fused with APOBEC1 and named AntiGCN4-BE3. The two were combined and named SunTag OEBE3, as shown in Figure 3. After construction, the plasmid sequences were confirmed to be correct and mutation-free by routine sequencing alignment. Single colonies with completely correct sequences were selected for amplification and plasmid extraction.

[0123] The sequence information for the SunTag OEBE3 system is as follows:

[0124] GCN4 amino acid sequence (SEQ ID NO: 21): EELLSKNYHLENEVARLKK;

[0125] GCN4 base sequence (SEQ ID NO: 22):

[0126] AntiGCN4 amino acid sequence (SEQ ID NO: 23):

[0127] AntiGCN4 base sequence (SEQ ID NO: 24):

[0128] As shown in Figure 2, the embedding sites of the antigenic epitopes within nCas9(D10A) are indicated by numbers, with those located on the α-helix or β-sheet marked after the number:

[0129] 1, 25, 54, 62 (α-spiral), 170, 203, 213, 231, 249, 257, 308, 355, 365, 400, 532, 573, 584, 674, 719, 765, 768, 776, 782, 790 (α-spiral), 808 (α-spiral), 819, 826, 831, 846, 868, 890, 910, 924 (α-spiral), 945, 975, 1010, 1020, 1033 (α-spiral), 10 50, 1051, 1055, 1059, 1068, 1072, 1102, 1110, 1120, 1130, 1227, 1246 (α-helix), 1248 (α-helix), 1252 (α-helix), 1260 (α-helix), 1276 (α-helix), 1290 (α-helix), 1300 (α-helix), 1302 (α-helix), 1327 (β-fold), 1332 (β-fold), 1340, 1368, 1048-1063 (β-fold). The numbers indicate that GCN4 is inserted after a specific amino acid in nCas9(D10A). For example, 532 means that GCN4 is inserted after the 532nd amino acid in nCas9(D10A) (i.e., GCN4 is inserted between the 532nd and 533rd amino acids), and this nCas9(D10A) variant is named D10A-532GCN4. 1048-1063 means that amino acids 1048 to 1063 of nCas9(D10A) are replaced with GCN4, and this variant is named D10A-1048-1063GCN4. The naming rules for other variants follow the same pattern.

[0130] Taking the D10A-532GCN4 variant as an example, its base sequence (SEQ ID NO: 25) is shown:

[0131] The uppercase base sequence is the GCN4 sequence, and "GCT(D10A)" marks the mutation of the 10th amino acid of nCas9(D10A) from D to A.

[0132] AntiGCN4-BE3 base sequence (SEQ ID NO: 26):

[0133] 4. Compare the editing characteristics of the original BE3 and SunTag OEBE3.

[0134] Human embryonic kidney 293T cell lines were co-transfected with sgRNA plasmids HEK3-sgRNA, FANCF-sgRNA, ZAP70-sgRNA, and VEGFA-sgRNA along with the SunTag OEBE3 double plasmid (Figure 3). A control group was established with transfections of sgRNA and BE3 plasmids. Editing characterization was accurately assessed using sequencing.

[0135] The specific steps are as follows:

[0136] (1) Cell culture: Human embryonic kidney cell line 293T was cultured in DMEM complete medium containing 10% serum at 37°C and 5% CO2. When the cell confluence reached 90%, it was digested with 0.25% trypsin and then digested with DMEM complete medium to stop the digestion. The cells were then seeded into 12-well plates and cultured for another 24 hours.

[0137] (2) Plasmid transfection: After 24 hours, once cell adhesion was confirmed to be good and cell confluence reached 80%, transfection was performed. Each well was transfected with 0.5 μg of sgRNA plasmid, 0.5 μg of D10A-GCN4 plasmid (Figure 3), and 0.5 μg of AntiGCN4-BE3 plasmid (Figure 3). Transfection was performed using Yeasen's Polyethylenimine Linear (PEI) MW40000 transfection reagent according to the manufacturer's instructions, with 1 μg of BE3 plasmid used as a control group. Transfected cells were then cultured at 37°C in a 5% CO2 incubator.

[0138] (3) Extraction of genomic DNA: 48 hours after transfection, cells were routinely digested with 0.25% trypsin and digestion was terminated with DMEM complete medium. Cells were collected into centrifuge tubes, centrifuged at 300g for 5 minutes, the medium was discarded, washed once with PBS, centrifuged again at 300g for 5 minutes, the PBS was discarded, and cell residue was obtained. Genomic DNA was extracted from the cells using a cell genomic DNA extraction kit (TransGen Biotech Ltd., catalog number: EE101-01), and the DNA concentration was measured.

[0139] (4) Genomic PCR: Based on the genomic sequences of HEK3, FANCF, ZAP70, and VEGFA, upstream and downstream primers were designed, and the specific sequences are shown in Table 2. After amplifying the target DNA product, the product was sent to Guangzhou Sangon Biotech Co., Ltd. for sequencing.

[0140] Table 2 lists the PCR primers.

[0141] Analysis of the Sanger sequencing results revealed the editing efficiency of each D10A-GCN4 variant at four target sites, as shown in Figure 4. In the SunTag OEBE3 group, multiple D10A-GCN4 variants significantly improved editing efficiency, outperforming the original BE3 control group, and also altered the editing window range.

[0142] To further and more intuitively compare the efficiency and window differences between SunTag OEBE3 and BE3, the target editing results were merged into a window plot, as shown in Figure 5. Classifying the 62 D10A-GCN4 variants revealed that some variants showed approximately 2 times the editing efficiency compared to the BE3 control group. Furthermore, variants in different groups exhibited significantly different sequence preferences and expanded the editing window to C15 positions.

[0143] (5) Design of amplicon library construction primers: Primers were designed based on the gene sequences of HEK3, FANCF, ZAP70 and VEGFA, with both ends of the primers spanning the target site to amplify the target fragment. The amplicon primer sequences are shown in Table 3.

[0144] Table 3. Amplicon primer sequences for target sites

[0145] (6) PCR reaction for amplicon library construction: Based on the Sanger sequencing results, DNA samples of the D10A-GCN4 variant with high editing efficiency were selected for amplicon library construction, and PCR reaction was performed using the primers described above. The high-fidelity Kapa polymerase used in this experiment was KAPA HiFi HotStart ReadyMix, purchased from Nanjing Novizan Pharmaceutical Co., Ltd., catalog number KK2602.

[0146] The conditions for the first round of PCR are shown in Table 4:

[0147] Table 4. Conditions for the first round of PCR

[0148] First round PCR program: 98℃ for 3 min; 25 cycles (98℃ for 20 s, 65℃ for 15 s, 72℃ for 15 s), 72℃ for 1 min.

[0149] The conditions for the second round of PCR are shown in Table 5:

[0150] Table 5. Conditions for the second round of PCR

[0151] The I7 and I5 primers use commercial Illumina sequencing adapter primers: Hieff NGS384 Dual Index Primer Kit for... Set1 (ILLUMINA, item number 12613ES02).

[0152] Second round PCR program: 98℃ for 3 min, 11 cycles (98℃ for 20 s, 65℃ for 15 s, 72℃ for 15 s), 72℃ for 1 min.

[0153] After the PCR reaction was completed, the PCR products were subjected to agarose gel electrophoresis. DNA products of the target fragment size were then recovered from the gel and analyzed using high-throughput sequencing to determine the target mutation efficiency. The amplicon analysis results are shown in Figure 6. The results demonstrate that multiple variants of SunTag OEBE3 significantly improved editing efficiency while significantly reducing byproducts and non-specific products such as indels, greatly enhancing product purity.

[0154] Example 2

[0155] Efficiency analysis of Example 1, including editing window, editing efficiency, and byproducts, revealed that the SunTag OEBE3 system exhibited powerful and excellent editing performance. To further expand the OEBE system, another base editor, ABE8eWQ, was introduced, and the SunTag OEABE8eWQ system was constructed.

[0156] 1. Constructing sgRNA plasmids targeting specific sites

[0157] The sgRNA plasmid construction protocol is the same as above. The constructed plasmids are named HEK3-sgRNA, HEK2-sgRNA, OCT4-sgRNA, and ZAP70-sgRNA.

[0158] 2. Positive control plasmid

[0159] The ABE8eWQ plasmid was purchased from Addgene, catalog number 161815. The ABE8eWQ plasmid (Figure 7) contains the nCas9 (D10A) sequence with DNA single-strand cutting activity and the TadA8eWQ editing enzyme sequence that enables the editing of base A to G.

[0160] 3. Construct the Suntag OEABE8eWQ plasmid

[0161] The construction scheme was the same as above. The TadA8eWQ sequence was derived from the ABE8eWQ plasmid (Figure 7). The antibody AntiGCN4 was fused with ABE8eWQ and expressed, named AntiGCN4-ABE8eWQ, thus establishing the SunTag OEABE8eWQ system. The results are shown in Figure 7. After construction, the correctness and absence of mutations in the constructed plasmid sequence were confirmed by routine sequencing alignment. Single colonies with completely correct sequences were selected for amplification and plasmid extraction.

[0162] The base sequence of AntiGCN4-ABE8eWQ (SEQ ID NO: 79):

[0163] 4. Compare the editing characteristics of the original ABE8eWQ and SunTag OEABE8eWQ.

[0164] Human embryonic kidney 293T cell lines were co-transfected with sgRNA plasmids HEK3-sgRNA, HEK2-sgRNA, OCT4-sgRNA, and ZAP70-sgRNA along with the SunTag OEABE8eWQ double plasmid. A control group was prepared by transfecting cells with both sgRNA and ABE8eWQ plasmids. Sequencing and amplicon sequencing protocols were used as described above to accurately assess the editing characterization. The results are shown in Figure 8.

[0165] Editing window plots show that the editing activity of SunTag OEABE8eWQ is primarily concentrated in A6 and A7, with the blue group (D10A-1 / 532 / 945 / 1068 / 1072GCN4) exhibiting comparable efficiency to ABE8eWQ. While TadA8eWQ is known for its high fidelity, amplicon sequencing revealed that specific D10A-GCN4 variants can further reduce indels. For example, D10A-532GCN4 reduced the indels at OCT4 from 1.4% of ABE8eWQ to 0.44%. On the other hand, the shrinking window leads to improved editing precision and reduced bystander editing. At E21, D10A-532GCN4 increased the indels only from A8 to G from 1% of ABE8eWQ to 42%, while D10A-945GCN4 increased the indels only from A4 to G from 26% of ABE8eWQ to 61%. At OCT4, compared to ABE8eWQ, D10A-1055GCN4 significantly improved only A7 to G from 1% to 90%, while D10A-945GCN4 improved only A5 to G from 0.7% to 38%. These improvements highlight the ability of SunTag OEABE8eWQ to enhance the purity of single-base mutation products, a key factor in correcting point mutation genetic diseases in humans.

[0166] Example 3

[0167] To further expand the development of the OEBE system, another base editor, miniCGBE1, was introduced, and the SunTag OEminiCGBE1 system was constructed.

[0168] 1. Constructing sgRNA plasmids targeting specific sites

[0169] The sgRNA plasmid construction protocol is the same as above. The constructed plasmids are named HEK3-sgRNA, FANCF-sgRNA, EMX1-sgRNA, and PPP1R12C-sgRNA.

[0170] 2. Positive control plasmid

[0171] The miniCGBE plasmid was purchased from Addgene, catalog number 140253. The miniCGBE1 plasmid (Figure 9) contains the nCas9 (D10A) sequence with DNA single-strand cutting activity and the R33A editing enzyme sequence that enables the editing of base C to G.

[0172] 3. Construct the Suntag OEminiCGBE1 plasmid

[0173] The construction scheme was the same as above. The R33A sequence was derived from the miniCGBE1 plasmid. The antibody AntiGCN4 was fused with miniCGBE1 for expression, named AntiGCN4-miniCGBE, thus establishing the SunTag OEminiCGBE1 system (see Figure 9). After construction, the plasmid sequence was confirmed to be correct and without mutations by routine sequencing alignment. Single colonies with completely correct sequences were selected for amplification and plasmid extraction.

[0174] The base sequence of AntiGCN4-miniCGBE1 (SEQ ID NO: 80):

[0175] 4. Compare the editing characteristics of the original miniCGBE1 and SunTag OEminiCGBE1.

[0176] Human embryonic kidney 293T cell lines were co-transfected with sgRNA plasmids HEK3-sgRNA, FANCF-sgRNA, EMX1-sgRNA, and PPP1R12C-sgRNA along with the SunTag OEminiCGBE1 double plasmid (Figure 9). A control group was prepared by transfecting cells with either the sgRNA or miniCGBE1 plasmid. The sequencing protocol was the same as above. Editing characterization was accurately assessed through sequencing. The results are shown in Figure 10.

[0177] The merge window for SunTag OEminiCGBE1 narrowed to C6, with variants in the blue and purple groups exhibiting the highest efficiency. Excessive nonspecific product from miniCGBE1 led to a significant overestimation of the C to G efficiency calculated using Sanger sequencing data. In contrast, the efficiency of the D10A-GCN4 variants identified by Sanger sequencing was closer to that of amplicon sequencing due to a significant reduction in nonspecific product. At FANCF, D10A-532 / 1020 / 1246 / 1248 / 1252 / 1260GCN4 showed better performance, with 35% to 40% editing efficiency and 30% to 40% nonspecific product, compared to miniCGBE1's 30% editing efficiency and 50% nonspecific product. At PPP1R12C, D10A-231 / 1068 / 1072GCN4 achieved 30% editing efficiency and 10% nonspecific product, compared to miniCGBE1's 23% editing efficiency and 40% nonspecific product. Furthermore, the D10A-GCN4 variant significantly improved the purity of the C to G conversion, reaching approximately 75% at PPP1R12C, compared to 25% for miniCGBE1. Given that the C to G conversion requires the formation of a purine-free / pyrimidine-free site, CGBEs tend to produce a high proportion of nonspecific products. In contrast, SunTag OEminiCGBE1 significantly reduced these nonspecific products, providing a narrow editing window and improving the editing accuracy and practical application value of CGBEs.

[0178] Example 4

[0179] The SunTag system contains a GCN4 consisting of 19 amino acids and an AntiGCN4 consisting of 272 amino acids. In contrast, the MoonTag system contains a smaller epitope gp41 consisting of 15 amino acids and an antibody Antigp41, which is half the size of the GCN41 and has 123 amino acids.

[0180] 1. Constructing sgRNA plasmids targeting the target site

[0181] The sgRNA plasmid construction protocol is the same as above. The sgRNA plasmid uses the same SunTag OEBE system.

[0182] 2. Positive control plasmid

[0183] Same as SunTag OEBE system.

[0184] 3. Build the Moontag OEBE system

[0185] The construction scheme was the same as above. To verify this hypothesis, three MoonTag OEBE systems were constructed: MoonTag OEBE3, MoonTag OEABE8eWQ, and MoonTag OEminiCGBE1, as shown in Figure 11. These systems maintained the same nCas9 (D10A) embedding position as the SunTag OEBE system, and their naming conventions were consistent. After construction, the plasmid sequences were confirmed to be correct and mutation-free by routine sequencing alignment. Completely correct colonies were selected for amplification and plasmid extraction.

[0186] gp41 amino acid sequence (SEQ ID NO: 81): KNEQELLELDKWASL;

[0187] gp41 base sequence (SEQ ID NO: 82):

[0188] Antigp41 amino acid sequence (SEQ ID NO: 83):

[0189] Antigp41 base sequence (SEQ ID NO: 84):

[0190] 4. Comparison of editing characterization between the original base editor and MoonTag OEBE3: The experimental protocol was the same as above. Human embryonic kidney 293T cell line was co-transfected with the sgRNA plasmid and the MoonTag OEBE3 plasmid, with the transfection group containing the sgRNA plasmid and the control plasmid serving as the control group. Editing characterization was accurately assessed by sequencing. The results are shown in Figure 12.

[0191] Notably, in the MoonTag OEBE3 system, Sanger sequencing confirmed that D10A-gp41 variants with the same embedding position generally exhibited higher efficiency than their D10A-GCN4 counterparts. MoonTag OEBE3 significantly altered the sequence preference and window width of APOBEC1 within the merge window. Unlike GCN4, the window width of D10A-gp41 variants fluctuated synchronously to some extent, suggesting that these variants may have similar motif preferences. Twenty-three variants containing D10A-532gp41, among others, were the most efficient, while thirteen variants containing D10A-1368gp41 showed higher efficiency at C13 / 14 / 15 than other groups. Amplicon sequencing further confirmed the superior performance of most efficient D10A-gp41 variants, both in terms of increased efficiency and reduced byproducts. At ZAP70, D10A-gp41 variants reduced C-to-G by approximately two-thirds while improving overall efficiency. At the FANCF site, variants such as D10A-945 / 1020 / 1033gp41 increased C-to-T efficiency from approximately 20% for BE3 to 35%-70%, while reducing nonspecific products by more than 50%. Similarly, at the HEK3 and VEGFA sites, the D10A-gp41 variant exhibited excellent editing efficiency while minimizing nonspecific products. The enhanced performance of MoonTag OEBE3 compared to SunTag OEBE3 is likely due to the smaller size of Antigp41, which may allow the editing enzyme to better access the DNA.

[0192] Surprisingly, MoonTag OEABE8eWQ is able to generate a series of highly efficient variants comparable to ABE8eWQ, as shown in Figure 13.

[0193] The merged window analysis clearly shows that MoonTag OEABE8eWQ exhibits a variety of window patterns. The purple group, including D10A-532 / 1020 / 1260gp41, displays a window similar to that of ABE8eWQ. The green group, consisting of 12 variants, has a narrow window, with the highest efficiency at A7, while the blue group, with 9 variants, shrinks its window to A5, A6, and A7. Amplicon sequencing revealed that while A-to-C / T byproducts of the highly efficient variants remained at very low levels, insertion / deletion variations were significant. For example, compared to ABE8eWQ, D10A-945gp41 reduced insertions / deletions by approximately 20% to 80% at four sites, while D10A-1020gp41 reduced insertions / deletions only at OCT4. Similar to the SunTag system, MoonTag OEABE8eWQ significantly increased the proportion of single-base mutations due to variations in the editing window. At HEK2, D10A-1048-1063gp41 improved the A7 to G only editing rate from 0.15% (ABE8eWQ) to 50%. At OCT4, D10A-945gp41 and D10A-1368gp41 improved the A5 to G only editing rate from 0.7% (ABE8eWQ) to 48% and 58%, respectively. Significant improvements in precise editing were also observed at E21 and ZAP70 sites. Therefore, MoonTag OEABE8eWQ not only maintains excellent efficiency but also achieves specific single-base editing by programming the gp41 insertion site.

[0194] In MoonTag OEminiCGBE1, as shown in Figure 14, Sanger sequencing revealed higher C6 to G efficiencies for the D10A-1246 / 1248 / 1252 / 1260gp41 variants and higher C8 to G efficiencies for the D10A-532gp41 variant compared to miniCGBE1. Similar to SunTag, the merge editing window of MoonTag OEminiCGBE1 primarily narrows to C6. The highest efficiency was observed at C6 for six variant groups (including variants such as D10A-1246gp41). Another six variant groups (including variants such as D10A-532gp41) showed rejection of the TC7A motif. The blue group with 11 variants showed significant efficiency only at C6. Furthermore, amplicon sequencing confirmed the superior performance of specific D10A-gp41 variants in improving efficiency and reducing nonspecific products. D10A-1246 / 1248 / 1260gp41 showed 10% to 15% higher efficiency at FANCF compared to miniCGBE1, with a nearly 50% reduction in byproducts C to T and Indels. At PPP1R12C, D10A-532gp41 increased C to G-only editing products from 17% to 35% of miniCGBE1, while reducing non-specific products from 50% to 25%. Furthermore, editing accuracy at FANCF improved from 38% (miniCGBE1) to 57%, and at PPP1R12C from 25% (miniCGBE1) to 70%.

[0195] Example 5

[0196] To further verify the compatibility of OEBE, another CBE (Target AID) was introduced, and the SunTag OETarget AID system and the MoonTag OETarget AID system were established.

[0197] 1. Constructing sgRNA plasmids targeting specific sites

[0198] The sgRNA plasmid construction protocol is the same as above. The sgRNA plasmid uses the same SunTag OEBE system.

[0199] 2. Positive control plasmid

[0200] The Target AID plasmid (Figure 15) was purchased from Addgene, catalog number 131300, and contains the nCas9 (D10A) sequence, the editing enzyme pmCDA1 that enables the C-to-T mutation, and a UGI sequence.

[0201] 3. Build the OETarget AID system

[0202] The APOBEC1 sequence in the SunTag OEBE3 system and the MoonTag OEBE3 system was replaced with the pmCDA1 sequence, while all other structures remained unchanged, as shown in Figure 15.

[0203] 4. Compare the editing characteristics of the original Target AIF and OETarget AID.

[0204] As shown in Figure 16, both SunTag OETarget AID and MoonTag OETarget AID exhibited a range of wider editing windows and more complex editing modes. Amplicon results further demonstrate that specific combinations of SunTag OETarget AID and MoonTag OETarget AID can significantly reduce nonspecific products.

[0205] Example 6

[0206] To further validate the scalability of OEBE as a general strategy for manipulating editing patterns, the ALFA system, widely used in applications such as protein purification, was integrated. Similar to the small size of MoonTag, the ALFA epitope consists of only 13 amino acids, while its antibody, AntiALFA, consists of 132 amino acids.

[0207] 1. Constructing sgRNA plasmids targeting specific sites

[0208] The sgRNA plasmid construction protocol is the same as above. The sgRNA plasmid uses the same SunTag OEBE system.

[0209] 2. Positive control plasmid

[0210] Same as SunTag OEBE system.

[0211] 3. Construct the ALFA OEBE system

[0212] Three ALFA OEBE systems were constructed: ALFA OEBE3, ALFA OEABE8eWQ, and ALFA OEminiCGBE1, as shown in Figure 17. These systems maintained the same nCas9(D10A) embedding position as the SunTag OEBE system, and their naming conventions were consistent. After construction, the plasmid sequences were confirmed to be correct and mutation-free by routine sequencing alignment. Completely correct colonies were selected for amplification and plasmid extraction.

[0213] ALFA amino acid sequence (SEQ ID NO: 85): SRLEEELRRRLTE;

[0214] ALFA base sequence (SEQ ID NO: 86):

[0215] AntiALFA amino acid sequence (SEQ ID NO: 87):

[0216] AntiALFA base sequence (SEQ ID NO: 88):

[0217] 4. Compare the editing characterization of the original base editor and ALFA OEBE.

[0218] The sgRNA plasmid and the ALFA OEBE3 double plasmid were co-transfected into the human embryonic kidney 293T cell line, and a control group was set up with the sgRNA plasmid and the control plasmid. Editing characterization was accurately evaluated by sequencing.

[0219] First, ALFA OEBE3 was constructed, and D10A-231ALFA and D10A-1246ALFA were tested. The results, shown in Figure 18, indicate that the editing windows for GCN4, gp41, and ALFA differ at the same embedding location. At C15, D10A-231ALFA achieved an efficiency of 35.8%, while D10A-231GCN4 / gp41 achieved 5% to 10%. At C3 and C4, D10A-1246ALFA showed a 20% improvement in efficiency compared to D10A-1246GCN4 / gp41. Amplicon sequencing revealed that D10A-231ALFA and D10A-1246ALFA significantly reduced nonspecific byproducts compared to BE3, but remained comparable to their GCN4 / gp41 counterparts.

[0220] Next, the ALFA OEABE8eWQ with embedding positions 945, 1246, and 1252 was studied, and the results are shown in Figure 19.

[0221] The editing window varies depending on the protein recruitment system used. At position 945, ALFA is less effective than GCN4 and gp41 at A5 and A6. At position 1246, ALFA performs similarly to gp41. At position 1252, ALFA improves efficiency from A3 to A6 compared to GCN4 and gp41. Amplicon sequencing further revealed that at E21, HEK2, and OCT4, D10A-945 / 1246 ALFA produces less nonspecific product than ABE8eWQ, while at ZAP70, they are comparable. Interestingly, changing the epitope type affects editing accuracy. At OCT4, D10A-945 ALFA primarily achieves A7 to G, while D10A-945 GCN4 and D10A-945 gp41 primarily produce A5 to G. At E21, D10A-1246ALFA showed approximately 60% A4 to G, exceeding the 38% and 26% observed in D10A-1246GCN4 and D10A-1246gp41, respectively. Therefore, the integration of the ALFA system further enhances the OEABE8eWQ's ability to reprogram editing modes and reduce bystander editing.

[0222] Furthermore, incorporating miniCGBE1 into the ALFA OEBE (see Figure 20) revealed that D10A-1246ALFA and D10A-1252ALFA exhibited significantly improved efficiency at C7 and C8 compared to their GCN4 and gp41 counterparts. Therefore, ALFA can compensate for the low efficiency of SunTag / MoonTag OEminiCGBE1 at C7 and C8.

[0223] Example 7

[0224] Since component reordering can alter the spatial location of deaminases and the final protein polymerization structure, the impact of reordering deaminases and antibodies on editing results was investigated. As shown in Figure 21, four different sequences of antibodies and deaminases were established within the OEABE8eWQ framework. Three embedding positions, 532, 945, and 1059, were selected for evaluation, and significant changes in efficiency and window width were observed. Nomenclature rules are defined as follows: 532gp41-ABE8eWQ indicates the use of the D10A-532gp41 variant and Antigp41-ABE8eWQ; ABE8eWQ-532GCN4 indicates the use of the D10A-532GCN4 variant and ABE8eWQ-AntiGCN4; other combinations follow the same nomenclature rules. Compared to 532gp41 / GCN4-ABE8eWQ, which contains D10A-532gp41 / GCN4 and Antigp41 / GCN4-ABE8eWQ, ABE8eWQ-532gp41, which contains D10A-532gp41 / GCN4 and ABE8eWQ-AntiGCN4, showed increased efficiency at A3, A4, and A5, but decreased efficiency at A7 and A8. ABE8eWQ-1059GCN4 widened the window width compared to 1059GCN4-ABE8eWQ. In poor OEBE combinations, strategic reordering of deaminases and antibodies can optimize editing efficiency. At the E21 and HEK2 sites, ABE8eWQ-1059GCN4 significantly improved efficiency from 5% of 1059GCN4-ABE8eWQ to 36%, and ABE8eWQ-1059gp41 improved efficiency by approximately 20% compared to 1059gp41-ABE8eWQ. Furthermore, the reordering of deaminases and antibodies affected editing precision. For example, at E21, ABE8eWQ-1059GCN4 further improved A8 to G efficiency from 40% to 56% compared to 1059GCN4-ABE8eWQ, while reducing other single-base mutations from 33% to 6%.

[0225] To further confirm whether this strategy applies to another editing enzyme, TadA9 from ABE9 (known for its minimal editing window) was integrated into OEBE and named OEABE9. The naming convention for the combinations was the same as above. The results, as shown in Figure 22, show that the reordering within the OEABE9 framework produced clearer window changes. At position 532, the window of ABE9-532gp41 is similar to that of ABE9, while the other three combinations shrink the window towards A7. At position 945, ABE9-945gp41 shows higher efficiency than the other combinations. ABE9-1059GCN4, ABE9-1059gp41, and 1059gp41-ABE9 provide narrower windows than ABE9. Efficiency optimization was also observed in the reordering of OEABE9. At HEK2, the ABE9-GCN4 / gp41 combination generally achieves higher efficiency than GCN4 / gp41-ABE9, with ABE9-532GCN4 improving by over 10%, ABE9-532gp41 by over 23%, and ABE9-1059gp41 by over 20%. Furthermore, at OCT4 and ZAP70, ABE9-945gp41 shows 32.5% and 24.5% higher efficiency than 945gp41-ABE9, respectively. Similarly, reordering in OEABE9 affects editing accuracy. At OCT4, compared to 945gp41-ABE9, ABE9-945gp41 improves A7 to G from 7% to 32% and reduces A5 to G from 79% to 6%. At E21, ABE9-1059GCN4 increases the A8 to G ratio from 29.5% for 1059GCN4-ABE9 to 56.5%.

[0226] In the OECGBE1 system integrating OEminiCGBE1 and eUNG (E. coli uracil DNA N-glycosylation enzyme), efficiency improvements were achieved by combining eUNG and R33A in different orders, as shown in Figure 23. The naming conventions were the same as above. At FANCF, compared to CGBE, 1246 / 1252GCN4-R33A, and 1246 / 1252GCN4-eUNG-R33A, 1246 / 1252GCN4-R33A-eUNG further improved efficiency by 5% to 10%. Conversely, at PPP1R12C, 1246CGN4 / gp41-eUNG-R33A showed a 5% to 15% improvement in efficiency compared to other groups.

[0227] In summary, reordering strategies for editing enzymes and antibodies can be used to improve OEBE performance and customize specific editing patterns.

[0228] Example 8

[0229] To further verify whether the OEBE system is compatible with other Cas proteins, it was further extended to FrCas9 and LbCas12 proteins.

[0230] 1. Constructing sgRNA plasmids targeting specific sites

[0231] The sgRNA plasmid construction scheme is the same as above, and the sgRNA sequence is shown in Table 1.

[0232] 2. To test OEBE compatibility, GCN4 was inserted after amino acid 670 of the FrCas9 (E796A) protein and after amino acid 1040 of the LbCas12 (dead form, D832A / E1006A / D1125A) protein. These were then used with the AntiGCN4-ABE8eWQ and AntiGCN4-BE3 plasmids (Figure 24), respectively. After construction, conventional sequencing alignment confirmed the correctness and absence of mutations in the constructed plasmid sequences. Completely correct colonies were selected for amplification and plasmid extraction.

[0233] The FrCas9(E796A)-610GCN4 sequence (SEQ ID NO: 89) contains the uppercase base sequence GCN4.

[0234] LbCas12-1040GCN4 sequence (SEQ ID NO: 90, where the uppercase base sequence is GCN4):

[0235] 3. The transfection and testing protocols are the same as above.

[0236] Cells were transfected with the double plasmid, and DNA was extracted and Sanger sequencing was performed. PCR primers are shown in Table 2. As shown in Figure 24, in the OEBE system, replacing nCas9(D10A) with FrCas9(E796A) and LbCas12(dead) both exhibited normal base editing capabilities.

[0237] Example 9

[0238] To test whether OEBE can be applied to tumor gene therapy, two targets, the E6 and E7 oncogenes, on the HPV18 tumor genome of the HeLa cell line, were selected.

[0239] 1. Constructing sgRNA plasmids targeting specific sites

[0240] The sgRNA plasmid construction scheme is the same as above, and the sgRNA sequence is shown in Table 1.

[0241] 2. To test the editing efficiency of OEBE on oncogenes, D10A-1246gp41 and AntiGCN4-BE3 were selected for use.

[0242] 3. The transfection and testing protocols are the same as above.

[0243] The plasmid was transfected into HeLa cells, and then DNA was extracted and Sanger sequencing was performed. The PCR primers are shown in Table 2.

[0244] As shown in Figure 25, the target base C of the two selected target sequences can be edited to T, thereby generating a stop codon. This causes premature termination of translation of the E6 and E7 oncogenes, resulting in incomplete E6 and E7 proteins. This is helpful for the development of tumor gene therapy strategies for the HPV18 oncogene. Furthermore, at the selected target sites, the target base C is located in a CCN background, which is unsuitable for traditional base editors such as BE3 and BE4, preventing editing, or the target base C exceeds the editing window of BE3 and BE4, making editing impossible. Conversely, in the OEBE system, the combination of D10A-1246gp41 and AntiGCN4-BE3 achieves efficient editing of the target base C of the target sequence. In contrast, the control group BE3 and BE4 failed to edit the target base C. More importantly, after OEBE editing, HeLa cells showed a significant decrease in growth rate due to premature termination of translation of E6 and E7, further demonstrating the value of the OEBE system in tumor gene therapy.

[0245] Example 10

[0246] To test whether OEBE could be used for live animal editing, it was applied to zebrafish embryo editing.

[0247] 1. In vitro transcription of sgRNA and related protein mRNA, as shown in Figure 26A. mRNA was prepared using the HiScribe T7ARCA mRNA Kit (NEB, E2060S) according to the manufacturer's instructions. sgRNA was synthesized by GenScript (Nanjing, China). The sgRNA sequence is listed in Table 1.

[0248] 2. Embryo injection: Zebrafish 1-cell stage embryos were selected and microinjected with a 2 nL mixture of mRNA (BE3 or ABE8eWQ, 100 ng / μL) and sgRNA (200 ng / μL), or OEBE mRNA (D10A-gp41 and Antigp41, 100 ng / μL) and sgRNA (200 ng / μL). Fluorescent embryos were sorted 6 to 8 h after microinjection, as shown in Figure 26B, and their genomic DNA was extracted and sequenced 72 h later. PCR primers and amplicon primers are listed in Tables 2 and 3.

[0249] 3. The in vivo editing capabilities of OEBE were explored by editing the zebrafish genes TWIST2 and TYR. These mRNAs contain sequences encoding fluorescent proteins to facilitate the identification of embryos expressing nCas9 (D10A) and Antigp41 (Fig. 26B). When the TWIST2 gene was edited using ABE8eWQ and 231gp41-ABE8eWQ (containing D10A-231gp41 and Antigp41-ABE8eWQ), juveniles in both groups exhibited slightly curved tails, smaller body size, and enlarged vacuoles in the yolk sac (Fig. 26C, left). Sequencing results of DNA extracted from juveniles showed that ABE8eWQ primarily edited A6 in TWIST2. Notably, 231gp41-ABE8eWQ avoided the CA6G sequence but showed high efficiency at A8, approximately 3.8 times that of ABE8eWQ (Fig. 26C, right). For further analysis, five embryos with the highest editing efficiency were selected from each group for subsequent amplicon sequencing. 231gp41-ABE8eWQ significantly improved the conversion rate from A8 to G only from 10% (ABE8eWQ) to an average of 59% (Figure 26D).

[0250] Notably, when editing the zebrafish TYR gene using BE3 and 1246gp41-BE3 (containing D10A-1246gp41 and Antigp41-BE3), only 1246gp41-BE3 produced the transparent juvenile phenotype (Fig. 26E, left). Sequencing revealed that due to the limited editing window, BE3 only achieved C7 to T editing, while 1246gp41-BE3 was able to achieve C11 to T editing in the TYR coding sequence, thus generating a stop codon (Fig. 26E, right). In amplicon sequencing analysis of the edited region, 1246gp41-BE3 improved the efficiency of C to T editing to nearly 90% while reducing the level of nonspecific products (Fig. 26F). Furthermore, in sequence type analysis, the major product of BE3 was C7 to G, while in 1246gp41-BE3, the C7 / 11 to T readings exceeded 75%.

[0251] In summary, the core concept of the OEBE technology in this application is that the base editing characterization pattern changes depending on the spatial location and protein polymerization structure of the editing enzyme on the nCas9(D10A) protein. The core principle of OEBE technology is that multiple sites within nCas9(D10A) can insert antigenic epitopes. Using a protein recruitment system, the deaminase is immobilized at specific spatial locations on the nCas9(D10A) protein, altering the protein polymerization structure and thus changing its base editing characterization pattern.

[0252] (1) The first key technical point is to fix the deaminase at different spatial positions on the surface of the Cas protein and achieve protein polymerization while retaining the functional activity of nCas9(D10A).

[0253] In protein recruitment systems, antibodies specifically recognize antigenic epitopes, thereby transporting the antibody-fused protein to the target protein site fused with the epitope. The epitope is short, resulting in less adverse effects from insertion into the protein. The Suntag system is a highly efficient protein recruitment system containing a 19-amino acid epitope, GCN4, and its antibody, AntiGCN4. Notably, GCN4 possesses dimerizing capabilities, enabling the dimerization of the fused protein. In this application, GCN4 was inserted into the irregular coil within SpCas9. Base editing experiments revealed that most insertion sites did not impair Cas9 activity. GCN4 was also inserted into nCas9, and AntiGCN4 fused with the editing enzyme APOBEC1 to form a modular plug-in. Through protein recruitment, the plug-in "AntiGCN4-APOBEC1" could be fixed at a specific spatial location, altering the protein dimerization structure. Experimental results show that as the position of GCN4 in nCas9(D10A) changes, the characterization of base editing also changes significantly, including editing efficiency, editing window, and product purity. This indicates that the Suntag system can effectively immobilize APOBEC1 in different spatial locations and alter the protein dimerization structure without affecting the function of nCas9.

[0254] (2) The second key technical point is that OEBE technology has strong scalability, that is, it can be compatible with other protein recruitment systems and other editing enzymes.

[0255] The Moontag system is also a highly efficient protein recruitment system. The antigenic epitope gp41 is 15 amino acids long, and the antibody Antigp41 is a nanobody. gp41 has trimerizing capabilities, enabling it to trim the protein it is fused with. Antigp41 is smaller and has a higher binding capacity, resulting in a shorter and more stable distance between the deaminase and nCas9; gp41 is also smaller, so its influence on protein conformation after insertion is less. Introducing another protein recruitment system, ALFA, into OEBE also achieves highly efficient editing. This indicates that OEBE technology is compatible with multiple protein recruitment systems. Furthermore, this application used other editing enzymes (TadA8eWQ, TadA9, and R33A, etc.) fused with AntiGCN4, Antigp41, and AntiALFA, respectively. The results show that OEBE can still exhibit superior editing performance by changing the type of deaminase, and its base editing characterization changes with the insertion position of the antigenic epitope. Other common protein recruitment systems include Flag and HA, and more and more base editing enzymes are being reported. Introducing these protein recruitment systems and base editing enzymes into the OEBE system can generate new editing characterization patterns, greatly increasing the ways to combine modules and significantly improving the scalability of the OEBE system. Therefore, modularity and scalability are important characteristics of OEBE.

[0256] (3) The third key technology is that the OEBE system can generate a series of new base editors.

[0257] Due to the existence of multiple protein recruitment systems and editing enzymes, OEBE technology possesses extremely strong plug-in scalability. Furthermore, the differences in size, conformation, and specificity among different protein recruitment systems mean that even with the same insertion site and the same deaminase type, their base editing characterization is not entirely the same, further expanding the diversity and flexibility of OEBE technology. OEBE technology can generate numerous modular combinations and multiple characterization modes, such as expanding or shrinking the editing window and altering sequence preferences. The OEBE system uses CRISPR / Cas as the insertion platform, multiple embedded antigenic epitopes on the Cas protein as slots, and protein recruitment systems and editing enzymes as plug-ins, thus generating a series of novel base editors without protein evolution. Therefore, OEBE is a modular, polymerizable, programmable, and scalable base editing platform. As the CRISPR platform, slots, and plug-ins are added, it can continuously expand, continuously generating novel base editors, bringing more possibilities and choices to the field of gene editing.

[0258] This application organically integrates the CRISPR / Cas system, protein recruitment system, and base editing technology. It develops an OEBE base editing system that can adjust the spatial location of the editing enzyme and the protein polymerization structure, as well as reprogram its editing characterization. This application modifies Cas by embedding antigenic epitopes at different positions within Cas, thereby generating a series of Cas variants. Furthermore, it fuses the editing enzyme with antibodies to produce a series of novel modular base editors. These variants and their combinations can significantly optimize the editing characterization of the original base editor.

[0259] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit the scope of protection of this application. Although this application has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of this application without departing from the substance and scope of the technical solutions of this application.

Claims

1. A fusion protein, characterized in that, The fusion protein is a Cas9 protein with an internally embedded or replaced antigenic epitope; the antigenic epitope is an antigenic epitope in a protein recruitment system.

2. The fusion protein according to claim 1, characterized in that, The Cas9 protein is any one of nCas9 (D10A) protein, FrCas9 (E796A) protein, or LbCas12 (dead) protein.

3. The fusion protein according to claim 1, characterized in that, The protein recruitment system includes any one of SunTag, MoonTag, ALFA, Flag, and HA.

4. The fusion protein according to claim 1, characterized in that, The antigenic epitopes include any one of GCN4, gp41, and Alfa.

5. The fusion protein according to claim 1, characterized in that, The amino acid sequence of the antigenic epitope is shown in any one of SEQ ID NO: 21, 81 and 85.

6. The fusion protein according to claim 2, characterized in that, The antigenic epitopes are embedded in the nCas9(D10A) protein at positions 1, 25, 54, 62, 170, 203, 213, 231, 249, 257, 308, 355, 365, 400, 532, 573, 584, 674, 719, 765, 768, 776, 782, 790, 808, 819, 826, 831, 846, 868, 890, 910, 924, 945, 975, 1010, 1020, and 10. The epitope is located after at least one of the following amino acids: 33, 1050, 1051, 1055, 1059, 1068, 1072, 1102, 1110, 1120, 1130, 1227, 1246, 1248, 1252, 1260, 1276, 1290, 1300, 1302, 1327, 1332, 1340, and 1368, or the antigenic epitope replaces amino acids 1048-1063 of the nCas9(D10A).

7. The fusion protein according to claim 2, characterized in that, The antigenic epitopes are embedded in the FrCas9 (E796A) protein at positions 1, 40, 51, 55, 108, 112, 120, 125, 170, 192, 200, 215, 234, 250, 265, 270, 275, 290, 295, 315, 320, 325, 363, 379, 428, 539, 566, 610, and 618. The following at least one amino acid from the following amino acids: 690, 702, 707, 727, 755, 761, 800, 830, 882, 980, 1002, 1043, 1055, 1095, 1111, 1117, 1134, 1184, 1188, 1224, 1248, 1321, 1329, 1333, 1347, and 1372.

8. The fusion protein according to claim 2, characterized in that, The antigenic epitope is embedded after at least one of the following amino acids in the LbCas12(dead) protein: 1, 85, 212, 270, 372, 406, 441, 478, 571, 625, 655, 713, 730, 760, 770, 774, 780, 807, 825, 847, 865, 965, 985, 1010, 1040, 1055, 1074, 1079, 1087, 1109, 1120, 1121, 1142, 1143, 1158, 1171, 1201, 1212, and 1228.

9. A nucleic acid molecule, characterized in that, Encodes the fusion protein as described in any one of claims 1-8.

10. The nucleic acid molecule according to claim 9, characterized in that, The nucleotide sequence of the antigenic epitope in the fusion protein is shown in any one of SEQ ID NO: 22, 82, and 86.

11. A plasmid, characterized in that, Includes the nucleic acid molecule as described in claim 9 or 10.

12. A base editor system, characterized in that, The invention comprises the fusion protein and editing enzyme fusion protein according to any one of claims 1-8; the editing enzyme fusion protein comprises an antibody against the antigenic epitope and an editing enzyme.

13. The base editor system according to claim 12, characterized in that, The antibody includes any one of AntiGCN4, Antigp41, and AntiALFA; the amino acid sequence of AntiGCN4 is shown in SEQ ID NO: 23; the amino acid sequence of Antigp41 is shown in SEQ ID NO: 83; the amino acid sequence of AntiALFA is shown in SEQ ID NO: 87; the editing enzyme includes at least one of APOBEC1, pmCDA1, TadA8eWQ, TadA9, and R33A.

14. The base editor system according to claim 13, characterized in that, The nucleotide sequence of the antibody is shown in any one of SEQ ID NO: 24, 84, and 88.

15. The base editor system according to claim 12, characterized in that, The functional elements of the editing enzyme fusion protein are sequentially connected in any of the following ways: i. Promoter, antibody, editing enzyme, UGI, GB1, NLS, BGH; ii. Promoter, antibody, editing enzyme, GB1, NLS, BGH; iii. Promoter, the editing enzyme, the antibody, GB1, NLS, BGH; iv. Promoter, antibody, editing enzyme, eUNG, GB1, NLS, BGH; v, promoter, antibody, eUNG, editing enzyme, GB1, NLS, BGH.

16. The base editor system according to claim 13, characterized in that, The functional elements of the editing enzyme fusion protein are sequentially connected in any of the following ways: i. Promoter, AntiGCN4, APOBEC1, UGI, GB1, NLS, BGH; ii. Promoter, AntiGCN4, TadA8eWQ, GB1, NLS, BGH; iii. Promoter, AntiGCN4, R33A, GB1, NLS, BGH; iv. Promoter, Antigp41, APOBEC1, UGI, GB1, NLS, BGH; v, promoter, Antigp41, TadA8eWQ, GB1, NLS, BGH; vi, promoter, Antigp41, R33A, GB1, NLS, BGH; vii, the promoter, the AntiALFA, the APOBEC1, UGI, GB1, NLS, BGH; viii, promoter, AntiALFA, TadA8eWQ, GB1, NLS, BGH; ix, promoter, AntiALFA, R33A, GB1, NLS, BGH; x, the promoter, AntiGCN4, pmCDA1, UGI, GB1, NLS, BGH; xi, promoter, Antigp41, pmCDA1, UGI, GB1, NLS, BGH; xii, promoter, TadA8eWQ, AntiGCN4, GB1, NLS, BGH; xiii, promoter, TadA8Ewq, Antigp41, GB1, NLS, BGH; xiv, promoter, AntiGCN4, TadA9, GB1, NLS, BGH; xv, promoter, TadA9, AntiGCN4, GB1, NLS, BGH; xvi, the promoter, the Antigp41, the TadA9, GB1, NLS, BGH; xvii, promoter, TadA9, Antigp41, GB1, NLS, BGH; xviii, promoter, AntiGCN4, R33A, eUNG, GB1, NLS, BGH; xix, promoter, AntiGCN4, eUNG, R33A, GB1, NLS, BGH; xx, promoter, Antigp41, R33A, eUNG, GB1, NLS, BGH; xxi, promoter, Antigp41, eUNG, R33A, GB1, NLS, BGH.

17. The base editor system according to claim 12, characterized in that, It also includes sgRNA; the sgRNA includes a target site recognition sequence.

18. A cell characterized in that, The base editor system includes any one of claims 12-17.

19. The fusion protein of any one of claims 1-8, the nucleic acid molecule of claim 9 or 10, the plasmid of claim 11, the base editor system of any one of claims 12-17, and the cell of claim 18 in gene editing.

20. The application according to claim 19, characterized in that, The gene editing includes at least one of editing the base C of the target site sequence to the base T, editing the base C of the target site sequence to the base G, and editing the base A of the target site sequence to the base G.

21. The use of the fusion protein of any one of claims 1-8, the nucleic acid molecule of claim 9 or 10, the plasmid of claim 11, the base editor system of any one of claims 12-17, and the cell of claim 18 in the preparation of a medicament for tumor gene therapy.

22. The fusion protein of any one of claims 1-8, the nucleic acid molecule of claim 9 or 10, the plasmid of claim 11, the base editor system of any one of claims 12-17, and the cell of claim 18 in agricultural breeding.