A fusion protein and use thereof

By fusing the R-loop binding domain with a single-gene base editor to form a fusion protein, the limitations of existing base editors in terms of targeting and editing precision are overcome, achieving more efficient base editing and wider applicability.

CN119431603BActive Publication Date: 2026-02-17SUN YAT SEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411402674.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-09
Publication Date
2026-02-17
Estimated Expiration
2044-10-09

AI Technical Summary

Technical Problem

Existing base editors have limitations in terms of targeting and editing precision. In particular, the dependence of Cas nucleases on PAM and the bystander effect affect their editing efficiency and window controllability, making them unable to adapt to more editing needs.

Method used

By fusing R-loop binding domains from different species with a single-gene base editor, a fusion protein was formed, including an R-loop binding domain, nucleoside deaminase, and nuclease, thereby optimizing editing efficiency and expanding the editing window.

Benefits of technology

It improves base editing efficiency, widens the editing window, reduces cytosine off-target editing activity and insertion/deletion frequency, and enhances the ability to edit target sites.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119431603B_ABST
    Figure CN119431603B_ABST
Patent Text Reader

Abstract

The application discloses a fusion protein and application thereof. The application discloses that the base editing efficiency of a fusion protein obtained by fusing and expressing R-loop binding domains from different species with a single-gene base editor is improved, and the base editing window is widened compared with the base editor before fusion. That is, the base editing efficiency of the fusion protein can be improved by fusing the R-loop binding domain with the single-gene base editor, and the editing window can be adjusted, so that the fusion protein can be applied to more base editing scenes. The application is beneficial to the popularization and application of single-gene editing technology, and has wide application prospects in the prevention and / or treatment of diseases, the construction of animal models or plant varieties.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of base editing. More particularly, it relates to a fusion protein and its application. BACKGROUND

[0002] Single nucleotide variants (SNVs) are the most common mutations in the human genome, accounting for 90% of genomic mutations. Among them, nearly 100,000 SNVs are related to the occurrence of human diseases, which has led to an urgent need for precise and efficient single-base editing technology. Single-base editing technology not only deepens the understanding of the genetic basis of human health, but also provides hope for correcting pathogenic SNVs.

[0003] Base editors (BEs) are a revolutionary genetic engineering tool that can achieve precise DNA base conversion without double-strand breaks or donor templates. Its basic principle is to locate the target site by means of the CRISPR-Cas9 system, and then achieve precise mutation of the target base by fused nucleoside deaminase. According to the difference of deaminase, base editors are divided into cytosine base editors (CBE) with cytosine deaminase as the core and adenine base editors (ABE) with adenosine deaminase as the core. CBE is composed of nCas9(D10A)-APOBEC fusion protein and uracil glycosylase inhibitor (UGI), which can deaminate cytosine to uracil at the target site, thereby realizing the conversion from cytosine (C) to thymine (T). ABE is composed of nCas9(D10A) and TadA (adenosine deaminase), which can effectively convert adenine (A) to guanine (G) at the target site. CBE and ABE are the most widely studied and applied base editors, and theoretically can correct 59% of pathogenic point mutations, and have broad application prospects in the field of gene therapy for human genetic diseases.

[0004] Although base editors have great potential in many fields, existing base editing tools still face many challenges in application. The requirement of Cas nuclease to recognize specific protospacer adjacent motifs (PAMs) limits targeting to all genetic sites. For example, the optimal editing window of typical SpCas9-based ABE and CBE systems is 4-8 bases at the end of NGG-PAM, which makes more than two-thirds of pathogenic SNVs untargetable. In addition, ABE and CBE can edit multiple variable nucleotides within the editing window, a phenomenon known as the bystander effect, which affects the precise editing of bases. Subsequent development of editors such as ABEmax, ABE8e and BE4max has improved editing activity compared to previous editors, but still has problems such as insufficient editing activity and uncontrollable editing window, which cannot be adjusted to adapt to more editing needs. Therefore, there is still a need to improve and optimize base editors to improve editing activity, adjust editing window, etc. SUMMARY

[0005] The present application is directed to the above technical problems, and provides a fusion protein which can improve the editing efficiency of single gene base editors and adjust (widen) the editing window thereof.

[0006] The first object of the present application is to provide a fusion protein.

[0007] The second object of the present application is to provide a gene encoding the fusion protein.

[0008] The third object of the present application is to provide a recombinant plasmid containing the gene.

[0009] The fourth object of the present application is to provide a recombinant cell or a recombinant bacteria containing the fusion protein or the gene.

[0010] The fifth object of the present application is to provide a single gene base editing system.

[0011] The sixth object of the present application is to provide the use of the fusion protein, the gene, the recombinant plasmid, the recombinant cell or the recombinant bacteria, or the single gene base editing system in the preparation of gene editing products, disease treatment and / or prevention products, animal models or plant varieties.

[0012] The above objects of the present application are achieved by the following technical solutions:

[0013] The application discloses a fusion protein, which comprises an R-loop binding domain, a nucleoside deaminase and a nuclease.

[0014] Specifically, the R-loop binding domain comprises a binding domain with an amino acid sequence having a similarity higher than 75% to the amino acid sequence shown in SEQ ID NO. 1 and / or a binding domain with an amino acid sequence shown in any one or more of SEQ ID NO. 1, 3, 4, 5 and 7; the fusion protein has a connection sequence that the nucleoside deaminase is located at the N-terminal of the nuclease, and the binding domain of the R-loop structure is located at the N-terminal of the nucleoside deaminase or is located between the nucleoside deaminase and the nuclease.

[0015] Optionally, the amino acid sequence of the binding domain with a similarity higher than 75% to the amino acid sequence shown in SEQ ID NO. 1 is shown in SEQ ID NO. 2 or SEQ ID NO. 6.

[0016] Preferably, when used for widening the editing window, the R-loop binding domain is located between the nucleoside deaminase and the nuclease.

[0017] Preferably, the R-loop binding domain comprises a binding domain with an amino acid sequence shown in any one or more of SEQ ID NO. 1 to 7.

[0018] Specifically, when the amino acid sequence of the R-loop binding domain is shown in SEQ ID NO. 1, the R-loop binding domain is located at the N-terminal of the nucleoside deaminase or is located between the nucleoside deaminase and the nuclease, which can improve the editing efficiency, widen the editing window, and reduce the cytosine off-target editing activity and the insertion-deletion (InDel) frequency.

[0019] When the amino acid sequence of the R-loop binding domain is shown in SEQ ID NO. 2 or SEQ ID NO. 6, the R-loop binding domain is located at the N-terminal of the nucleoside deaminase, which can improve the editing efficiency and widen the editing window.

[0020] Specifically, the nucleoside deaminase comprises an adenosine deaminase and / or a cytosine deaminase.

[0021] Optionally, the adenosine deaminase is selected from the adenosine deaminase of ABEmax or / and ABE8e.

[0022] Optionally, the cytosine deaminase is selected from cytosine deaminase derived from a rat.

[0023] Optionally, the nuclease is selected from one or more of the following nucleases: Cpf1, C2C1, C2C2, C2C3, Cas9, Cas12, Cas13, Cas14, TnpB, IscB, Argonaute, or a mutant thereof.

[0024] Specifically, the Cas9 includes but is not limited to dCas9, GeoCas9, CjCas9, nCas9.

[0025] Specifically, the Cas12 includes but is not limited to Cas12a, Cas12b, Cas12c, Cas12d, Cas12f, Cas12g, Cas12h, Cas12i, Cas12m.

[0026] Specifically, the Cas13 includes but is not limited to Cas13d.

[0027] Specifically, the Cas9 is selected from Cas9 derived from Streptococcus pneumoniae, Staphylococcus aureus, Streptococcus pyogenes, or Streptococcus thermophilus.

[0028] More specifically, the Cas9 is nSpCas9 (carrying D10A or mutant SpCas9) derived from Streptococcus pyogenes.

[0029] In a specific embodiment of the present application, the adenosine deaminase contained in the fusion protein is TadA and / or TadA* (TadA* is a mutant of TadA); the amino acid sequence of the TadA is shown in SEQ ID NO. 8, and the amino acid sequence of the TadA* is shown in SEQ ID NO. 9.

[0030] In a specific embodiment of the present application, the cytosine deaminase contained in the fusion protein is cytosine deaminase APOBEC1 selected from BE4max; the amino acid sequence of the APOBEC1 is shown in SEQ ID NO. 10.

[0031] In a specific embodiment of the present application, the nuclease contained in the fusion protein is SpCas9 (D10A); the amino acid sequence of the SpCas9 (D10A) is shown in SEQ ID NO. 11.

[0032] Specifically, the fusion protein further comprises a nuclear localization signal (NLS) at at least one end of the fusion protein, i.e., the nuclear localization signal is located at the N-terminus and / or C-terminus of the fusion protein.

[0033] Specifically, the amino acid sequence of the NLS is shown as SEQ ID NO. 12.

[0034] Specifically, the fusion protein further comprises two copies of a uracil glycosidase inhibitor (UGI) at at least one end of the fusion protein, i.e., the uracil glycosidase inhibitor is located at the N-terminus and / or the C-terminus of the fusion protein.

[0035] Specifically, the amino acid sequence of the UGI is shown as SEQ ID NO. 13.

[0036] The application also provides a gene, which is the fusion protein of the application.

[0037] The application also provides a recombinant plasmid containing the gene of the application.

[0038] Optionally, the vector used in the recombinant plasmid comprises a viral vector and / or a non-viral vector; wherein the viral vector comprises an adeno-associated viral vector, an adenoviral vector, a lentiviral vector, a retroviral vector and / or an oncolytic viral vector; and the non-viral vector comprises a cationic polymer, a plasmid vector and / or a liposome.

[0039] The application also provides a recombinant cell or a recombinant bacterium containing the fusion protein of the application or containing the gene of the application.

[0040] The cell can be a target cell to be edited.

[0041] The application also provides a single-gene base editing system comprising one or more of the fusion protein of the application, the gene, the recombinant plasmid, the recombinant cell or the recombinant bacterium and an sgRNA.

[0042] Specifically, the sgRNA guides the fusion protein to perform single-base gene editing on a target sequence in a target cell.

[0043] More specifically, the target sequence of the sgRNA comprises at least one of the sequences shown in SEQ ID NO. 14-28.

[0044] In view of the fact that the fusion protein of the application can improve the editing efficiency of single-base gene and adjust the base editing window, the application also claims the use of the fusion protein, the gene, the recombinant plasmid, the recombinant cell or the recombinant bacterium or the single-gene base editing system in the preparation of a gene editing product, a disease treatment and / or prevention product, an animal model or a plant variety.

[0045] Specifically, the gene editing includes but is not limited to base editing from A to G or C to T, which can be applied to edit the splice acceptor / donor site to regulate RNA splicing, and can also be used for model (such as disease model, cell model, animal model, etc.) construction or human disease treatment, etc.

[0046] The application further provides a method for improving the performance of a base editor, comprising the steps of introducing the fusion protein of any one of the above into a cell and performing gene editing on a target gene.

[0047] The application has the following beneficial effects:

[0048] The application has the following beneficial effects: BRIEF DESCRIPTION OF DRAWINGS

[0049] Figure 1 Figure 1 is a schematic diagram of an ABE system based on an R-loop binding domain (RHBD); the system components consist of nCas9, RHBD, TadA and sgRNA; the protein linker is represented by a brown curve, and the thick black line represents the DNA strand; the red ten character symbol marks the cutting position of nCas9 near the protospacer adjacent motif (PAM) site.

[0050] Figure 2 Figure 5 is a structural schematic diagram of the fusion expression of the R-loop binding domain (RHBD1) of human RNaseH1 and its mutant (mRHBD1) with ABEmax, and the editing efficiency detection results of the fusion protein of ABEmax and RHBD1 or mRHBD1 as a base editor; A and C in the figure are structural schematic diagrams of the fusion expression of RHBD1 or mRHBD1 with ABEmax, respectively; B and D in the figure are editing efficiency detection results of the fusion protein of ABEmax and RHBD1 or mRHBD1 as a base editor, respectively.

[0051] Figure 3Structure diagram of RHBD1 and ABE8e fusion and editing efficiency detection results of fusion proteins of R-loop binding domains of different sources and ABE8e as base editors; A in the figure is a structure diagram of RHBD1 and ABE8e fusion expression; B in the figure is the editing efficiency detection results of the fusion proteins obtained by inserting the R-loop binding domains of different sources into the N-terminus of the ABE8e editor for fusion expression as base editors; C in the figure is the editing efficiency detection results of the fusion proteins obtained by inserting the R-loop binding domains of different sources into the middle position of the ABE8e editor for fusion expression as base editors.

[0052] Figure 4 Editing performance detection results of the protein ABE8e-M-RHBD1 obtained by RHBD1 and ABE8e fusion as a base editor; A in the figure is the editing efficiency detection results of ABE8e-M-RHBD1 on different target bases; B in the figure is the base editing efficiency statistical results of ABE8e-M-RHBD1 on A10-A15 sites of different targets; C in the figure is the cytosine off-target editing activity detection results of ABE8e-M-RHBD1; D in the figure is the InDel frequency detection results of ABE8e-M-RHBD1.

[0053] Figure 5 Editing performance detection results of the protein ABE8e-M-RHBD1 obtained by RHBD1 and ABE8e fusion as a base editor at disease-related gene sites; A and B in the figure are the editing efficiency detection results of ABE8e-M-RHBD1 on different target bases; C and D in the figure are the editing frequency detection results of ABE8e-M-RHBD1 on the base mutation genotypes of different targets.

[0054] Figure 6 Structure diagram of RHBD1 and mRHBD1 fusion expression with BE4max and editing efficiency detection results of the fusion proteins of RHBD1 or mRHBD1 and BE4max as base editors; A in the figure is a structure diagram of RHBD1 or mRHBD1 and BE4max fusion expression; B in the figure is a schematic diagram of CBE-HEK293T-stop-EGFP reporter cell line for detecting CBE editing activity and editing window; C in the figure is the detection results of the CBE-HEK293T-stop-EGFP reporter cell line detecting the editing efficiency of the fusion proteins of RHBD1 or mRHBD1 and BE4max as base editors at different sites; D in the figure is the editing efficiency detection results of the N-terminal fusion proteins of RHBD1 or mRHBD1 and BE4max as base editors.

[0055] *P<0.05; **P<0.01; ***P<0.001; ****P<0.0001; ns, no significant difference. DETAILED DESCRIPTION

[0056] The present application will be further described by the following description of the drawings and specific examples, but the examples do not limit the present application in any form. Unless otherwise specified, the reagents, methods and equipment used in the present application are conventional reagents, methods and equipment in the technical field.

[0057] Unless otherwise specified, the reagents and materials used in the following examples are commercially available.

[0058] The synthetic sequences described in the embodiments of the present application are synthesized by GENEWIZ company.

[0059] Example 1 Effect of RHBD1 on the editing performance of ABEmax

[0060] The performance of base editors is affected by many factors, including substrate accessibility, intrinsic properties of deaminases, and intracellular environment. Among them, substrate accessibility is closely related to the ability of Cas protein to unwind double-stranded DNA to form an R-loop with sgRNA. After the formation of an R-loop structure between sgRNA and DNA target site, it provides an opportunity for deaminases to bind to single-stranded DNA substrates. Therefore, the editing efficiency is likely to be determined by the interaction between the substrate nucleotides in the R-loop and the deaminases. The present application accidentally found that the fusion of R-loop binding domain (RHBD) with base editor system can improve the editing efficiency or adjust the editing window of base editor. As an example, the schematic diagram of ABE system based on R-loop binding domain is shown in Figure 1 As can be seen from Figure 1 , the system components consist of nCas9, RHBD, TadA and sgRNA; the protein linker is represented by a brown curve, the thick black line represents the DNA strand; the red ten symbol marks the cutting position of nCas9 near the protospacer adjacent motif (PAM) site.

[0061] In this embodiment, the R-loop binding domain (RHBD1) of human (Homo sapiens) ribonuclease H1 (RNaseH1) is inserted into different positions of ABEmax base editor, and different ABEmax and RHBD1 fusion proteins are constructed and the editing performance of the obtained fusion proteins as base editors is detected. The present application also carries out inactivation mutation on RHBD1 (the obtained mutant is named mRHBD1), and also carries out fusion expression with ABEmax, and detects the editing performance of the obtained fusion protein as base editor.

[0062] The specific process is as follows:

[0063] 1. Construction of sgRNA expression plasmid and fusion protein recombinant expression plasmid

[0064] The artificially synthesized DNA sequence targeting endogenous gene site NIBAN1 (ACACACACACTTAGAATCTGTGG, as shown in SEQ ID NO. 14, the PAM sequence is shown in bold) and HEK2-site1 (GAACACAAAGCATAGACTGCGGG, as shown in SEQ ID NO. 15, the PAM sequence is shown in bold) were cloned into the BbsI cleavage site of pUC19-U6-sgRNA (SpCas9) vector (for expressing sgRNA of the corresponding target), obtaining sgRNA expression plasmids pUC19-U6-sgNIBAN1 and pUC19-U6-HEK2-site1.

[0065] ABEmax base editor plasmid (#112095) was purchased from Addgene, and the structure of ABEmax base editor is shown as A in Figure 2 On the basis of the plasmid, the DNA sequences encoding RHBD1 (the amino acid sequence of RHBD1 is shown as SEQ ID NO. 1) and mRHBD1 (compared with RHBD1, the 17th amino acid W is mutated to A (W17A), and the 33rd and 34th amino acids K are both mutated to A (K33A, K34A), as shown in C in Figure 2 Figure 2 The RHBD1 was inserted into the N-terminal of ABEmax base editor (the N-terminal of TadA, named ABEmax-N-RHBD1), the middle position (between TadA* and SpCas9 (D10A), named ABEmax-M-RHBD1) and the C-terminal of SpCas9 (D10A) (the C-terminal of SpCas9 (D10A), named ABEmax-C-RHBD1) for fusion expression using MultiF Seamless Assembly Mix (ABclonal); the mRHBD1 was inserted into the N-terminal of ABEmax base editor (the N-terminal of TadA, named ABEmax-N-mRHBD1), and four different recombinant expression plasmids of ABEmax and RHBD1 fusion protein were constructed, and the structure of RHBD1 or mRHBD1 fusion expression with ABEmax is shown as A and C in Figure 2 respectively; all plasmids were verified by Sanger sequencing.

[0066] 2. Cell transfection

[0067] HEK293T cells were plated in 24-well plates (15000 cells / well), and when the cell density was about 80%, the transfection operation was performed. 3 μL of polyethyleneimine (PEI) with a concentration of 1 mg / mL and 1 μg of plasmid (containing 750 ng of base editor expression plasmid and 250 ng of sgRNA expression plasmid) were mixed, diluted with opti-MEM transfection buffer, and then incubated at room temperature for 15 min, and then added to the cells for transfection; the medium was changed after 6 hours of transfection, and the cells were cultured in the cell incubator. Cells without transfection were set as the control group (None), and each experiment was repeated three times.

[0068] 3. Extraction of cell genomes and construction of amplicon sequencing library

[0069] After 72 hours of transfection, the genomic DNA of the transfected / non-transfected cells was extracted using the blood cell genomic kit (Axygen), and the specific method was referred to the instruction manual. The construction of the amplicon sequencing library was entrusted to GENEWIZ Company. 50-100 ng of extracted genomic DNA was used as a template, and specific primers containing a bridging sequence were used for the first round of PCR amplification of the target site. Then, 1 μL of the first round of PCR amplification product was used for the second round of PCR amplification using primers containing a bridging sequence and a barcode sequence. After the second round of PCR amplification product was subjected to end repair, A tailing reaction and sequencing adapter, the DNA sequencing library was constructed by magnetic bead purification and gel recovery, and high-throughput sequencing was performed using Illumina Hiseq 2500 PE150.

[0070] The specific primers containing a bridging sequence (the bridging sequence is shown in bold) are as follows:

[0071] Forward primer:

[0072] 5'-GGAGTGAGTACGGTGTGCGCCTTGTAAGTGTTGCTGTCC-3';

[0073] Reverse primer:

[0074] 5'-GAGTTGG ATGCTGGATGGTCTCGCCTGCAGAAAGGTAT-3'.

[0075] 4. Analysis of high-throughput sequencing results

[0076] First, the sequencing file was split using MATLAB software, then the base conversion ratio was analyzed using BE-Analyzer, and the editing efficiency at the target site was calculated; CRISPResso2 software was used to analyze the frequency of InDel in each sample and Graphpad Prism 8 was used for plotting and display.

[0077] The editing efficiency detection results of the fusion protein obtained by fusing ABEmax with RHBD1 as a base editor are shown in B of FIG. 6. Figure 2 Figure 2 As can be seen from B of FIG. 6, compared with the ABEmax group, the editing efficiency of ABEmax-N-RHBD1 at position A5 is slightly reduced, but the editing efficiency at positions A7 and A9 is increased by 1.6 times and 6.5 times, respectively, that is, RHBD1 can affect the editing efficiency of ABEmax within the editing window, and effectively improve the editing efficiency of the base (A9) outside the editing window of ABEmax, which shows that RHBD1 has the effect of widening the editing window of ABEmax. However, ABEmax-M-RHBD1 and ABEmax-C-RHBD1 cannot effectively improve the editing efficiency of ABEmax.

[0078] The editing efficiency detection results of the fusion protein obtained by fusing ABEmax with mRHBD1 as a base editor are shown in D of FIG. 7. Figure 2

[0079] Example 2 Influence of R-loop binding domains from different sources on editing performance of ABEmax

[0080] In addition to RHBD1, the present application also artificially synthesizes DNA sequences encoding R-loop binding domains from different sources, and respectively fuses and expresses them with ABE8e editor to test the editing efficiency of the obtained fusion protein as a base editor. The amino acid sequences of the R-loop binding domains from different sources used are shown in Table 1.

[0081] Table 1 R-loop binding domains from different sources

[0082]

[0083] The specific process is as follows:

[0084] 1. Construction of fusion protein recombinant expression plasmid

[0085] The ABE8e base editor plasmid (#138889) is purchased from Addgene, and the structure of the ABE8e base editor is as shown in FIG. 8. Figure 3 ​​ABE8e and R-loop binding domain fusion expression schematic diagram as shown in FIG. 1A, and the synthetic DNA sequences encoding each R-loop binding domain shown in Table 1 were inserted into the N-terminal (N-terminal of TadA*) and middle position (between TadA* and SpCas9(D10A), denoted as M) of ABE8e editor for fusion expression, respectively, to construct different recombinant expression plasmids of ABE8e and R-loop binding domain. The structure schematic diagram of ABE8e and R-loop binding domain fusion expression is shown in FIG. 1A. Figure 3 All plasmids were verified by Sanger sequencing.

[0086] 2. Cell transfection

[0087] The same as Example 1.

[0088] 3. Extraction of cell genome and construction of amplicon sequencing library

[0089] The same as Example 1, except that the specific primers containing bridging sequences (the bridging sequences are shown in bold) used are as follows:

[0090] Forward primer:

[0091] 5'-GGAGTGAGTACGGTGTGCAGGACGTCTGCCCAATATGT-3';

[0092] Reverse primer:

[0093] 5'-GAGTTGG ATGCTGGATGGAGCCCCATCTGTCAAACTGT-3'.

[0094] 4. Analysis of high-throughput sequencing results

[0095] The same as Example 1.

[0096] The editing efficiency detection results of the fusion proteins obtained by inserting R-loop binding domains of different sources into the N-terminal of ABE8e editor for fusion expression as base editors are shown in FIG. 1B; the editing efficiency detection results of the fusion proteins obtained by inserting R-loop binding domains of different sources into the middle position of ABE8e editor for fusion expression as base editors are shown in FIG. 1C. Figure 3 The editing efficiency detection results of the fusion proteins obtained by inserting R-loop binding domains of different sources into the N-terminal of ABE8e editor for fusion expression as base editors are shown in FIG. 1B; the editing efficiency detection results of the fusion proteins obtained by inserting R-loop binding domains of different sources into the middle position of ABE8e editor for fusion expression as base editors are shown in FIG. 1C. Figure 3As shown in FIG. 1, the R-loop binding domain can affect the editing activity and editing window of ABE8e, and the RHBD of different species has different effects on the editing activity of ABE8e, and the fusion method also affects the editing performance of the editor. Compared with ABE8e, the N-terminal fusion of RHBD slightly reduces the editing activity at the A3 position (except for ABE8e-N-RHBD2), and has little effect on the editing activity at the A5 position. The fusion proteins ABE8e-N-RHBD1, ABE8e-N-RHBD2 and ABE8e-N-RHBD6 significantly enhance the editing efficiency at the A8 and A9 positions. The conventional editing window of ABE8e is A2 to A11, and the M-terminal fusion of RHBD can effectively edit the bases outside the editing window (such as A12), which shows that the RHBD widens the editing window of ABE8e.

[0097] Example 3 Detection of base editing performance of ABE8e-M-RHBD1

[0098] In order to further characterize the editing efficiency of ABE8e-M-RHBD1 and its cytosine off-target editing activity, the present application selects eight target endogenous gene sites (referred to as target sites) containing no adenine at the distal end of PAM (the sequence of the target sites is shown in Table 2, and the part in bold in the sequence is the PAM sequence, and the high-throughput sequencing primers corresponding to the target sites are shown in Table 3). The experiment and result analysis are carried out according to the method described in Example 1, and the editing efficiency and InDel frequency of the fusion protein obtained in this example are analyzed.

[0099] Table 2 Eight target endogenous gene sites containing no adenine at the distal end of PAM

[0100] Target name Sequence (5’-3’) SEQ ID NO VEGFA_site2-g2 GCCCGCGCCCGGAGGCGGGGTGG 16 EGFR-g2 TCTCTCTGTCATAGGGACTCTGG 17 ABE_site16-g2 CTGTCCTTCAAACCTTGTCCAGG 18 PDCD1_site1-g2 GGTGCCGCTGTCATTGCGCCGGG 19 PPP1R12C_site4-g2 GCCCCTCTGAGGCTCCTGTGTGG 20 HEK293_site3-g2 CTGCTTCTCCAGCCCTGGCCTGG 21 HEK293_site4-g2 GGTGCTGTGTGACTACAGTGGGG 22 AAVS1-g2 CTGTCCCCTCCACCCCACAGTGG 23

[0101] Table 3 High-throughput sequencing primers for target sites

[0102]

[0103]

[0104] The results of the editing performance detection of the ABE8e-M-RHBD1 fusion protein of RHBD1 and ABE8e as a base editor are shown in Table 4. Figure 4 Table 4 ABE8e-M-RHBD1 fusion protein of RHBD1 and ABE8e as a base editor Figure 4 A in Table 4 is the base editing efficiency detection result of ABE8e-M-RHBD1 on different target sites; Figure 4 B in Table 4 is the base editing efficiency statistical result of ABE8e-M-RHBD1 on the A10-A15 sites of different target sites; Figure 4 C in Table 4 is the cytosine off-target editing activity detection result of ABE8e-M-RHBD1; Figure 4D in the figure represents the InDel frequency detection result of ABE8e-M-RHBD1.

[0105] Combination Figure 4 Analysis of the A-to-G editing efficiency of the fusion protein ABE8e-M-RHBD1 revealed that, except for HEK293_site3-g2, ABE8e-M-RHBD1 exhibited higher editing efficiency than ABE8e at all other tested endogenous gene sites, with increases ranging from 1.8 to 16.6 times. Figure 4 (A in the text). Statistical data shows that the A to G editing efficiency at sites A10–A15 increased by 3.6 times ( Figure 4 (B in the text). Importantly, the cytosine off-target editing activity of ABE8e-M-RHBD1 was significantly lower than that of ABE8e, with conversion rates at C4, C5, C6, and C7 positions reduced by 49%, 52%, 44%, and 56%, respectively. Figure 4 (C in the text). The above results indicate that RHBD1 can effectively reduce the off-target editing activity of ABE8e within the editing window of cytosine, while simultaneously improving its adenine editing efficiency. Furthermore, at the eight endogenous gene loci tested, the InDel frequency of ABE8e-M-RHBD1 was slightly lower than that of ABE8e (C ​​in the text). Figure 4 (D in the middle).

[0106] Example 4: Potential applications of ABE8e-M-RHBD1 in building disease models

[0107] To further characterize the application prospects of ABE8e-M-RHBD1 in constructing disease models, this invention selected two gene targets related to genetic diseases (hereinafter referred to as targets). The target sequences are shown in Table 4, with the bolded portion indicating the PAM sequence. The corresponding high-throughput sequencing primers are shown in Table 5. Specifically, the SERPINA1 c.1226T>C point mutation leads to α-1-antitrypsin deficiency; the CASR c.518T>C point mutation leads to familial hypocalcemic hypercalcemia. Experiments and results analysis were performed according to the method described in Example 1 to analyze the editing efficiency of the fusion protein obtained in this example.

[0108] Table 4. Two disease-related targeted endogenous gene loci.

[0109] Target name Sequence (5’-3’) SEQ ID NO SERPINA1 c.1226 ACCACTTTTCCCATGAAGAGGGG 24 CASR c.518 GTTGCTGAGGAGTCTGCTGGAGG 25

[0110] Table 5 lists the high-throughput sequencing primers used for the target sites.

[0111]

[0112] The results of the performance test of the RHBD1 and ABE8e fusion protein ABE8e-M-RHBD1 as a base editor on disease-related targets are as follows:Figure 5 As shown; Figure 5 A and B in the figure represent the results of the base editing efficiency of ABE8e-M-RHBD1 on different target sites. Figure 5 C and D in the figure represent the results of the base mutation genotype editing of ABE8e-M-RHBD1 at different target sites.

[0113] Combination Figure 5 Analysis of the editing efficiency from A to G revealed that ABE8e-M-RHBD1 outperformed ABE8e in editing target bases at both disease-related endogenous gene loci. For the target base A13 at the SERPINA1 c.1226 locus, ABE8e-M-RHBD1 demonstrated a 3.2-fold increase in editing efficiency compared to ABE8e, and also reduced the editing efficiency of the bystander base A4 by 38%. Figure 5 (A in the original text). For A13 of CASR c.518, the target base editing efficiency of ABE8e-M-RHBD1 is 1.8 times that of ABE8e ( ). Figure 5 (B in the text). At the SERPINA1 c.1226 site, compared with ABE8e, the proportion of cells treated with ABE8e-M-RHBD1 showing only A13 editing events increased by 5.7-fold, while the proportion of only A4 editing events decreased by 66%. Figure 5 (C in the text). At the CASR c.518 site, compared with ABE8e, the proportion of cells treated with ABE8e-M-RHBD1 showing only A11 editing events increased by 7.5-fold, while the proportion of only A8 editing events decreased by 92%. Figure 5 (D in the original text). These results demonstrate that RHBD1 significantly enhances the editing activity of ABE8e on target bases within the editing window at disease-related gene loci, while reducing the editing efficiency on non-target bases. These findings highlight the potential of RHBD1-based adenine base editors in constructing disease models.

[0114] Example 5: The impact of RHBD1 on BE4max editing performance

[0115] In this embodiment, RHBD1 was inserted into different positions of BE4max to construct a BE4max-RHBD1 fusion protein, and the editing performance of the obtained fusion protein as a base editor was detected by mEGFP fluorescent reporter cell line and high-throughput sequencing.

[0116] The specific process is as follows:

[0117] 1. Construction of the CBE-HEK293T-stop-EGFP reporter cell line for detecting CBE editing activity

[0118] HEK293T cells were seeded in 6-well plates, 24 hours after seeding, the cell confluence rate was about 60-80%, at this time, the lentivirus packaging plasmids containing 300 ng pMD2.G (Addgene, #12259), 900 ng psPAX2 (plasmid #12260) and 1200 ng pLenti-stop-EGFP-puro were transfected into HEK293T cells by PEI; the supernatant of the exchanged virus was collected 48 hours and 72 hours after transfection, filtered with a 0.45 μm membrane (Millipore), and stored at -80℃. After 48 hours of HEK293T cell infection with the collected supernatant, 1 μg / mL puromycin was used for screening for 7 days; monoclonal cells were picked up by limiting dilution, and after two weeks, the cells that grew well and expressed mGFP were amplified, which were CBE-HEK293T-stop-EGFP reporter cell lines.

[0119] 2. Construction of sgRNA expression plasmid and fusion protein recombinant expression plasmid

[0120] The DNA sequences targeting mEGFP sites (gGFP-4, gGFP-5, gGFP-6) were artificially synthesized, and the target sequences are shown in Table 6, and the bold part in the sequence is the PAM sequence. The DNA sequence of mEGFP is shown in SEQ ID NO. 29. The gGFP-4, gGFP-5 and gGFP-6 sequences were cloned into the Bbs I cleavage site of the pUC19-U6-sgRNA (SpCas9) vector (for expressing sgRNA of the corresponding target), obtaining sgRNA expression plasmids pUC19-U6-sgGFP-4, pUC19-U6-sgGFP-5 and pUC19-U6-sgGFP-6.

[0121] Table 6 Three sgRNAs targeting EGFP sites

[0122] Target name Sequence (5’-3’) SEQ ID NO gGFP-4 AACGGTGAGCAAGGGGGAGGAGG 26 gGFP-5 GATAACGGTGAGCAAGGGGGAGG 27 gGFP-6 TCGGGATAACGGTGAGCAAGGGG 28

[0123] The BE4max base editor plasmid (#112093) was purchased from Addgene, and the structure of the BE4max base editor is as follows Figure 6The structure of the fusion protein of RHBD1 and BE4max is shown in A of FIG. 1. RHBD1 was inserted into the N-terminal of BE4max base editor (N-terminal of APOBEC1, named BE4max-N-RHBD1), the middle position (between APOBEC1 and SpCas9 (D10A), named BE4max-M-RHBD1) and the C-terminal of SpCas9 (D10A) (C-terminal of SpCas9 (D10A), named BE4max-C-RHBD1) for fusion expression using MultiF Seamless Assembly Mix (ABclonal), and mRHBD1 was inserted into the N-terminal of BE4max base editor (N-terminal of TadA, named BE4max-N-mRHBD1), to construct four different recombinant expression plasmids of BE4max and RHBD1 fusion protein, and the structure of the fusion expression of RHBD1 or mRHBD1 and BE4max is shown in FIG. 1. Figure 6 A of FIG. 1; all plasmids were verified by Sanger sequencing.

[0124] 3. CBE-HEK293T-stop-EGFP reporter cell line was used to detect the editing activity of CBE

[0125] CBE-HEK293T-stop-EGFP cells were plated in a 24-well plate (15000 cells / well), and when the cell density was about 80%, transfection was performed. 3 μL of polyethyleneimine (PEI) with a concentration of 1 mg / mL and 1 μg of plasmid (containing 750 ng of cytosine base editor expression plasmid and 250 ng of gGFP expression plasmid) were mixed, diluted with opti-MEM transfection buffer, and then incubated at room temperature for 15 min, and then added to the cells for transfection; the medium was changed after 6 hours of transfection, and the cells were cultured in a cell incubator. Cells without transfection were set as a control group (None), and each experiment was repeated three times. 72 hours after transfection, the cells were collected and resuspended in PBS. The proportion of GFP positive cells was analyzed using a CytoFLEX flow cytometer (Beckman Coulter) equipped with a 488 nm laser and a FITC filter, and the proportion of GFP positive cells can reflect the editing efficiency of CBE at the site.

[0126] 4. The experiment and result analysis were carried out according to the method described in embodiment 1, and the editing efficiency of the fusion protein obtained in this embodiment at the endogenous gene site was analyzed.

[0127] The CBE-HEK293T-stop-EGFP reporter cell line was constructed to quickly evaluate the editing activity and editing window of BE4max. Figure 6B in FIG. 1 is a schematic diagram of the reporter cell line used to detect CBE editing activity and editing window. When BE4max successfully edits the mutated start codon ACG to normal ATG, the cell expresses GFP, thereby producing fluorescence. The present application designs three gGFPs targeting the bases at C3, C6 and C10 positions, respectively. BE4max and gGFP are transfected into CBE-HEK293T-stop-EGFP reporter cell lines, and GFP expression is used to evaluate the editing activity of BE4max.

[0128] The results of detecting the editing efficiency of the fusion protein of RHBD1 or mRHBD1 and BE4max as a base editor at different positions by the CBE-HEK293T-stop-EGFP reporter cell line are shown in FIG. C in Figure 6 As can be seen from the figure, compared with BE4max, the editing efficiency of BE4max-C-RHBD1 at the three positions is reduced. The efficiency of BE4max-M-RHBD1 at C10 position is improved, but the editing efficiency at C3 and C6 positions is reduced. BE4max-N-RHBD1 improves the editing efficiency at C3 and C10 positions, but the efficiency at C6 position is reduced. Unlike BE4max-N-RHBD1, the fusion of mRHBD1 and BE4max does not enhance the editing activity at C3 and C10 positions, and partially alleviates the reduction of efficiency at C6 position.

[0129] The results of detecting the editing efficiency of the fusion protein of RHBD1 or mRHBD1 and BE4max as a base editor at different positions by the CBE-HEK293T-stop-EGFP reporter cell line are shown in FIG. C in Figure 6 As can be seen from the figure, compared with BE4max, the editing efficiency of BE4max-C-RHBD1 at the three positions is reduced. The efficiency of BE4max-M-RHBD1 at C10 position is improved, but the editing efficiency at C3 and C6 positions is reduced. BE4max-N-RHBD1 improves the editing efficiency at C3 and C10 positions, but the efficiency at C6 position is reduced. Unlike BE4max-N-RHBD1, the fusion of mRHBD1 and BE4max does not enhance the editing activity at C3 and C10 positions, and partially alleviates the reduction of efficiency at C6 position.

[0130] The above embodiments are the preferred embodiments of the present application, but the embodiments of the present application are not limited by the above embodiments, and any changes, modifications, substitutions, combinations, simplifications made without departing from the spirit and principles of the present application shall be equivalent replacement modes, and shall be included in the protection scope of the present application.

Claims

1. A fusion protein, characterized in that, The fusion protein consists of an R-loop binding domain, a nucleoside deaminase, a nuclease and a nuclear localization signal, and the connection order is that the nucleoside deaminase is located at the N-terminal of the nuclease, the R-loop binding domain is located at the N-terminal of the nucleoside deaminase, and the nuclear localization signal is located at both ends of the fusion protein; wherein the amino acid sequence of the R-loop binding domain is shown as SEQ ID NO. 1; the nucleoside deaminase is adenosine deaminase TadA and TadA*, the TadA* is located at the N-terminal of the nuclease, and the TadA is located at the N-terminal of the TadA*, and the amino acid sequence of the TadA is shown as SEQ ID NO. 8, and the amino acid sequence of the TadA* is shown as SEQ ID NO. 9; the amino acid sequence of the nuclease is shown as SEQ ID NO. 11; and the amino acid sequence of the nuclear localization signal is shown as SEQ ID NO. 12; Or the fusion protein consists of an R-loop binding domain, a nucleoside deaminase, a nuclease, a nuclear localization signal and two copies of a uracil glycosylase inhibitor, and the connection order is that the nuclease is located at the N-terminal of the uracil glycosylase inhibitor, the nucleoside deaminase is located at the N-terminal of the nuclease, the R-loop binding domain is located at the N-terminal of the nucleoside deaminase, and the nuclear localization signal is located at both ends of the fusion protein; wherein the amino acid sequence of the R-loop binding domain is shown as SEQ ID NO. 1; the nucleoside deaminase is cytosine deaminase APOBEC1, and the amino acid sequence of the APOBEC1 is shown as SEQ ID NO. 10; the amino acid sequence of the nuclease is shown as SEQ ID NO. 11; the amino acid sequence of the nuclear localization signal is shown as SEQ ID NO. 12; and the amino acid sequence of the uracil glycosylase inhibitor is shown as SEQ ID NO. 13; Or the fusion protein consists of an R-loop binding domain, a nucleoside deaminase, a nuclease and a nuclear localization signal, and the connection order is that the nucleoside deaminase is located at the N-terminal of the nuclease, the R-loop binding domain is located at the N-terminal of the nucleoside deaminase, and the nuclear localization signal is located at both ends of the fusion protein; wherein the amino acid sequence of the R-loop binding domain is shown as SEQ ID NO. 1, 2 or 6; the nucleoside deaminase is adenosine deaminase TadA*, and the amino acid sequence of the TadA* is shown as SEQ ID NO. 9; the amino acid sequence of the nuclease is shown as SEQ ID NO. 11; and the amino acid sequence of the nuclear localization signal is shown as SEQ ID NO. 12; or the fusion protein consists of an R-loop binding domain, a nucleoside deaminase, a nuclease and a nuclear localization signal, and the connection order is: the nucleoside deaminase is located at the N-terminal of the nuclease, the R-loop binding domain is located in the middle of the nucleoside deaminase and the nuclease, and the nuclear localization signal is located at both ends of the fusion protein; wherein the amino acid sequence of the R-loop binding domain is any one of SEQ ID NO. 1-7; the nucleoside deaminase is adenosine deaminase TadA*, the amino acid sequence of the TadA* is shown in SEQ ID NO. 9; the amino acid sequence of the nuclease is shown in SEQ ID NO. 11; and the amino acid sequence of the nuclear localization signal is shown in SEQ ID NO.

12.

2. A gene, characterized in that, The gene encodes the fusion protein of claim 1.

3. A recombinant plasmid, characterized in that, The gene of claim 2.

4. A recombinant cell, characterized in that, The fusion protein of claim 1 or the gene of claim 2.

5. A recombinant bacterium, characterized in that, The fusion protein of claim 1 or the gene of claim 2.

6. A single gene base editing system, comprising: The fusion protein of claim 1 or the gene of claim 2. The fusion protein of claim 1, the gene of claim 2, the recombinant plasmid of claim 3, the recombinant cell of claim 4 or the recombinant bacteria of claim 5, and sgRNA; the sgRNA is used to guide the fusion protein to perform single-base gene editing on the target sequence in the target cell.

7. The fusion protein of claim 1, the gene of claim 2, the recombinant plasmid of claim 3, the recombinant cell of claim 4 or the recombinant bacteria of claim 5, the single-base gene editing system of claim 6 in the preparation of a gene editing product.