Small-molecule-targeted RNA recognition technique based on base editing
By using a base editor and tag protein in the fusion protein to identify the binding site of small molecule drugs to RNA, the off-target problem of small molecule drugs targeting RNA in existing technologies has been solved, and precise RNA targeting recognition and editing has been achieved.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- LINGANG LAB
- Filing Date
- 2025-08-04
- Publication Date
- 2026-07-23
AI Technical Summary
Existing small molecule drugs exhibit an off-target effect when targeting RNA, making it difficult to accurately identify and map mRNA binding sites, thus affecting treatment efficacy.
Develop a fusion protein comprising a base editor and a tag protein linked together by a linker. The base editor will be used to edit the binding sites of small molecule compounds to RNA, and the tag protein will be used for specific recognition to detect the target sites of small molecule drugs.
This technology enables precise identification of the binding sites of small molecule drugs and RNA, reducing off-target effects and improving the accuracy and efficiency of treatment.
Smart Images

Figure PCTCN2025112392-FTAPPB-I100001 
Figure PCTCN2025112392-FTAPPB-I100002 
Figure PCTCN2025112392-FTAPPB-I100003
Abstract
Description
Small molecule targeted RNA recognition technology based on base editing
[0001] This invention claims priority to patent application No. CN 202510082979.X, filed on January 20, 2025; the entire contents of which are incorporated herein by reference. Technical Field
[0002] This invention belongs to the field of biotechnology, and more specifically, it relates to a small molecule targeted RNA recognition technology based on base editing. Background Technology
[0003] Base editing is a revolutionary genome editing technology that allows scientists to perform precise single-base editing without creating double-strand breaks in DNA. Since its initial report in 2016, this technology has demonstrated enormous potential and application value in fields such as basic research, gene therapy, and plant and animal breeding improvement.
[0004] Base editors are mainly divided into two categories: cytosine base editors (CBEs) and adenine base editors (ABEs). CBEs can convert cytosine (C) in DNA to thymine (T), while ABEs can convert adenine (A) to guanine (G). These editors typically consist of a deaminase and a modified CRISPR / Cas9 system, in which the Cas9 protein loses its cleavage activity and is only used to recognize and bind to specific DNA sequences. In recent years, base editing technology has made continuous breakthroughs. For example, in 2023, Yang Hui's team developed a novel DNA base editor, AYBE, which achieved efficient adenine base transversion editing for the first time, which has significant potential value for the field of gene therapy. In addition, in early 2024, Bi Changhao and Zhang Xueli's team developed novel base editors DAF-CBE and DAF-TBE that do not rely on deaminases. These editors achieved C-to-G and T-to-G base transversion editing in E. coli and mammalian cells, providing a new direction for the development of base editing technology.
[0005] The development of base editing technology also faces several challenges, including reducing off-target effects, improving editing efficiency and accuracy, and developing new base editors to expand the types of bases that can be edited. To improve therapeutic efficacy, researchers are working to optimize the design of base editors to reduce non-target effects and increase editing efficiency. Furthermore, how to effectively deliver base editors to target tissues within the body is also an important research area. Despite these challenges, advances in base editing technology offer new hope for treating genetic diseases, cancer, and other ailments. With continued technological maturation and optimization, base editing is expected to play an even more significant role in the future of medicine and biotechnology.
[0006] SNAP tagging technology is an innovative protein engineering tool that allows scientists to precisely label and track specific proteins in living cells. At the heart of this technology is the SNAP tag, a mutant based on O6-alkylguanine-DNA alkyltransferase (AGT) that specifically undergoes an irreversible covalent reaction with benzylguanine (BG) derivatives. This property allows the SNAP tag to bind to various BG derivatives carrying different functional groups, such as fluorophores and biotin, thereby enabling the visual labeling and functional study of proteins.
[0007] The research background and application prospects of SNAP tagging technology are very broad. In the field of basic research, it is used to study the subcellular localization, dynamic behavior, protein-protein interactions, and the role of proteins in diseases. For example, by using SNAP tagging technology, researchers can observe protein behavior in living cells in real time, which is of great significance for understanding the mechanisms of protein action within cells and changes in protein behavior in disease states. Furthermore, SNAP tagging technology can also be used to develop novel biomedical diagnostic tools and treatment methods, such as improving the targeting of cancer treatment by tagging specific therapeutic agents or imaging agents. Technically, SNAP tag ligands are diverse, easily enabling both low-throughput and high-throughput protein-protein interaction studies. In addition, SNAP tagging technology can also be used to label fusion proteins in living cells; its ligands are small molecule compounds that easily enter cells and are non-toxic, making ligands with various fluorescent groups highly suitable for cell imaging.
[0008] Recent advances in SNAP tagging technology include the development of novel fluorescent probes that enable real-time monitoring of protein distribution within cells without washing, improving the convenience and real-time nature of experiments. Furthermore, SNAP tagging technology is also being used for site-specific conjugation of antibodies, equipping them with effector molecules to generate homogeneous antibody conjugates with customized properties, which plays a crucial role in medicine, particularly in cancer management. In summary, SNAP tagging technology, due to its specificity, speed, and versatility, demonstrates enormous potential and value in protein science research and biomedical applications. With continuous technological optimization and the development of new applications, SNAP tagging technology is expected to play an even more important role in the future.
[0009] The development of small molecule drugs targeting RNA is an emerging field full of challenges and opportunities. RNA, as a key molecule in the transmission of genetic information, plays a crucial role in various biological processes through its structure and function. By recognizing and binding to specific domains of RNA, small molecules can modulate RNA function, such as regulating gene expression, influencing RNA splicing, and stability. For example, small molecules targeting viral RNA can interfere with the viral infection process; the structures of HIV TAR and RRE RNAs have been studied for the development of antiviral therapies. Small molecules can also modulate RNA function by promoting RNA degradation, such as the ribonucleic acid-targeting chimera (RIBOTAC) technology, which recruits intracellular nucleases to specifically degrade target RNA. Traditional medicinal chemistry methods, structure-guided approaches, and modular assembly methods have been used to optimize lead compounds for RNA-targeting small molecules to improve their efficacy and selectivity. In this field, researchers are exploring new strategies and technologies to identify and optimize novel small molecules capable of targeting RNA structure and function. These studies not only contribute to a deeper understanding of the biological role of RNA but also provide new opportunities for the development of novel therapeutic drugs. With technological advancements and in-depth research, small molecule drugs targeting RNA are expected to play an even more important role in the future of medicine. However, determining whether existing small molecules have off-target effects and exploring the structural locations of the mRNAs they actually target remain significant challenges. Therefore, developing novel base-editing-based small molecule-targeted RNA recognition technologies for mapping small molecule-targeted mRNAs is crucial. Summary of the Invention
[0010] The purpose of this invention is to provide a small molecule targeted RNA recognition technology based on base editing.
[0011] In a first aspect of the invention, a fusion protein is provided, comprising a base editor and a tag protein.
[0012] In one or more embodiments, the base editor is an adenine base editor (ABE) or a cytosine base editor (CBE), wherein the adenine base editor or cytosine base editor also includes mutants thereof.
[0013] In one or more embodiments, the adenine base editor comprises: rABE, ABE7.10, ABE7.10(F148A), HuADAR2DD, HuADAR2DD(E488Q), tADAR1q, ABE7.7, ABE3.2, ABE5.3, ABE7.2, ABE6.3, ABE6.4, ABE7.8, ABE7.9, ABEMax, ABE8e, ABE8e-V106W, SaABE8e, SaKKH-ABE8e, NG-ABE 8e, ABE-xCas9, CP1028-ABE8e, ABE7.10-CP1041, CP1041-ABE8e, ABE8e-NRTH, ABE8e-NRRH, ABE8e-NRCH, NG-CP1041, ABE8e-VRQR-CP1041, ABE8e-VRQR, ABE8e-LbCasl2a(LbABE8e), ABE8e-AsCasl2a(enAsABE8e), ABE8e-SpyMac, ABE8e(TadA-8e V106W), ABE8e(K20A,R21A), ABE8e(TadA-8e V82G), or a combination of two or more of these.
[0014] In one or more embodiments, the amino acid sequence of rABE is shown in SEQ ID NO:1, the amino acid sequence of ABE7.10 is shown in SEQ ID NO:27, the amino acid sequence of ABE7.10(F148A) is shown in SEQ ID NO:28, the amino acid sequence of HuADAR2DD is shown in SEQ ID NO:29, the amino acid sequence of HuADAR2DD(E488Q) is shown in SEQ ID NO:30, and the amino acid sequence of tADAR1q is shown in SEQ ID NO:31.
[0015] In one or more embodiments, the mutant of the adenine base editor is the rABE mutant.
[0016] In one or more embodiments, the amino acid sequence of the rABE mutant, compared to the amino acid sequence shown in SEQ ID NO:1, has a mutation at any one or more of the following positions: positions 154, 23, 36, 48, 51, 106, 109, 111, 119, 122, 123, 146, 152, 157, and 166 of SEQ ID NO:1; more preferably, the mutation is an amino acid addition, amino acid deletion, or amino acid substitution.
[0017] In one or more embodiments, the amino acid at position 154 is substituted with Q, which is substituted with R, H, I, P or S, more preferably with Q (Q154R).
[0018] In one or more embodiments, the amino acid at position 23 is substituted with R, which is substituted with W, D, F, G, L, S, T or V, more preferably with R substituted with W (R23W).
[0019] In one or more embodiments, the amino acid at position 36 is substituted with H, I, P, R or S, more preferably with H (L36H).
[0020] In one or more embodiments, the amino acid at position 48 is substituted with P, H, I, R or S, more preferably with P (A48P).
[0021] In one or more embodiments, the amino acid at position 51 is substituted with L, R, H, I, P, or S, more preferably with R (L51R).
[0022] In one or more embodiments, the amino acid at position 106 is substituted with V, which is substituted with A, C, D, G, K, R, S or T, more preferably with V substituted with A (V106A).
[0023] In one or more embodiments, the amino acid at position 109 is substituted with S, D, F, G, L, T, V or W, more preferably with S (A109S).
[0024] In one or more embodiments, the amino acid at position 111 is substituted with T, R, H, I, P or S, more preferably with T (T111R).
[0025] In one or more embodiments, the amino acid at position 119 is substituted with D to N or Q, more preferably with D to N (D119N).
[0026] In one or more embodiments, the amino acid at position 122 is substituted with H to N or Q, more preferably with N (H122N).
[0027] In one or more embodiments, the amino acid at position 123 is substituted with Y, which is substituted with H, I, P, R or S, more preferably with H (Y123H).
[0028] In one or more embodiments, the amino acid at position 146 is substituted with C, A, D, G, K, R, T or V, more preferably with C (S146C).
[0029] In one or more embodiments, the amino acid at position 152 is substituted with P, or R, H, I, P, or S, more preferably with P substituted with R (P152R).
[0030] In one or more embodiments, the amino acid at position 157 is substituted with K, A, C, D, G, R, S, T or V, more preferably with K (N157K).
[0031] In one or more embodiments, the amino acid at position 166 is substituted with T, which is substituted with I, H, P, R or S, more preferably with T substituted with I (T166I).
[0032] In one or more embodiments, the rABE mutant comprises a protein having a substitution mutation at one or more of the following amino acid sites corresponding to the amino acid sequence shown in SEQ ID NO:1: R23W, L36H, A48P, L51R, V106A, A109S, T111R, D119N, H122N, Y123H, S146C, P152R, Q154R, N157K, and T166I.
[0033] In one or more embodiments, the amino acid sequence of the rABE mutant is shown in any one of SEQ ID NO: 2, 3, 4, 5, 7, 11, 12, 13, 14, 15, 16, 19, 20, 23 and 24.
[0034] In one or more embodiments, the tagged protein refers to a protein that can specifically bind to its substrate.
[0035] In one or more embodiments, the tag protein includes: SNAP-tag, CLIP-tag, HaloTag, TMP-tag, and ACP-tag.
[0036] In one or more embodiments, the tag protein is a SNAP-tag protein.
[0037] In one or more embodiments, the amino acid sequence of the tag protein is shown in SEQ ID NO:25.
[0038] In one or more embodiments, the fusion protein further comprises a linker; preferably, the linker comprises a linker containing G and S; more preferably, the amino acid sequence of the linker is shown in SEQ ID NO:26.
[0039] In one or more embodiments, the fusion protein is selected from:
[0040] (a) A protein with the amino acid sequence shown in SEQ ID NO:36;
[0041] (b) A protein derived from (a) that has the function of the protein in (a) and is formed by substitution, deletion or addition of one or more amino acid residues of the amino acid sequence shown in SEQ ID NO:36.
[0042] (c) A protein whose amino acid sequence is more than 80% identical to the amino acid sequence defined in (a) and which has the function of the protein in (a); or
[0043] (d) A fragment of SEQ ID NO:36 having the protein function of (a); preferably comprising a base editor and a tag protein; more preferably, the amino acid sequence of the base editor is as shown in SEQ ID NO:20, and the amino acid sequence of the tag protein is as shown in SEQ ID NO:25.
[0044] A second aspect of the present invention provides a polynucleotide comprising:
[0045] 1) A polynucleotide encoding the fusion protein described in any embodiment of the present invention; or
[0046] 2) A polynucleotide complementary to the polynucleotide described in 1).
[0047] A third aspect of the present invention provides an expression vector or host cell comprising the polynucleotides described in any embodiment of the present invention, or expressing the fusion protein described in any embodiment of the present invention.
[0048] A fourth aspect of the present invention provides a composition comprising:
[0049] 1) The fusion protein described in any embodiment of the present invention, or the polynucleotide described in any embodiment of the present invention, or the expression vector or host cell described in any embodiment of the present invention, wherein the fusion protein comprises a base editor and a tag protein; and
[0050] 2) A modified small molecule compound, wherein the modified small molecule compound comprises: the small molecule compound and the substrate of the tag protein described in 1).
[0051] In one or more embodiments, the tag protein is a SNAP-tag protein, and the substrate of the tag protein is benzylguanine (BG).
[0052] A fifth aspect of the present invention provides a kit comprising a fusion protein as described in any embodiment of the present invention, or a polynucleotide as described in any embodiment of the present invention, or an expression vector or host cell as described in any embodiment of the present invention, or a composition as described in any embodiment of the present invention.
[0053] A sixth aspect of the invention provides the application of the fusion protein described in any embodiment of the invention, or the polynucleotide described in any embodiment of the invention, or the expression vector or host cell described in any embodiment of the invention, or the composition described in any embodiment of the invention, or the kit described in any embodiment of the invention, selected from:
[0054] (1) Application in detecting or predicting the binding sites of small molecule compounds with nucleic acids;
[0055] (2) Application in screening nucleic acid sites that bind to small molecule compounds.
[0056] In one or more embodiments, the nucleic acid is DNA or RNA.
[0057] In one or more embodiments, the RNA is mRNA.
[0058] A seventh aspect of the present invention provides a method for detecting or predicting the binding site of a small molecule compound to a nucleic acid, the method comprising: contacting a fusion protein, a polynucleotide, an expression vector or host cell, a composition, or a kit with a modified small molecule compound as described in any embodiment of the present invention.
[0059] In one or more embodiments, the modified small molecule compound comprises a small molecule compound and a substrate of a tag protein, the substrate of which specifically binds to the tag protein in the fusion protein; preferably, the concentration of the modified small molecule compound is 0.001–1000 nM, more preferably 0.01–800 nM, 0.1–700 nM, 1–600 nM, 10–600 nM, 50–500 nM, or 200–400 nM.
[0060] In one or more embodiments, the contact includes simultaneous, sequential, or sequential contact.
[0061] In one or more embodiments, the method further includes: collecting nucleic acids before and after contact with the modified small molecule compound, detecting the nucleic acid sequence, and obtaining the nucleic acid editing site by comparing the nucleic acid sequences before and after contact with the modified small molecule compound, wherein the site is the binding site between the small molecule compound and the nucleic acid.
[0062] An eighth aspect of the present invention provides a method for screening nucleic acid sites that bind to small molecule compounds, the method comprising:
[0063] 1) Construct an expression vector, fuse the ABE editor and SNAP tag with the GS linker for expression, and then transiently transfect the plasmid into HEK293T cells via PEI. After 6 hours of transfection, replace with fresh culture medium.
[0064] 2) Add the modified small molecule compound to the system in 1), use the fusion protein to edit the nucleic acid bound to the modified small molecule compound, and collect the edited nucleic acid;
[0065] 3) Sequencing the nucleic acids in 1) and 2) respectively, comparing the sequences of the nucleic acids in 1) with those in 2) to obtain the sites where nucleic acid editing occurred. These editing sites are adjacent to the nucleic acid sites that bind to small molecule compounds.
[0066] Other aspects of the invention will be apparent to those skilled in the art from the disclosure herein. Attached Figure Description
[0067] Figure 1. Small molecule editing system used to screen for the optimal ABE editor for mRNA editing. Here, A1 represents the theoretical editing site (its editing can change TAG to TGG), and A2, A3, etc., indicate the base positions after A1. For example, A2 represents the position of the second A after the theoretical editing site; the A shown in the shaded area in Figure 8 is A2. The numbers presented in the figure represent the editing efficiency for sites A1 to A20, respectively. These values were obtained through analysis using EditR software (https: / / moriaritylab.shinyapps.io / editr_v10 / ).
[0068] Figure 2. Flow cytometry analysis of the editing ability of the mutant rABE on specific bases.
[0069] Figure 3 shows the optimal mutant rABE for mRNA editing in the BG-JQ1-mediated small molecule editing system. The numbers presented in the figure represent the editing efficiency of the sites described in Figure 8, obtained through analysis using EditR software (https: / / moriaritylab.shinyapps.io / editr_v10 / ).
[0070] Figure 4 shows the editing efficiency of srABE under the action of BG-JUN binder small molecules at 0 nM, 10 nM, 50 nM, 200 nM, 600 nM, and 1000 nM, respectively. The numbers presented in the figure represent the editing efficiency of the sites described in Figure 9, which were obtained through analysis using EditR software (https: / / moriaritylab.shinyapps.io / editr_v10 / ).
[0071] Figure 5 shows the editing efficiency after 12h, 24h, 36h, or 48h of editing with 200nM BG-JUN binder added together with the plasmid, with no BG-JUN binder added (final concentration 0nM) as a control. The numbers shown in the figure represent the editing efficiency of the sites described in Figure 9, which were obtained through analysis using EditR software (https: / / moriaritylab.shinyapps.io / editr_v10 / ).
[0072] Figure 6 shows the editing efficiency of specific sites after plasmid transfection and expression for 0h, 6h, 12h, 18h, or 24h, with the addition of 200nM BG-JUN binder for 48h of editing. The control group (without BG-JUN binder, final concentration 0nM) uses this method. The numbers presented in the figure represent the editing efficiency of the sites described in Figure 9, obtained through analysis using EditR software (https: / / moriaritylab.shinyapps.io / editr_v10 / ).
[0073] Figure 7. Schematic diagram of small molecule targeted RNA recognition technology based on base editing. This technology includes a fusion protein of an adenine base editor (ABE) and a SNAP protein, and a small molecule compound modified with benzylguanine (BG). The fusion protein is expressed by transfecting plasmids, and a BG-modified small molecule drug is added exogenously. The adenine base editor (ABE) in the fusion protein is used to edit the bases of the RNA bound to the small molecule. Cells are harvested, total RNA is extracted, reverse transcribed, and then PCR is performed using specific primers. The PCR product is then subjected to NGS sequencing. By comparing the RNA sequence information before and after editing, the location of the small molecule drug binding to the mRNA can be detected.
[0074] Figure 8. Editing sites of the BG-JQ1-mediated small molecule editing system. The shaded area A is the editing site.
[0075] Figure 9. Editing sites of the BG-JUN binder-mediated small molecule editing system. The shaded area A represents the editing site. Detailed Implementation
[0076] Through in-depth research, the inventors have developed a method for recognizing small molecule targeted RNA based on base editing. They discovered that linking a base editor (BE) to a tag protein (e.g., SNAP-tag) via a linker creates a fusion protein that specifically recognizes small molecule compounds coupled to the tag protein substrate (e.g., BG). By using the base editor (BE) to edit specific bases, it is possible to detect the binding sites of small molecule compounds to target nucleic acids (e.g., RNA), determining whether small molecule drugs exhibit off-target effects, demonstrating significant application value.
[0077] Fusion protein
[0078] This invention provides a fusion protein comprising a base editor and a tag protein. In some embodiments, the fusion protein further includes a linker that connects the base editor and the tag protein together.
[0079] 1. Base Editor
[0080] In this invention, the term "base editor" generally refers to a protein capable of base editing a target sequence. The base editing includes: base editing from adenine (A) to guanine (G) and / or base editing from cytosine (C) to thymine (T).
[0081] The base editor that edits adenine (A) to guanine (G) is often called an adenine base editor (ABE). An adenine base editor (ABE) typically contains an adenosine deaminase domain, which deaminates adenine from adenosine residues in nucleic acids (DNA or RNA) and converts it to inosine residues (which typically pair with cytosine residues), resulting in a point mutation. This process is known as adenine (A) to guanine (G) base editing. In this invention, the ABE includes, but is not limited to: rABE, ABE7.10, ABE7.10(F148A), HuADAR2DD, HuADAR2DD(E488Q), tADAR1q, ABE7.7, ABE3.2, ABE5.3, ABE7.2, ABE6.3, ABE6.4, ABE7.8, ABE7.9, ABEMax, ABE8e, ABE8e, ABE8e-V106W, SaABE8e, SaKKH-ABE8e, NG-ABE8e, ABE-xCas9, CP1028-ABE8e, ABE7.10-CP1041, CP1041-ABE8e, ABE8e-NRTH, ABE8e-NRRH, ABE8e-NRCH, NG-CP1041, ABE8e-VRQR-CP1041, ABE8e-VRQR, ABE8e-LbCasl2a(LbABE8e), ABE8e-AsCasl2a(enAsABE8e), ABE8e-SpyMac, ABE8e(TadA-8e V106W), ABE8e(K20A, R21A), ABE8e(TadA-8e V82G), or combinations of two or more of these.
[0082] A base editor that edits cytosine (C) to thymine (T) is commonly called a cytosine base editor (CBE). A cytosine base editor (CBE) typically contains a cytosine deaminase domain, which deaminates cytosine from cytidine residues in nucleic acids (DNA or RNA), converting them to thymine residues, thus causing a point mutation. This process is known as cytosine (C) to thymine (T) base editing. In this invention, the CBE includes, but is not limited to, ApoBEC1, ApoBEC3G, hA3A, eA3A, etc., or combinations of two or more of these.
[0083] In this invention, the adenine base editor or cytosine base editor also includes its mutants. The "mutant" refers to a protein formed by substituting, deleting, and / or adding one or more amino acids of the adenine base editor or cytosine base editor. Substitution means replacing an amino acid occupying a position with a different amino acid. Deletion means removing an amino acid occupying a position. Addition (or insertion) means adding an amino acid adjacent to and immediately following the amino acid occupying a position.
[0084] In some embodiments, the adenine base editor is an rABE mutant, i.e., a protein obtained by mutagenesis based on the amino acid sequence of rABE (shown in SEQ ID NO:1). In some embodiments, the rABE mutant of the present invention has at least 80%, at least 82%, at least 84%, at least 86%, at least 88%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO:1. Sequence identity is used herein to describe the correlation between two amino acid sequences or two nucleotide sequences. Sequence identity can be calculated using methods known in the art. For example, the Needleman-Wunsch algorithm (Needleman and Wunsch, 1970, Journal of Molecular Biology, 48:443-453) implemented in the Needle program of the EMBOSS package (EMBOSS: European Open Software Suite for Molecular Biology, Rice et al., 2000, Trends in Genetics, 16:276-277) can be used to determine sequence identity between two amino acid sequences. Alternatively, BLASTP on NCBI can be used to calculate sequence identity between two amino acid sequences.
[0085] In some embodiments, the mutation in the rABE mutant is a substitution mutation. In some embodiments, the mutation includes a mutation occurring at any one or more positions selected from positions 23, 36, 48, 51, 106, 109, 111, 119, 122, 123, 146, 152, 154, 157, and 166 of SEQ ID NO:1. Compared to the wild-type SEQ ID NO:1, these rABE mutants exhibit superior base editing capabilities.
[0086] In some embodiments, the amino acid at position 23 of SEQ ID NO:1 is substituted with R, which is substituted with W, D, F, G, L, S, T, or V, preferably with R substituted with W (R23W). In some embodiments, the amino acid at position 36 of SEQ ID NO:1 is substituted with L, which is substituted with H, I, P, R, or S, preferably with L substituted with H (L36H). In some embodiments, the amino acid at position 48 of SEQ ID NO:1 is substituted with A, which is substituted with P, H, I, R, or S, preferably with A substituted with P (A48P). In some embodiments, the amino acid at position 51 of SEQ ID NO:1 is substituted with L, which is substituted with R, H, I, P, or S, preferably with L substituted with R (L51R). In some embodiments, the amino acid at position 106 of SEQ ID NO:1 is substituted with V, which is substituted with A, C, D, G, K, R, S, or T, preferably with V substituted with A (V106A). In some embodiments, the amino acid at position 109 of SEQ ID NO:1 is substituted with S, D, F, G, L, T, V, or W, preferably with S (A109S). In some embodiments, the amino acid at position 111 of SEQ ID NO:1 is substituted with R, H, I, P, or S, preferably with R (T111R). In some embodiments, the amino acid at position 119 of SEQ ID NO:1 is substituted with N or Q, preferably with N (D119N). In some embodiments, the amino acid at position 122 of SEQ ID NO:1 is substituted with N or Q, preferably with N (H122N). In some embodiments, the amino acid at position 123 of SEQ ID NO:1 is substituted with H, I, P, R, or S, preferably with H (Y123H). In some embodiments, the amino acid at position 146 of SEQ ID NO:1 is substituted with S, which is substituted with C, A, D, G, K, R, T, or V, preferably S is substituted with C (S146C). In some embodiments, the amino acid at position 152 of SEQ ID NO:1 is substituted with P, which is substituted with R, H, I, P, or S, preferably P is substituted with R (P152R). In some embodiments, the amino acid at position 154 of SEQ ID NO:1 is substituted with Q, which is substituted with R, H, I, P, or S, preferably Q is substituted with R (Q154R). In some embodiments, the amino acid at position 157 of SEQ ID NO:1 is substituted with N, which is substituted with K, A, C, D, G, R, S, T, or V, preferably N is substituted with K (N157K). In some embodiments, the amino acid at position 166 of SEQ ID NO:1 is substituted with T, which is substituted with I, H, P, R, or S, preferably T is substituted with I (T166I).
[0087] In some specific embodiments, the rABE mutant comprises a protein having a substitution mutation at one or more of the following amino acid sites corresponding to the amino acid sequence shown in SEQ ID NO:1: R23W, L36H, A48P, L51R, V106A, A109S, T111R, D119N, H122N, Y123H, S146C, P152R, Q154R, N157K, and T166I.
[0088] In some specific embodiments, the amino acid sequence of the rABE mutant is shown in any one of SEQ ID NO:2-24. In some more specific embodiments, the amino acid sequence of the rABE mutant is shown in any one of SEQ ID NO:2,3,4,5,7,11,12,13,14,15,16,19,20,23 and24.
[0089] In some embodiments, the adenine base editor is selected from: rABE or its mutants, ABE7.10, ABE7.10(F148A), HuADAR2DD, HuADAR2DD(E488Q), or tADAR1q. In some specific embodiments, the amino acid sequence of the adenine base editor is shown in any one of SEQ ID NO: 1, 2, 3, 4, 5, 7, 11, 12, 13, 14, 15, 16, 19, 20, 23, 24, 27, 28, 29, 30, and 31. In some preferred embodiments, the amino acid sequence of the adenine base editor is shown in SEQ ID NO: 20.
[0090] 2. Tag proteins and their substrates
[0091] In this invention, the term "tag protein" refers to a protein capable of specifically binding to its substrate, thereby labeling a target protein. A detectable marker can be used to couple the tag protein to its substrate, and the presence and quantity of the target protein can be detected by utilizing the specific binding of the tag protein to the substrate. The "detectable marker" includes, but is not limited to, fluorescent groups, biotin, radionuclides, enzymes, quantum dots, colloidal gold, etc. More specifically, the detectable marker can be a fluorescent protein, such as green fluorescent protein (Green), EGFP, red fluorescent protein, rhodamine, etc.
[0092] In this invention, the tag proteins include, but are not limited to, SNAP-tag, CLIP-tag, HaloTag, TMP-tag, and ACP-tag. SNAP-tag refers to a protein that specifically binds to its substrate benzylguanine (BG) or its derivatives. CLIP-tag refers to a protein that specifically binds to its substrate benzylcytosine (BC) or its derivatives. HaloTag refers to a protein that specifically binds to its substrate chloroalkane. TMP-tag refers to a protein that specifically binds to its substrate trimethoprim (TMP). ACP-tag refers to a protein that specifically binds to its substrate acyl carrier protein (ACP).
[0093] In some embodiments, the tag protein is a SNAP-tag, the substrate of which is benzylguanine (BG) or a derivative thereof. The cysteine residue in the SNAP-tag protein, serving as the reaction site, can nucleophilically attack O. 6 The benzyl group of modified benzylguanine, after the guanine structure is removed, allows cysteine to form a stable thioether covalent bond with the benzyl group, thereby achieving specific binding of the SNAP-tag to BG or its derivatives. In some specific embodiments, the tag protein is a SNAP-tag protein, preferably a SNAP-tag protein with an amino acid sequence as shown in SEQ ID NO:25.
[0094] In this invention, the term "BG derivative" generally refers to a substance derived from, containing, or coupled with a benzylguanine (BG) group. In this invention, the BG derivative includes, but is not limited to: BG-JQ1, BG-JUN binder, BG-MYC, BG-Lev, etc.
[0095] In this invention, the benzylguanine (BG), also known as O-6-benzylguanine, has the CAS number 19916-73-5. The benzylguanine (BG) group typically refers to a group having the following structure:
[0096] In some embodiments, the BG derivative is a small molecule compound coupled with a benzylguanine (BG) group. In this invention, "small molecule compound" generally refers to an organic compound with a molecular weight between several hundred and several thousand Daltons (Da). Typically, the small molecule compounds in this invention bind to nucleic acids (DNA or RNA) when they function. Exemplary small molecule compounds include, but are not limited to, JQ1, JUN-binder, MYC, and Lev. The benzylguanine (BG) group can be directly coupled to the small molecule compound, or it can be coupled to the small molecule compound via a linker.
[0097] 3. Linker
[0098] In the fusion protein described in this invention, the linker connects the base editor to the tag protein. The linker can be 2–18, 3–27, or 5–45 amino acids long. The linker can be a flexible linker. The linker can be a linker containing both G and S, such as, but not limited to, (GS). n (GSG) n (GGGGS) n More specifically, the amino acid sequence of the linker may be as shown in SEQ ID NO:26.
[0099] In some embodiments, the fusion protein comprises an adenine base editor and a tag protein. In some embodiments, the fusion protein comprises an adenine base editor, a tag protein, and a linker. In some embodiments, the fusion protein consists of an adenine base editor, a tag protein, and a linker.
[0100] In some specific embodiments, the amino acid sequence of the adenine base editor is shown in any one of SEQ ID NO: 1, 2, 3, 4, 5, 7, 11, 12, 13, 14, 15, 16, 19, 20, 23, 24, 27, 28, 29, 30 and 31.
[0101] In some specific embodiments, the tag protein is a SNAP-tag protein, preferably a SNAP-tag protein with the amino acid sequence shown in SEQ ID NO:25. In some specific embodiments, the linker has the amino acid sequence shown in SEQ ID NO:26. In some more specific embodiments, the fusion protein has the amino acid sequence shown in SEQ ID NO:36.
[0102] Nucleic acid, expression vector, host cell
[0103] The present invention also provides isolated nucleic acids encoding the fusion protein, which may be its complementary strand. The DNA sequence encoding the fusion protein of the present invention can be synthesized artificially in its entirety, or the DNA sequences encoding TAT and LPTS amino acids can be obtained separately by PCR amplification and then spliced together to form the DNA sequence encoding the fusion protein of the present invention.
[0104] After obtaining the DNA sequence encoding the fusion protein of the present invention, it is ligated into a suitable expression vector and then transformed into a suitable host cell. Therefore, the present invention also provides a vector comprising a nucleic acid molecule encoding the fusion protein. The vector may also contain an expression regulatory sequence operatively linked to the sequence of the nucleic acid molecule to facilitate the expression of the fusion protein. A variety of suitable vectors can be used, such as those used for cloning and expression in bacterial, fungal, yeast, and mammalian cells. Various vectors known in the art, such as commercially available vectors, can be used. For example, a commercially available vector can be used, and the nucleotide sequence encoding the novel fusion protein of the present invention can be operatively linked to the expression regulatory sequence to form a protein expression vector. The expression vector may also contain elements such as promoters (e.g., pCMV), terminators (e.g., bGH polyA), and enhancers. It should be understood that the present invention is not limited to the expression vectors defined in the embodiments.
[0105] Furthermore, recombinant cells containing nucleic acid sequences encoding the fusion protein are also included in this invention. In this invention, the term "host cell" includes both prokaryotic and eukaryotic cells. Commonly used prokaryotic host cells include *Escherichia coli*, *Bacillus subtilis*, etc. Commonly used eukaryotic host cells include yeast cells, insect cells, and mammalian cells.
[0106] Preparation methods of fusion proteins
[0107] The present invention also provides a method for producing the fusion protein described herein, the method comprising: culturing the host cell described herein under conditions suitable for expressing the fusion protein, expressing and isolating the fusion protein.
[0108] Specifically, the method may include:
[0109] 1) Provide the nucleic acid sequence encoding the fusion protein;
[0110] 2) Insert the nucleic acid sequence from 1) into a suitable expression vector to obtain a recombinant expression vector;
[0111] 3) Introduce the recombinant expression vector from 2) into a suitable host cell;
[0112] 4) Culture and transform host cells under suitable expression conditions;
[0113] 5) Collect the supernatant and purify the fusion protein product.
[0114] The coding sequence can be introduced into host cells using a variety of known techniques in the art, such as, but not limited to: calcium phosphate precipitation, protoplast fusion, liposome transfection, electroporation, microinjection, reverse transcription, phage transduction, and alkali metal ion transfection.
[0115] For information on host cell culture and expression, see Olander RM Dev Biol Stand 1996; 86: 338. Cells and residues in the suspension can be removed by centrifugation, and the supernatant can be collected. Identification can be performed using agarose gel electrophoresis.
[0116] The fusion protein obtained through the above preparation can be purified to achieve essentially homogeneous properties, such as appearing as a single band on SDS-PAGE electrophoresis. For example, when the recombinant protein is secreted, a commercially available ultrafiltration membrane can be used to separate the protein, and the expression supernatant can be concentrated first. The concentrate can be further purified by gel chromatography or by ion exchange chromatography, such as anion exchange chromatography (DEAE, etc.) or cation exchange chromatography. The gel matrix can be agarose, dextran, polyamide, or other commonly used protein purification matrices. Q- or SP- groups are ideal ion exchange groups. Finally, the purified product can be further refined using methods such as hydroxyapatite adsorption chromatography, metal chelate chromatography, hydrophobic interaction chromatography, and reversed-phase high-performance liquid chromatography (RP-HPLC). All the above purification steps can be combined in different ways to ultimately achieve essentially homogeneous purity of the fusion protein.
[0117] Composition
[0118] The present invention also provides a composition comprising: 1) a fusion protein or its nucleic acid, expression vector or host cell, wherein the fusion protein comprises a base editor and a tag protein; and 2) a modified small molecule compound comprising a substrate of the small molecule compound and the tag protein of 1).
[0119] In this invention, the "modified small molecule compound" generally refers to a small molecule compound linked to the tag protein substrate described in 1). The linking can be direct, such as covalent or non-covalent bonding, or indirect, such as via a linker, through conjugation or coupling. The linker is typically a flexible linker. The linker may contain polyethylene glycol.
[0120] It should be understood that the composition may not contain sgRNA. The specific binding of the tag protein to its substrate allows the fusion protein to accumulate near the modified small molecule compound. The composition may also contain sgRNA, for example, sgRNA modified with BG, to further enhance the targeting of the composition. In some preferred embodiments, the composition does not contain sgRNA.
[0121] In some embodiments, the composition comprises: 1) a fusion protein or its nucleic acid, expression vector or host cell, wherein the fusion protein comprises a base editor (preferably an adenine base editor) and a SNAP-tag protein; and 2) a modified small molecule compound comprising a small molecule compound and benzylguanine (BG) or a derivative thereof.
[0122] Since the modified small molecule compound contains benzylguanine (BG) or its derivatives, the specific binding of benzylguanine (BG) or its derivatives to SNAP-tag proteins allows fusion proteins containing SNAP-tag proteins to accumulate near the binding site of the small molecule compound. Using a base editor, the nucleic acids bound to the small molecule compound can be genetically edited. By comparing the sequence information before and after base editing, the location information of the nucleic acids bound to the small molecule compound can be determined.
[0123] In some embodiments, the composition comprises: 1) a fusion protein or its nucleic acid, expression vector or host cell, wherein the fusion protein comprises a base editor, a SNAP-tag protein and a linker; and 2) a modified small molecule compound comprising a small molecule compound and benzylguanine (BG) or a derivative thereof.
[0124] In the composition, the fusion protein or its nucleic acid, expression vector or host cell can be the fusion protein or its nucleic acid, expression vector or host cell described in any embodiment of the present invention.
[0125] In the composition, the small molecule compound can be the small molecule compound described in any embodiment of the present invention.
[0126] Reagent test kit
[0127] The present invention also provides a kit for detecting the binding site of small molecule compounds to nucleic acids, comprising: a fusion protein as described in any embodiment of the present invention, or a nucleic acid as described in any embodiment of the present invention, or an expression vector or host cell as described in any embodiment of the present invention, or a composition as described in any embodiment of the present invention.
[0128] Applications and methods
[0129] The present invention also provides the application of the fusion protein or its nucleic acid, expression vector or host cell, or the composition, or the kit, for detecting or predicting the binding sites of small molecule compounds with nucleic acids, or for screening nucleic acid sites that bind to small molecule compounds.
[0130] The present invention also provides a method for detecting or predicting the binding site of a small molecule compound to a nucleic acid, the method comprising the step of contacting the fusion protein or its nucleic acid, expression vector or host cell, or the kit of the present invention with the modified small molecule compound.
[0131] In this invention, the nucleic acid can be DNA or RNA. Specifically, the RNA can include mRNA.
[0132] In this invention, the contact can be simultaneous, sequential, or in turn.
[0133] In some embodiments, the fusion protein, its nucleic acid, or expression vector described in this invention can be simultaneously added to cells along with the modified small molecule compound, and the fusion protein can be used to edit the nucleic acid bound to the small molecule compound. In this method, the editing time of the nucleic acid bound to the small molecule compound by the fusion protein can be 0–48 h, for example, 6–36 h, 12–24 h, or 12–18 h.
[0134] In some embodiments, the modified small molecule compound can also be added to cells expressing or containing the fusion protein or its nucleic acid or expression vector of the present invention, or added to host cells containing the fusion protein of the present invention, so that the fusion protein can edit the nucleic acid bound to the small molecule compound. In this method, the nucleic acid or expression vector containing the fusion protein of the present invention can first be transferred into cells to express the fusion protein, and then the modified small molecule compound can be added to the cells. In this method, the expression time of the fusion protein can be 0–24 h, for example 6–24 h, 12–24 h, or 12–18 h, and the editing time of the nucleic acid bound to the small molecule compound by the fusion protein can be 0–48 h, for example 6–36 h, 12–24 h, or 12–18 h.
[0135] In the method, the concentration of the modified small molecule compound can be 0.001 to 1000 nM, for example 0.01 to 800 nM, 0.1 to 700 nM, 1 to 600 nM, 10 to 600 nM, 50 to 500 nM or 200 to 400 nM.
[0136] This invention also provides a method for screening nucleic acid sites that bind to small molecule compounds, the method comprising:
[0137] 1) Provide a system capable of expressing the fusion protein described in any embodiment of the present invention, and collect the nucleic acids in the system;
[0138] 2) Add the modified small molecule compound to the system in 1), use the fusion protein to edit the nucleic acid bound to the modified small molecule compound, and collect the edited nucleic acid;
[0139] 3) Sequencing the nucleic acids in 1) and 2) respectively, comparing the sequences of the nucleic acids in 1) with those in 2) to obtain the sites where nucleic acid editing occurred. These sites are adjacent to the nucleic acid sites that bind to small molecule compounds.
[0140] Compared with the prior art, the advantages of the present invention include:
[0141] 1. It can detect the target sites of small molecule compounds binding to mRNA within cells.
[0142] 2. Highly efficient and convenient, accurately identifying potential mRNA binding sites for small molecule compounds.
[0143] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. Experimental methods in the following embodiments that do not specify specific conditions are generally performed according to conventional conditions such as those described in J. Sambrook et al., Molecular Cloning: A Laboratory Manual, 4th Edition, Science Press, or according to the manufacturer's recommendations.
[0144] Example 1: Construction of a small molecule-mediated mRNA editing system
[0145] 1.1 Carrier Construction
[0146] Effect plasmids pCMV-SNAP-rABE, pCMV-SNAP-ABE7.10, pCMV-SNAP-ABE7.10(F148A), pCMV-SNAP-HuADAR2DD, pCMV-SNAP-HuADAR2DD(E488Q), and pCMV-SNAP-tADAR1q, as well as binding plasmid pCMV-BRD4-BD-Lamda N-NES, and reporter plasmids pCMV-Green-TAG-Boxb-mCherry and pCMV-EGFP-TAG-Boxb-mCherry vectors were constructed.
[0147] 1.1.1 Effect plasmids
[0148] The pCMV-SNAP-rABE effector plasmid consists of the following components from N to C: promoter CMV, SNAP (SEQ ID NO:25), linker (SEQ ID NO:26), rABE (SEQ ID NO:1), and terminator bGH poly(A).
[0149] The pCMV-SNAP-ABE7.10 effector plasmid, from N-terminus to C-terminus, consists of: promoter CMV, SNAP (SEQ ID NO:25), linker (SEQ ID NO:26), ABE7.10 (SEQ ID NO:27), and terminator bGH poly(A).
[0150] The pCMV-SNAP-ABE7.10(F148A) effector plasmid consists of the following components from N-terminus to C-terminus: promoter CMV, SNAP (SEQ ID NO:25), linker (SEQ ID NO:26), ABE7.10(F148A)(SEQ ID NO:28), and terminator bGH poly(A).
[0151] The pCMV-SNAP-HuADAR2DD effector plasmid consists of the following components from N to C: promoter CMV, SNAP (SEQ ID NO:25), linker (SEQ ID NO:26), HuADAR2DD (SEQ ID NO:29), and terminator bGH poly(A).
[0152] The pCMV-SNAP-HuADAR2DD(E488Q) effector plasmid consists of the following components from N-terminus to C-terminus: promoter CMV, SNAP (SEQ ID NO:25), linker (SEQ ID NO:26), HuADAR2DD(E488Q) (SEQ ID NO:30), and terminator bGH poly(A).
[0153] The pCMV-SNAP-tADAR1q effector plasmid consists of the following components from N to C: promoter CMV, SNAP (SEQ ID NO:25), linker (SEQ ID NO:26), tADAR1q (SEQ ID NO:31), and terminator bGH poly(A).
[0154] 1.1.2 Binding plasmids
[0155] The pCMV-BRD4-BD-Lamda N-NES binding plasmid consists of, from N-terminus to C-terminus: promoter CMV, BRD4-BD-Lamda N-NES sequence (SEQ ID NO:32), and terminator bGH poly(A).
[0156] 1.1.3 Reporting plasmids
[0157] The pCMV-Green-TAG-Boxb-mCherry reporter plasmid, from N-terminus to C-terminus, consists of: promoter CMV, Green protein sequence (SEQ ID NO:33), TAG-Boxb-mCherry sequence (SEQ ID NO:34), and terminator bGH poly(A).
[0158] The pCMV-EGFP-TAG-Boxb-mCherry reporter plasmid, from N-terminus to C-terminus, consists of: promoter CMV, EGFP protein sequence (SEQ ID NO:35), TAG-Boxb-mCherry sequence (SEQ ID NO:34), and terminator bGH poly(A).
[0159] All components in the vector were synthesized by Shanghai Qingke Biotechnology. Following the experimental methods for vector construction in the fourth edition of *Molecular Cloning*, plasmids were constructed via homologous recombination. Positive clones were selected for sequencing. After confirming correct sequencing results, plasmids were extracted using the OMEGA endotoxin-free plasmid extraction kit (OMEGA, D6228-01) and used for subsequent cell transfection.
[0160] 1.2 Chemical Synthesis of BG-JQ1 Small Molecule
[0161] BG-JQ1 crafting route:
[0162] Reaction conditions: (a) HATU, DMF, DIPEA, 0℃-rt; (b) TFA, DCM, rt; (c) HATU, DIPEA, DMF, 0℃-rt
[0163] Synthesis method:
[0164] In a dried 100 mL two-necked flask, JQ-1 (1.00 g, 2.50 mmol) was added sequentially, dissolved in 25 mL of anhydrous DMF, followed by NH2-PEG3-COO-tBu (0.83 g, 3.00 mmol) and HATU (1.43 g, 3.75 mmol). Under argon protection and an ice bath, DIPEA (0.49 g, 3.75 mmol) was diluted with 10 mL of DMF and added to the reaction solution. The solution turned pale green, and the reaction was carried out overnight at room temperature. After the reaction was completed by TLC, DMF was removed by vacuum distillation, the solution was diluted with water, and extracted with EA. The EA layer was washed sequentially with saturated brine, dried over anhydrous sodium sulfate, and ethyl acetate was removed by vacuum distillation. The residue was purified by silica gel column chromatography to give 1.20 g of intermediate 2, a white solid, in 72.8% yield.
[0165] Intermediate 2 (1.10 g, 1.67 mmol) was added to a 100 mL round-bottom flask and dissolved in 50 mL of DCM. 10 mL of TFA was added at room temperature, and the mixture was stirred for 2 h at room temperature. After the reaction was complete as detected by TLC, saturated sodium bicarbonate solution was added to adjust the pH to 7-8. DCM was then added for extraction, and the DCM layer was washed sequentially with saturated brine, dried over anhydrous sodium sulfate, and concentrated to obtain 0.96 g of intermediate 3, a white, foamy solid, which was directly proceeded to the next step without purification.
[0166] In a dried 100 mL two-necked flask, intermediate 3 (0.92 g, 1.53 mmol) was added sequentially, dissolved in 30 mL of anhydrous DMF, followed by 6-((4-(aminomethyl)benzyl)oxy)-7H-purine-2-amine (0.50 g, 1.84 mmol) and HATU (0.88 g, 2.30 mmol). Under argon protection and an ice bath, DIPEA (0.30 g, 2.30 mmol) was diluted with 5 mL of DMF and added to the reaction solution. The ice bath was removed, and the reaction was allowed to proceed overnight at room temperature. After the reaction was completed by TLC, DMF was removed under reduced pressure, the mixture was diluted with water, extracted with EA, and the EA layer was washed sequentially with saturated brine, dried over anhydrous sodium sulfate, and ethyl acetate was removed under reduced pressure. The residue was purified by silica gel column chromatography to give 1.16 g of JQ-Peg3-BG, a white solid, in 88.6% yield. 11H NMR (400 MHz, DMSO-d6) δ 8.35 (t, J = 5.9 Hz, 1H), 8.29 (t, J = 5.6 Hz, 1H), 7.87 (s, 1H), 7.48 (d, J = 8.4 Hz, 2H), 7.45 (s, 1H), 7.42 (d, J = 8.7 Hz, 3H), 7.26 (d, J = 7.9 Hz, 2H), 6.33 (s, 2H), 5.45 (s, 2H), 4.51 (dd, J = 8.0, 6.1 Hz, 1H), 4.27 (d, J = 5.9 Hz, 2H), 3.63 (t, J = 6.4 Hz, 2H), 3.50 (d, J = 6.9 Hz, 9H), 3.44 (t, J = 5.9 Hz, 4H), 3.29–3.21 (m, 5H), 2.58 (s, 3H), 2.39 (s, 3H), 2.37 (d, J = 6.4 Hz, 2H), 1.61 (s, 3H). 13 13C NMR (151 MHz, DMSO-d6) δ 170.15, 169.72, 163.04, 159.56, 155.13, 149.84, 139.40, 136.77, 135.24, 135.14, 132.28, 130.72, 130.17, 129.84, 129.57, 128.50, 128.48, 127.23, 69.78, 69.74, 69.63, 69.56, 69.21, 66.88, 66.61, 53.84, 41.83, 38.65, 37.52, 36.15, 14.07, 12.69, 11.31.
[0167] 1.3 Cell transfection
[0168] The effector plasmid, pCMV-BRD4-BD-Lamda N-NES, and the reporter plasmid were selected and co-transfected at a mass ratio of 1:1:1, with a total plasmid amount of 2 μg per well. Transfection was performed using 1 mg / mL polyethyleneimine (PEI), with the PEI volume being twice the plasmid mass (i.e., 2 μg plasmid plus 4 μL PEI). The transfection procedure was as follows: the plasmid was placed in 50 μL of antibiotic- and serum-free DMEM medium (mixed by pipetting), and PEI was simultaneously added to 50 μL of antibiotic- and serum-free DMEM medium. Immediately afterward, the plasmid was added to the PEI solution over 5 minutes, mixed, and allowed to stand for 15 minutes before being dropwise into 293T cells with 85% confluence. The effect plasmid pCMV-SNAP-ABE7.10, along with the binding plasmid pCMV-BRD4-BD-Lamda N-NES and the reporter plasmid pCMV-Green-TAG-Boxb-mCherry, were mixed at a mass ratio of 1:1:1, with a total volume of 2 μg, and transfected. Only the effect plasmid needs to be changed to screen for the optimal small molecule-mediated mRNA editor. This system utilizes BG-JQ1 to bind to the BRD4-BD protein, and the fusion protein Lamda N to bind to the Boxb mRNA sequence ggccctgaaaaagggcc. Through BG-JQ1 mediation, the ABE editor reaches the vicinity of the Boxb sequence, thereby editing the A bases near the Boxb sequence.
[0169] 1.4 Screening of rABE using BG-JQ1 small molecule
[0170] After 18 hours of expression, the medium was changed to complete DMEM (Gibco, C11995500BT) containing BG-JQ1 (final concentration 600 nM). After 48 hours of editing, RNA was extracted using a kit (Promega, LS1040). A control was used without BG-JQ1 (final concentration 0 nM). Subsequently, reverse transcription was performed, followed by PCR using specific primers. The PCR products were then subjected to NGS sequencing.
[0171] Sequencing revealed that the final small molecule-mediated mRNA editing system was the effector plasmid pCMV-SNAP-rABE.
[0172] 1.5 Analysis of Screening Results
[0173] The results of screening rABE using BG-JQ1 small molecules are shown in Figure 1. The results show that rABE has low editing efficiency when treated with 0 nM BG-JQ1, and rABE exhibits the highest editing ability (29%) for A2 (the position of the second A after the theoretical editing site) when treated with 600 nM BG-JQ1. Therefore, this study selected rABE as a template for subsequent evolution.
[0174] Example 2: Flow cytometry analysis of the editing ability of the mutant rABE on specific bases.
[0175] In this embodiment, point mutations were performed on rABE, and the editing effect of the mutant rABE was evaluated by flow cytometry analysis.
[0176] 2.1. Selecting different point mutation sites for plasmid construction
[0177] The selected point mutation sites included R23W, L36H, A48P, L51R, F84L, V106A, V106W, N108Q, N108D, A109S, T111R, D119N, H122N, Y123H, S146C, F148A, F149Y, P152R, Q154R, V155E, F156I, N157K, and T166I. The sequence of the point-mutated rABE protein is shown in SEQ ID NO:2-24. The pCMV-Lamda N-mutation rABE plasmid was constructed using the point-mutated rABE. The specific vector construction procedure was referenced in the fourth edition of *Molecular Cloning*.
[0178] The plasmid construction method is as follows: primers with corresponding mutation sites are designed, and high-fidelity PCR is performed using the original plasmid pCMV-SNAP-rABE as a template. The PCR product is then purified and digested with DpnI. After digestion, the plasmid is transformed, and positive clones are selected for sequencing. Once the sequencing results are correct, the plasmid can be extracted using an endotoxin-free plasmid extraction kit (OMEGA, D6228-01) and used for cell transfection in subsequent steps.
[0179] 2.2. Select 293T cells for transfection
[0180] 293T cells were cultured and seeded in 24-well plates for transfection. The pCMV-Lamda N-mutation rABE plasmid and the pCMV-EGFP-TAG-Boxb-mCherry plasmid were co-transfected at a mass ratio of 1:1, with a total plasmid amount of 2 μg per well. PEI (1 mg / mL) was used for transfection, with the PEI amount being twice the plasmid mass (2 μg plasmid plus 4 μL PEI). The transfection procedure was as follows: the plasmid was placed in 50 μL of antibiotic- and serum-free DMEM medium (mixed by pipetting), and PEI was simultaneously added to 50 μL of antibiotic- and serum-free DMEM medium. Immediately after 5 minutes, the plasmid was added to the PEI solution, mixed, and allowed to stand for 15 minutes before being dropwise into 293T cells with 85% confluence.
[0181] 2.3. Flow cytometry analysis of red fluorescence expression levels to detect editing efficiency.
[0182] After transfection, cells were edited for 48 hours, followed by cell digestion and flow cytometry analysis. Editing ability was indirectly assessed by analyzing the expression level of red fluorescence (Figure 2). The effector plasmid pCMV-Lamda N-rABE (R23W) and reporter plasmid pCMV-EGFP-TAG-Boxb-mCherry were transfected at a mass ratio of 1:1, with a total amount of 2 μg. After 48 hours, the proportion of red fluorescence was analyzed by flow cytometry to determine the editing ability of the mutant rABE. A higher proportion of red fluorescence indicates stronger rABE editing activity. Green fluorescence represents transfection efficiency.
[0183] Example 3: Screening of the mutant rABE using a small molecule editing system
[0184] In this embodiment, the pCMV-SNAP-mutation rABE plasmid (point mutation sites include R23W, L36H, A48P, L51R, F84L, V106A, V106W, N108Q, N108D, A109S, T111R, D119N, H122N, Y123H, S146C, F148A, F149Y, P152R, Q154R, V155E, F156I, N157K, T166I, and the point mutation rABE protein sequence is shown in SEQ ID NO:2-24) obtained in Example 2 was used as the effector plasmid. 293T cells were transfected using the same method as in Example 1, and screening was performed using BG-JQ1 small molecule (final concentration of 600 nM).
[0185] The results of screening the mutant rABE using the BG-JQ1 small molecule are shown in Figure 3. The results showed that, based on statistical analysis of specific A molecules, rABE(Q154R) exhibited an editing efficiency of 0 nM BG-JQ1 treatment, while reaching 29% efficiency at 600 nM BG-JQ1 treatment, making it the most active mutant among the selected mutants. Therefore, the rABE(Q154R) mutation was chosen for further research and named srABE. Sequencing revealed that the optimal small molecule-mediated mutant mRNA editing system was pCMV-SNAP-Linker-rABE(Q154R), and the sequence of this fusion protein is shown in SEQ ID NO:36.
[0186] Example 4: Editing efficiency of srABE under different concentrations of BG-JUN treatment
[0187] 4.1 Constructing the srABE vector
[0188] First, we analyzed the JUN-binder small molecule (refer to Nature. 2023 Jun; 618(7963):169-179.), and found that its targeted mRNA sequence is ggcaguauaguccgaacugcaaaucuuauuuucuuuucaccuucucucuaacugcc.
[0189] 4.2 Chemical Synthesis of BG-JUN Small Molecule
[0190] BG-JUN small molecule synthetic route:
[0191] Reaction conditions: (a) THF, DIPEA, 60℃, 2h; (b) THF, DIPEA, 60℃, overnight; (c) MeOH, NaOH, H2O, Reflux; (d) HATU, DMF, DIPEA, 0℃-rt ; (e) TFA, DCM, rt; (c) HATU, DIPEA, DMF, 0℃-rt (f) HATU, DMF, DIPEA, 0℃-rt; (g) S8, Ethylenediamine, 100℃, 2h.
[0192] Synthesis method:
[0193] In a 100 mL single-necked flask, trichlorazine (4.00 g, 21.6 mmol) was added sequentially, dissolved in 35 mL of THF, followed by p-aminobenzonitrile (6.40 g, 52.8 mmol) and DIPEA (12.8 mL). The mixture was heated at 60 °C for 2 h. After the reaction was completed by TLC, THF was removed under reduced pressure, the mixture was diluted with water, extracted with EA, and the EA layer was washed sequentially with saturated brine, dried over anhydrous sodium sulfate, and ethyl acetate was removed under reduced pressure. The residue was purified by silica gel column chromatography to give 6.55 g of intermediate J1, a white solid, in 87.4% yield. ¹H NMR (400 MHz, DMSO-d6) δ 10.79 (s, 2H), 7.84 (d, J = 7.8 Hz, 8H).
[0194] In a 100 mL single-necked flask, J1 (2.40 g, 6.91 mmol) was added sequentially, dissolved in 35 mL of THF, followed by methyl 3-aminopropionate hydrochloride (1.20 g, 8.52 mmol) and DIPEA (2.13 g, 16.50 mmol). The mixture was heated to 60 °C overnight. After the reaction was completed by TLC, THF was removed under reduced pressure, the mixture was diluted with water, extracted with EA, and the EA layer was washed sequentially with saturated brine, dried over anhydrous sodium sulfate, and ethyl acetate was removed under reduced pressure. The residue was purified by silica gel column chromatography to give 2.60 g of intermediate J2, a white solid, in 90.8% yield. 1H NMR (400MHz, DMSO-d6) δ9.95 (d, J = 13.3Hz, 2H), 8.02 (d, J = 8.5Hz, 4H), 7.75 (d, J = 8.5Hz, 5H), 3.65 (s, 3H), 3.56 (t, J = 7.2Hz, 2H), 2.62 (t, J = 7.1Hz, 2H).
[0195] In a 100 mL single-necked flask, J2 (2.40 g, 5.80 mmol), 20 mL MeOH, 2 mL water, and sodium hydroxide (463 mg, 11.58 mmol) were added sequentially, and the mixture was refluxed for 5 h. After the reaction was completed by TLC, 1 M dilute hydrochloric acid was added to adjust the pH to 3-4, and a solid precipitated. The solid was filtered, the filter cake was washed, and dried to give 2.51 g of grayish-white solid intermediate J3, with a yield of 92.3%. ¹H NMR (400 MHz, DMSO-d6) δ 9.98 (d, J = 13.3 Hz, 2H), 8.01 (d, J = 8.4 Hz, 4H), 7.73 (d, J = 8.4 Hz, 5H), 3.55 (t, J = 7.1 Hz, 2H), 2.58 (t, J = 7.1 Hz, 2H).
[0196] In a dried 100 mL two-necked flask, J3 (1.50 g, 3.75 mmol) was added sequentially, dissolved in 20 mL of DMF, followed by NH2-PEG3-COO-tBu (1.04 g, 3.75 mmol) and HATU (1.79 g, 4.69 mmol). Under argon protection and an ice bath, DIPEA (0.62 g, 4.69 mmol) was diluted with 2 mL of DMF and added to the reaction solution. The solution turned pale green, and the reaction was carried out overnight at room temperature. After the reaction was completed by TLC, DMF was removed by vacuum distillation, the solution was diluted with water, and extracted with EA. The EA layer was washed sequentially with saturated brine, dried over anhydrous sodium sulfate, and ethyl acetate was removed by vacuum distillation. The residue was purified by silica gel column chromatography to give 2.06 g of intermediate J4, a white foamy solid, with a yield of 82.3%. 1H NMR (400MHz, DMSO-d6) δ9.80(s,1H),9.69(s,1H),8.03(d,J=7.4Hz,4H),7.95(t,J=5.7Hz,1H),7.71(d,J=8.4Hz,4H),7.42(s,1H),3 .55(t,J=6.3Hz,4H), 3.47(d,J=8.8Hz,9H), 3.40(t,J=5.9Hz,3H), 3.21(q,J=5.8Hz,2H), 2.41(dt,J=15.7,6.7Hz,4H), 1.37(s,9H).
[0197] Intermediate J4 (1.22 g, 1.85 mmol) was added to a 100 mL round-bottom flask and dissolved in 50 mL of DCM. 10 mL of TFA was added at room temperature, and the mixture was stirred for 2 h at room temperature. After the reaction was complete as detected by TLC, saturated sodium bicarbonate solution was added to adjust the pH to 7-8. DCM was then added for extraction, and the DCM layer was washed sequentially with saturated brine, dried over anhydrous sodium sulfate, and concentrated to obtain 1.05 g of intermediate J5, a white, foamy solid, which was directly proceeded to the next step without purification. 1H NMR (400MHz, DMSO-d6) δ9.83(s,1H),9.72(s,1H),8.03(d,J=8.4Hz,4H),7.96(t,J=5.6Hz,1H),7.71(d,J=8.4Hz,4H),7.46(s ,1H),3.61–3.51(m,4H),3.48(s,8H),3.40(d,J=11.7Hz,2H),3.21(q,J=5.8Hz,2H),2.42(td,J=6.6,3.5Hz,4H),1.90(s,2H).
[0198] In a dried 100 mL two-necked flask, intermediate J5 (0.93 g, 1.53 mmol) was added sequentially to dissolve in 15 mL of anhydrous DMF, followed by 6-((4-(aminomethyl)benzyl)oxy)-7H-purine-2-amine (0.50 g, 1.84 mmol) and HATU (0.88 g, 2.30 mmol). Under argon protection and an ice bath, DIPEA (0.30 g, 2.30 mmol) was diluted with 2 mL of DMF and added to the reaction solution. The ice bath was removed, and the reaction was allowed to proceed overnight at room temperature. After the reaction was completed by TLC, DMF was removed under reduced pressure, the mixture was diluted with water, and extracted with EA. The EA layer was washed sequentially with saturated brine, dried over anhydrous sodium sulfate, and ethyl acetate was removed under reduced pressure. The residue was purified by silica gel column chromatography to give 0.95 g of compound J6 as a white solid, with a yield of 72.3%. 1H NMR (400MHz, DMSO-d6) δ12.42(s,1H),9.81(s,1H),9.69(s,1H),8.35(t,J=5.9Hz,1H),8.03(d,J=8.4Hz,4 H),7.96(q,J=4.5,3.5Hz,1H),7.82(s,1H),7.71(d,J=8.4Hz,4H),7.43(d,J=7.7Hz,3H),7.25(d,J=7.8Hz ,2H),6.27(s,2H),5.44(s,2H),4.26(d,J=5.9Hz,2H),3.61(t,J=6.4Hz,2H),3.53(q,J=6.9Hz,2H),3.46( d, J=2.2Hz, 8H), 3.39 (t, J=5.9Hz, 4H), 3.20 (q, J=5.8Hz, 2H), 2.43 (t, J=7.1Hz, 2H), 2.36 (t, J=6.4Hz, 2H).
[0199] Intermediate J6 (0.32 g, 0.37 mmol) was added to a dried 50 mL pressure-resistant flask, dissolved in 10 mL of ethylenediamine, and sulfur (80 mg, 2.50 mmol). The mixture was reacted at 100 °C for 2 h, cooled to room temperature, and the reaction solution was poured into ice water and stirred. The pH was adjusted to 6-7 with 1 M dilute hydrochloric acid, and a solid precipitated. The solid was filtered, washed with water, dried, and separated by silica gel column chromatography to obtain 77 mg of gray solid Jun-PEG3-BG, with a yield of 22.3%. NMR (400MHz, DMSO-d6) δ9.49(s,1H),9.36(s,1H),8.35(dt,J=11.4,6.0Hz,1H),7.97(t,J=5.7Hz,1H),7.90(d,J=9.3Hz,4H),7.8 2(s,1H),7.74(dd,J=9.1,3.0Hz,4H),7.69(qd,J=5.4,2.9Hz,2H),7.43(d,J=7.8Hz,2H),7.25(d,J=7.9Hz,2H),7.20(dd,J=14.9 ,7.3Hz,1H),6.27(s,1H),4.25(dd,J=10.3,5.8Hz,2H),4.18–4.09(m,3H),3.63(s,10H),3.60(d,J=6.4Hz,7H),3.54(q,J=6.7Hz,9H),3.38(q,J=6.2Hz,5H),3.21(q,J=5.9Hz,3H),2.43(t,J=7.2Hz,2H),2.36(t,J=6.4Hz,3H).LC-MS(ESI):m / z:942.40[M+H]+. The final synthesized BG-JUN binder structure is:
[0200] 4.3 293T cell transfection
[0201] pCMV-SNAP-srABE was transfected, with a total plasmid volume of 2 μg per well. Transfection was performed using PEI (1 mg / mL), with the PEI volume being twice the plasmid mass (2 μg plasmid plus 4 μL PEI). The transfection procedure was as follows: the plasmid was placed in 50 μL of antibiotic- and serum-free DMEM medium (mixed by pipetting), and PEI was simultaneously added to 50 μL of antibiotic- and serum-free DMEM medium. Immediately afterward, the plasmid was added to the PEI solution over 5 minutes, mixed, and allowed to stand for 15 minutes before being dropwise into 293T cells with 85% confluence.
[0202] 4.4 Screening for the optimal BG-JUN small molecule concentration
[0203] After transfection and expression for 18 hours, the sample was replaced with complete DMEM containing BG-JUN binder molecules at wavelengths of 0 nm, 10 nm, 50 nm, 200 nm, 600 nm, and 1000 nm. RNA extraction was performed 48 hours after editing using a kit. Subsequently, reverse transcription was performed, followed by PCR using specific primers. The PCR products were then subjected to NGS sequencing.
[0204] 4.5 Screening Results
[0205] The results of screening using BG-JUN small molecules are shown in Figure 4. The results show that the editing efficiency reached 37 and 34 under treatment with 50 nM and 200 nM BG-JUN binder small molecules, respectively, indicating that both concentrations can induce the strongest editing effect of srABE. Considering the metabolic damage of small molecules in cells, 200 nM was selected as the treatment concentration for 293T cells.
[0206] Example 5: Determining the optimal editing level of srABE on the JUN-binder target within cells.
[0207] 5.1 Small molecules were added together with plasmids to explore the effect of editing duration on editing specific sites.
[0208] Before transfection, the plasmid was replaced with a small molecule containing the appropriate concentration of BG-JUN binder (200 nM), followed by the addition of the plasmid and PEI complex. Editing times were set at 12 h, 24 h, 36 h, and 48 h. After the editing time was completed, RNA extraction was performed using an RNA kit (Promega, LS1040). A control group without BG-JUN binder (final concentration 0 nM) was used. Subsequently, reverse transcription was performed, followed by PCR using specific primers, and the PCR products were sequenced using NGS. The results are shown in Figure 5. It can be observed that the overall editing effect gradually increased with the increase in editing time, and the control group also showed a corresponding trend. rABE (Q154R) showed the same results at 36 h as rABE at 48 h.
[0209] 5.2 After plasmid transfection, expression is preferentially performed for a certain duration. Subsequently, the effect of adding small molecules to the editing duration on the editing of specific sites is explored.
[0210] After transfection of the plasmid with the PEI complex, expression was performed for a certain period of time at time gradients of 0h, 6h, 12h, 18h, and 24h. After expression, the expression was replaced with a small molecule containing the corresponding concentration of BG-JUN binder (200nM), followed by another 48h of editing. After the editing time was completed, RNA extraction was performed using an RNA kit (Promega, LS1040). A control was used with no BG-JUN binder (final concentration 0nM). Subsequently, reverse transcription was performed, followed by PCR using specific primers, and the PCR products were sequenced by NGS. The results are shown in Figure 6. With the increase of the total editing time, both rABE and rABE(Q154R) showed an increasing trend, with the most significant difference between the two at 6h. rABE edited to 20, and rABE(Q154R) edited to 29.
[0211] Combining the results presented in Figures 5 and 6, the optimal results were achieved when the pCMV-SNAP-rABE(Q154R) plasmid was preferentially expressed for 6 hours after transfection into the cells, followed by replacement with a small molecule containing the corresponding concentration of BG-JUN binder (200 nM), and then edited for another 30 hours, for a total editing time of 36 hours.
[0212] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims. Furthermore, all documents mentioned in this invention are incorporated herein by reference as if each document were individually incorporated by reference.
[0213] sequence
Claims
1. A fusion protein comprising a base editor and a tag protein.
2. The fusion protein as described in claim 1, characterized in that, The base editor is an adenine base editor or a cytosine base editor, wherein the adenine base editor or cytosine base editor also includes its mutants; Preferably, the adenine base editor comprises: rABE, ABE7.10, ABE7.10(F148A), HuADAR2DD, HuADAR2DD(E488Q), tADAR1q, ABE7.7, ABE3.2, ABE5.3, ABE7.2, ABE6.3, ABE6.4, ABE7.8, ABE7.9, ABEMax, ABE8e, ABE8e, ABE8e-V106W, SaABE8e, SaKKH-ABE8e, NG-ABE8e, ABE-xCas9, CP1028-ABE 8e, ABE7.10-CP1041, CP1041-ABE8e, ABE8e-NRTH, ABE8e-NRRH, ABE8e-NRCH, NG-CP1041, ABE8e-VRQR-CP1041, ABE8e-VRQR, ABE8e-LbCasl2a(LbABE8e), ABE8e-AsCasl2a(enAsABE8e), ABE8e-SpyMac, ABE8e(TadA-8eV106W), ABE8e(K20A,R21A), ABE8e(TadA-8e V82G), or a combination of two or more of these; More preferably, the amino acid sequence of rABE is shown in SEQ ID NO:1, the amino acid sequence of ABE7.10 is shown in SEQ ID NO:27, the amino acid sequence of ABE7.10(F148A) is shown in SEQ ID NO:28, the amino acid sequence of HuADAR2DD is shown in SEQ ID NO:29, the amino acid sequence of HuADAR2DD(E488Q) is shown in SEQ ID NO:30, and the amino acid sequence of tADAR1q is shown in SEQ ID NO:
31.
3. The fusion protein as described in claim 2, characterized in that, The mutant of the adenine base editor is the rABE mutant; Preferably, the amino acid sequence of the rABE mutant, compared to the amino acid sequence shown in SEQ ID NO:1, has a mutation at any one or more of the following positions: positions 154, 23, 36, 48, 51, 106, 109, 111, 119, 122, 123, 146, 152, 157, and 166; more preferably, the mutation is an amino acid addition, amino acid deletion, or amino acid substitution. More preferably, The amino acid at position 154 is substituted with R, H, I, P or S, more preferably with R (Q154R). The amino acid at position 23 is substituted with R, which can be W, D, F, G, L, S, T or V, more preferably with R substituted with W (R23W); The amino acid at position 36 is substituted with H, I, P, R or S, more preferably with H (L36H). The amino acid at position 48 is substituted with P, H, I, R or S, more preferably with P (A48P). The amino acid at position 51 is substituted with R, H, I, P or S, more preferably with R (L51R). The amino acid at position 106 is substituted with V, which can be A, C, D, G, K, R, S or T, more preferably V is substituted with A (V106A); The amino acid at position 109 is substituted with S, D, F, G, L, T, V or W, more preferably with S (A109S). The amino acid at position 111 is substituted with T, which is substituted with R, H, I, P or S, more preferably with T substituted with R (T111R); The amino acid at position 119 is substituted with N or Q, more preferably with N (D119N); The amino acid at position 122 is substituted with H to N or Q, more preferably with N (H122N); The amino acid at position 123 is substituted with Y, which can be H, I, P, R or S, more preferably with H (Y123H); The amino acid at position 146 is substituted with C, A, D, G, K, R, T or V, more preferably with C (S146C). The amino acid at position 152 is substituted with P, which can be replaced with R, H, I, P or S, more preferably with R (P152R). The amino acid at position 157 is substituted with N, and can be substituted with K, A, C, D, G, R, S, T or V, more preferably with N substituted with K (N157K); The amino acid at position 166 is substituted with T, and the substitution is I, H, P, R or S, more preferably T is substituted with I (T166I); More preferably, the rABE mutant comprises: a protein corresponding to the amino acid sequence shown in SEQ ID NO:1, having a substitution mutation at one or more amino acid sites among R23W, L36H, A48P, L51R, V106A, A109S, T111R, D119N, H122N, Y123H, S146C, P152R, Q154R, N157K, and T166I; More preferably, the amino acid sequence of the rABE mutant is shown in any one of SEQ ID NO: 2, 3, 4, 5, 7, 11, 12, 13, 14, 15, 16, 19, 20, 23 and 24.
4. The fusion protein according to any one of claims 1-3, characterized in that, The tagged protein refers to a protein that can specifically bind to its substrate; Preferably, the tag protein includes: SNAP-tag, CLIP-tag, HaloTag, TMP-tag, and ACP-tag; More preferably, the tag protein is a SNAP-tag protein; More preferably, the amino acid sequence of the tag protein is shown in SEQ ID NO:
25.
5. The fusion protein according to any one of claims 1-4, characterized in that, The fusion protein also includes a linker; Preferably, the connector includes a connector containing G and S; More preferably, the amino acid sequence of the linker is shown in SEQ ID NO:
26.
6. The fusion protein according to any one of claims 1-5, characterized in that, The fusion protein is selected from: (a) A protein with the amino acid sequence shown in SEQ ID NO:36; (b) A protein derived from (a) that has the function of the protein in (a) and is formed by substitution, deletion or addition of one or more amino acid residues of the amino acid sequence shown in SEQ ID NO:
36. (c) A protein whose amino acid sequence is more than 80% identical to the amino acid sequence defined in (a) and which has the function of the protein in (a); or (d) A fragment of SEQ ID NO:36 having the protein function of (a); preferably comprising a base editor and a tag protein; more preferably, the amino acid sequence of the base editor is as shown in SEQ ID NO:20, and the amino acid sequence of the tag protein is as shown in SEQ ID NO:
25.
7. A polynucleotide comprising: 1) A polynucleotide encoding the fusion protein according to any one of claims 1-6; or 2) A polynucleotide complementary to the polynucleotide described in 1).
8. An expression vector or host cell comprising the polynucleotide of claim 7, or expressing the fusion protein of any one of claims 1-6.
9. A composition comprising: 1) The fusion protein according to any one of claims 1-6, or the polynucleotide according to claim 7, or the expression vector or host cell according to claim 8, wherein the fusion protein comprises a base editor and a tag protein; and 2) A modified small molecule compound, wherein the modified small molecule compound comprises: the small molecule compound and the substrate of the tag protein described in 1); Preferably, the tag protein is a SNAP-tag protein, and the substrate of the tag protein is benzylguanine.
10. A kit comprising the fusion protein of any one of claims 1-6, or the polynucleotide of claim 7, or the expression vector or host cell of claim 8, or the composition of claim 9.
11. The application of the fusion protein according to any one of claims 1-6, or the polynucleotide according to claim 7, or the expression vector or host cell according to claim 8, or the composition according to claim 9, or the kit according to claim 10, selected from: (1) Application in detecting or predicting the binding sites of small molecule compounds with nucleic acids; (2) Application in screening nucleic acid sites that bind to small molecule compounds; Preferably, the nucleic acid is DNA or RNA; more preferably, the RNA is mRNA.
12. A method for detecting or predicting the binding site of a small molecule compound to a nucleic acid, the method comprising: The step of contacting the fusion protein of any one of claims 1-6, or the polynucleotide of claim 7, or the expression vector or host cell of claim 8, or the composition of claim 9, or the kit of claim 10 with the modified small molecule compound; Preferably, the modified small molecule compound comprises a small molecule compound and a substrate of a tag protein, wherein the substrate of the tag protein specifically binds to the tag protein in the fusion protein; more preferably, the concentration of the modified small molecule compound is 0.001–1000 nM, even more preferably 0.01–800 nM, 0.1–700 nM, 1–600 nM, 10–600 nM, 50–500 nM, or 200–400 nM; And / or, Preferably, the contact includes simultaneous, sequential, or sequential contact.
13. The method as described in claim 12, characterized in that, The method further includes: collecting nucleic acids before and after contact with the modified small molecule compound, detecting the nucleic acid sequence, and obtaining the nucleic acid editing site by comparing the nucleic acid sequences before and after contact with the modified small molecule compound, which is the binding site between the small molecule compound and the nucleic acid.
14. A method for screening nucleic acid sites that bind to small molecule compounds, the method comprising: 1) Construct an expression vector, fuse the ABE editor and SNAP tag with the GS linker for expression, and then transiently transfect the plasmid into HEK293T cells via PEI. After 6 hours of transfection, replace with fresh culture medium. 2) Add the modified small molecule compound to the system in 1), use the fusion protein to edit the nucleic acid bound to the modified small molecule compound, and collect the edited nucleic acid; 3) Sequencing the nucleic acids in 1) and 2) respectively, comparing the sequences of the nucleic acids in 1) with those in 2) to obtain the sites where nucleic acid editing occurred. These editing sites are adjacent to the nucleic acid sites that bind to small molecule compounds.