A fusion protein and a method of using the same
By using a fusion protein to bind RNA editing enzyme, the problem of identifying RNA-binding protein targets and providing transcriptome information in existing technologies has been solved, achieving high-sensitivity and high-temporal-resolution dual-omics analysis, which is suitable for RBP studies of single-cell and tissue samples.
Patent Information
- Application Number
- CN202411632711.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-15
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2044-11-15
AI Technical Summary
Existing technologies struggle to efficiently and sensitively identify RNA-binding protein (RBP) targets and provide transcriptomic information in low-input samples, especially in single-cell and tissue samples, and lack high temporal resolution and dual-omics analysis capabilities.
A fusion protein, comprising an antibody-binding domain and an RNA-editing domain, was expressed and purified using a baculovirus expression vector system. After binding to cell or tissue samples, an in situ deamination reaction of RNA-editing enzymes was performed. Bioinformatics analysis was then used to identify RBP targets and provide transcriptome information.
It enables highly sensitive RBP target identification and transcriptome analysis in single-cell and tissue samples, possesses high temporal resolution and dual-omics analysis capabilities, reduces material input, and is suitable for studying RNA functional mechanisms.
Smart Images

Figure CN119505015B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of RNA-binding protein (RBP) analysis technology, in particular to a sequencing method for simultaneously identifying RBP targets and providing transcriptome information and application thereof. BACKGROUND
[0002] Gene expression regulation at the RNA level is the basis of many biological processes, which accurately controls the life cycle of RNA molecules in a complex and precise manner. Such post-transcriptional regulation is mainly achieved by the interaction of intracellular proteins and RNA in a complex manner, assembling dynamic ribonucleoprotein particles (RNPs) to jointly coordinate a series of processes such as cell fate and environmental changes. RNA-binding proteins (RBPs) can be involved in different biological stages after RNA transcription, and form molecular machines with rRNA, tRNA, etc., playing a pivotal role in the gene expression regulation network. However, the functions and molecular mechanisms of many RBPs are still blank, which leads to slow progress in the study of the role of RBP in development and disease. Nowadays, more and more studies have shown that many human diseases, including neurological diseases, metabolic diseases and cancer, are related to RBP dysfunction. At present, the number of disease-related RBPs identified has exceeded 1000, but except for a few exceptions, people lack understanding of the physiological regulation system and pathological mechanism of RBP. When the disease phenotype cannot be explained by the known functions of RBP, it is crucial to find out which targets it specifically recognizes, and to detect the abnormal interaction caused by pathogenic mutations centered on RBP, in order to understand the process of disease occurrence.
[0003] The current frontier technology for studying the binding of RBP and RNA can be divided into two categories: the first is the cross-linking and antibody-dependent immunoprecipitation method; the second is the in situ detection method dependent on RNA processing enzymes.
[0004] RNA immunoprecipitation (RIP) and cross-linking immunoprecipitation (CLIP) are the most fundamental techniques for studying RNA-protein interactions in cells, which use specific antibodies to precipitate RNA-binding proteins (RBPs) and their associated RNA targets. However, these methods have significant limitations, including time-consuming and laborious experimental procedures, non-specific interactions identified in low-complexity libraries, and the need for large amounts of starting material. These challenges make them difficult to apply in certain contexts, particularly when dealing with low-input or rare samples, or when a large number of samples need to be processed in parallel. More importantly, these methods still face challenges in resolving targets at the level of transcript isoforms and achieving single-cell resolution. In addition, none of the above methods support parallel transcriptome analysis in the same cellular environment, limiting their ability to directly establish the connection between RBP binding and gene expression regulation.
[0005] Recently, RNA-editing-based TRIBE and STAMP have been developed, which use ectopically expressed RBP-deaminase fusion proteins to label RNA substrates. RNA deamination is mainly mediated by two types of enzymes: adenosine deaminases acting on RNA (ADARs) and apolipoprotein B mRNA-editing enzyme cytosine deaminases (also known as AID / APOBEC family), which are responsible for A-to-I and C-to-U editing, respectively. The ADAR structure is modular, consisting of double-stranded RNA binding motifs (dsRBMs) and catalytic deamination domains (ADARcd), which can work independently. ADARcd deaminates adenosine on RNA to inosine, which is recognized as guanosine by reverse transcriptase, so A-G mutations can be determined by RNA sequencing to identify ADAR editing. Endogenous ADAR editing sites are usually in non-coding regions Alu repeat elements, as they are biased to form double-stranded RNA features. AOPBEC1 has weak RNA binding ability and cannot edit its substrates without the presence of its partner protein APOBEC1 coupling factor (ACF), and its editing may depend on C-terminal-mediated homodimerization. There are structural and sequence preferences for AOPBEC1 editing in cells, almost all of which occur in 3'UTRs, and are biased to bind AU-rich regions, and are more biased to edit 5'-AC-3' in sequence. The methods of these editing enzymes avoid target enrichment steps and retain complete transcriptome information. Although these IP-free techniques allow the discovery of substrates from low-input materials, including single cells, the use of a single enzyme can cause false-negative results, and their dependence on genetic manipulation limits their application in primary cells and clinical samples. In addition, ectopic expression of fusion proteins can introduce artifacts or interfere with the function of endogenous proteins. Finally, the lack of temporal resolution severely limits their ability to study dynamic biological processes.
[0006] Therefore, in order to fully understand the regulatory role of RBP function under different environments, the field urgently needs a technology that is sensitive, universal and convenient, can perform dual-omics co-analysis, has high time resolution, requires only a small amount of sample input, and can be applied to tissue samples. SUMMARY
[0007] In view of the deficiencies of the prior art, the present application provides a fusion protein and an application method thereof, which can simultaneously identify RBP targets and provide transcriptome information, solving the problems existing in the prior art.
[0008] In one aspect of the present application, a fusion protein is provided, which comprises an antibody binding domain and an RNA editing domain, the antibody binding domain being a pAG domain comprising functional domains of Protein A and Protein G, and the RNA editing domain comprising a rat APOBEC1 enzyme deamination domain and / or a human ADAR2 enzyme deamination domain.
[0009] In one specific embodiment of the present application, the fusion protein comprises the pAG domain, and the rat APOBEC1 enzyme deamination domain and the human ADAR2 enzyme deamination domain.
[0010] In another aspect of the present application, an application method for simultaneously identifying RBP targets and providing transcriptome information using the fusion protein disclosed in the present application is provided, the method comprising the following steps:
[0011] 1. Expression and purification of the fusion protein: the fusion protein of the present application is expressed and purified using a baculovirus expression vector system;
[0012] 2. Cell / tissue section fixation: cell / tissue samples are fixed using a formaldehyde solution or 3,3'-dithiodipropionic acid di(N-hydroxysuccinimidyl ester) (DSP);
[0013] 3. Cell binding ConA magnetic beads: resuspend the fixed cells with DPBS, add ConA magnetic beads, rotate incubate at room temperature, and transfer the cells bound with the magnetic beads to a PCR tube;
[0014] 4. Antibody incubation: prepare antibody incubation buffer, add RBP antibody or IgG antibody buffer in the experimental group and the control group, respectively, resuspend the cells, mix gently using vortex oscillation, and resuspend the cells after washing;
[0015] 5. Deaminase binding: prepare a buffer containing the fusion protein of the present application in advance, and wash and resuspend the cells obtained in step four after incubation with the secondary antibody in the buffer containing the fusion protein;
[0016] 6. Editing reaction initiation and termination: prepare the deamination reaction buffer containing zinc ions for initiating the deamination reaction, and resuspend the cells co-incubated with the fusion protein of the present application in the reaction buffer after washing to perform the editing reaction;
[0017] 7. RNA extraction and sequencing library construction: extract total RNA from cells, construct a library, and after detecting the library concentration, sequence the cDNA in the library;
[0018] 8. Bioinformatics analysis: the bioinformatics analysis includes data quality control, two rounds of unique alignment, fine-tuning alignment, identification of single nucleotide variations (SNVs) and differential editing analysis.
[0019] In an embodiment of the present application, 3,3'-dithiodipropionic acid di(N-hydroxysuccinimidyl ester) (DSP) is used in step 2 described above to fix the cells.
[0020] In an embodiment of the present application, the use of the fusion protein of the present application for identifying RBP targets and providing transcriptome information is provided, characterized in that it is applied in tissue samples, frozen sections or single cell samples.
[0021] Advantages
[0022] The present application provides a sequencing method for simultaneously identifying RBP targets and providing transcriptome information and applications.
[0023] Compared with the prior art, the following advantages are achieved:
[0024] 1. The fusion protein and its application method disclosed in the present application can simultaneously identify RBP targets and provide transcriptome information. Using the technical solutions disclosed in the present application, the present inventors successfully identified the RNA targets of RBP YTHDF2 and G3BP1 in HeLa cells, and detected the transcripts bound by PTBP1 and G3BP1 and the binding motifs in HEK293T cells. The results of the endogenous proteins obtained by the technical solutions of the present application were compared with the gold standard CLIP-seq, and the results fully verified the practicability and sensitivity of the technical solutions of the present application.
[0025] 2、Since the fusion protein and the application method thereof disclosed by the application can realize RBP target detection with single cell resolution and also provide transcriptome information, the transcriptome information can be used to distinguish cells at different stages in the cell cycle and reveal G3BP1 targets at different cell phases, the fusion protein and the application method thereof disclosed by the application are a multifunctional dual-omics platform, which can characterize the RNA-protein interaction spectrum and has many advantages compared with other methods, effectively reduces the material input, does not need to construct a heterologous expression model, is suitable for studying the RBP function in tissue sections and single cells, and has a wide application prospect in the functional mechanism research of RNA in the development and disease process. BRIEF DESCRIPTION OF DRAWINGS
[0026] Figure 1 It is a method flowchart of the application.
[0027] Figure 2 It is a schematic diagram of the cumulative number of the editing signal of the different processing groups of MAPIT-seq in the 400 nucleotide window near the CLIP enrichment peak region.
[0028] Figure 3 It is a schematic diagram of the editing signal of the experimental group and the control group on the G3BP1 binding target.
[0029] Figure 4 It is a schematic diagram of the comparison result of the G3BP1 target identified by the MAPIT-seq of the application and the PAR-CLIP.
[0030] Figure 5 It is a schematic diagram of the overlapping proportion obtained by comparing the G3BP1 target identified by the MAPIT-seq of the application with CLIP-seq and its variants, RIP-seq and HyperTRIBE.
[0031] Figure 6 It is a schematic diagram of the specific editing event on the PTBP1 and YTHDF2 target mRNA identified by the MAPIT-seq of the application.
[0032] Figure 7 It is a schematic diagram of the overlap of the PTBP1 and YTHDF2 target identified by the MAPIT-seq of the application and the target detected by CLIP.
[0033] Figure 8 It is a schematic diagram of the high-quality transcriptome spectrum corresponding to the cell environment studied by the RBPs of the application.
[0034] Figure 9 It is a schematic diagram of the PTBP1 binding motif detected by the MAPIT-seq of the application.
[0035] Figure 10Schematic of YTHDF2 binding motif for MAPIT-seq detection of the present invention.
[0036] Figure 11 Schematic of enriched editing events for MAPIT-seq application on fixed frozen sections and fresh frozen sections of the present invention.
[0037] Figure 12 Schematic of G3BP1 editing signal and gene expression reproducibility for MAPIT-seq application on mouse embryo tissue sections of the present invention.
[0038] Figure 13 Schematic of gene expression for the transition from neurogenesis to gliogenesis between E12.5 and E16.5 of the present invention.
[0039] Figure 14 Schematic of G3BP1-MAPIT and IgG-MAPIT enriched editing events comparison for MAPIT-seq application on two time period mouse embryo brain tissue sections of the present invention.
[0040] Figure 15 Schematic of specific editing events generated by MAPIT-seq on G3BP1 targets Mapt and Cadm2 of the present invention.
[0041] Figure 16 Schematic of G3BP1 binding motif content identified by CLIP in identified targets for MAPIT-seq application on mouse embryo brain tissue sections of the present invention.
[0042] Figure 17 Schematic of gene expression profile of G3BP1-MAPIT library fixed with DSP and correlation with untreated group of the present invention.
[0043] Figure 18 Schematic of signal to noise ratio for bulk MAPIT-seq generated by DSP-MAPIT of the present invention.
[0044] Figure 19 Schematic of G3BP1 binding substrates compared to formaldehyde-MAPIT and ARTR-seq results of the present invention.
[0045] Figure 20 Schematic of gene expression profile of scMAPIT-seq library performed with DSP fixation and correlation with untreated group of the present invention.
[0046] Figure 21 Schematic of double editing events enrichment on G3BP1 targets in single cell and single cell aggregated G3BP1-MAPIT datasets of the present invention.
[0047] Figure 22 G3BP1 target identified for scMAPIT-seq of the present application after all single cells were aggregated. Comparison with bulk MAPIT-seq and ARTR-seq results.
[0048] Figure 23 Grouping of cells into three cell cycle phases using gene expression for scMAPIT-seq of the present application.
[0049] Figure 24 Differences of G3BP1 targets in three phases (G1, S and G2 / M) for the present application.
[0050] Figure 25 Distribution of the correlation between G3BP1 binding strength and RNA expression level in cell cycle progression detected by scMAPIT-seq of the present application.
[0051] Figure 26 Comparison of sequencing read length for DSP-G3BP1-MAPIT and FA-G3BP1-MAPIT samples of the present application.
[0052] Figure 27 Comparison of editing signal strength for DSP-G3BP1-MAPIT and FA-G3BP1-MAPIT samples of the present application.
[0053] Figure 28 Comparison of all targets for long read MAPIT-seq data and short read MAPIT-seq of the present application.
[0054] Figure 29 Comparison of editing signal strength for long read MAPIT-seq data on two transcript isoforms of YTHDF2, GDAP2 and ING3 of the present application. DETAILED DESCRIPTION
[0055] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.
[0056] The application provides a fusion protein, which comprises an antibody binding domain and an RNA editing domain, the antibody binding domain is a pAG domain, comprising the functional domains of Protein A and Protein G, and the RNA editing domain comprises a rat APOBEC1 enzyme deamination domain and / or a human ADAR2 enzyme deamination domain. On the basis of the above-mentioned fusion protein, the application provides a technical scheme based on an RNA editing enzyme and in-situ capturing of intracellular RBP targets. The core principle is different from that of the TRIBE method, and genetic operation is not required, but depends on the antibody-guided recombination editing enzyme (i.e. the fusion protein of the application) to target the RBP binding position to perform in-situ deamination reaction. Therefore, the technology can not only realize the combined detection of RBP targets and the transcriptome, but also can draw dynamic RBP interaction.
[0057] The present application designs key experimental procedures based on the characteristics of RBP: 1) The sample is fixed with 0.5% formaldehyde, because the binding of RBP on RNA is transient and dynamic, so efficient and reversible cross-linking is a key step to capture RBP interaction and preserve high-quality transcriptome in subsequent operations. 2) One of the cores of the present application is the construction of RNA editing enzyme components, including antibody binding domains and RNA editing domains. For the antibody binding domain, the prior art usually uses Protein A, which binds well with rabbit, guinea pig and other antibody constant regions, but binds poorly with mouse-derived IgG1 subtype, so Protein G with higher affinity to mouse antibodies is needed. In order to improve the universality of the fusion protein, the present application uses pAG-ERH constructed by Henikoff et al. (see Meers MP, Bryson TD, Henikoff JG, Henikoff S. Improved CUT&RUN chromatin profiling tools. Elife. 2019 Jun 24;8:e46314. doi: 10.7554 / eLife.46314. PMID: 31232687; PMCID: PMC6598765. The entire contents of this document are incorporated into the present application), that is, a Protein G domain containing three point mutations is added to Protein A, which is sufficient to ensure the affinity with common host antibodies. Secondly, as described in the background art, the TRIBE and STAMP technologies are limited by the efficiency of a single editing enzyme, and APOBEC and ADAR both have strong substrate and sequence preferences. The present application connects the deamination domain of rat APOBEC1 (rAPOBEC1) and the enzyme activity-enhanced mutant of human ADAR2 (E488Q) (hADAR2dd) through an XTEN linker and pAG fusion, and the double editing event produced can combine to reduce false negative detection and improve the sensitivity of the technology. After testing different combinations of vectors, the present inventors determined that rAPOBEC1-pAG-hADAR2 is the best combination. In order to ensure the biological activity and yield of the fusion protein, the present inventors expressed rAPOBEC1-pAG-hADAR2 in eukaryotic insect cells sf21, and purified it through Strep magnetic beads and molecular sieve.
[0058] Based on the above ideas, the present inventors developed the MAPIT-seq technology (Modification Added to RBP Interacting Transcript-sequencing), the specific principle of which is as follows Figure 1As shown, in situ fixation of intracellular molecules with formaldehyde covalently crosslinks RBP-bound RNA. Addition of digitonin permeabilizes the cell membrane, and 4°C incubation of antibodies (RBP-specific antibodies as experimental group, IgG as control group to remove non-specific editing) and fusion protein rAPOBEC1-pAG-hADAR2dd, which is guided to target RBP antibodies by the pAG domain in the fusion protein after entering the cell. Excess antibodies and enzymes are washed away at the corresponding steps, and then the fusion protein undergoes a deamination reaction at the RBP-bound RNA region at room temperature. Finally, the reaction is terminated and proteinase K is added to uncrosslink, RNA is extracted, and sequencing analysis identifies mutation sites to identify RBP substrates.
[0059] Example 1: MAPIT-seq experimental steps
[0060] 1. Expression and purification of fusion proteins
[0061] The fusion protein rAPOBEC1-pAG-hADAR2dd used in the present application is expressed and purified using a baculovirus expression vector system with authentic protein modifications to achieve maximum enzymatic activity. The general procedure: Recombinant pFastBac plasmid is transformed into DH10Bac cells (Biomed, BC112) to generate bacmids. Bacmid DNA is transfected into Sf21 cells using X-tremeGENE HP DNA Transfection Reagent (06366236001, Roche) and cultured for 4 days. Low titer recombinant baculovirus is collected and amplified in Sf21 cells to produce high titer virus. Next, baculovirus is used to infect Hi5 cells at a ratio of 1:200 and cultured for 48-60 hours. Cells are lysed in lysis buffer containing 50 mM Tris-HCl (pH 8.0) and 200 mM NaCl. The lysate is sonicated using an ultrasonic processor (SCIENTZ) at a setting of 7 seconds on and 5 seconds off for 20 minutes at a power of 300W. After sonication, the lysate is centrifuged at 16,000g at 4°C for 40 minutes. The deaminase fusion protein is then purified using Strep-Tactin beads (SA053, Smart Lifesciences). Briefly, the supernatant is incubated with Strep-Tactin beads for 1 hour at 4°C. After washing the beads with 30 volumes of lysis buffer, the deaminase fusion protein is eluted with lysis buffer containing 5 mM desthiobiotin. Further purification is performed by a Superdex 200 Increase column (GE Healthcare). The deaminase protein is eluted into a buffer containing 25 mM HEPES (pH 7.5) and 150 mM NaCl, which is then supplemented with 5% glycerol. The deaminase protein is aliquoted and quickly frozen in liquid nitrogen and stored at -80°C.
[0062] 2 Cell / tissue section fixation
[0063] 500-5 million cells are collected per experiment, 0.1% Trypsin digestion for 5 min, after stopping with culture medium, collected into centrifuge tubes pre-washed with 1% BSA. The following are described for one set of cells or tissue sections; 250 x g centrifugation at room temperature for 5 min, washed once with DPBS; prepare 0.2% (tissue) or 0.5% (cells) formaldehyde solution; add formaldehyde solution to each set of cells, incubate at room temperature for 5 min; add Glycine or Tris-HCl pH 7.4 to stop crosslinking, incubate at room temperature for 5 min; centrifuge in a non-fixed centrifuge at 450 x g for 5 min at 4°C. Wash the cells with pre-cooled DPBS for 3 times.
[0064] 3 Cell binding ConA magnetic beads
[0065] Pipette 5 μL of ConA beads and add 10× binding buffer, mix well, place on a magnetic stand for 30 seconds to 2 minutes, and remove the supernatant. Repeat this process and resuspend with the original volume of binding buffer. Resuspend the fixed cells in DPBS, add ConA magnetic beads, and incubate at room temperature with rotation for 10 to 20 minutes. Transfer the cells bound to the magnetic beads to an eight-tube strip for PCR.
[0066] 4. Antibody Incubation
[0067] Prepare antibody incubation buffer (for two groups as an example, freshly prepared): (1mM PMSF, 1× protease inhibitor cocktail, 1U / μL RiboLock RNase Inhibitor (EO0382, ThermoFisher), 0.01% Digitonin, 1mM DTT, 2mM EDTA, and 0.1% BSA, make up to 100μL with DPBS). Resuspend cells in the experimental and control groups with RBP antibody or IgG antibody buffer (100:1 ratio), gently mix by vortexing, and incubate at 4°C for 3-4 hours. Mix thoroughly during the incubation period, wash cells once with DPBS, resuspend in secondary antibody incubation buffer (antibody buffer-EDTA), and incubate at 4°C for 1 hour.
[0068] 5. Binding of deaminase
[0069] Prepare enzyme buffer in advance: 20mM HEPES-KOH pH 7.5 + 300mM NaCl, replacing the DPBS in the secondary antibody incubation buffer. Add 1μg of the rAPOBEC1-pAG-hADAR2dd fusion protein to 50μL of buffer. After secondary antibody incubation, wash cells once with DPBS, resuspend cells in enzyme buffer, and incubate at 4°C for 1 hour.
[0070] 6 Editing reaction initiation and termination
[0071] Prepare deamination reaction buffer (15 mM HEPES pH 7.9, 60 mM KCl, 15 mM NaCl, 5% Glycerol, 0.5 mM DTT, and 0.1 μM ZnCl2). Add proteinase inhibitor, RNAase inhibitor, and DTT to the deamination reaction buffer freshly before use. After deaminase incubation, wash cells three times with 200 μL of wash buffer (20 mM HEPES-KOH pH 7.5 + 300 mM NaCl, 0.005% Digitonin). Resuspend cells in 40 μL of reaction buffer and incubate at 30°C for 3 hr.
[0072] 7RNA extraction and sequencing library construction
[0073] Add 60 μΐ of Proteinase K digestion buffer and 1 μΐ of Proteinase K, digest at 56°C for 1 hr. Extract total RNA using Magzol RNA extraction reagent, add microcell with glycogen to help precipitation. Take 500 ng of RNA for library construction, library construction uses Universal V6 RNA-seq Library Prep Kit for After detecting the library concentration by Qubit, take 100 ng of cDNA to Illumina NovaSeq 6000 platform for sequencing.
[0074] 8 Bioinformatics analysis
[0075] After quality control of sequencing data, TrimGalore was used to remove adapter sequences and low-quality bases. For MAPIT-seq data, BWA was used to remove high-abundance RNAs in cells, including ribosomal RNA, mitochondrial RNA and transfer RNA, before alignment. The data alignment reads used the HISAT2 and BWA "two-round unique alignment" strategy, and then the GATK tool was used to correct the alignment file, identify SNV sites using HaplotypeCaller and remove known SNP sites. Finally, A-to-G and C-to-U editing sites were screened. bedtools was used for gene annotation of sites, such as transcript elements and repeat sequences. Differential editing analysis divided transcripts into small blocks according to a given window size (default 50bp), and calculated the editing index of each block window, which was defined as the cumulative value of the editing rate of the editing site in the specified region. Finally, the Wilcoxon signed rank test was used to screen the transcripts with significantly higher editing in the experimental group than in the control group (MAPITscore greater than 0.5, P-value less than 0.001 and editing fold > 2), i.e. RBP-targeted RNA targets, where MAPITscore was defined as the editing index of a transcript in the experimental group minus the value in the control group. For Motif analysis, high-confidence editing sites were screened using Poisson, and editing clusters were clustered within a neighborhood distance of 150 nt to obtain editing clusters, and editing clusters containing at least 3 sites and higher editing rates in the experimental group than in the control group were retained as high-confidence RNA editing segments, and homer was used to predict RBP binding motifs. For single-cell MAPIT-seq, STARsolo module was used for alignment, and after read de-duplication and cell quality control, the alignment results were split by single cell, and RNA editing recognition analysis was performed for each single cell. For long-read MAPIT-seq, Minimap2 was used for alignment, and after quantification and quality control using isoquant, the alignment results were split by transcript isoform, and RNA editing recognition analysis was performed for each transcript isoform. After quality control of sequencing data, TrimGalore was used to remove adapter sequences and low-quality bases. For MAPIT-seq data, BWA was used to remove high-abundance RNAs in cells, including ribosomal RNA, mitochondrial RNA and transfer RNA, before alignment.
[0076] The MAPIT-seq method for tissue sections is a modified version of the standard MAPIT-seq. A brief description is as follows:
[0077] The sections stored at -80°C were thawed at room temperature for 10 minutes.
[0078] A hydrophobic circle was drawn around the tissue using a PAP pen.
[0079] Slices were fixed with 0.2% formaldehyde.
[0080] Cell permeabilization and deamination process was essentially the same as standard MAPIT-seq.
[0081] After deamination reaction, samples were digested with 0.4 mg / mL proteinase K at 55 °C for 2 hours.
[0082] Finally, the mixture was transferred from the glass slide to a 1.5 mL tube using a pipette for subsequent RNA extraction, library construction and sequencing.
[0083] The MAPIT-seq method for high-throughput single cell was performed using DSP fixation combined with the 10x Genomics platform:
[0084] Cells were fixed with 1 mg / mL DSP solution at room temperature for 30 min.
[0085] Throughout the experiment, DSP-fixed cells were collected by centrifugation at 650 g for 5 minutes.
[0086] After deamination reaction, 50 mM DTT was added for reverse crosslinking.
[0087] Cells were washed once and then resuspended in DPBS containing 0.05% BSA and 1 U / μL RiboLock RNase inhibitor.
[0088] Subsequently, samples were sorted on a MoFlo Astrios EQ (Beckman) to eliminate cell debris and aggregates.
[0089] Finally, samples were loaded onto the 10x Genomics Chromium platform.
[0090] Example Two: MAPIT-seq enables specific editing of RBP targets
[0091] To preliminarily verify the feasibility of the system, the inventors first selected the RNA binding protein YTHDF2, which has clear functions and CLIP data sets, and they are m 6A plays a key role in the regulation. And considering the efficiency of antibody, the present application constructs a cell line overexpressing RBP-Flag, and 250,000 cells are collected for MAPIT-seq experiment. YTHDF2-Flag overexpressing HeLa cells are used as experimental group by targeting editing with Flag antibody. IgG / no antibody treatment is used as simple control group, and Flag antibody treatment of cells expressing only Flag is used as more reliable non-specific editing control. First, the present inventors counted the cumulative number of MAPIT-seq editing signals in the 400-nucleotide window near the CLIP enrichment peak. Within 400 bp around YTHDF2 binding sites, the total number of editing events in the YTHDF2-Flag experimental group was significantly higher than that in the efficiency and three control groups Figure 2 ) It is worth noting that the increase in editing events requires both YTHDF2 binding to RNA and the corresponding antibody recruiting the recombinase to edit, which proves the specificity of MAPIT technology for RBP.
[0092] Example Three MAPIT-seq combined detection of endogenous RBP binding motifs and gene expression
[0093] After verifying the feasibility of the system and optimizing the experimental conditions, the present inventors applied MAPIT-seq to endogenous protein G3BP1 in HEK293T cells. The present inventors first looked at the G3BP1 target genes previously reported by PAR-CLIP, and IGV (Integrative Genomic Viewer) results showed that the experimental group produced rich double editing signals on the target, which had similar distribution to the G3BP1 binding sites detected by PAR-CLIP, while the control group had almost no editing in the binding region Figure 3 ) These results show that MAPIT-seq can enrich target-specific editing to indicate RBP binding substrates. To evaluate the effect of MAPIT-seq in identifying G3BP1 interacting RNAs across the whole transcriptome, the present inventors applied a set of stringent criteria to define G3BP1 targets. Briefly, if the fold enrichment of editing of a transcript is >2, the MAPIT score is >0.5, and the P value is <0.001, the present inventors consider it as a G3BP1 target. To minimize false negatives, the present inventors combined A to I and C to U editing to calculate the MAPIT score of each gene. The results show that the G3BP1 targets identified by MAPIT-seq significantly overlap with the PAR-CLIP results Figure 4 ), and most of the G3BP1 targets identified by MAPIT-seq are cross-validated by multiple independent methods, including CLIP-seq and its variants, RIP-seq, HyperTRIBE and TRIBE-ID Figure 5), which can reflect the accuracy and reliability of MAPIT-seq for detecting RBP target genes.
[0094] To further explore the universality of MAPIT-seq, the inventors also detected PTBP1 protein RNA substrates in HEK293T and YTHDF2 protein RNA substrates in HeLa cells. Similar to the discovery of G3BP1, MAPIT-seq also revealed specific editing events on these RBP target mRNAs Figure 6 ). In addition, the targets identified by MAPIT-seq all significantly overlapped with those detected by CLIP-like methods Figure 7 ), highlighting its reliability and broad applicability in different RBPs. Moreover, MAPIT-seq provided high-quality transcriptome profiles corresponding to the cellular environment of these RBPs under study Figure 8 ).
[0095] Finally, MAPIT-seq can also be used to detect the binding motifs of RBPs. The inventors focused on PTBP1, an RBP known to prefer CU-rich sequences. Using PTBP1-MAPIT editing clusters, the inventors identified a significant enrichment of the "CUUCUUUC" motif by the HOMER suite Figure 9 ). Similarly, the inventors found that YTHDF2 enriched a conserved "GGAC" motif Figure 10 ). Overall, these data demonstrate the effectiveness of MAPIT-seq in detecting RBP binding targets and revealing their associated sequence motifs.
[0096] Example Four MAPIT-seq detects RBP targets in tissue sections
[0097] Having fully validated the advanced nature and reliability of the MAPIT-seq technology, the inventors further explored whether MAPIT-seq could be applied in tissues to provide RBP binding sets and gene expression information in tissue environments. Therefore, the inventors tested the MAPIT-seq method in mouse tissues. The inventors first tested the MAPIT-seq experimental conditions for G3BP1 on fixed and fresh-frozen mouse embryo sections at embryonic day 10.5 (E10.5). The results showed that enriched editing events occurred on fresh-frozen sections but not on fixed-frozen sections Figure 11 ), and both the editing signals of MAPIT-seq and gene expression had high reproducibility in brain tissues Figure 12Based on this, the inventors decided to use fresh frozen sections in subsequent experiments. To further investigate the function of G3BP1 in brain development, the inventors performed MAPIT-seq on fresh frozen sections of E12.5 and E16.5 mouse embryonic brains. Gene expression confirmed the transition from neurogenesis to gliogenesis between E12.5 and E16.5 ( Figure 13 The present inventors found that G3BP1-MAPIT enriched more editing events than IgG-MAPIT in mouse brain slices at all stages ( Figure 14 G3BP1-MAPIT discovered specific editing events on the G3BP1 targets Mapt and Cadm2, previously identified by CLIP ( Figure 15 Overall, the G3BP1 targets identified by the inventors in E12.5 and E16.5 mouse brain slices were significantly enriched in the G3BP1 binding motifs identified by CLIP ( Figure 16 ), confirming the feasibility and specificity of MAPIT-seq for detecting RBP binding in tissue sections. GO analysis of these targets can be performed to explore the regulatory function of G3BP1.
[0098] Example 5 MAPIT-seq detects RBP targets in single cells at a high-throughput level
[0099] The inventors first explored the possibility of MAPIT-seq to reliably detect in situ RBP-RNA interactomes and transcriptomes at the single-cell level. The present invention first considered that detecting mutations in formaldehyde-fixed samples requires a decrosslinking step, which is incompatible with most high-throughput single-cell RNA-seq (scRNA-seq) platforms. To solve this problem, the inventors adopted an alternative fixative that does not form covalent bonds and is compatible with scRNA-seq: 3,3'-dithiodipropionic acid bis(N-hydroxysuccinimide ester) (DSP), a new crosslinker that can be easily reversed by dithiothreitol (DTT). The inventors first evaluated whether DSP can simultaneously capture protein-RNA interactions and gene expression profiles in bulk MAPIT-seq. The results showed that the gene expression profile of the G3BP1-MAPIT library fixed with DSP was highly correlated with that of untreated cells and very similar to the expression profile of formaldehyde-fixed samples ( Figure 17 Notably, DSP-MAPIT yielded a significantly higher signal-to-noise ratio compared to formaldehyde-MAPIT, with a clear enhancement around the PAR-CLIP peak ( Figure 18 The G3BP1 binding substrates obtained by DSP-MAPIT were strongly correlated with the results of formaldehyde-MAPIT and ARTR ( Figure 19 ).
[0100] Next, the inventors performed scMAPIT-seq using DSP fixation and confirmed that it successfully preserved transcriptome integrity in single cells while not affecting RNA quality Figure 20 ). These findings indicate that DSP is an effective fixative for MAPIT-seq and scMAPIT-seq, making it suitable for expanding scMAPIT-seq to high-throughput applications. The inventors sequenced MAPIT-seq cells using the 10x Genomics single-cell workflow, capturing 3,400 G3BP1-MAPIT cells and 3,945 IgG-MAPIT cells. Results showed that doublet editing events were significantly enriched on G3BP1 targets in both single-cell and single-cell aggregated G3BP1-MAPIT datasets compared to IgG-MAPIT Figure 21
[0101] Furthermore, the inventors aggregated signals from all individual cells and found that scMAPIT-seq preserved comparable sensitivity and precision to bulk MAPIT-seq, showing strong agreement with ARTR-seq Figure 22 ). These data support the inventors’ high-throughput scMAPIT-seq can effectively reveal RBP targets and gene expression simultaneously at single-cell resolution. scMAPIT-seq preserved comparable sensitivity and precision to bulk MAPIT-seq.
[0102] Taking advantage of the dual-omics capability of scMAPIT-seq, the inventors further identified cell cycle stage-specific G3BP1 target repertoire Figure 23 ). The inventors observed significant differences in G3BP1 targets between three stages (G1, S, and G2 / M) Figure 24 ). And there were positive and negative correlations between G3BP1 binding strength and RNA expression levels during cell cycle progression, with positive correlations being more common Figure 25 ). Altogether, these data demonstrate the dual-omics capability of scMAPIT-seq to reveal stage- and subpopulation-specific RBP interactions and highlight the specific roles of G3BP1 in the cell cycle.
[0103] Example Six Long-read MAPIT-seq reveals differential binding of RBPs to RNA isoforms
[0104] The present inventors further improved the molecular resolution of MAPIT-seq to identify RBP binding events specific to individual transcript isoforms. The present inventors input G3BP1-MAPIT samples from formaldehyde-fixed and DSP-fixed HEK293T cells into the PacBio HiFi sequencing platform. Compared with the formaldehyde-fixed samples, the DSP-G3BP1-MAPIT samples produced longer reads and higher editing signals ( Figure 26 and 27 ), indicating that DSP fixation enhances target detection efficiency in long-read sequencing. Based on these results, the inventors used DSP-MAPIT data for further analysis.
[0105] When comparing long-read MAPIT-seq data with short-read MAPIT-seq data, the inventors observed strong consistency in target gene identification between the two datasets ( Figure 28 ). The accuracy of long-read MAPIT-seq was demonstrated. Further analysis of editing events revealed that G3BP1 binding is isoform-specific for some genes, including YTHDF2, GDAP2, and ING3 ( Figure 29 ).
[0106] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.
[0107] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A fusion protein, wherein the connection order of the fusion protein is as follows: a first RNA editing domain-XTEN linker-antibody binding domain-XTEN linker-second RNA editing domain, wherein the first RNA editing domain is a rat APOBEC1 enzyme, the antibody binding domain is a functional domain of Protein A and Protein G connected in sequence, and the second RNA editing domain is hADAR2dd, which is a mutant of the human ADAR2 enzyme with E488Q.
2. A method for simultaneously identifying RBP targets and providing transcriptome information using the fusion protein of claim 1, wherein the method is for the purpose of non-disease diagnosis and treatment, and the method comprises the following steps: 1) Expression and purification of the fusion protein: using a baculovirus expression vector system to express and purify the fusion protein of claim 1; 2) Fixation of cell / tissue sections: Fix the cell / tissue samples using formaldehyde solution or 3,3'-dithiodipropionic acid (N-hydroxysuccinimide ester); 3) Cell binding to ConA magnetic beads: Resuspend the fixed cells in DPBS, add ConA magnetic beads, incubate with rotation at room temperature, and transfer the bound cells to a PCR tube. 4) Antibody incubation: Prepare antibody incubation buffer, add RBP antibody or IgG antibody buffer to the experimental group and control group respectively, resuspend the cells, and gently mix by vortexing. After washing, cells were resuspended in secondary antibody incubation buffer; 5) Binding of deaminase: Prepare in advance a buffer solution containing the fusion protein of claim 1, wash the cells incubated with the secondary antibody obtained in step 4) and resuspend them in a buffer solution containing the fusion protein; 6) Initiation and termination of the editing reaction: prepare a deamination reaction buffer containing zinc ions for initiating the deamination reaction, wash the cells co-incubated with the fusion protein of claim 1, and resuspend them in the reaction buffer to perform the editing reaction; 7) RNA extraction and sequencing library construction: Extract total cellular RNA, build a library, test the library concentration, and sequence the cDNA in the library; 8) Bioinformatics analysis: The bioinformatics analysis includes data quality control, two rounds of unique alignment, fine-tuning alignment, identification of single nucleotide variants (SNVs), and differential editing analysis.
3. The method according to claim 2, characterized in that In step 2), 3,3'-dithiodipropionic acid bis(N-hydroxysuccinimide ester) was used to fix the cells.
4. Use of the fusion protein according to claim 1 for identifying RBP targets and providing transcriptome information, wherein the use is for purposes other than disease diagnosis and treatment, and is characterized in that: It is applied to tissue samples, frozen sections or single cell samples.