A high-efficiency rna interactome identification method based on hybrid proximity labeling tool
By designing an engineered peroxidase APEX2 and linking it to a complementary nucleic acid sequence, and utilizing APEX2 to catalyze the formation of biotin-phenoxy radicals from biotin and phenol to covalently label neighboring proteins, combined with affinity purification and mass spectrometry identification, the problem of low efficiency in the identification of RNA-protein interactions in existing technologies has been solved, achieving efficient and specific RNA-interacting proteome identification.
Patent Information
- Application Number
- CN202411827460.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-12
- Publication Date
- 2026-02-17
- Estimated Expiration
- 2044-12-12
AI Technical Summary
Existing technologies are insufficient for efficiently identifying RNA-protein interactions, especially for atypical RBPs, and traditional methods are either inefficient or unsuitable for rare cell samples.
An engineered peroxidase, APEX2, was designed and linked to complementary nucleic acid sequences via click chemistry. APEX2 catalyzes the formation of biotin-phenoxy radicals from biotin and phenol to covalently label neighboring proteins. Combined with affinity purification and mass spectrometry identification techniques, this enables the efficient identification of RNA-interacting proteomes.
It enables efficient and specific identification of RNA-interacting proteomes, can identify low-abundance RNA-binding proteins, is applicable to rare cell samples, and improves identification efficiency and accuracy.
Smart Images

Figure CN119736268B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of biotechnology and relates to a highly efficient method for identifying RNA-interacting proteomes based on a heterozygous proximity marker tool. Background Technology
[0002] The genome enables the complex expression of cellular functions through the dynamic expression of different RNAs. RNA does not function in isolation within the busy cellular environment but interacts with RNA-binding proteins (RBPs), which either influence RNA fate and function or are influenced by RNA. Furthermore, RNA-RBP interactions can trigger liquid-liquid phase separation and the formation of biomolecular condensates, thereby delineating cellular processes. Given the large number of RNAs and RBPs within the cell, combined interaction networks often drive complex regulatory outcomes.
[0003] Many discovered and studied RNA-binding proteins (RBPs) share two common characteristics: they contain structural features that explain their RNA binding, called RNA-binding domains (RBDs), and they have functional annotations relevant to RNA biology, such as involvement in RNA processing or expression, or as part of ribonucleoprotein (RNP) complexes (such as ribosomes, spliceosomes, or telomerases). These RBPs are often referred to as “conventional” or “orthodox” RBPs. Conventional RBPs also include proteins with enzymatic functions, those used for chemical modification or structural remodeling of RNA, and those that simply bind to RNA. The latter typically recruit other effector molecules to interact with RNA. RBP-RNA interactions are mediated through structurally well-defined RBDs, such as RNA recognition motifs (RRMs) or DEAD-DEAH helicase domains. Depending on their function, conventional RBPs exhibit broad substrate affinity and specificity. Overall, these RBPs can control the entire life cycle of their substrates, including mRNA processing, transport, translation, and degradation.
[0004] However, hundreds of newly discovered RNA-binding proteins (RBPs) lack both an RNA-binding domain (RBD) and functional annotations relevant to RNA biology. Instead, these RBPs are often well-studied proteins with highly diverse functional annotations, including metabolic enzymes, ion channels or solute carriers, cytoskeletal proteins, or effector molecules transported intracellularly. Some of these RBPs are highly abundant cellular proteins, while others are not. These RBPs (or candidate RBPs before independent validation) are often referred to as unconventional, atypical, or non-classical RBPs. This suggests that RNA-protein interactions are more complex than expected. Given the importance of messenger ribonucleoprotein (mRNP) complexes in almost all biological processes, their dysregulation can lead to a variety of diseases, including cancer and neurodegenerative diseases. For example, the human antigen R protein (HuR) binds to UDP-glucose, thus affecting its binding to specific mRNAs; mutations in HuR can lead to increased mRNA stability in cancer cells, promoting cancer cell invasiveness. Fusion sarcoma protein (FUS) is involved in multiple aspects of RNA metabolism. Mutations can lead to protein mislocalization and aggregation, which in turn can trigger neurodegenerative diseases such as amyotrophic lateral sclerosis (ALS). Therefore, in-depth research into the diversity and mechanisms of RNA-protein interactions is of great significance for understanding gene expression regulation, cell fate determination, and disease mechanisms.
[0005] Currently, there are two main strategies for detecting RNA-protein interactions: a protein-centric strategy for studying RNA binding to the target protein, and an RNA-centric strategy for studying proteins binding to the target RNA. The UV-crosslinking method for RNA and protein can only crosslink directly interacting RNA and protein, failing to effectively identify indirectly interacting RNA-protein interactions (RBPs). Furthermore, due to the low efficiency of UV crosslinking, it requires large amounts of cell material, making it unsuitable for rare cell samples. Methods using embedded nucleotide analogs to label newly transcribed RNA require long labeling times, making them unsuitable for studying transient RNA-protein interactions and dynamic changes. Phase-separated RNPs are unsuitable for low-abundance RNAs, and unstable RNA-protein interactions may be disrupted during the separation process.
[0006] In conclusion, how to achieve efficient identification of RNA-interacting proteomes remains one of the urgent problems to be solved in the field of RNA-protein interaction research. Summary of the Invention
[0007] To address the problem of how to achieve high-efficiency identification of RNA-interacting proteomes, this invention provides a high-efficiency RNA-interacting proteome identification method based on a hybrid proximity labeling tool. This invention designs an engineered peroxidase APEX2 and further develops proximity labeling complexes and RNA-interacting protein labeling technologies to establish a highly efficient and widely applicable method for studying RNA-interacting proteins.
[0008] To achieve this objective, the present invention adopts the following technical solution:
[0009] In a first aspect, the present invention provides an engineered peroxidase APEX2, the structural formula of which is shown in Formula I;
[0010] AL-APEX2 formula I;
[0011] In this context, A represents a non-natural amino acid containing an azide group, L represents a linker, and APEX2 represents peroxidase APEX2.
[0012] In this invention, the peroxidase APEX2 (Ascorbate Peroxidase 2) is redesigned by linking a non-natural amino acid containing an azide group to its N-terminus via a linker, thus obtaining engineered peroxidase APEX2. The azide group can undergo a click reaction with substances such as dibenzocyclooctyne (DBCO). Therefore, substances modified with dibenzocyclooctyne can be linked to engineered peroxidase APEX2 to expand its applications, such as linking it to complementary sequences of mRNA. This allows APEX2 to be linked to mRNA, and under the action of hydrogen peroxide, APEX2 catalyzes the formation of biotin-phenoxy radicals from biotin-phenol to biotinylate adjacent proteins of mRNA.
[0013] It is understood that in this invention, peroxidase APEX2 refers to a proximity labeling technology tool enzyme. APEX2 catalyzes the conversion of substrates (such as biotin and phenol) into active free radicals, which can then covalently bind to neighboring proteins, thereby achieving labeling of neighboring proteins. Peroxidases with similar functions in the art, such as APEX2, are also applicable to this invention; for example, their amino acid sequences are shown in SEQ ID NO. 1.
[0014] SEQ ID NO.1:
[0015] GKSYPTVSADYQDAVEKAKKKLRGFIAEKRCAPLMLRLAFHSAGTFDKGTKTGGPFGTIKHPAELAHSANNGLDIAVRLLEPLKAEFPILSYADFYQLAGVVAVEVTGGPKVPFHPGREDKPEPP PEGRLPDPTKGSDHLRDVFGKAMGLTDQDIVALSGGHTIGAAHKERSGFEGPWTSNPLIFDNSYFTELLSGEKEGLLQLPSDKALLSDPVFRPLVDKYAADEDAFFADYAEAHQKLSELGFADA.
[0016] Preferably, the non-natural amino acid containing an azide group includes phenylalanine analogues containing an azide group (p-azido-phenylalanine).
[0017] Preferably, the phenylalanine analog containing an azide group includes at least one of p-azidophenylalanine, 3-azido-L-phenylalanine, or 4-azido-L-phenylalanine.
[0018] Preferably, the connector comprises (Gly-Gly-Gly-Gly-Ser)n, where Gly is glycine, Ser is serine, and n is a positive integer, such as 1, 2, 3, or 4.
[0019] In a second aspect, the present invention provides a method for preparing the engineered peroxidase APEX2 described in the first aspect, the method comprising:
[0020] A succinate stop codon was introduced into the APEX2 peroxidase encoding gene, and an expression vector was inserted to obtain a recombinant vector. The recombinant vector was introduced into a host cell, which also expressed an orthogonal aminoacyl-tRNA synthetase and tRNA to obtain recombinant cells. The recombinant cells were cultured, and non-natural amino acids containing azide groups were added to the culture medium. After the culture was completed, the culture medium was collected, separated, and purified to obtain the engineered peroxidase APEX2.
[0021] In this invention, an amber stop codon (TAG) is first introduced into the APEX2 gene sequence at a specific site. Based on the technique of inserting non-natural amino acids at a specific site, an engineered aminoacyl-tRNA synthetase and a system of corresponding tRNA orthogonal pairs (such as products developed by Ambrx) are used to precisely insert non-natural amino acids containing azide groups at the TAG position of the APEX2 protein.
[0022] Preferably, the culture temperature is 23-26°C, for example, 23.5, 24, 24.5, 25 or 25.5°C, preferably 24.5-25.5°C, and the culture time is 22-26 hours, for example, 22.5, 23, 23.5, 24, 24.5, 25 or 25.5 hours, preferably 23.5-24.5 hours.
[0023] Thirdly, the present invention provides a proximity marker complex comprising the engineered peroxidase APEX2 described in the first aspect and a complementary nucleic acid sequence covalently linked thereto; the complementary nucleic acid sequence is complementary to the target RNA.
[0024] In this invention, a nucleic acid sequence complementary to the target RNA is linked to an engineered peroxidase APEX2, which can bind the engineered peroxidase APEX2 to the RNA and label proteins that interact with the RNA.
[0025] It is understood that the length and specific composition of the complementary nucleic acid sequence can be selected according to actual needs, as long as it can be complementary to and bind to the target RNA.
[0026] Preferably, the length of the complementary nucleic acid sequence is 20 to 40 nt, for example, it can be 21, 22, 23, 24, 25, 30, 35, 36, 37, 38 or 39 nt.
[0027] Preferably, the complementary nucleic acid sequence includes a polydT sequence.
[0028] In this invention, the polydT sequence can be complementary to the polyA tail of the mRNA.
[0029] Fourthly, the present invention provides a method for preparing the proximity-labeled complex described in the third aspect, the method comprising:
[0030] The engineered peroxidase APEX2 described in the first aspect is mixed with the complementary nucleic acid sequence modified with dibenzocyclooctyne to obtain the neighbor-labeled complex.
[0031] Preferably, the molar ratio of the engineered peroxidase APEX2 to the complementary nucleic acid sequence modified with dibenzocyclooctyne is 1:(1-20), for example, it can be 1:1, 1:2, 1:3, 1:4, 1:5, 1:6, 1:10, 1:12, 1:14, 1:15, 1:18 or 1:19, etc., preferably 1:(9-11).
[0032] Preferably, the mixing temperature is 4 to 40°C, for example, it can be 5, 6, 7, 8, 10, 15, 20, 25, 30, 35, 36, 37, 38 or 39°C, preferably 36 to 38°C, and the time is 1 to 5 hours, for example, it can be 1.5, 2, 2.5, 3, 3.5, 4 or 4.5 hours, preferably 1.5 to 2.5 hours.
[0033] In this invention, specific conditions are designed to further improve the binding efficiency of engineered peroxidase APEX2 to complementary nucleic acid sequences.
[0034] Fifthly, the present invention provides a method for labeling RNA-interacting proteins, the method comprising:
[0035] The cells were pretreated, and the pretreated cells were mixed with bovine serum albumin and incubated. The incubated cells were then mixed with the proximity labeling complex described in the third aspect, and biotin, phenol and hydrogen peroxide were added to carry out the reaction.
[0036] Preferably, the pretreatment includes formaldehyde cross-linking treatment and cell perforation treatment.
[0037] This invention further designs a method for labeling RNA-interacting proteins. Cells in a culture dish are cross-linked with formaldehyde to fix RNA-protein interactions under physiological conditions. Cells are perforated by breaking the cell membrane using methods such as Triton, and then bovine serum albumin (BSA) is added for incubation to block non-specific binding sites. A neighboring labeling complex is added for incubation to make it specifically bind to mRNA. After incubation, the cells are washed to remove unbound complexes, and H2O2 is added to activate biotin-phenol free radicals, causing them to covalently bind to neighboring proteins, thereby completing the specific labeling of RNA-binding proteins.
[0038] Preferably, the reaction temperature is 2 to 30°C, for example, 3, 4, 5, 6, 7, 8, 10, 15, 20, 25, 26, 27, 28 or 29°C, and more preferably 3 to 5°C, and the reaction time is 1 to 12 minutes, for example, 2, 3, 4, 5, 6, 7, 8 or 9 minutes, and more preferably 11 to 12 minutes.
[0039] The present invention designs specific marking conditions that can further improve marking efficiency.
[0040] Sixthly, the present invention provides a method for identifying RNA-interacting proteomes, the method comprising:
[0041] The RNA-interacting proteins were labeled using the method described in the fifth aspect, the labeled RNA-interacting proteins were then isolated and purified, and the isolated and purified RNA-interacting proteins were identified.
[0042] Preferably, the separation and purification method includes biotin-labeled purification methods, such as affinity purification using streptavidin agarose beads.
[0043] Preferably, the identification method includes Western blotting detection, liquid chromatography-mass spectrometry (LC-MS), and liquid chromatography-mass spectrometry (LC-MS / MS).
[0044] Preferably, based on the identified interacting protein spectral data, the mass spectrometry data can be analyzed using MaxQuant and ProteomeDiscoverer software to obtain protein identification (ID) and quantitative information; differential protein statistical analysis and screening can be performed using the Limma package in R / Bioconductor, and proteins interacting with mRNA can be inferred through functional and interaction network analysis.
[0045] Compared with the prior art, the present invention has at least the following beneficial effects:
[0046] This invention designs a novel peroxidase APEX2 complex, which binds peroxidase APEX2 to mRNA. Under the action of hydrogen peroxide, peroxidase APEX2 catalyzes the formation of biotin-phenoxy radicals from biotin-phenol to covalently label neighboring proteins. This allows for biotinylation of proteins bound to mRNA, followed by efficient and specific identification of mRNA-interacting proteins using affinity purification and mass spectrometry. Subsequently, bioinformatics analysis is used to obtain candidate mRNA-interacting proteins, and finally, techniques such as CLIP and WB are combined to verify the interaction between mRNA and candidate proteins, achieving highly efficient identification of the RNA-interacting proteome. Attached Figure Description
[0047] Figure 1 This is a schematic diagram illustrating the labeling principle of the adjacent labeling complex of the present invention;
[0048] Figure 2 A schematic diagram of the RNA-interacting proteome identification method;
[0049] Figure 3 This is a schematic diagram of the engineered peroxidase APEX2 expression vector;
[0050] Figure 4 Figure showing the protein expression results of (G4S)2-APEX2 and (G4S)3-APEX2;
[0051] Figure 5 Figure showing the protein expression results of (G4S)1-APEX2 and (G4S)4-APEX2;
[0052] Figure 6 The image shows the results of imidazole elution of (G4S)4-APEX2 protein onto a nickel column;
[0053] Figure 7A The image shows the labeling efficiency of Azf-APEX2 with linkers of different lengths. The left image is the fluorescence imaging image, and the right image is the Coomassie Brilliant Blue staining image.
[0054] Figure 7B This is a graph showing the results of engineered peroxidase activity detection;
[0055] Figure 8 The binding effects of engineered peroxidase with DBCO-modified polydT under different conditions are shown in the figure.
[0056] Figure 9 Gel electrophoresis image of the elution buffer from the adjacent labeled complex anion exchange column;
[0057] Figure 10 Gel electrophoresis images of protein samples from each stage of the ultrafiltration concentration process for adjacent labeled complexes;
[0058] Figure 11 The results of Western blotting analysis of biotin levels under different labeling conditions are shown in the figure.
[0059] Figure 12 The figure shows the results of Western blotting analysis to detect the effect of BSA on labeling performance.
[0060] Figure 13 The labeling effect of adjacent labeling complexes with different addition amounts is shown in the figure.
[0061] Figure 14 The image shows the results of fluorescent staining for RNA and biotin in cells.
[0062] Figure 15 A bar chart showing the distribution of RNA-interacting proteins;
[0063] Figure 16A This is a graph showing the results of hierarchical clustering analysis.
[0064] Figure 16B The two-dimensional planar plot of PCA analysis is shown, with the horizontal axis representing the first principal component and the vertical axis representing the second principal component. The distance between these points is approximately the sum of the differences in protein abundance among the samples.
[0065] Figure 17A A graph showing the results of identifying 1224 proteins using PrimAPEX;
[0066] Figure 17B Comparison of Venn diagrams for protein identification and annotation of RBP databases using PrimAPEX;
[0067] Figure 18 GO analysis diagram of proteins identified by PrimAPEX. Detailed Implementation
[0068] The technical solution of the present invention will be further described below with reference to the accompanying drawings and specific embodiments. However, the following examples are merely simplified examples of the present invention and do not represent or limit the scope of protection of the present invention. The scope of protection of the present invention is determined by the claims.
[0069] Where specific techniques or conditions are not specified in the examples, they shall be performed in accordance with the techniques or conditions described in the literature in this field, or in accordance with the product instructions. Reagents or instruments whose manufacturers are not specified are all conventional products that can be purchased from legitimate channels.
[0070] To achieve highly efficient identification of RNA-interacting proteomes, this invention first utilizes codon expansion technology to introduce non-natural amino acids containing azide groups into the peroxidase APEX2. Then, APEX2 is bound to a DBCO-modified complementary sequence (such as polydT) via click chemistry. The DBCO-modified complementary sequence then pairs with the mRNA sequence to specifically bind APEX2 to the vicinity of the mRNA, completing biotinylation labeling. Subsequently, affinity purification and mass spectrometry are used to efficiently and specifically identify mRNA-interacting proteins. Following this, bioinformatics analysis is used to obtain candidate mRNA-interacting proteins, and finally, CLIP and WB techniques are combined to verify the interaction between mRNA and candidate proteins.
[0071] (1) Development of the RNA interaction protein labeling technology "PrimAPEX"
[0072] To achieve specific labeling of APEX2 on mRNA-interacting proteins, this invention uses *E. coli* strain BL21 as the chassis cell and constructs a system comprising an engineered aminoacyl-tRNA synthetase and a corresponding orthogonal pair of tRNAs. Amber stop codons (TAGs) are introduced site-specifically into the APEX2 gene sequence using homologous recombination technology. The engineered aminoacyl-tRNA synthetase specifically links the azide-containing phenylalanine analog Azf (p-azido-phenylalanine, a non-natural amino acid) to its corresponding tRNA, thereby precisely embedding Azf at the TAG position in the APEX2 protein. Since the azide group in Azf can serve as a reactive group in bioorthogonal chemistry, this invention utilizes click chemistry to cycloaddite Azf-APEX2 with the complementary sequence modified by DBCO, thereby coupling them together. Then, based on the base pairing principle, this coupling complex binds to mRNA, allowing APEX2 to be efficiently and specifically guided to the vicinity of the polyA tail of the mRNA. Under the action of hydrogen peroxide, APEX2 catalyzes the formation of biotin-phenoxy radicals from biotin-phenol to covalently label neighboring proteins (technical principle diagram shown). Figure 1 As shown in the image, it is named "PrimAPEX".
[0073] APEX2 protein containing Azf modification was obtained through protein purification, and the peroxidase activity of Azf-APEX2 and the Azf embedding efficiency were tested. The binding efficiency of DBCO-modified polydT to form the polydT-APEX2 complex with Azf-APEX2 and the complementary pairing efficiency with mRNA were tested.
[0074] (2) Establishing a method for identifying RNA-interacting proteomes based on PrimAPEX technology
[0075] To efficiently identify RNA-interacting proteins, this invention uses the commonly used laboratory cell line HELA as a model to establish a PrimAPEX-based method. The technical roadmap is as follows: Figure 2As shown, firstly, cells in culture dishes were cross-linked with formaldehyde to fix RNA-protein interactions under physiological conditions. Cells were then perforated using a Triton assay to disrupt the cell membrane, followed by incubation with BSA to block non-specific binding sites. Then, a polydT-APEX complex was added for incubation to induce specific binding to mRNA. After incubation, cells were washed to remove unbound polydT-APEX complex, and H2O2 was added to activate biotin-phenol radicals, causing them to covalently bind to neighboring proteins, thus completing the specific labeling of RNA-binding proteins. Cells successfully labeled with biotin were added to lysis buffer, sonicated, and decross-linked at high temperature. Protein concentration was then determined using the BCA method. Streptavidin agarose beads were used to affinity purify potential mRNA-interacting proteins. A portion of the agarose beads was analyzed by Western blotting; the other portion was digested with trypsin buffer to extract polypeptide fragments, which were then identified by LC-MS / MS.
[0076] Based on the identified protein interaction profiles, this invention utilizes MaxQuant and ProteomeDiscoverer software to analyze LC-MS / MS data, obtaining protein identification (ID) and quantification information. Differential protein statistical analysis and screening are performed using the Limma package in R / Bioconductor, and proteins interacting with mRNA are inferred through functional and interaction network analysis.
[0077] Example 1
[0078] In this embodiment, the peroxidase APEX2 complex was prepared, and the RNA interaction protein labeling technology "PrimAPEX" was designed.
[0079] To precisely insert the phenylalanine analogue Azf into the APEX2 protein, the highly efficient *E. coli* BL21 strain was selected as the chassis cell, and the pET28b(+) plasmid was used as the expression vector. Through homologous recombination technology, a series of expression vectors containing TAG-linker-APEX2 of different lengths were constructed. Figure 3Using ATG as a control, a six-histidine tag was fused to the end to facilitate subsequent protein purification. Simultaneously, a plasmid containing an aminoacyl-tRNA synthetase / tRNA orthogonal pair (aminoacyl-tRNA synthetase nucleic acid sequence SEQ ID NO.2, amino acid sequence as shown in SEQ ID NO.3; tRNA nucleic acid sequence SEQ ID NO.4) was transformed into BL21 cells to form a strain expression system capable of specifically recognizing and inserting pAzf (para-azidophenylalanine). A single colony of the protein expression vector was inoculated into 5 mL of LB liquid medium (containing 50 μg / mL spectinomycin and 100 μg / mL ampicillin) and cultured overnight at 37°C and 200 rpm. 4 mL of the overnight culture was transferred to 500 mL of LB liquid medium containing 1 mM pAzf for expansion. The strain was cultured until it reached OD... 600 When the concentration reached 0.6, IPTG was added to a final concentration of 1 mM, and the cells were cultured further at different temperatures and times to promote protein expression. The cells were collected by centrifugation, resuspended in lysis buffer (NaCl 0.5 mol / L, Tris-HCl 20 mmol / L, pH 8.0), and sonicated on ice. The mixture was then filtered through a 0.45 μm filter. APEX2 with a 6×His tag was purified using a nickel affinity column with imidazole gradient elution. The purity of the purified protein was determined by SDS-PAGE. Protein concentration was determined by BCA.
[0080] SEQ ID NO.2:
[0081] 。
[0082] SEQ ID NO.3:
[0083] MDEFEMIKRNTSEIISEEELREVLKKDEKSAVIGFEPSGKIHLGHYLQIKKMIDLQNAGFDIIIYLADLHAYLNQKGELDEIRKIGDYNKKVFEAMGLKAKYVYGSEHGLDKDYTLNVYRLALKTTLKRARRSMELIAREDENPKVAEVIYPI MQVNGIHYEGVDVAVGGMEQRKIHMLARELLPKKVVCIHNPVLTGLDGEGKMSSSKGNFIAVDDSPEEIRAKIKKAYCPAGVVEGNPIMEIAKYFLEYPLTIKRPEKFGGDLTVNSYEELESLFKNKELHPMRLKNAVAEELIKILEPIRKRL.
[0084] SEQ ID NO.4:
[0085] ccggcggtagttcagcagggcagaacggcggactctaaatccgcatggcaggggttcaaatcccctccgccggacca.
[0086] The expression levels of APEX2 protein at different temperatures and times are as follows: Figure 4 and Figure 5 As shown, the expression level was higher when cultured at 25℃ for 24 hours.
[0087] The results of imidazole gradient elution are shown in Figure 6, yielding relatively pure pAzF-APEX2 protein.
[0088] The embedding efficiency of pAzf in APEX2 was detected using the fluorescent dye DBCO-Cy5. 10 μM DBCO-Cy5 was added to purified pAzf-APEX2, and after incubation at room temperature for 30 min, SDS-PAGE was performed. HEK293T cell protein lysis buffer (Lysate) was used as a negative control for fluorescence imaging. No fluorescence was observed with HEK293T cell protein lysis buffer, while Azf-APEX2 with four linkers of different lengths all showed fluorescence. Figure 7A This indicates that APEX2 has been successfully embedded in Azf.
[0089] The catalytic activity of Azf-APEX2 with different linker lengths was detected: HEK293T cell protein lysates were collected, and 1 μM pAzf-APEX2, 500 μM biotin-phenol, and 1 mM hydrogen peroxide were added to 160 μg of cell lysates for 1 min. The reaction was then terminated by adding 50 mM sodium ascorbate. The reaction was then analyzed by streptavidin-horseradish peroxidase on Western blotting. The catalytic activity of the four Azf-APEX2 proteins was also detected. The results showed that the APEX2 labeling effect of the four different linkers was basically consistent. Figure 7B Based on the theoretical premise that a longer linker length results in a larger label space, (G4S)4 Azf-APEX2 can be selected for subsequent experiments.
[0090] Optimization of the binding of Azf-APEX2 and DBCO-modified polydT: By co-incubating Azf-APEX2 and DBCO-modified polydT, different molar ratios, incubation times, and temperatures were designed to select the conditions with the best binding effect. The results are as follows: Figure 8 As shown, the pAzf-APEX2 protein exhibits the highest binding efficiency to the DBCO-modified primers under the incubation conditions of 37℃, 1h, and a molar ratio of 1:10.
[0091] To remove proteins and primers that did not participate in the formation of the polydT-APEX2 complex, an anion exchange column was used to bind the polydT-APEX2 complex. Unbound proteins were eluted with a low salt concentration, followed by elution with an increased salt concentration to remove the polydT-APEX2 complex. SDS-PAGE was then used to select proteins eluted with 500-900 mM NaCl. Figure 9 Finally, ultrafiltration was used to remove salt ions from the protein solution and concentrate the protein to obtain the final polydT-APEX2 complex. Figure 10 ).
[0092] Example 2
[0093] This embodiment is based on the PrimAPEX method designed in Example 1 for the identification of RNA-interacting proteomes.
[0094] (1) Cell Culture
[0095] Cultured in 100mm petri dishes for approximately 8 × 10⁸ cm⁻¹. 6 One HeLa cell (per sample) was incubated at 37°C and 5% CO2 until the cells reached 80% confluence (usually 1–2 days).
[0096] (2) Formaldehyde cross-linking and cell perforation
[0097] Perform the procedure in a fume hood, remove the DMEM medium, add 5 mL of 4% (v / v) formaldehyde, incubate on a shaker at room temperature for 10 min, remove the formaldehyde, add 5 mL of 50 mM glycine, incubate on a shaker at room temperature for 5 min, stop cross-linking, remove the formaldehyde, gently wash three times with PBS, add 5 mL of 0.3% Triton-X100, incubate at room temperature for 10 min, punch wells in the cell membrane, discard the Triton-X100, and gently wash three times with PBS.
[0098] (3) APEX2 tag
[0099] Add 5 mL of 3% BSA to PBS and incubate on a shaker at room temperature for 1 h. Remove the BSA and add 1 μg of polydT-APEX2 complex (diluted with 5 mL of 3% BSA in PBS) to each sample. Incubate gently on a shaker at room temperature for 1 h, discard the supernatant, wash three times with PBST, and twice with PBS. Dissolve biotin and phenol in PBS to a final concentration of 250 μM. Incubate the cells at room temperature for 30 min, then add H2O2 to a final concentration of 1 mM to induce biotinylation. Test incubation at different temperatures (4℃, 25℃) and times (1 min, 5 min, 10 min).
[0100] Pour off the culture medium, add 50 mM ascorbic acid, and incubate at room temperature for 5 min to terminate the reaction. Discard the supernatant, wash twice with PBS, add 150 μL of 4℃ RIPA lysis buffer (containing 50 mM Tris-HCl (pH 7.5), 150 mM NaCl, 1% (wt / vol) SDS, 0.5% (wt / vol) sodium deoxycholate, 1% (vol / vol) Triton X-100, 1 mM EDTA and 1× protease inhibitor), and incubate at 4℃ for 10 min. Collect cells using a cell scraper.
[0101] (4) High-temperature decrosslinking of proteins
[0102] The sample was sonicated at 30% power for 5 min (3 seconds on, 3 seconds off), and then incubated at 95°C with shaking for 1 h to decrosslink. It was then centrifuged at 10000×g for 10 min at 4°C, and the supernatant was collected (30 μL of supernatant from each sample was used as "input" for Western blot analysis). The concentration of each sample was determined using a BCA kit.
[0103] (5) Streptavidin agarose beads enrich biotinylated proteins
[0104] Add streptavidin agarose beads (“beads”, 30 μL per sample) to centrifuge tubes, centrifuge at 300 × g for 2 min at 4 °C, discard the supernatant, and wash twice with RIPA buffer 2 (50 mM Tris-HCl (pH 7.5), 50 mM NaCl, 0.25% (w / v) SDS, 0.125% (w / v) sodium deoxycholate, 0.25% (v / v) Triton, 1 mM EDTA, and 1 × protease inhibitor). Take 1 mg of protein from each sample, dilute 1% SDS in the RIPA lysis buffer to 0.25% with TBS, and place on ice. Add 30 μL of streptavidin agarose beads to each sample, incubate at 4 °C for 2 h by rotation, centrifuge at 300 × g for 2 min at 4 °C, wash twice with RIPA buffer 2, wash three times with 4 M Urea (dissolved in 50 mM Tris-HCl, pH 7.5), and 30% of the solution is added. Wash twice with ACN, then three times with 50mM TEAB, using 1mL for each wash. Take a portion of the agarose beads for Western blotting, and add the other portion to trypsin buffer for digestion to cut into polypeptide fragments. Identify the fragments using LC-MS / MS.
[0105] (6) Quality control analysis
[0106] For streptavidin Western blotting, add 5× protein loading buffer to the “beads” container, boil to elute biotinylated protein, run 4-20% SDS-PAGE on the “Input” and “bead” containers, transfer the protein to a 0.22 μm PVDF (Millipore) membrane, block in TBST with 5% BSA at room temperature for 1 h, add streptavidin-HRP (Beyotime, A0303, 1:5000 dilution) to TBST with 5% BSA, incubate at room temperature for 1 h, wash 3 times with TBST buffer for 5 min each time, develop using Clarity Western ECL substrate, and image using a ChemiDoc MP imaging system (Bio-Rad).
[0107] (7) Proteomics analysis
[0108] For LC-MS / MS analysis, after enrichment and washing, the beads were resuspended in 200 μL of bead digestion buffer, and 10 mM Tris(2-carboxyethyl)phosphine (TCEP) and 40 mM chloroacetamide (CAA) were added. The mixture was incubated at room temperature for 30 min, and the beads were washed with 1 mL of bead digestion buffer. Then, a digestion buffer containing 1 μg of mass spectrometry-grade trypsin (50 mM Tris-HCl, 1 mM CaCl2, and 2% ACN, pH 8.0) was added, and the mixture was digested at 37 °C for 16 h. After digestion, the mixture was centrifuged at 600 × g for 2 min, and the supernatant was collected in a new centrifuge tube. Buffer A (0.1% formic acid) was added to a final concentration of 0.01% to terminate the reaction. Before LC-MS / MS analysis, the peptide samples were desalted using StageTips. The C18 material was loaded into 200 μL pipettes to prepare StageTips (6 layers).
[0109] (8) Processing of biological mass spectrometry data
[0110] Maxquant software was used to search libraries and quantify proteins based on mass spectrometry identification results. The reference database was the Swiss-Prot Human fasta database. Variable modifications were set as methionine oxidation and N-terminal acetylation. Fixed modifications were set according to the experimental design, such as iodoacetamide alkylation of cysteine. The error ranges for primary and secondary mass spectrometry were set according to the type of mass spectrometer used (e.g., 20 ppm and 50 mmu for QE-HF data). The minimum peptide length was 7, and the minimum number of independent peptides was 2. A mixed positive and negative library search strategy was used to assess the false positive rate of peptides. Peptides and proteins with an FDR of less than 0.01 at both the peptide and protein levels were retained. Protein quantification was performed using the iBAQ method.
[0111] (9) Utilizing biostatistics and bioinformatics techniques for omics data analysis
[0112] First, the data obtained from the database search was filtered for contaminating proteins, missing values were filled in, and intensity-based absolute quantification (iBAQ) was used as the quantitative basis. Differential protein statistical analysis and screening were performed using the Limma package in R / Bioconductor to infer protein complexes that specifically interact with specific RNAs. Second, the differentially expressed proteins were annotated with functional, pathway, and network information: using the DAVID tool (http: / / david.abcc.ncifcrf.gov / ), extraction was performed based on GO (Gene Ontology) annotation information for biological processes, molecular functions, and cellular components. Pathways and known regulatory networks were annotated using the Ingenuity Pathway Analysis (IPA) tool. Cluster analysis was performed using the K-means algorithm, implemented in Matlab; co-expression analysis was performed using the WGCNA algorithm, implemented in the R toolkit.
[0113] Western blotting analysis of biotinylated proteins revealed a significant increase in biotinylated protein levels at 25°C for 10 min, followed by the 4°C, 10 min group. The labeling effects at the remaining conditions were largely consistent. Figure 11 To avoid non-specific labeling due to excessively long labeling time while ensuring labeling efficiency, a labeling condition of 4℃ for 10 minutes can be selected.
[0114] To investigate the effect of adding BSA for incubation before labeling to block non-specific sites on the labeling effect, control groups were further designed, including groups with pAzf-APEX2 alone or polydT modified with DBCO alone. The experimental group was the polydT-APEX2 complex group. The effects of adding or not adding BSA were tested separately. Western blotting results are shown below. Figure 12 As shown, "-" indicates that no corresponding substance was added, and "+" indicates that the corresponding substance was added. This indicates that the non-specific labeling was significantly reduced after the addition of BSA blocking, and the labeling effect of the polydT-APEX2 group was enhanced.
[0115] To verify the minimum amount of polydT-APEX2 complex added, 0.5 μg, 1 μg, 2 μg, and 3 μg of polydT-APEX2 were added for labeling. The biotinylated protein was then bound with Streptavidin-HRP antibody. Western blotting results were obtained. Figure 13 This indicates that adding 1 μg of the complex can achieve a good labeling effect.
[0116] To investigate whether RNA and biotinylated proteins co-localize, fluorescent staining was used to stain RNA and biotin, respectively. The results are as follows: Figure 14 As shown, the RNA and biotinylated protein sites overlap, and the polydT-APEX2 group has significantly more biotinylated protein than the control group.
[0117] Further analysis of the proteomic identification results was conducted.
[0118] 1) Identification of biotinylated proteins using mass spectrometry
[0119] RNA-interacting proteins were labeled using the PrimAPEX method. The experimental group consisted of polydT-APEX (two replicates), while the control group consisted of either APEX2 or polydT alone. Each group had three biological replicates, totaling 12 samples. Maxquant software was used to search the library and quantify proteins based on the mass spectrometry identification results. The minimum peptide length was 7, and the minimum number of independent peptides was 2. Peptides and proteins with an FDR (Free Path Difference) less than 0.01 were retained. Contaminating proteins were filtered from the library search data, and missing values were filled in. A total of 631 proteins were identified in the APEX2 group, 814 proteins in the polydT group, 2436 proteins in polydT-APEX2Rep1, and 2477 proteins in polydT-APEX2 Rep2. The protein distribution for each sample is shown in the figure. Figure 15 .
[0120] 2) Proteomics data analysis
[0121] Unsupervised clustering was used to analyze the identified proteins. Samples from the same group could be clustered together, and the differences between groups were significant. Figure 16A To assess inter-group differences and intra-group sample replication, we performed unsupervised principal component analysis (PCA) on all identified proteins, allowing groups to cluster together. Figure 16B ).
[0122] 3) PrimAPEX identified 1224 RNA-interacting proteins.
[0123] Proteins with differential upregulation were screened based on a p-value < 0.05 and a fold change (log2) > 1 between the experimental and control groups, resulting in 307 differentially upregulated proteins. Proteins identified in every replicate of the experimental group but not in any replicate of the control group were defined as unique proteins of the experimental group, totaling 917. In total, 1224 proteins were identified and defined as RNA-interacting proteins identified by PrimAPEX. Figure 17ABased on the RBP2GO database (https: / / rbp2go.dkfz.de / ), literature on RBP research from the past five years was added, and proteins identified in more than two publications were selected and compiled into a new database called Annotated RBPs. This database contains 4659 proteins. The 1224 RNA-interacting proteins identified by PrimAPEX technology were compared with the Annotated RBPs database; 87% of the proteins identified by PrimAPEX overlapped with this database. Figure 17B ).
[0124] 4) Most of the proteins identified by PrimAPEX are related to RNA biological processes.
[0125] GO analysis was performed on 1224 PrimAPEX-identified RNA-interacting proteins using Metascape. The results showed that most of the 1224 PrimAPEX-identified proteins were related to RNA biological processes. Figure 18 This indicates that the PrimAPEX technology has good specificity and can be used for RNA interaction proteome identification.
[0126] In summary, this invention designs a novel peroxidase APEX2 complex, introducing a non-natural amino acid containing an azide group into APEX2 at a specific site. Through click chemistry, APEX2 binds to a DBCO-modified complementary sequence. This complementary sequence can bind to the target RNA, thereby binding APEX2 to the mRNA. Under the action of hydrogen peroxide, APEX2 catalyzes the formation of biotin-phenoxy radicals from biotin-phenol to covalently label neighboring proteins, enabling biotinylation of proteins bound to mRNA. Subsequently, affinity purification and mass spectrometry are used to efficiently and specifically identify mRNA-interacting proteins. Following this, bioinformatics analysis is used to obtain candidate mRNA-interacting proteins, and finally, CLIP and WB techniques are combined to verify the interaction between mRNA and candidate proteins, achieving highly efficient identification of the RNA-interacting proteome.
[0127] The applicant declares that the above description is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Those skilled in the art should understand that any changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention fall within the protection and disclosure scope of the present invention.
Claims
1. A method for labeling RNA-interacting proteins, characterized in that, The method for labeling RNA-interacting proteins includes: Cells were pretreated, and the pretreated cells were mixed with bovine serum albumin and incubated. The incubated cells were then mixed with a neighboring labeling complex, and biotin, phenol, and hydrogen peroxide were added to initiate the reaction. The pretreatment consists of formaldehyde cross-linking treatment and cell perforation treatment; The proximity marker complex comprises engineered peroxidase APEX2 and a complementary nucleic acid sequence covalently linked thereto; the complementary nucleic acid sequence is complementary to the target RNA. The structural formula of the engineered peroxidase APEX2 is shown in Formula I. AL-APEX2 Formula I; In this context, A represents a non-natural amino acid containing an azide group, L represents a linker, and APEX2 represents peroxidase APEX2.
2. The method for labeling RNA-interacting proteins according to claim 1, characterized in that, The non-natural amino acids containing azide groups include phenylalanine analogs containing azide groups.
3. The method for labeling RNA-interacting proteins according to claim 2, characterized in that, The phenylalanine analogs containing azido groups include p-azidophenylalanine or 3-azido-L-phenylalanine.
4. The method for labeling RNA-interacting proteins according to claim 1, characterized in that, The connector comprises (Gly-Gly-Gly-Gly-Ser)n, where n is a positive integer.
5. The method for labeling RNA-interacting proteins according to claim 1, characterized in that, The preparation method of the engineered peroxidase APEX2 includes: A succinate stop codon was introduced into the APEX2 peroxidase encoding gene, and an expression vector was inserted to obtain a recombinant vector. The recombinant vector was introduced into a host cell, which also expressed an orthogonal aminoacyl-tRNA synthetase and tRNA to obtain recombinant cells. The recombinant cells were cultured, and non-natural amino acids containing azide groups were added to the culture medium. After the culture was completed, the culture medium was collected, separated, and purified to obtain the engineered peroxidase APEX2.
6. The method for labeling RNA-interacting proteins according to claim 5, characterized in that, The culture temperature is 23~26℃ and the time is 22~26 h.
7. The method for labeling RNA-interacting proteins according to claim 1, characterized in that, The length of the complementary nucleic acid sequence is 20-40 nt.
8. The method for labeling RNA-interacting proteins according to claim 1, characterized in that, The complementary nucleic acid sequence includes a polydT sequence.
9. The method for labeling RNA-interacting proteins according to claim 1, characterized in that, The method for preparing the proximity-labeled complex includes: The engineered peroxidase APEX2 was mixed with the complementary nucleic acid sequence modified with dibenzocyclooctyne to obtain the neighbor-labeled complex.
10. The method for labeling RNA-interacting proteins according to claim 9, characterized in that, The molar ratio of the engineered peroxidase APEX2 to the complementary nucleic acid sequence modified with dibenzocyclooctylene is 1:(1~20).
11. The method for labeling RNA-interacting proteins according to claim 9, characterized in that, The mixing temperature is 4~40℃, and the time is 1~5 h.
12. The method for labeling RNA-interacting proteins according to claim 1, characterized in that, The reaction is carried out at a temperature of 2-30°C for 1-12 minutes.
13. A method for identifying RNA-interacting proteomes, characterized in that, The method for identifying RNA-interacting proteomes includes: The RNA-interacting proteins are labeled using the method described in claim 1, the labeled RNA-interacting proteins are then isolated and purified, and the isolated and purified RNA-interacting proteins are identified.