Methods for extracting RNA-binding proteins

The method of forming RNA-protein adducts and using enzymatic cleavage with DNA oligonucleotides and RNase H addresses the inefficiencies of current RNA-centric methods, providing efficient and reproducible RNA-protein interaction profiling with region-specific mapping.

JP2026508293APending Publication Date: 2026-03-10QUEEN MARY UNIV OF LONDON
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-28
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Current RNA-centric methods for profiling protein partners of specific RNAs suffer from low efficiency, reproducibility, and off-target capture, making comprehensive discovery of RNA-protein interactions challenging.

Method used

A method involving cross-linking agents to form RNA-protein adducts, followed by phase separation and enzymatic cleavage using DNA oligonucleotides and RNase H to specifically release and identify RNA-binding proteins, enabling region-specific mapping of interactions.

Benefits of technology

Achieves efficient, reproducible, and specific profiling of RNA-protein interactions with enhanced region-specific mapping capabilities, surpassing existing pull-down-based assays in accuracy and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026508293000010
    Figure 2026508293000010
  • Figure 2026508293000011
    Figure 2026508293000011
  • Figure 2026508293000012
    Figure 2026508293000012
Patent Text Reader

Abstract

The present invention relates to novel methods for isolating RNA-binding proteins. The present invention also provides methods for identifying RNA-binding proteins and for identifying binding sites of RNA-binding proteins in RNA transcript sequences.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to novel methods for isolating RNA-binding proteins. The present invention also provides methods for identifying RNA-binding proteins and for identifying binding sites of RNA-binding proteins in RNA transcript sequences. [Background technology]

[0002] The recent success of COVID mRNA vaccines has demonstrated the great potential of the field of RNA biology in bringing transformative innovations to the biomedical and biotech sectors. A key aspect of RNA biology research is the comprehensive study of RNA-protein interactions. Comprehensive discovery and mapping of RNA molecules that bind to proteins of interest in cellulo can be easily performed thanks to recent advances in methods such as RNA immunoprecipitation sequencing (RIP-seq), individual-nucleotide resolution crosslinking and immunoprecipitation (iCLIP), or enhanced crosslinking and immunoprecipitation (eCLIP). These "protein-centric" methods utilize the sensitivity of next-generation sequencing to reveal and map transcriptome-wide targets of a given protein (Hafner et al., 2021).

[0003] In contrast, comprehensive discovery of proteins that bind to endogenous RNAs of interest is much more challenging, requiring purification of these RNAs from cells followed by mass spectrometry (MS) of the associated proteins. Such "RNA-centric" methods typically rely on antisense oligo-based pulldown of target RNAs from RNA-protein cross-linked cell lysates, followed by quantitative MA analysis to reveal the identity of specifically interacting proteins (Hafner et al., 2021). However, despite significant interest in this field, all available RNA-centric approaches suffer from several significant drawbacks. These include low efficiency of antisense oligo-based RNA pulldown, a significant lack of reproducibility between different approaches, and a significant risk of off-target capture resulting in low specificity.

[0004] Thus, there is currently a major gap in the field of RNA biology for new methods that can enable efficient, robust, and scalable profiling of the protein partners of specific RNAs. The present invention fills this major gap. As described in more detail herein, in brief, the present invention provides the field of RNA biology with a set of revolutionary new RNA-centric methods for exploring RNA-protein interactions that are far more specific, efficient, and reproducible than existing methods, and that additionally enable region-specific mapping of RNA-protein interactions. Summary of the Invention

[0005] The present inventors have identified a novel method for isolating RNA-binding proteins from a sample, which in turn makes it possible to identify the RNA-binding proteins and further to identify the binding sites of the RNA-binding proteins in the RNA transcript sequence.

[0006] The method involves the use of multiple DNA oligonucleotides that hybridize to specific target regions in an RNA molecule of interest. The method subsequently involves the use of an enzyme to cleave the DNA-RNA hybrid, thereby releasing proteins previously crosslinked to the RNA molecule of interest. The released proteins can then be identified by known protein identification techniques, such as quantitative mass spectrometry. The use of an enzyme with specific activity against DNA-RNA hybrids, along with sequence-specific probe design, results in the method of the present invention achieving unprecedented activity: only the specific target RNA of interest undergoes cleavage and release of its associated proteins. Such an enzyme-based approach is significantly more efficient than alternative methods using RNA pull-down. Furthermore, because enzyme-based cleavage can be specifically targeted to specific regions of a given RNA, the method of the present invention enables region-specific mapping of interacting proteins from endogenous samples, a capability not possible with other existing pull-down-based assays.

[0007] The present invention provides a method comprising the steps of: (a) contacting a sample with a cross-linking agent, thereby providing RNA-protein adducts; (b) performing phase separation on the sample containing the RNA-protein adducts to produce a separate phase comprising the RNA-protein adducts but at least substantially free of free protein and free RNA; (c) isolating the RNA-protein adducts from the separated phases of part (b); (d) contacting the isolated RNA-protein adduct with a plurality of DNA oligonucleotides containing complementary nucleotide sequences to the target RNA molecule of interest, thereby forming one or more DNA-RNA hybrids; (e) contacting the sample from part (d) with an enzyme that cleaves DNA-RNA hybrids, thereby liberating proteins from the RNA molecules of interest.

[0008] The method optionally further comprises isolating the RNA binding protein from the sample, the method comprising the steps of: (f) subjecting the sample to a further step of phase separation to produce a separate phase containing free RNA-binding protein but at least substantially free of free RNA; and (g) Isolating the released RNA-binding protein from the separated phases of part (f).

[0009] The present invention further provides a method for identifying an RNA binding protein, which method comprises the steps of the method for isolating an RNA binding protein from a sample according to the present invention, and further comprises subjecting the released RNA binding protein to a technique for protein identification.

[0010] The present invention further provides a method for identifying binding sites of an RNA-binding protein in an RNA transcript sequence, the method comprising the steps of the method for isolating an RNA-binding protein from a sample according to the present invention together with the method for identifying an RNA-binding protein according to the present invention, and further comprising mapping a plurality of DNA oligonucleotides comprising complementary nucleotide sequences to a target RNA molecule of interest to the RNA transcriptome, and inferring binding sites of the released RNA-binding protein in the RNA transcript sequence based on the DNA oligonucleotides that have been mapped to the RNA transcriptome. [Brief explanation of the drawings]

[0011] [Figure 1] Schematic diagram of TREX. [Figure 2] FIG. 1 shows an example of poor overlap across three different antisense-based transcript-specific RNA capture studies aimed at uncovering interactors of U1 RNA, a well-characterized non-coding RNA component of the spliceosome core. [Figure 3]FIG. 1 shows a comparison of the TREX method with the most common antisense-based transcript-specific RNA capture methods (RAP, CHART, CHIRP). [Figure 4] FIG. 1 shows that retrospective assessment of efficiency and specificity is possible for TREX by including an additional QC step (red box marked "QC") performed on fractions after the RNase H treatment reaction (RNA is first extracted by proteinase K treatment, followed by RT-qPCR analysis (to test for degradation efficiency) or RNA sequencing (to test for degradation specificity). [Figure 5] Figure 1 shows the evaluation of the efficiency of 18S degradation by RT-qPCR from fractions of +RNase H-treated and -RNase H-treated samples. The abundance of 18S rRNA transcripts relative to GAPDH (a housekeeping gene) mRNA was quantified in each sample. An almost complete loss of 18S rRNA was evident in the +RNase H sample. [Figure 6] Figure 1 shows the interactome of 18S as revealed by TREX. Volcano plot of t-test results (FDR<0.05) comparing +RNase H-treated and -RNase H-treated 18S TREX samples. The vast majority of 40S ribosomal proteins, as well as many known 40S ribosomal subunit-associated factors, are among the confidently identified interactors of 18S (right-hand side of graph). [Figure 7] Figure 1 shows category enrichment analysis of 18S interacting proteins from TREX. The list of identified interactors of 18S from Figure 6 was subjected to category enrichment analysis using Fisher's exact test (FDR<0.02), and protein annotations were taken from the Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) databases. The top 20 enriched categories were selected and depicted. [Figure 8]Analysis of known protein-protein interactions of 18S binding proteins from TREX using the STRING database. The vast majority of identified proteins are part of the network of proteins around the 40S core ribosomal subunit (highlighted by the bold circle). [Figure 9] Venn diagram of overlap between the 18S TREX and RAP-MS datasets (McHugh et al., 2015). While most 40S ribosomal proteins are among the shared targets of both methods (overlap), TREX identifies several other 40S proteins, as well as many more known 18S interactors (e.g., dozens of 18S biogenesis factors, many translation initiation factors, the SRP-dependent ER targeting machinery, and several large (60S) ribosomal proteins that primarily contact the small subunit). On the other hand, RAP-MS exclusively reveals several large subunit biogenesis factors, as well as several nucleolar proteins. Because the majority of both small and large subunit RNAs begin as joint, unprocessed transcripts (45S rRNA), it is possible that these transcripts are co-captured in RAP-MS prior to their processing and separation from each other. Because TREX acts in a region-specific manner, large subunit biogenesis factors are not expected to copurify, even before processing of the 45S rRNA (simply because these factors are not directly associated with the 18S portion). [Figure 10] Figure 1 shows the evaluation of the efficiency of U1 RNA degradation by RT-qPCR from fractions of +RNase H-treated and -RNase H-treated samples. The abundance of U1 RNA transcript relative to GAPDH (a housekeeping gene) mRNA was quantified in each sample. Near-complete loss of U1 RNA was evident in the +RNase H sample. [Figure 11]Figure 1 shows the interactome of U1 RNA as revealed by TREX. Volcano plot of t-test comparison of +RNase H-treated and -RNase H-treated U1 TREX samples (FDR<0.05). Proteins specifically extracted upon U1 degradation are on the right-hand side of the graph. The vast majority of U1 snRNP proteins and several splicing factors were among the confidently identified interactors of U1. [Figure 12] Figure 1 shows an analysis of known protein-protein interactions of U1-binding proteins from TREX using the STRING database. As expected, the vast majority of identified proteins are part of the known network of U1 snRNP proteins and splicing-related factors, but a small independent complex of redox regulators (PRDX5 and TXN) was also identified, suggesting a possible novel role for U1 in redox regulation. [Figure 13] Venn diagram of overlap between the U1 TREX and RAP-MS datasets (McHugh et al., 2015). Most U1 snRNP proteins are among the shared targets of both methods (overlap), but TREX identifies additional splicing factors and one additional U1 snRNP (SNRPE). [Figure 14] Schematic diagram of the different regions of the 45S pre-rRNA. The 18S, 5.8S, and 28S regions ultimately constitute the majority of the rRNA on the ribosome, while the 5' ETS, ITS1, ITS2, and 3' ETS regions are the last to be processed and removed. [Figure 15] Figure 1 shows the region-specific interactome of the 5'ETS as revealed by TREX. Volcano plot of t-test comparison of +RNase H-treated and -RNase H-treated 5'ETS TREX samples (FDR<0.05). Proteins specifically extracted upon 5'ETS degradation are on the right-hand side of the graph. Eight known ribosome biogenesis factors were among the confidently identified interactors of the 5'ETS region. [Figure 16] Figure 1 shows an analysis of known protein-protein interactions of 5'ETS-binding proteins from TREX using the STRING database. As expected, 8 of the 16 identified proteins are part of a known network of ribosome biogenesis factors, including UTP14A, UTP15, and WDR75, known to be important for early biogenesis events involving the 5'ETS.

[0012] [Table 1-1] [Table 1-2]

[0013] [Table 2]

[0014] [Table 3-1] [Table 3-2] [Table 3-3] DETAILED DESCRIPTION OF THE INVENTION

[0015] The present invention will be described with respect to particular embodiments and with reference to certain drawings, but the disclosure is not limited thereto, but only by the claims. Any reference signs in the claims should not be construed as limiting the scope. Of course, it should be understood that not necessarily all aspects or advantages may be achieved in accordance with any particular embodiment. Thus, for example, one skilled in the art will recognize that the disclosed embodiments may be embodied or performed to achieve or optimize one advantage or group of advantages as taught herein, without necessarily achieving other aspects or advantages as taught or suggested herein.

[0016] The disclosure, relating to both the organization and method of operation, along with its features and advantages, can best be understood by reference to the following detailed description when read in conjunction with the accompanying drawings. Aspects and advantages of the present invention will be apparent from and elucidated with reference to the embodiment(s) described below. References throughout this specification to "one embodiment" or "an embodiment" mean that a particular feature, structure, or characteristic described in connection with that embodiment is included in at least one disclosed embodiment. Thus, the appearances of the phrases "in one embodiment" or "in an embodiment" in various places throughout this specification may apply to that embodiment, although not necessarily all of them. Similarly, it should be recognized that in the description of exemplary disclosed embodiments, various features may be grouped together in a single embodiment, figure, or description thereof for the purpose of streamlining the disclosure and aiding in the understanding of one or more of the various inventive aspects. However, this method of disclosure should not be interpreted as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive aspects lie in less than all features of a single foregoing disclosed embodiment.

[0017] It is to be recognized that "embodiments" of the present disclosure may be specifically combined together unless the context dictates otherwise. Any specific combination of disclosed embodiments is an additional disclosed embodiment of the claimed invention (unless the context implies otherwise).

[0018] Additionally, as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the content clearly dictates otherwise. Thus, for example, reference to "a polynucleotide" includes two or more polynucleotides; reference to "a protein" includes two or more proteins, etc.

[0019] Unless otherwise indicated, nucleic acid sequences herein are written left to right in 5' to 3' orientation.

[0020] All publications, patents, and patent applications cited herein, whether supra or infra, are hereby incorporated by reference in their entirety.

[0021] Methods for isolating RNA-binding proteins The present disclosure provides a method comprising the steps of: (a) contacting a sample with a cross-linking agent, thereby providing RNA-protein adducts; (b) performing phase separation on the sample containing the RNA-protein adducts to produce a separate phase comprising the RNA-protein adducts but at least substantially free of free protein and free RNA; (c) isolating the RNA-protein adducts from the separated phases of part (b); (d) contacting the isolated RNA-protein adduct with a plurality of DNA oligonucleotides containing complementary nucleotide sequences to the target RNA molecule of interest, thereby forming one or more DNA-RNA hybrids; (e) contacting the sample from part (d) with an enzyme that cleaves DNA-RNA hybrids, thereby liberating proteins from the RNA molecules of interest.

[0022] The method includes the step of providing an RNA-protein adduct. Preferably, the protein is an RNA-binding protein. The method includes the step of isolating the RNA-protein adduct, and thus preferably, isolating the RNA-binding protein.

[0023] The method may further comprise isolating the RNA binding protein from the sample, the method comprising the steps of: (f) subjecting the sample to a further step of phase separation to produce a separate phase containing free RNA-binding protein but at least substantially free of free RNA; and (g) Isolating the released RNA-binding protein from the separated phases of part (f).

[0024] "Sample" is used herein to refer to any material containing at least one protein bound to RNA. A sample preferably contains multiple proteins bound to different RNA molecules and / or different regions of the RNA molecules.

[0025] Exemplary samples can be soil samples, or any material or tissue sample obtained from a plant, animal, or microorganism. The sample material can be tissue. Preferred animal materials include hair follicles and bodily fluids such as blood, saliva, semen, vaginal fluid, mucus, urine, or any other bodily fluid material. The sample can include material derived from cells cultured in vitro or material derived from human or animal tissue, optionally the material includes a lysate of cells or human or animal tissue. The sample can include cells undergoing in vitro culture, optionally the cells are attached to a cell culture plate or suspended in a solution.

[0026] The material may comprise a lysate of cells or human or animal tissue, which lysate is further subjected to subcellular fractionation. Subcellular fractionation methods are well known in the art. The cellular fraction may be, for example, a whole cell fraction, a cytoplasmic fraction, or a nuclear fraction. Preferably, the lysate is prepared under conditions that are free of active ribonuclease enzymes.

[0027] The number of cells in the sample can be any suitable amount, provided that the cross-linking procedure is routinely adapted to ensure that a sufficient number of cells are cross-linked and RNA-protein adducts are successfully delivered. A sample should contain at least about 1 x 10 cells. 6 pieces, optionally at least about 2.5 x 10 6 pieces, 5×10 6 pieces, 10×10 6 pieces, or 100 x 10 6 Most preferably, the sample contains at least about 5 x 10 cells. 6 It contains cells.

[0028] The "RNA-binding protein" isolated in the methods disclosed herein can be any RNA-binding protein. The protein may bind to RNA directly, preferably through a non-covalent interaction. The protein may bind to RNA indirectly, for example, through an interaction with another protein, preferably, the interaction is non-covalent. Most preferably, the protein binds to RNA directly through a non-covalent interaction.

[0029] The sample may contain any one or more types of RNA. RNAs include messenger RNA (mRNA), ribosomal RNA (rRNA), signal recognition particle RNA (7SL RNA or SRP RNA), transfer RNA (tRNA), transfer-messenger RNA (tmRNA), small nuclear RNA (snRNA), small nucleolar RNA (snoRNA), SmY RNA (SmY), small Cajal body-specific RNA (scaRNA), guide RNA (gRNA), ribonuclease P (RNase P), ribonuclease MRP (RNase MRP), Y RNA, telomerase RNA component (TERC), splice leader RNA (SL RNA), antisense RNA (aRNA, asRNA), cis-natural antisense transcript (cis-NAT), CRISPR RNA (crRNA), long non-coding RNA (lncRNA), microRNA (miRNA), piwi-interacting RNA (piwi-interacting The sample may be a ribosomal RNA (ribosomal RNA), a small interfering RNA (piRNA), a small interfering RNA (siRNA), a small hairpin RNA (shRNA), a trans-acting siRNA (tasiRNA), a repeat associated siRNA (rasiRNA), a 7SK RNA (7SK), an enhancer RNA (eRNA), a retrotransposon, a viral genome, a viroid, or a satellite RNA. Most preferably, the sample contains rRNA, mRNA, and / or ncRNA.

[0030] The RNA bound by the RNA-binding protein can be any known RNA sequence. The RNA can be a wild-type sequence, including splice variants. Alternatively, the RNA sequence can be a mutant sequence. The known sequence can be the entire RNA molecule or a region contained within the RNA molecule. Knowing the sequence of a particular RNA means that a DNA oligonucleotide can be designed containing a nucleotide sequence complementary to the sequence within the RNA, thereby enabling isolation of the RNA-binding protein by the methods disclosed herein.

[0031] The method includes contacting a sample with a cross-linking agent, thereby providing RNA-protein adducts. The cross-linking agent functions to induce cross-links between RNA and protein molecules that are in close proximity to one another. Typically, the proximity is such that the RNA and protein molecules are determined to be bound to one another via non-covalent interactions. The cross-linking agent forms one or more covalent bonds between such RNA and protein molecules, thereby "locking" the non-covalent interactions and thereby providing one or more RNA-protein adducts.

[0032] Therefore, the crosslinking agent can be any agent capable of inducing crosslinks between interacting proteins and RNA molecules. Suitable crosslinking agents are well known in the art. For example, the sample can be irradiated and / or contacted with one or more chemicals to induce crosslink formation and, accordingly, RNA-protein adducts. The crosslinking agent preferably comprises ultraviolet radiation or formaldehyde. The ultraviolet radiation can be of any suitable wavelength, but is preferably 365 nm, or more preferably 254 nm. When UV 365 nm is used to contact the sample to form RNA-protein adducts, the sample is preferably also contacted with 4-thiouridine.

[0033] The method includes performing phase separation on a sample containing RNA-protein adducts to produce a separate phase that contains the RNA-protein adducts but is at least substantially free of free protein and free RNA.

[0034] The phase separation step is preferably a phase separation method comprising separating a liquid sample into an organic phase and an aqueous phase. In the method of the present invention, in which the sample is contacted with a cross-linking agent, thereby providing RNA-protein adducts, an intermediate phase additionally forms along with the organic and aqueous phases.

[0035] Preferably, all or substantially all of the free RNA, i.e., RNA that is not cross-linked to protein and therefore not contained within RNA-protein adducts, is contained within the aqueous phase. Preferably, all or substantially all of the free protein, i.e., protein that is not cross-linked to RNA, is contained within the organic phase. The RNA-protein adducts are contained within a separate phase that contains the RNA-protein adducts but is at least substantially free of free protein and free RNA, and preferably, the separate phase corresponds to an interphase.

[0036] Liquid phase separation methods are generally known in the art. The phase separation step preferably includes the use of one or more of acidic phenol, guanidine isothiocyanate, and chloroform. Even more preferably, the phase separation step preferably includes the use of all of acidic phenol, guanidine isothiocyanate, and chloroform. The use of one or more of these agents includes contacting the sample with the one or more agents, thereby inducing phase separation. General liquid phase separation methods including the use of the one or more agents are known in the art, and are particularly known to include one or more steps of mixing and centrifugation, as described in the examples herein.

[0037] The method includes isolating the RNA-protein adducts from the separated phases in the phase separation step. Such isolation can be achieved by conventional methods of liquid extraction from a liquid sample. The isolation step can include removing the aqueous and organic phases from a vessel containing the sample, leaving an interphase contained in said vessel, thereby allowing the RNA-protein adducts to be isolated from the interphase.

[0038] The phase separation and isolation steps together can be repeated one or even more times in sequence before beginning to contact the RNA-protein adduct with multiple DNA oligonucleotides.

[0039] Prior to contacting the RNA-protein adducts with the plurality of DNA oligonucleotides, the sample may also be contacted with a ribonuclease inhibitor and an agent that digests genomic DNA (preferably, the agent is a deoxyribonuclease).

[0040] The method includes contacting the isolated RNA-protein adduct with a plurality of DNA oligonucleotides comprising complementary nucleotide sequences to a target RNA molecule of interest, thereby forming one or more DNA-RNA hybrids, wherein one or more, or all, of the plurality of DNA oligonucleotides can consist of a nucleotide sequence that is complementary to the target RNA molecule of interest.

[0041] The complementary nucleotide sequence allows the DNA oligonucleotide to anneal with the complementary RNA sequence to form a DNA-RNA hybrid. The complementary nature of the DNA oligonucleotides means that the method can be used to target the DNA oligonucleotides to one or more specific regions of the RNA that the user wishes to examine. This allows the user to isolate and thereby determine the RNA-binding proteins that bind to the one or more specific regions. The DNA oligonucleotides preferably do not overlap in sequence within the targeted RNA molecule of interest, and therefore do not compete with each other in annealing to the one or more specific regions.

[0042] A plurality of DNA oligonucleotides can be tiled to one or more specific regions of a target RNA molecule, or to the entire, or optionally substantially the entire, of one or more specific RNA transcripts.When one or more specific regions of a target RNA molecule are to be investigated, it is preferred that a plurality of DNA oligonucleotides be tiled to the entire, or optionally substantially the entire, of one or more specific regions, so that the entire one or more specific regions anneal with the DNA oligonucleotides.When the entire one or more specific RNA transcripts are the target of interest, it is preferred that a plurality of DNA oligonucleotides be tiled to the entire, or optionally substantially the entire, of one or more specific RNA transcripts, so that the entire RNA transcripts anneal with the DNA oligonucleotides.

[0043] The DNA oligonucleotide may be of any appropriate length so that specific annealing with the target RNA can be achieved. Preferably, the DNA oligonucleotide is at least about 20 nucleotides long, at least about 30 nucleotides long, at least about 40 nucleotides long, at least about 50 nucleotides long, or at least about 60 nucleotides long. More preferably, the DNA oligonucleotide is about 30 to about 90 nucleotides long, and most preferably, the DNA oligonucleotide is about 60 nucleotides long. Preferably, the DNA oligonucleotide is unmodified.

[0044] The method includes contacting a sample with an enzyme that cleaves DNA-RNA hybrids, thereby liberating proteins from the RNA molecule of interest. Enzymes capable of cleaving DNA-RNA hybrids are known in the art. Furthermore, means for testing the efficiency of such cleavage are well known in the art. Preferably, the enzyme that cleaves DNA-RNA hybrids in the method of the present invention is RNase H or a functionally active fragment or analog thereof. Preferably, the RNase H or a functionally active fragment or analog thereof is thermostable, thereby capable of cleaving DNA-RNA hybrids at a temperature of at least 50°C. Alternatively, the enzyme can be a recombinant enzyme comprising the catalytic domain of RNase H or a functionally active fragment or analog thereof. The effect of cleavage preferably results in degradation of the DNA-RNA hybrid sequence.

[0045] The method may include a step of testing cleavage efficiency and / or specificity. This step allows the user to determine whether the RNA targeted by multiple DNA oligonucleotides and subsequently targeted by an enzyme that cleaves DNA-RNA hybrids remains intact after performing the steps of the method. Intact RNA indicates that one or both of the steps of forming a DNA-RNA hybrid and cleaving with an enzyme that cleaves DNA-RNA hybrids lack specificity and / or efficiency. Intact RNA further indicates that no protein was released from the RNA molecule of interest, in accordance with the objectives of the present invention. Testing cleavage efficiency and / or specificity may involve comprehensive techniques such as whole-transcriptome RNA-seq. Testing cleavage efficiency / specificity may involve targeted techniques such as quantitative PCR.

[0046] The method may further comprise isolating the RNA binding protein from the sample, the method comprising the steps of: (f) subjecting the sample to a further step of phase separation to produce a separate phase containing free RNA-binding protein but at least substantially free of free RNA; and (g) Isolating the released RNA-binding protein from the separated phases of part (f).

[0047] Liquid phase separation methods are generally known in the art. After the protein is released from the target RNA molecule, a further step of phase separation is performed on the sample, resulting in a separate phase containing free RNA-binding proteins but at least substantially free of free RNA. The separated phase is preferably at least substantially free of free DNA. The separated phase is preferably at least substantially free of RNA-protein adducts. Such RNA-protein adducts may remain present in the sample in step (f), for example, because in some embodiments, not all of the RNA-protein adducts are in contact with multiple DNA oligonucleotides, and therefore, in some embodiments, not all of the RNA-binding proteins are released from the RNA to which they are crosslinked when the sample is contacted with an enzyme that cleaves DNA-RNA hybrids.

[0048] The further step of phase separation preferably includes the use of one or more of acidic phenol, guanidine isothiocyanate, and chloroform. Even more preferably, the step of phase separation preferably includes the use of all of acidic phenol, guanidine isothiocyanate, and chloroform. The use of one or more of these agents includes contacting the sample with the one or more agents, thereby inducing phase separation. General liquid phase separations involving the use of the one or more agents are known in the art, and are particularly known to include one or more steps of mixing and centrifugation, as described in the Examples herein.

[0049] The further step of phase separation preferably results in the production of an organic phase, an interphase, and an inorganic phase, the separated phase containing the free RNA-binding protein but at least substantially free of free RNA corresponding to the organic phase.

[0050] The method includes isolating the liberated RNA-binding protein from the separated phases in a further step of phase separation. Such isolation can be achieved by conventional methods of liquid extraction from a liquid sample.

[0051] Methods for identifying RNA-binding proteins The present disclosure also provides a method for identifying an RNA binding protein, the method comprising the steps of a method for isolating an RNA binding protein from a sample according to the present invention, and further comprising subjecting the released RNA binding protein to a technique for protein identification.

[0052] Techniques for protein identification can be targeted techniques, optionally involving the use of antibodies to target the protein of interest. Targeted techniques can be antibody-based pull-down or immunoblot assays.

[0053] The technique for protein identification may be a comprehensive technique, and optionally the technique may comprise any suitable form of quantitative mass spectrometry, preferably the quantitative mass spectrometry comprises a shotgun proteomics method such as label-free quantitative mass spectrometry.

[0054] Methods for identifying binding sites of RNA-binding proteins The present disclosure also provides a method for identifying binding sites of an RNA-binding protein in an RNA transcript sequence, the method comprising the steps of the method for isolating an RNA-binding protein from a sample according to the present invention together with the steps of the method for identifying an RNA-binding protein according to the present invention, and further comprising mapping a plurality of DNA oligonucleotides, each comprising a nucleotide sequence complementary to a target RNA molecule of interest, to the RNA transcriptome, and inferring binding sites of the released RNA-binding protein in the RNA transcript sequence based on the DNA oligonucleotides mapped to the RNA transcriptome.

[0055] The multiple DNA oligonucleotides can be mapped to the RNA transcriptome using any suitable technique or software for sequence mapping.

[0056] The transcriptome can be that of any organism for which some or all of the RNA transcriptome is known.

[0057] While specific embodiments, specific configurations, as well as materials and / or molecules, have been discussed herein for methods according to the present disclosure, it should be understood that various changes or modifications in form and detail may be made without departing from the scope and spirit of the invention. The foregoing embodiments and the following examples are provided by way of illustration only and should not be considered as limiting the present application, which is limited only by the claims. [Example]

[0058] A novel method for the purification and identification of proteins associated with any selected endogenous RNA of interest from cells and tissues is disclosed. This method, termed Targeted RNase H-mediated Extraction of X-linked proteins (TREX), utilizes organic phase separation to isolate protein-RNA adducts from cells or tissues after RNA-protein crosslinking, followed by specific degradation of a given RNA target region of interest using RNase H, resulting in the specific release of proteins covalently associated with that RNA region. The released proteins can then be identified by quantitative mass spectrometry or other protein identification techniques (Figure 1).

[0059] Thanks to the highly specific nature of RNase H activity, TREX achieves unprecedented specificity, ensuring that only the specific target RNA of interest undergoes complete degradation and release of its associated proteins. Furthermore, as an enzymatic reaction, RNase H-based release of associated proteins is far more efficient than alternative methods using RNA pull-down. Finally, because RNase H degradation can be specifically targeted to specific regions of a given RNA, TREX enables region-specific mapping of interacting proteins from endogenous samples, a capability not possible with other existing pull-down-based assays. As proof-of-concept, we have validated TREX with different known RNA targets, demonstrating that TREX can robustly and accurately characterize interacting proteins of RNA targets. Crucially, this is achieved from significantly smaller amounts of starting material compared to previously published pull-down-based methods, meaning that TREX can be easily scaled and used in diverse experimental settings.

[0060] There have been several iterations of TREX. In some iterations, the approach uses whole-cell or whole-tissue lysis to capture all proteins bound to the RNA of interest. In other iterations, subcellular fractionation can be applied first to focus on complexes in a given subcellular compartment. Different crosslinking methods can also be used to capture either direct RNA-protein interactions (using UV-C crosslinking) or indirect interactions (using formaldehyde crosslinking). Finally, in some iterations, RNase H-based degradation can be performed on the full length of the target RNA, while in other iterations, only specific regions of the target RNA can be targeted, allowing for region-specific mapping of interactions. Here, we describe a detailed step-by-step protocol for one iteration of TREX and use it to generate the example data shown for 18S RNA and U1 RNA.

[0061] Materials - Biological Materials The biological material for our TREX testing studies was human HCT116 cells, but any other cell line can be used as well.

[0062] It is important to check cell lines regularly to ensure they are authentic and free of mycoplasma infection. Handle cell lines according to the supplier's instructions. Work in a biosafety hood, use sterile equipment, and wear gloves to minimize the risk of contamination.

[0063] Materials and Reagents RNaseZAP™ Surface Decontamination (Sigma, R2020) Nuclease-free water (Thermo, AM9932) Cytoplasmic lysis buffer (1% Triton X-100, 10 mM Tris-HCl pH 7.5, 150 mM NaCl, cOmplete™ protease inhibitor cocktail (Sigma), SUPERase·In™ 1 / 500 v / v (Invitrogen, AM2694)) (use ribonuclease-free reagents and nuclease-free water to prepare this buffer) Caution! Triton is toxic and an irritant. Triton is harmful to the environment. Handle solutions containing Triton with care and dispose of waste according to your facility's regulations. Caution! cOmplete™ protease inhibitor cocktail is an irritant. Handle solutions containing the protease inhibitor cocktail with care and dispose of waste according to facility regulations. TRIzol™ LS Reagent (Thermo, 10296028) TRIzol™ Reagent (Fisher Scientific UK Ltd, 12034977) CAUTION! TRIzol™ Reagent is toxic if inhaled and a potential carcinogen. Always handle chloroform-containing solutions with care in a chemical fume hood and wearing personal protective equipment, and dispose of waste according to facility regulations. Chloroform CAUTION! Chloroform is volatile and toxic. Chloroform is an irritant. Handle solutions containing chloroform with care and dispose of chloroform waste according to facility regulations. TE (Tris-Cl pH 7, 10 mM, EDTA 1 mM) (Use ribonuclease-free reagents and nuclease-free water to prepare this buffer) TE + SDS 0.1% (Tris-Cl pH 7, 10 mM, EDTA 1 mM, SDS 0.1%) (Use ribonuclease-free reagents and nuclease-free water to prepare these buffers) TE + SDS 0.5% (Tris-Cl pH 7, 10 mM, EDTA 1 mM, SDS 0.5%) (Use ribonuclease-free reagents and nuclease-free water to prepare these buffers) 1x Hybridization Buffer: (50mM NaCl, 1mM EDTA, 100mM TrisHCl pH 7.0) (Use ribonuclease-free reagents and nuclease-free water to prepare this buffer) Thermostable ribonuclease H (NEB, M0523S) Heat-stable RNase H buffer (NEB, M0523S) Isopropanol Proteinase K (Thermo, AM2546) Proteinase K buffer (0.1 M NaCl, 10 mM Tris HCl pH 8.0, 1 mM EDTA, 0.5% SDS) (use ribonuclease-free reagents and nuclease-free water to prepare this buffer) acetone ethanol Dithiothreitol (DTT) ≥ 99.4%, proteomics grade (VWR, M109) Iodoacetamide (IAA) ≥98%, proteomics grade (Sigma, I1149) Urea≧99.5% (Sigma, U1250) ABC buffer (50 mM ammonium bicarbonate, pH 7.8) Trypsin, LC-MS / MS grade (Sigma, 650279) Stop buffer (4% acetonitrile, 1% TFA) 5M NaCl (Invitrogen REF AM9760G) Invitrogen GlycoBlue coprecipitant 15mg / mL (Fisher Scientific UK Ltd, 10301575) 1M Tris HCl pH7.0 (Invitrogen, AM9850G) 0.5M EDTA pH8.0 (Invitrogen, AM9260G) 1M MgCl2 (Invitrogen, AM9530G) 10% SDS (Invitrogen, 15553-035) TURBO DNase 10x Buffer (Fisher Scientific, 10646175) TURBO deoxyribonuclease (Fisher Scientific, 10646175) RNasin® Plus Ribonuclease Inhibitor (Promega, N2611)

[0064] Materials-Equipment UV crosslinker (Hoefer, UVC500) Cell scraper (Corning, catalog number 353087) 150mm TC-treated culture dish Ribonuclease-free microcentrifuge tubes, 2 ml (Thermo, AM12425) Ribonuclease-free microcentrifuge tubes, 1.5 ml 0.2mL 8-strip PCR tubes (Thermo, AM12425) Food for Trizol Veriti™ Thermocycler (Thermo, 4375786) Thermomixer Comfort (Eppendorf, No. 5355) Barrier Pipette Tips Tabletop centrifuge (Eppendord, 5415R) C18 Stage Chip mass spectrometer Vortexer (Scientific Industries, SI-0236) Qubit 2.0 Fluorometer (Invitrogen, Q32866) QuantStudio 7 Flex Real-Time PCR (Thermo, 4485701) RTqPCR plates

[0065] Materials-Software Q7 RTqPCR Software Maxquant Perseus

[0066] Protocol Steps: 1) In vivo crosslinking - 2 minutes HCT116 cells are seeded in 150 mm TC-treated culture dishes and allowed to attach for 16 hours in a 37° C. incubator with 5% CO 2 before crosslinking. 1. Aspirate the medium and add 15 ml cold PBS. 2. Aspirate the PBS and add another 15 ml of cold PBS. 3. Aspirate the PBS. 4. Transfer the dish onto ice or onto the ice container that is installed inside the crosslinker device (the crosslinker device sensor must not be covered as it continuously measures UV energy). 5. On ice, 200-400mJ / cm 2Crosslink the cells with 0.5% COOH (the dish lid should be removed from the plate to ensure crosslinking efficiency). NOTE: To perform TREX on cellular total RNA, proceed with 2A. If prior subcellular fractionation is required, proceed with 2B.

[0067] Protocol Step: 2A) Whole Cell Lysis - 10 minutes 6. Add 1 ml of TRIzol™ per 20 million cells to the plate. 7. Scrape the cells thoroughly. While scraping, make sure that the TRIzol™ covers the entire plate. 8. Homogenize the lysate through repeated pipetting. If necessary, store the lysate at -80°C for up to 3 months.

[0068] Protocol Step: 2B) Subcellular Fractionation - Time Required: 20 minutes 9. Scrape the cells into 1 mL of RNase-free ice-cold PBS per 20 million cells. Transfer each 1 mL of lysate to a 2 mL RNase-free Eppendorf tube. 10. Pellet the cells by centrifugation at 500g for 1 minute at 4°C. Discard the PBS supernatant. 11. Dissolve the cell pellet in 150 μl cytoplasmic lysis buffer. 12. Spin at 10,000 g for 15 minutes at 4° C. and collect the supernatant, which contains the cytoplasmic fraction. Note: Alternatively, if a nuclear fraction is required, retain the remaining pellet. If necessary, store fractions at -80°C for up to 3 months.

[0069] Protocol Step: 3A) TRIzol™ Phase Separation for Whole Cell Lysates (continued from 2A) Duration: 1 hour 13. Incubate the homogenized lysate at room temperature (20-25°C) for 5 minutes to dissociate any uncrosslinked RNA-protein interactions. 14. Add 200 μl chloroform / 1 ml TRIzol™ and vortex the sample until the contents are homogenous. 15. Centrifuge at 12,000g for 15 minutes at 4°C. At this point, three phases should be visible. 16. Remove the upper aqueous phase until the interface collapses. Remove as much of the lower organic phase as possible without disturbing the interface. 17. Resolubilize the interface in 1 ml of TRIzol™ per 20 million cells. 18. Add 200 μl chloroform / 1 ml TRIzol™ and vortex the sample until the contents are homogenous. 19. Centrifuge at 12,000g for 15 minutes at 4°C. 20. Remove the upper aqueous phase until the interface collapses. Remove as much of the lower organic phase as possible without touching the interface. 21. Resolubilize the interface in 1 ml of TRIzol™ per 20 million cells. Pause If necessary, store the lysate at -80°C. 22. Add 200 μl chloroform / 1 ml TRIzol™ and vortex the sample until the contents are homogenous. 23. Centrifuge at 12,000g for 15 minutes at 4°C. 24. Remove the upper aqueous phase until the interface collapses. Remove as much of the lower organic phase as possible without touching the interface. NOTE: If it is necessary to work with more than 20 million cells per replicate, the following steps are required. 25. Combine the interface from up to 100 million cells in 1 ml of TRIzol™ and transfer to a 2 ml tube. 26. Add 200 μl of chloroform and vortex the sample until the contents are homogenous. 27. Centrifuge at 12,000g for 15 minutes at 4°C. 28. Remove the upper aqueous phase until the interface collapses. Remove as much of the lower organic phase as possible without touching the interface.

[0070] Protocol Step: 3B) TRIzol™ LS Phase Separation for Intracellular TREX (continued from 2B) Duration: 1 hour 29. Add 1 ml of TRIzol™ LS to the sample and homogenize through repeated pipetting. Note: If nuclear fractionation is to follow, standard TRIzol™ is used. 30. Incubate the homogenized lysate at room temperature (20-25°C) for 5 minutes to dissociate any uncrosslinked RNA-protein interactions. 31. Add 200 μl of chloroform and vortex the sample until the contents are homogenous. 32. Centrifuge at 12,000g for 15 minutes at 4°C. At this point, three phases should be visible. 33. Remove the upper aqueous phase until the interface collapses. Remove as much of the lower organic phase as possible without touching the interface. 34. Resolubilize the interface in 1 ml of TRIzol™. 35. Add 200 μl of chloroform and vortex. 36. Centrifuge at 12,000 g for 15 minutes at 4° C. Remove the upper and lower phases as before. 37. Resolubilize the interface in 1 ml of TRIzol™. Pause If necessary, store the lysate at -80°C. 38. Add 200 μl of chloroform and vortex. 39. Centrifuge at 12,000 g for 15 minutes at 4° C. Remove the upper and lower phases as before, leaving a maximum of 50 μl. NOTE: When working with more than 20 million cells per replicate, the following steps are required. 40. Combine the interface from up to 100 million cells in 1 ml of TRIzol™ and transfer to a 2 ml tube. 41. Add 200 μl of chloroform and vortex the sample until the contents are homogenous. 42. Centrifuge at 12,000g for 15 minutes at 4°C. 43. Remove the upper aqueous phase until the interface collapses. Remove as much of the lower organic phase as possible without touching the interface.

[0071] Protocol step: 4) Interface solubilization and DNA removal. Time required: 4 hours. Interface cleaning: 44. Without disturbing the interface, add 1 ml of TE buffer to the tube by slowly pipetting it along the wall. Do not resuspend or disturb the interface. 45. Gently invert the tube three times, taking special care not to disturb the interface too much. 46. ​​Centrifuge at 5000g for 1 minute at 4°C. 47. Remove the supernatant. 48. If necessary, repeat steps 44-47 one or two more times until the red dye from TRIzol™ is no longer visible. NOTE: After centrifugation, a small amount of TRIzol™ may appear trapped below the interfacial pellet. Repeated washing (steps 44-47) will remove this.

[0072] Interface failure: 49. Resuspend the interface in 1 ml of TE + SDS 0.1% by pipetting up to 30-60 times. 50. When the interface is no longer disrupted, centrifuge at 5000 g for 2 minutes at room temperature. 51. Transfer the supernatant to a new 2 ml tube (Collection Tube 1). Note: From this point onwards the interface should become increasingly translucent. Do not transfer any of the interface to the collection tube. 52. Resuspend the remaining interface in 1 ml of TE + SDS 0.1% by repetitive pipetting up to 30-60 times. 53. Centrifuge at 5000g for 2 minutes at room temperature. 54. Transfer the supernatant to a new 2 ml tube (Collection Tube 2). 55. Resuspend the remaining interface in 1 ml of TE + SDS 0.5% by repetitive pipetting up to 30-60 times. 56. Centrifuge at 5000g for 2 minutes at room temperature. 57. Transfer the supernatant to a new 2 ml tube (Collection Tube 3). 58. Resuspend the remaining interface in 1 ml of TE + SDS 0.5% by pipetting up to 30-60 times. 59. Centrifuge at 5000g for 2 minutes at room temperature. 60. Transfer the supernatant to a new 2 ml tube (Collection Tube 4).

[0073] Precipitation 1: 61. To each collection tube, add 60 μl NaCl 5M, 1 μl Glycoblue, and 1 ml isopropanol. 62. Invert the tube multiple times. If necessary, samples can be stored at -20°C overnight. 63. Centrifuge at 18000g (or faster) and 4°C (or colder) for 15 minutes.

[0074] Ethanol wash: 64. Carefully remove the supernatant from all collection tubes. 65. Add 1 ml of 70% ethanol to the first collection tube and disrupt the pellet through repeated pipetting. Then transfer the contents to the next collection tube and repeat the process for all four tubes until all pellets are combined in the same ethanol solution. 66. Use another 1 ml of 70% ethanol to rinse any remaining pellet from the now empty collection tube. Combine this solution with the collection from step 35, leaving 2 ml of ethanol suspension containing all the pellet. 67. Centrifuge at 18,000 g for 1 minute at room temperature. 68. Discard the supernatant and use a brief spin to remove any remaining ethanol.

[0075] Hydration: NOTE: Skip this step if TREX is to be performed on the cytoplasmic fraction. Note: The volumes below are for up to 100 million cells. 69. Add 1.8 ml of nuclease-free water to the pellet. 70. Invert the tube multiple times or vortex briefly to loosen the pellet from the bottom of the tube. 71. Hydrate the pellet on ice for 1 hour, inverting occasionally. NOTE: The pellet will become translucent. 72. Using a 1 ml pipette, resuspend the pellet repeatedly until the pellet is completely dissolved.

[0076] Deoxyribonuclease treatment: NOTE: Skip this step if TREX is to be performed on the cytoplasmic fraction. 73. Add 200 μl of 1×TURBO DNase buffer and mix the solution by pipetting. 74. Add 2 μl of ribonuclease inhibitor and 100 μl of TURBO deoxyribonuclease. 75. Invert the tube 4-5 times and incubate at 37°C with 700 rpm shaking for 50 minutes.

[0077] Precipitation 2: NOTE: Skip this step if TREX is to be performed on the cytoplasmic fraction. 76. Split the sample into two 2 ml tubes (1 ml each). 77. To each collection tube, add 60 μl of NaCl 5M and 1 ml of isopropanol. 78. Invert the tube multiple times. If necessary, samples can be stored at -20°C overnight. 79. Centrifuge at 18000g (or faster) at 4°C (or colder) for 15 minutes.

[0078] Ethanol Wash 2: NOTE: Skip this step if TREX is to be performed on the cytoplasmic fraction. 80. Carefully remove the supernatant from each tube. 81. Add 1 ml of 70% ethanol to the first tube and disrupt the pellet by pipetting. Transfer the contents to the second tube and repeat the process until the pellet is combined in the same ethanol solution. 82. Use another 1 ml of 70% ethanol to rinse any remaining pellet from the now empty tube. Combine this solution with the collection from step 81, leaving 2 ml of ethanol containing the combined pellet. 83. Centrifuge at 18,000 g for 1 minute at room temperature. 84. Discard the supernatant and use a brief spin to remove any remaining ethanol.

[0079] Protocol step: 5) Hybridization (annealing) of probe to RNA. Time required: 1 hour. 85. Dissolve the pellet in 110 μl of hybridization buffer per 20 million cells by repeated pipetting (e.g., if starting with 100 million cells per replicate, dissolve in 550 μl of hybridization buffer). Pipette up and down until the pellet no longer dissolves. 86. Add 6 μl of 1-200 μM tiling oligo* per 20 million cells** (For example: for U1 RNA, we used 6 μl of 10 μM oligo per 20 million cells. Note that oligo is added to both control and experimental samples). 87. Invert the sample 4-5 times and give the sample a brief rotation.

[0080] Sample summary table:

[0081] [Table 4]

[0082] *We use antisense tiling DNA oligos designed against a target RNA of interest. We use our design software to generate a list of approximately 60 nt long tiling oligos for any input sequence.

[0083] TREX antisense DNA oligonucleotides can be designed by converting a target RNA sequence into its reverse complementary DNA counterpart and dividing the resulting sequence into non-overlapping tiling sections of any selected length. In our example, the oligo lengths were designed to be approximately 60 nucleotides, with the last oligo at the 3' end being flexible, between 30 and 90, allowing complete coverage of any target RNA length.

[0084] **The amount of cells and oligos used depends on each target RNA and its copy number and must be determined empirically through qPCR for each set of oligos. 88. Annealing is performed in a Thermomixer. It is important to perform the following steps correctly to ensure efficient annealing. a) Heat to 95°C for 2 minutes while shaking at 1100 RPM. b) While shaking, the temperature is manually reduced by 2°C per minute until 50°C is reached. c) Stop shaking, but do not remove the sample or turn off the heat. Go to step 89. The sample must not be covered during annealing, as this would prevent a uniform drop in temperature.

[0085] Protocol step: 6) Digestion with thermostable ribonuclease H. Time required: 1 hour. 89. Without removing the samples from the thermomixer, add MgCl2 to the samples to a final concentration of 0.55 mM. This is to neutralize the EDTA in the hybridization buffer. 90. Without removing the sample from the thermomixer, add the following components to the experimental or control sample:

[0086] [Table 5]

[0087] [Table 6]

[0088] 91. Close the cap and turn on the shaker (1100 RPM). 92. Incubate the tube in a thermomixer at 50°C for 1 hour. 93. Centrifuge the sample at 12000g for 3 minutes. 94. Transfer the supernatant to a new tube. Be sure not to transfer any insoluble particles that may have collected at the bottom. 95. Take 1 / 10 of the sample for analysis of RNA degradation by RTqPCR or RNA sequencing (see Appendix). If necessary, this sample can be flash frozen and stored at -80°C for up to 3 months. 96. Proceed to step 97 with the remaining 9 / 10 of the sample. Do not stop at this step in the protein preparation.

[0089] Protocol step: 7) RBP extraction, required time: 20 hours 97. Add 900 μl of TRIzol™ LS to each sample and vortex. 98. Add 200 μl of chloroform and vortex the sample until the contents are homogenous. 99. Centrifuge at 12,000 g for 15 minutes at 4° C. The sample should separate into three phases: aqueous, interface, and organic. After phase separation, it is important to check that an interface is still forming: the lack of a visible interface indicates nonspecific degradation of the RNA. 100. Recover the released RBP by pipetting the organic phase into a new tube. Recover only the organic phase without any interface. If it is difficult to pipette all of the organic phase while eliminating the interface, some of the organic phase may remain to ensure no interface carryover. 101. To precipitate the released RBP, add 100% cold acetone to a final concentration of 80% acetone. For example, if 300 μl of organic phase is recovered, add 1200 μl of 100% acetone. Vortex the mixture thoroughly. Suspend the released RBP in acetone by incubating at -20°C overnight. 102. Centrifuge at 16,000g for 20 minutes at 4°C. 103. Discard the supernatant. The RBPs should have precipitated and therefore be present as a pellet after centrifugation (as shown in the image). 104. Add 1 ml of 80% acetone to wash the pellet. 105. Vortex and centrifuge at 16,000g for 15 minutes at 4°C. 106. Repeat steps 35-37. 107. Discard the supernatant and air dry the pellet for 5 minutes.

[0090] Protocol Step: 8) Sample preparation for shotgun proteomics (Time required: 2 hours) 108. Recover the RBP pellet in 100 μl 8M urea in ABC buffer. Prepare a fresh urea solution on the day. Pipette the sample up and down thoroughly to resuspend the pellet as thoroughly as possible. Any remaining insoluble material should dissolve during trypsin digestion (step 112). 109. Add DTT from a 1 M stock to reach a final concentration of 10 mM DTT. Incubate the sample at room temperature (20-25°C) for 30 minutes to reduce the protein. 110. Add IAA from a 0.55M stock to reach a final concentration of 55 mM. Incubate the sample in the dark at room temperature (20-25°C) for 30 minutes to alkylate the reduced residues. IAA is unstable and light sensitive. Prepare the solution immediately before use and ensure that the alkylation is carried out in the dark. 111. Dilute the urea concentration to 2M by adding ABC buffer to the sample. 112. Add 1 μg trypsin per sample and incubate the samples at room temperature (20-25°C) for 16 hours. 113. Add 1 volume of Stop Buffer to the digested sample. 114. Proceed with a C18-based cleanup (e.g., stage tip) of the sample to remove salts. 115. Inject the purified peptide into the LC-MS / MS (procedures vary depending on the MS instrument used). 116. Identified RBPs are searched and quantified using label-free quantification (LFQ). Multiple search and quantification computational pipelines can be used. We routinely use Maxquant and Perseus software packages for the data analysis step.

[0091] Protocol Steps: Appendix) RTqPCR and RNA-seq analysis of efficiency and specificity of RNase H digestion To retrospectively analyze the efficiency and specificity of RNase H-mediated target degradation, a 1 / 10 aliquot taken after RNase H treatment can be used to purify RNA and analyze it in a targeted (RT-qPCR) or non-targeted (RNA-seq) manner. To this end, the following steps are followed: 110 μl of proteinase K buffer is added to each collected aliquot. Proteins are digested by adding 10 μl proteinase K and incubating at 25° C. for 1 hour. An additional phase separation is performed by adding TRIzol™ LS to the Proteinase K reaction. The free RNA (top phase) is collected. The RNA is purified according to the TRIzol™ LS manual. RNA quantification is performed using Qubit. Pause: If necessary, store the extracted RNA at -80°C for up to 3 months. To assess degradation efficiency and specificity, RT-qPCR quantification of the target transcripts, as well as independent control transcripts, is performed. Alternatively, purified RNA can be subjected to whole-transcriptome RNA-seq to comprehensively assess the efficiency and specificity of RNase H-mediated degradation.

[0092] Results: Analysis of 18s rRNA protein interactors by TREX To evaluate the performance of TREX, we first applied TREX to analyze the direct interactome of 18S ribosomal RNA (rRNA), the RNA component of the small (40S) ribosomal subunit. Because the 18S interactome is well characterized, e.g., by other transcript target capture methods, this allows us to evaluate the performance of TREX against existing technologies.

[0093] A total of four experimental samples (+ RNase H) and four control samples (- RNase H) (each derived from 10 million HCT116 human colorectal cancer cells) were processed and subjected to TREX as described in the detailed stepwise protocol above. Using RT-qPCR, we first confirmed efficient degradation of target transcripts in the RNase H-treated samples (Figure 5). Indeed, we found that 18S was efficiently degraded in all + RNase H-treated samples but not in the - RNase H-treated samples. Next, we analyzed the digested protein extracts by LC-MS / MS using a Thermo Orbitrap Q Exactive Plus MS instrument. We identified a total of 480 proteins that were significantly more abundant in the + RNase H samples (Figure 6). Crucially, most of the core 40S small ribosomal proteins, plus several other known 40S-associated proteins, were among the significant interactors identified (Fig. 6).

[0094] Next, we assessed the types of proteins overrepresented in our 18S interactome by performing a category enrichment analysis (Tyanova et al., 2016). As expected, the top enriched categories belonged to small ribosomal protein annotations, but other protein categories involved in 18S biogenesis and maturation, as well as categories of proteins involved in protein targeting to the ER (known to be associated with the 40S ribosome), were also among the top enriched categories (Figure 7). Analysis of existing protein-protein interactions using the STRING protein-protein interaction database (https: / / string-db.org / ) revealed that the vast majority of the identified 18S interactors were known to be related to each other, further confirming the specificity of the TREX results (Figure 8).

[0095] Finally, we compared our TREX results with a previous 18S RAP-MS analysis (McHugh et al., 2015). More than 50% of the hits from the 18S RAP-MS analysis were also identified by TREX (Figure 9). These included the majority of 40S ribosomal proteins. However, TREX identified significantly more proteins, including several 40S ribosomal proteins missed by RAP-MS, translation initiation factors, small ribosomal subunit biogenesis factors, and many other known 40S-associated proteins, such as the ER targeting machinery (known to be associated with the 40S) (Figure 9). In contrast, 52 proteins not found in TREX were identified by RAP-MS. These were primarily factors involved in large ribosomal subunit biogenesis, as well as many other nucleolar proteins (Figure 9). Such differences may be due to the fact that TREX is region-specific, with probes acting only on 18S-specific interactors, whereas RAP-MS potentially captures preprocessed rRNA containing both large and small ribosomal subunit RNA.

[0096] In conclusion, this example demonstrates that TREX can accurately and efficiently reveal the 18S rRNA interactome from at least 10 million cells per sample. This means that the total number of cells used in TREX (both -RNase H and +RNase H) was only 80 million. For comparison, a total of 200 million or 800 million cells were used in RAP-MS analysis (McHugh et al., 2015).

[0097] Results: Analysis of U1 RNA-protein interactors by TREX Next, we applied TREX to analyze the direct interactome of U1 RNA, a well-characterized RNA component of the U1 small nuclear ribonucleoprotein (U1 snRNP) complex, a key core spliceosome complex that recognizes 5' exon-intron junction sites (Kondo et al., 2015). Similar to 18S, the interactome of U1 is well characterized, e.g., by other target capture methods, allowing for the evaluation of TREX against existing technologies. Furthermore, U1 is much smaller than 18S (<200 nt vs. >3000 nt), thus providing a comparative test of whether TREX can also be successfully applied to small RNAs.

[0098] A total of five experimental samples (+RNase H) and five control samples (-RNase H) (each derived from 20 million HCT116 human colorectal cancer cells) were processed and subjected to TREX as described in the detailed stepwise protocol above. Using RT-qPCR, we first confirmed the efficient degradation of U1 in the RNase H-treated samples (Figure 10). Indeed, we found that U1 was efficiently degraded in all +RNase H-treated samples but not in the -RNase H-treated samples. Next, the digested protein extracts were analyzed by LC-MS / MS using a Thermo Orbitrap Q Exactive Plus MS instrument. A total of 30 proteins were found to be significantly enriched in the +RNase H samples (Figure 11). While this number is much lower than the number of interacting proteins identified for 18S (Figure 6), this is expected because U1 is much smaller and RNA is less abundant. Crucially, 8 of the 10 known U1 snRNP proteins, plus several other known splicing factors known to associate with the U1 snRNP complex, were among the significant interactors identified (Fig. 11 ).

[0099] Next, we analyzed existing protein-protein interactions among the identified U1-binding proteins using the STRING protein-protein interaction database (https: / / string-db.org / ). This analysis revealed one major network of predicted U1 snRNP and splicing-related proteins, as well as a complex of thioredoxin (TXN) and peroxiredoxin-5 (PRDX5). These two proteins are involved in redox regulation, suggesting a potential novel interrelationship between U1 and redox regulatory components in these cells (Figure 12).

[0100] Finally, we compared our TREX results with a previous U1 RAP-MS analysis (McHugh et al., 2015). Over 50% (8 of 14) of the hits in the U1 RAP-MS analysis were also identified by our TREX (Figure 13). Crucially, these were all known U1 snRNP protein components and splicing factors. However, TREX also identified additional splicing factors, as well as additional U1 snRNP proteins (SNRPEs), that were missed by RAP-MS. In addition, a small number of metabolic enzymes known to function as unconventional RBPs (Castello et al., 2012) were also identified as U1 interactors by TREX.

[0101] Results: Analysis of region-specific protein interactions in pre-rRNA by TREX Next, we applied TREX to region-specific mapping of RNA-protein interactions. For this purpose, we selected the 45S pre-rRNA, a precursor RNA synthesized in the nucleolus as a >13 kb-long transcript. This pre-rRNA then undergoes extensive processing, including cleavage and removal of several spacer regions, ultimately resulting in three fully processed rRNA transcripts (18S, 5.8S, and 28S) that form the major RNA components of eukaryotic ribosomes (Moraleva et al., 2022). The regions of the 45S that undergo cleavage and removal are composed of the 5' external transcribed region (5'ETS), internal transcribed spacers 1 and 2 (ITS1 and ITS2), and the 3' external transcribed region (3'ETS) (Figure 14). Until now, little is known about the intrinsic protein interactions of these specific spacer regions.

[0102] As a proof-of-concept for the region-specific mapping capabilities of TREX, we applied it to specifically map the protein interactome of the 5'ETS region, the first spacer region processed during ribosome biogenesis. A total of five experimental samples (+ RNase H) and five control samples (- RNase H), each derived from 20 million HCT116 human colorectal cancer cells, were processed and subjected to TREX as described in the stepwise protocol detailed above. As previously described, we first confirmed efficient degradation of the 5'ETS in the RNase H-treated samples by RT-qPCR (data not shown). Next, trypsin-digested protein extracts were analyzed by LC-MS / MS using a Thermo Orbitrap Q Exactive Plus MS instrument. A total of 16 proteins were found to be significantly enriched in the + RNase H samples. Crucially, half of these proteins are known to be involved in ribosome biogenesis, and several other unknown factors have also been shown to localize to the nucleolus, the exclusive site for the 5′ ETS (Figure 15).

[0103] We also analyzed existing protein-protein interactions among the identified 5' ETS-binding proteins using the STRING protein-protein interaction database (https: / / string-db.org / ). This analysis revealed one major network of ribosome biogenesis factors (Figure 16). Crucially, while most of these factors are known to be important for ribosome biogenesis, particularly the early steps involving 5' ETS processing, their direct interaction with RNA within the 5' ETS region is a novel and potentially important finding.

[0104] In conclusion, this example demonstrates that TREX can accurately and efficiently reveal the interactome of a small RNA (U1 RNA) from at least 20 million cells per sample, revealing not only the majority of its known interactors but also several novel interactors.

[0105] References Castello, A., Fischer, B., Eichelbaum, K., Horos, R., Beckmann, BM, Strein, C., Davey, NE, Humphreys, DT, Preiss, T., Steinmetz, LM, et al. (2012). Insights into RNA biology from anatlas of mammalian mRNA-binding proteins. Cell 149, 1393-1406.10.1016 / j.cell.2012.04.031. Chu, C., Zhang, QC, da Rocha, ST,Flynn, RA, Bharadwaj, M., Calabrese, JM, Magnuson, T., Heard, E., andChang, HY (2015). Systematic discovery of Xist RNA binding proteins. Cell161, 404–416. 10.1016 / j.cell.2015.03.025. Desideri , F. , Cipriano , A. ,Petrezselyova , S. , Buonaiuto , G. , Santini , T. , Kasparek , P. , Prochazka , J. ,Janson , G. , Paiardini , A. , Calicchio , A. , et al. (2020). Intronic DeterminantsCoordinate Charm lncRNA Nuclear Activity through the Interaction with MATR3and PTBP1. Cell reports 33 , 108548 . Hafner , M. , Katsantoni , M. , Koester , T. ,Marks , J. , Mukherjee , J. , Staiger , D. , Ule , J. , and Zavolan , M. (2021). CLIPand complementary methods. Nature Reviews Methods Primers 1 , 20.10.1038 / s43586-021-00018-1 . Kondo, Y., Oubridge, C., van Roon,A.-M.M., and Nagai, K. (2015). Crystal structure of human U1 snRNP, a smallnuclear ribonucleoprotein particle, reveals the mechanism of 5' splice siterecognition. eLife 4, e04986. 10.7554 / eLife.04986. McHugh, C.A., Chen, C.K., Chow, A.,Surka, C.F., Tran, C., McDonel, P., Pandya-Jones, A., Blanco, M., Burghard, C.,Moradian, A., et al. (2015). The Xist lncRNA interacts directly with SHARP tosilence transcription through HDAC3. Nature 521, 232-236. 10.1038 / nature14443. Moraleva, A.A., Deryabin, A.S., Rubtsov,Y.P., Rubtsova, M.P., and Dontsova, O.A. (2022). Eukaryotic RibosomeBiogenesis: The 40S Subunit. ActaNaturae 14, 14-30. 10.32607 / actanaturae.11540. Queiroz, R.M.L., Smith, T., Villanueva,E., Marti-Solano, M., Monti, M., Pizzinga, M., Mirea, D.M., Ramakrishna, M.,Harvey, R.F., Dezi, V., et al. (2019). Comprehensive identification ofRNA-protein interactions in any organism using orthogonal organic phaseseparation (OOPS). Nat Biotechnol 37, 169-178. 10.1038 / s41587-018-0001-2. Trendel, J., Schwarzl, T., Horos, R.,Prakash, A., Bateman, A., Hentze, M.W., and Krijgsveld, J. (2019). The HumanRNA-Binding Proteome and Its Dynamics during Translational Arrest. Cell 176, 391-403e319. 10.1016 / j.cell.2018.11.004. Tyanova, S., Temu, T., Sinitcyn, P.,Carlson, A., Hein, M.Y., Geiger, T., Mann, M., and Cox, J. (2016). The Perseuscomputational platform for comprehensive analysis of (prote)omics data. NatureMethods 13, 731-740. 10.1038 / nmeth.3901. Urdaneta, E.C., Vieira-Vieira, C.H.,Hick, T., Wessels, H.-H., Figini, D., Moschall, R., Medenbach, J., Ohler, U.,Granneman, S., Selbach, M., and Beckmann, B.M. (2019). Purification ofcross-linked RNA-protein complexes by phenol-toluol extraction. NatureCommunications 10, 990. 10.1038 / s41467-019-08942-3.

Claims

1. (a) contacting a sample with a cross-linking agent, thereby providing RNA-protein adducts; (b) performing phase separation on the sample containing the RNA-protein adducts to produce a separate phase comprising the RNA-protein adducts but at least substantially free of free protein and free RNA; (c) isolating the RNA-protein adducts from the separated phases of part (b); (d) contacting the isolated RNA-protein adduct with a plurality of DNA oligonucleotides comprising complementary nucleotide sequences to a target RNA molecule of interest, thereby forming one or more DNA-RNA hybrids; (e) contacting the sample from part (d) with an enzyme that cleaves DNA-RNA hybrids, thereby liberating proteins from said RNA molecules of interest.

2. further comprising isolating the RNA binding protein from the sample: (f) subjecting the sample to a further step of phase separation to produce a separate phase comprising free RNA-binding proteins but at least substantially free of free RNA; and 10. The method of claim 1, comprising: (g) isolating the released RNA binding protein from the separated phases of part (f).

3. 3. The method of claim 1 or 2, wherein the RNA binding protein binds to any one or more of mRNA, rRNA, 7 SL RNA, tRNA, snRNA, snoRNA, lncRNA, miRNA, and eRNA.

4. 4. The method of any one of claims 1 to 3, wherein the sample comprises material derived from cells cultured in vitro or from human or animal tissue, optionally wherein said material comprises a lysate of cells or human or animal tissue.

5. 5. The method of claim 4, wherein the material comprises a lysate of cells or human or animal tissue, and the lysate is further subjected to subcellular fractionation.

6. 6. The method of claim 5, wherein the lysate is prepared under conditions that are free of active ribonuclease enzymes.

7. 7. The method of any one of claims 1 to 6, wherein the phase separation of part (b) results in the production of an organic phase, an interphase, and an inorganic phase, the separated phase comprising RNA-protein adducts but at least substantially free of free protein and free RNA corresponding to the interphase.

8. The method of any one of claims 1 to 7, wherein the enzyme that cleaves DNA-RNA hybrids is RNase H or a functionally active fragment or analogue thereof.

9. 9. The method of any one of claims 2 to 8, wherein the further step of phase separation of part (f) results in the production of an organic phase, an interphase, and an inorganic phase, and wherein the separated phase comprising free RNA binding protein but at least substantially free of free RNA corresponds to the organic phase.

10. 10. The method of any one of claims 1 to 9, wherein the cross-linking agent induces covalent bonds between interacting protein and RNA molecules in the sample, and optionally the cross-linking agent is UV irradiation.

11. 11. The method of any one of claims 1 to 10, wherein the phase separation comprises the use of one or more of acidic phenol, guanidine isothiocyanate, and chloroform.

12. 12. A method for identifying an RNA binding protein, comprising the steps of the method for isolating an RNA binding protein from a sample according to any one of claims 2 to 11, and further comprising subjecting the released RNA binding protein to a technique for protein identification.

13. 13. The method of claim 12, wherein the technique for protein identification is a targeted technique, optionally comprising the use of antibodies to target proteins of interest.

14. 13. The method of claim 12, wherein the technique for protein identification is an exhaustive technique, optionally comprising quantitative mass spectrometry.

15. 15. A method for identifying a binding site of an RNA-binding protein in an RNA transcript sequence, the method comprising the steps of the method for isolating an RNA-binding protein from a sample according to any one of claims 2 to 11 and the steps of any one of claims 12 to 14, and further comprising mapping a plurality of DNA oligonucleotides, each of which comprises a nucleotide sequence complementary to a target RNA molecule of interest, to an RNA transcriptome, and inferring a binding site of the released RNA-binding protein in the RNA transcript sequence based on the DNA oligonucleotides mapped to the RNA transcriptome.