Method for extracting RNA binding proteins

EP4673454A1Pending Publication Date: 2026-01-07QUEEN MARY UNIV OF LONDON
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
EP2024708418
Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-02-28
Filing Date
2024-02-28
Publication Date
2026-01-07

AI Technical Summary

Technical Problem

Current RNA-centric methods for identifying RNA-binding proteins are inefficient, lack reproducibility, and suffer from off-target capture, making it challenging to achieve specific and scalable profiling of protein partners of specific RNAs.

Method used

A method involving cross-linking agents, phase separation, and the use of DNA oligonucleotides that hybridize to specific RNA regions, followed by enzyme-mediated cleavage of DNA-RNA hybrids to release and isolate RNA-binding proteins, allowing for region-specific mapping and identification.

Benefits of technology

This approach significantly enhances specificity and efficiency in identifying RNA-binding proteins and their binding sites, enabling robust and scalable profiling with unprecedented precision compared to existing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000007_0001
    Figure IMGF000007_0001
  • Figure IMGF000008_0001
    Figure IMGF000008_0001
  • Figure IMGF000009_0001
    Figure IMGF000009_0001
Patent Text Reader

Abstract

The present invention relates to novel methods for isolating RNA-binding proteins. The present invention also provides methods for identifying RNA-binding proteins and for identifying binding sites of RNA-binding proteins in an RNA transcript sequence.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] METHOD

[0002] Field of the Invention

[0003] The present invention relates to novel methods for isolating RNA-binding proteins. The present invention also provides methods for identifying RNA-binding proteins and for identifying binding sites of RNA-binding proteins in an RNA transcript sequence.

[0004] Background of the Invention

[0005] The recent success of COVID mRNA vaccines has showcased the great potential of RNA-biology field in delivering transformative innovations in the biomedical and biotech sectors. A key aspect of RNA-biology research is the unbiased study of RNA-Protein interactions. Unbiased discovery and mapping of RNA molecules which bind to a protein of interest in cellulo can be readily performed, thanks to recent advances in methods such as RNA Immunoprecipitation-sequencing (RIP-seq), individual-nucleotide resolution crosslinking and Immunoprecipitation (iCLIP), or enhanced crosslinking and Immunoprecipitation (eCLIP). These ‘protein-centric’ methods take advantage of nextgeneration sequencing sensitivity to reveal and map the transcriptome-wide targets of a given protein (Hafner et al. 2021).

[0006] In contrast, unbiased discovery of proteins that bind to an endogenous RNA of interest is much more challenging, requiring the purification of these RNAs from cells, followed by mass-spectrometry (MS) analysis of the associated proteins. Such ‘RNA-centric’ methods typically rely on antisense oligo based pulldown of a target RNA from RNA-protein crosslinked cell lysates, followed by quantitative MS analysis to reveal the identity of specifically interacting proteins (Hafner et al. 2021). Despite the huge interest in the field, however, all available RNA-centric approaches suffer from several key shortcomings. These include low efficiency of antisense oligo-based RNA pulldowns, drastic lack of reproducibility between different approaches, and significant risk of off-target capture, resulting in low specificity.

[0007] Therefore, there is currently a major gap in the RNA biology field for new methods that can enable efficient, robust, and scalable profiling of proteins partners of specific RNAs. The invention fills this major gap. In short, and as described in more detail herein, the invention provides the RNA-biology field with a set of revolutionary new RNA-centric methods for probing RNA-protein interactions that are far more specific, efficient, and reproducible than existing methods, as well as enabling region-specific mapping of RNA- protein interactions.

[0008] Summary of the Invention

[0009] The present inventors have identified a novel method for isolating RNA-binding proteins from a sample. On the basis of this method, it is in turn possible to identify RNA- binding proteins and furthermore identify a binding site of an RNA-binding protein in an RNA transcript sequence.

[0010] The methods involve the use of a plurality of DNA oligonucleotides that hybridise to specific target regions in an RNA molecule of interest. The methods subsequently involve the use of an enzyme that cleaves DNA-RNA hybrids thereby releasing proteins previously cross-linked to the RNA molecule of interest. The released proteins may then be identified by known protein identification techniques such as quantitative mass spectrometry. Sequence-specific probe design in conjunction with the use of an enzyme having specific activity towards DNA-RNA hybrids results in the method of the invention achieving unprecedented activity, with only the specific target RNA of interest undergoing cleavage and release of its associated proteins. Such an enzyme-based approach is vastly more efficient than alternative methods which use RNA pulldown. Furthermore, since enzyme-based cleavage can be specifically targeted to a particular region of a given RNA, region-specific mapping of interacting proteins from endogenous samples is possible by the method of the invention, a capability that is not conceivable via other existing pulldownbased assays.

[0011] The invention provides a method comprising:

[0012] (a) contacting a sample with a cross-linking agent thereby providing RNA- protein adducts; (b) performing phase separation on the sample containing the RNA-protein adducts to produce a distinct phase that comprises RNA-protein adducts but is at least substantially free from free protein and free RNA;

[0013] (c) isolating the RNA-protein adducts from the distinct phase of part (b);

[0014] (d) contacting the isolated RNA-protein adducts with a plurality of DNA oligonucleotides that comprise complementary nucleotide sequence to a target RNA molecule of interest to thereby form one or more DNA-RNA hybrids;

[0015] (e) contacting the sample from part (d) with an enzyme that cleaves DNA-RNA hybrids thereby releasing proteins from the RNA molecule of interest.

[0016] The method optionally further comprises isolating an RNA-binding protein from the sample, the method comprising:

[0017] (f) performing a further step of phase separation on the sample to produce a distinct phase that comprises released RNA-binding proteins but is at least substantially free from free RNA; and

[0018] (g) isolating the released RNA-binding proteins from the distinct phase of part (f).

[0019] The invention further provides a method for identifying RNA-binding proteins, the method comprising the steps of the method for isolating an RNA-binding protein from a sample according to the invention, and further comprising subjecting the released RNA- binding protein to a technique for protein identification.

[0020] The invention further provides a method for identifying a binding site of an RNA- binding protein in an RNA transcript sequence, the method comprising the steps of the method for isolating an RNA-binding protein from a sample according to the invention together with the method for identifying RNA-binding proteins according to the invention, further wherein the method comprises mapping the plurality of DNA oligonucleotides that comprise complementary nucleotide sequence to a target RNA molecule of interest to an RNA transcriptome, and inferring the binding site of the released RNA-binding protein in the RNA transcript sequence on the basis of the DNA oligonucleotides that are mapped to the RNA transcriptome. Brief Description of the Figures

[0021] FIGURE 1 shows a schematic representation of TREX.

[0022] FIGURE 2 shows an example of poor overlap across three different Antisense based transcript-specific RNA capture studies, aimed at revealing U1 RNA interactors, a well characterised non-coding RNA component of the splicosome core.

[0023] FIGURE 3 shows a method comparison of TREX with the most common antisense based transcript-specific RNA capture methods (RAP, CHART, CHIRP).

[0024] FIGURE 4 shows a retrospective assessment of efficiency and specificity is possible with TREX, by including an additional QC step (red box, marked ‘QC’), performed on a fraction of the post-RNase-H treatment reaction, in which RNA is first extracted by proteinase K treatment, followed by RT-qPCR analysis (to test for degradation efficiency) or RNA-sequencing (to test for degradation specificity).

[0025] FIGURE 5 shows assessment of efficiency of 18S degradation by RT-qPCR, from a fraction of + and - RNase-H treated samples. The abundance of 18S rRNA transcript relative to GAPDH (a housekeeping gene) mRNAs was quantified in each sample. A near complete loss of 18S rRNA was evident in the +RNase-H samples.

[0026] FIGURE 6 shows the interactome of 18S, as revealed by TREX. Volcano plot of the t-test results comparing + and - RNase-H treated 18S TREX samples (FDR < 0.05). The vast majority of the 40S ribosomal proteins, as well as many known 40S ribosomal subunit associated factors are amongst the confidently identified interactors of 18S (right hand side of the graph).

[0027] FIGURE 7 shows the category enrichment analysis of 18S interacting proteins from TREX. The list of identified interactors of 18S from figure 6 were subjected to Category Enrichment analysis, using Fisher’s Exact test (FDR <0.02), with protein annotations taken from Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) databases. The top 20 enriched categories were selected and depicted.

[0028] FIGURE 8 shows analysis of known protein-protein interactions of the 18S binding proteins from TREX, using the STRING database. The vast majority of the identified proteins are part of a network of proteins around the 40S core ribosomal subunit (highlighted with the bold circle). FIGURE 9 shows a Venn diagram of the overlap between 18S TREX and RAP-MS (McHugh et al. 2015) datasets. Most 40S ribosomal proteins are amongst the shared targets of both methods (overlap), while TREX identifies several other 40S proteins as well as many more known 18S interactors (e.g. dozens of 18S biogenesis factors, many translation initiation factors, SRP-dependent ER targeting machinery, and several large (60S) ribosomal proteins that are primarily in contact with the small subunit). RAP-MS, on the other hand, exclusively reveals several large subunit biogenesis factors as well as a number of nucleolar proteins. As the majority of both small and large subunit RNAs initiate as a joint unprocessed transcript (45 S rRNA), it is possible that these transcripts get co-captured in RAP -MS before their processing and separation from each other. Since TREX acts in a region-specific manner, large subunit biogenesis factors will not be expected to be co-purified, even before processing of 45 S rRNA, as these factors simply do not directly bind to the 18S part.

[0029] FIGURE 10 shows an assessment of efficiency of U1 RNA degradation by RT-qPCR, from a fraction of + and -RNase-H treated samples. The abundance of U1 RNA transcript relative to GAPDH (a housekeeping gene) mRNAs was quantified in each sample. A near complete loss of U1 RNA was evident in the +RNase-H samples.

[0030] FIGURE 11 shows the interactome of U1 RNA, as revealed by TREX. Volcano plot of the t-test comparison of + and - RNase-H treated U1 TREX samples (FDR < 0.05). Specifically extracted proteins upon U1 degradation are on the right hand side of the graph. The vast majority of the U1 snRNP proteins and several splicing factors were amongst the confidently identified interactors of Ul.

[0031] FIGURE 12 shows an analysis of known protein-protein interaction of the Ul binding proteins from TREX, using the STRING database. As expected, the vast majority of the identified proteins are part of a known network of Ul snRNP proteins and splicing related factors, but a small independent complex of redox regulators (PRDX5 and TXN) is also identified, suggestive of a potentially novel role for Ul in redox regulation.

[0032] FIGURE 13 shows a Venn diagram of the overlap between Ul TREX and RAP-MS (McHugh et al. 2015) datasets. Most Ul snRNP proteins are amongst the shared targets of both methods (overlap), while TREX identifies another splicing factor and one additional Ul snRNP (SNRPE). FIGURE 14 shows a schematic diagram of different regions of 45 S pre-rRNA. The 18S, 5.8S and 28S regions are those that finally make up the majority of rRNA in ribosomes, while the 5’ETS, ITS1, ITS2 and 3’ETS regions are those that are ultimately processed and removed. FIGURE 15 shows the region-specific interactome of 5’ETS, as revealed by TREX.

[0033] Volcano plot of the t-test comparison of + and - RNase-H treated 5’ETS TREX samples (FDR < 0.05). Specifically extracted proteins upon 5’ETS degradation are on the right hand side of the graph. 8 known ribosome biogenesis factors were amongst the confidently identified interactors of the 5’ETS region FIGURE 16 shows an analysis of known protein-protein interaction of the 5’ETS binding proteins from TREX, using the STRING database. As expected, 8 out of the 16 identified proteins are part of a known network of ribosome biogenesis factors including UTP14A, UTP15, and WDR75, which are known to be important for early biogenesis events that involves 5’ETS.

[0034] Brief Description of the Sequences

[0035] Oligonucleotides for hybridising to 18S rRNA

[0036] Oligonucleotides for hybridising to U1 RNA

[0037] Oligonucleotides for hybridising to the 5 ’External Transcribed Region (5 TS) of the 45S pre-rRNA

[0038] Detailed Description of the Invention

[0039] The present invention will be described with respect to particular embodiments and with reference to certain drawings but the disclosure is not limited thereto but only by the claims. Any reference signs in the claims shall not be construed as limiting the scope. Of course, it is to be understood that not necessarily all aspects or advantages may be achieved in accordance with any particular embodiment. Thus, for example those skilled in the art will recognize that the disclosed embodiments may be embodied or carried out in a manner that achieves or optimizes one advantage or group of advantages as taught herein without necessarily achieving other aspects or advantages as may be taught or suggested herein.

[0040] The disclosure, both as to organization and method of operation, together with features and advantages thereof, may best be understood by reference to the following detailed description when read in conjunction with the accompanying drawings. The aspects and advantages of the invention will be apparent from and elucidated with reference to the embodiment(s) described hereinafter. Reference throughout this specification to "one embodiment" or "an embodiment" means that a particular feature, structure or characteristic described in connection with the embodiment is included in at least one disclosed embodiment. Thus, appearances of the phrases "in one embodiment" or "in an embodiment" in various places throughout this specification are not necessarily all referring to the same embodiment, but may. Similarly, it should be appreciated that in the description of exemplary disclosed embodiments, various features are sometimes grouped together in a single embodiment, figure, or description thereof for the purpose of streamlining the disclosure and aiding in the understanding of one or more of the various inventive aspects. This method of disclosure, however, is not to be interpreted as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive aspects lie in less than all features of a single foregoing disclosed embodiment.

[0041] It should be appreciated that “embodiments” of the disclosure can be specifically combined together unless the context indicates otherwise. The specific combinations of all disclosed embodiments (unless implied otherwise by the context) are further disclosed embodiments of the claimed invention.

[0042] In addition as used in this specification and the appended claims, the singular forms “a”, “an”, and “the” include plural referents unless the content clearly dictates otherwise. Thus, for example, reference to “a polynucleotide” includes two or more polynucleotides, reference to “a protein” includes two or more proteins, and the like.

[0043] Unless otherwise indicated, nucleic acid sequences herein are written in the 5’-to-3’ direction from left to right.

[0044] All publications, patents and patent applications cited herein, whether supra or infra, are hereby incorporated by reference in their entirety.

[0045] Method of isolating of isolating RNA-binding proteins

[0046] The disclosure provides a method comprising:

[0047] (a) contacting a sample with a cross-linking agent thereby providing RNA- protein adducts;

[0048] (b) performing phase separation on the sample containing the RNA-protein adducts to produce a distinct phase that comprises RNA-protein adducts but is at least substantially free from free protein and free RNA;

[0049] (c) isolating the RNA-protein adducts from the distinct phase of part (b);

[0050] (d) contacting the isolated RNA-protein adducts with a plurality of DNA oligonucleotides that comprise complementary nucleotide sequence to a target RNA molecule of interest to thereby form one or more DNA-RNA hybrids;

[0051] (e) contacting the sample from part (d) with an enzyme that cleaves DNA- RNA hybrids thereby releasing proteins from the RNA molecule of interest. The method comprises a step of providing RNA-protein adducts. Preferably said proteins are RNA-binding proteins. The method comprises a step of isolating RNA- protein adducts, and thus preferably the isolation of RNA-binding proteins.

[0052] The method may further comprise isolating an RNA-binding protein from the sample, the method comprising:

[0053] (f) performing a further step of phase separation on the sample to produce a distinct phase that comprises released RNA-binding proteins but is at least substantially free from free RNA; and

[0054] (g) isolating the released RNA-binding proteins from the distinct phase of p art (f).

[0055] The “sample” is used herein to refer to any material containing at least one protein bound to RNA. The sample preferably comprises a plurality of proteins bound to different RNA molecules and / or different regions of the same RNA molecule.

[0056] An exemplary sample may be a soil sample, or a sample of any material or tissue obtained from a plant, animal or microorganism. The sample material may be tissue. Preferred animal materials include hair follicles and body fluids such as blood, saliva, semen, vaginal fluids, mucus, urine or any other humoral material. The sample may comprise material derived from cells that have been cultured in vitro or derived from human or animal tissue, optionally wherein the material comprises a lysate of cells or human or animal tissue. The sample may comprise cells being subject to in vitro culture, optionally wherein the cells are adhered to a cell culture plate or are suspended in solution.

[0057] The material may comprise a lysate of cells or human or animal tissue, and wherein the lysate is further subject to subcellular fractionation. Subcellular fractionation methods are well known in the art. The cell fraction may for example be a whole cell fraction, a cytosolic fraction, or a nuclear fraction. Preferably the lysate has been prepared under conditions free from active ribonuclease enzymes.

[0058] The number of cells in the sample may be any suitable quantity, provided that the means for cross-linking is subject to routine adaptation to ensure that a sufficient number of cells are cross-linked and RNA-protein adducts are successfully provided. There may be at least about 1 x 106cells comprised in the sample, optionally at least about 2.5 x 106, 5 x 106, 10 x 106, or 100 x 106. Most preferably there are at least about 5 x 106cells comprised in the sample.

[0059] The “RNA-binding protein” isolated in the methods disclosed herein may be any RNA-binding protein. The protein may bind directly to RNA, preferably wherein the binding is via a non-covalent interaction. The protein may be bound indirectly to RNA, for example via an interaction with another protein, preferably wherein said interaction is non-covalent. Most preferably the protein is bound directly to RNA via a non-covalent interaction.

[0060] The sample may comprise any one or more types of RNA. The RNA may be messenger RNA (mRNA), ribosomal RNA (rRNA), signal recognition particle RNA (7SL RNA or SRP RNA), transfer RNA (tRNA), transfer-messenger RNA (tmRNA), small nuclear RNA (snRNA), small nucleolar RNA (snoRNA), SmY RNA (SmY), small Cajal body-specific RNA (scaRNA), guide RNA (gRNA), ribonuclease P (RNase P), ribonuclease MRP (RNase MRP), Y RNA, telomerase RNA component (TERC), spliced leader RNA (SL RNA), antisense RNA (aRNA, asRNA), cis-natural antisense transcript (cis-NAT), CRISPR RNA (crRNA), long noncoding RNA (IncRNA), microRNA (miRNA), piwi-interacting RNA (piRNA), small interfering RNA (siRNA), short hairpin RNA (shRNA), trans-acting siRNA (tasiRNA), repeat associated siRNA (rasiRNA), 7SK RNA (7SK), enhancer RNA (eRNA), retrotransposon, viral genome, viroid, or satellite RNA. Most preferably the sample comprises rRNA, mRNA and / or ncRNA.

[0061] The RNA bound by the RNA-binding protein may be any known RNA sequence. The RNA may be a wild type sequence, including splice variants. The RNA sequence may alternatively be a mutant sequence. The known sequence may be an entire RNA molecule or a region comprised within an RNA molecule. Knowledge of the sequence of a particular RNA means that DNA oligonucleotides can be designed the comprise complementary nucleotide sequence to a sequence within RNA, thereby allowing isolation of RNA-binding proteins by the methods disclosed herein.

[0062] The method comprises contacting a sample with a cross-linking agent thereby providing RNA-protein adducts. The cross-linking agent functions to induce cross-linking between RNA and protein molecules that are in close proximity to one another. Typically, the proximity is such that the RNA and protein molecules are deemed bound to one another via a non-covalent interaction. A cross-linking agent induces one or more covalent bonds to form between such RNA and protein molecules, thereby “fixing” the non-covalent interaction and consequently providing one or more RNA-protein adducts.

[0063] The cross-linking agent can therefore be any agent that is capable of inducing cross-linking between an interacting protein and RNA molecule. Suitable cross-linking agents are well known in the art. For example, the sample may be contacted with radiation and / or one or more chemicals in order to induce cross-link formation and thus RNA- protein adducts. The cross-linking agent preferably comprises ultraviolet radiation or formaldehyde. Ultraviolet radiation may be any suitable wavelength, although preferably 365 nm, or more preferably 254 nm. When UV 365 nm is used to contact a sample to form RNA-protein adducts, the sample is preferably also contacted with 4-Thiouridine.

[0064] The method comprises performing phase separation on the sample containing the RNA-protein adducts to produce a distinct phase that comprises RNA-protein adducts but is at least substantially free from free protein and free RNA.

[0065] The step of phase separation is preferably a method of phase separation which involves separation of a liquid sample into an organic phase and an aqueous phase. In the method of the invention whereby the sample has been contacted with a cross-linking agent to thereby provide RNA-protein adducts, an interphase will additionally form together with the organic phase an aqueous phase.

[0066] Preferably, all free RNA, or substantially all free RNA, i.e. RNA that is not crosslinked to a protein and is therefore not comprised within an RNA-protein adduct, is comprised within the aqueous phase. Preferably, all free proteins, or substantially all free proteins, i.e. proteins that are not crosslinked to RNA, are comprised within the organic phase. RNA-protein adducts are comprised within a distinct phase that comprises RNA- protein adducts but is at least substantially free from free protein and free RNA, and preferably wherein the distinct phase corresponds to the interphase.

[0067] Methods of liquid phase separation are generally known in the art. The step of phase separation preferably comprises the use of one or more of acidic phenol, guanidine isothiocyanate, and chloroform. Even more preferably, the step of phase separation preferably comprises the use all of acidic phenol, guanidine isothiocyanate, and chloroform. The use of one of more of these agents comprises contacting the sample with the one or more agents, and thereby inducing the phase separation. General liquid phase separation involving the use of said one or more agents is known in the art, and is particularly known to involve one or more steps of mixing and centrifugation as described in the Examples herein.

[0068] The method comprises isolating the RNA-protein adducts from the distinct phase in the phase separation step. Such isolation can be achieved by routine methods of liquid extraction from liquid samples. The step of isolation may comprise removing the aqueous and organic phases from a vessel containing the sample, and leaving the interphase contained within said vessel, thereby allowing the RNA-protein adducts to be isolated from the interphase.

[0069] The phase separation and isolation steps may together be repeated in turn on one or more further occasions before proceeding to contact the RNA-protein adducts with a plurality of DNA oligonucleotides.

[0070] Prior to contacting the RNA-protein adducts with a plurality of DNA oligonucleotides, the sample may also be contacted with a ribonuclease inhibitor and an agent that digests genomic DNA, preferably wherein the agent is a deoxyribonuclease.

[0071] The method comprises contacting the isolated RNA-protein adducts with a plurality of DNA oligonucleotides that comprise complementary nucleotide sequence to a target RNA molecule of interest to thereby form one or more DNA-RNA hybrids. One or more, or all, of the plurality of DNA oligonucleotides may consist of a nucleotide sequence that is complementary to a target RNA molecule of interest.

[0072] The complementary nucleotide sequence allows the formation of DNA-RNA- hybrids by the annealing of the DNA oligonucleotides to the complementary RNA sequence. The complementarity of the plurality of DNA oligonucleotides means that the method may be used to target the plurality of DNA oligonucleotides to one or more specific regions of RNA which the user wishes to investigate. This allows the user to isolate and thereby determine the RNA-binding proteins that are bound at said one or more specific regions. The DNA oligonucleotides are preferably non-overlapping in terms of sequence within the RNA molecule of interest that is being target and therefore do not compete with one another for annealing to the one or more specific regions.

[0073] The plurality of DNA oligonucleotides may tile the one or more specific regions of the RNA molecule of interest or the entirety of, or optionally substantially the entirety of, one or more specific RNA transcripts. When one or more specific regions of the RNA molecule of interest are subject to investigation, it is preferable that the plurality of DNA oligonucleotides tile the entirety of, or optionally substantially the entirety of, the one or more specific regions such that all of said one or more specific regions are annealed to a DNA oligonucleotide. When the entirety of one or more specific RNA transcripts are of interest, it is preferable that the plurality of DNA oligonucleotides tile the entirety of, or optionally substantially the entirety of, the one or more specific RNA transcripts such that all of said RNA transcripts are annealed to a DNA oligonucleotide.

[0074] The DNA oligonucleotides may be any suitable length such that specific annealing to a target RNA sequence can be achieved. Preferably the DNA oligonucleotides are at least about 20, at least about 30, at least about 40, at least about 50 or at least about 60 nucleotides in length. More preferably the DNA oligonucleotides are about 30 nucleotides in length to about 90 nucleotides in length, but most preferably the DNA oligonucleotides are about 60 nucleotides in length. Preferably the DNA oligonucleotides are unmodified.

[0075] The method comprises contacting the sample with an enzyme that cleaves DNA- RNA hybrids thereby releasing proteins from the RNA molecule of interest. Enzymes capable of cleaving DNA-RNA hybrids are known in the art. Furthermore, means for testing the efficiency of such cleavage are well known in the art. Preferably, the enzyme that cleaves the DNA-RNA hybrids in the method of the invention is RNaseH, or a functionally active fragment or analog thereof. Preferably, the RNaseH, or functionally active fragment or analog thereof, is thermostable and thereby capable of cleaving DNA- RNA hybrids at temperatures of at least 50° C. The enzyme may alternatively be a recombinant enzyme that comprises a catalytic domain of RNaseH, or a functionally active fragment or analog thereof. The effect of the cleavage preferably results in the degradation of DNA-RNA hybrid sequence. The method may comprise a step of testing cleavage efficiency and / or specificity. This step allows the user to determine whether the RNA targeted by the plurality of DNA oligonucleotides, and subsequently by the enzyme that cleaves DNA-RNA hybrids, remains intact after performing the steps of the method. Intact RNA would indicate that one or both of the steps of forming DNA-RNA hybrids and cleavage using an enzyme that cleaves DNA-RNA hybrids lack specificity and / or efficiency. Intact RNA would further indicate that proteins have not been released from the RNA molecule of interest in accordance with the objective of the method. The step of testing cleavage efficiency and / or specificity may comprise an unbiased technique, such as whole transcriptome RNA-seq. The step of testing cleavage efficiency and / or specificity may comprise a biased technique, such as quantitative PCR.

[0076] The method may further comprise isolating an RNA-binding protein from the sample, the method comprising:

[0077] (f) performing a further step of phase separation on the sample to produce a distinct phase that comprises released RNA-binding proteins but is at least substantially free from free RNA; and

[0078] (g) isolating the released RNA-binding proteins from the distinct phase of part (f).

[0079] Methods of liquid phase separation are generally known in the art. The performing of a further step of phase separation on the sample to produce a distinct phase that that comprises released RNA-binding proteins but is at least substantially free from free RNA is following the release of the proteins from the RNA molecule of interest. The distinct phase is preferably at least substantially free from free DNA. The distinct phase is preferably at least substantially free from RNA-protein adducts. Such RNA-protein adducts may remain present in the sample in step (f) for example because in some aspects not all RNA-protein adducts are contacted by the plurality of DNA oligonucleotides, and hence in some aspects not all RNA-binding proteins are released from the RNA to which they are crosslinked upon contacting the sample with an enzyme that cleaves DNA-RNA hybrids. The further step of phase separation preferably comprises the use of one or more of acidic phenol, guanidine isothiocyanate, and chloroform. Even more preferably, the step of phase separation preferably comprises the use all of acidic phenol, guanidine isothiocyanate, and chloroform. The use of one of more of these agents comprises contacting the sample with the one or more agents, and thereby inducing the phase separation. General liquid phase separation involving the use of said one or more agents is known in the art, and is particularly known to involve one or more steps of mixing and centrifugation as described in the Examples herein.

[0080] The further step of phase separation preferably leads to the production of an organic phase, an interphase and an inorganic phase, and wherein the distinct phase that comprises released RNA-binding proteins but is at least substantially from free RNA corresponds to the organic phase.

[0081] The method comprises isolating the released RNA-binding proteins from the distinct phase in the further step of phase separation. Such isolation can be achieved by routine methods of liquid extraction from liquid samples.

[0082] Method for identifying RNA-binding proteins

[0083] The disclosure also provides a method for identifying RNA-binding proteins, the method comprising the steps of the method for isolating an RNA-binding protein from a sample according to the invention, and further comprising subjecting the released RNA- binding protein to a technique for protein identification.

[0084] The technique for protein identification may be a biased technique, optionally wherein the technique comprises the use of an antibody for targeting a protein of interest. The biased technique may be an antibody-based pull-down or an immunoblot assay.

[0085] The technique for protein identification may be an unbiased technique, optionally wherein the technique may comprise any suitable form of quantitative mass spectrometry, preferably wherein the quantitative mass spectrometry comprises a method of shotgun proteomics such as label-free quantification mass spectrometry.

[0086] Method for identifying a binding site of an RNA-binding protein The disclosure also provides a method for identifying a binding site of an RNA- binding protein in an RNA transcript sequence, the method comprising the steps of the method for isolating an RNA-binding protein from a sample according to the invention together with the steps of the method for identifying RNA-binding proteins according to the invention, further wherein the method comprises mapping the plurality of DNA oligonucleotides that comprise complementary nucleotide sequence to a target RNA molecule of interest to an RNA transcriptome, and inferring the binding site of the released RNA-binding protein in the RNA transcript sequence on the basis of the DNA oligonucleotides that are mapped to the RNA transcriptome.

[0087] The plurality of DNA oligonucleotides can be mapped to an RNA transcriptome using any suitable technique or software for sequence mapping.

[0088] The transcriptome can be the transcriptome of any organism for which some or all of the RNA transcriptome is known.

[0089] It is to be understood that although particular embodiments, specific configurations as well as materials and / or molecules, have been discussed herein for methods according to the present disclosure, various changes or modifications in form and detail may be made without departing from the scope and spirit of this invention. The preceding embodiment and following examples are provided for illustration only, and should not be considered limiting the application. The application is limited only by the claims

[0090] Examples

[0091] Disclosed are novel methods for purification and identification of proteins that are associated with any chosen endogenous RNA of interest, from cells and tissues. Named Targeted RNase H mediated Extraction of X-linked proteins (TREX), the methods utilise organic phase separation to isolate protein-RNA adducts from cells or tissues following RNA-protein crosslinking, followed by specific degradation of a given RNA target region of interest using RNase-H, which in turn results in specific release of proteins that were covalently associated with that RNA region. The released proteins can subsequently be identified by quantitative mass spectrometry, or other protein identification techniques (Figure 1).

[0092] Thanks to the highly specific nature of RNase-H activity, TREX achieves unprecedented specificity, with only the specific target RNA of interest undergoing complete degradation and release of its associated proteins. Moreover, as an enzymatic reaction, RNase-H based release of associated proteins is vastly more efficient than alternative methods which use RNA pulldown. Finally, since RNase-H degradation can be specifically targeted to a particular region of a given RNA, region-specific mapping of interacting proteins from endogenous samples is possible by TREX, a capability that is not conceivable via other existing pulldown-based assays. As proof of concept, we have validated TREX on different well-known RNA targets, showing that it can robustly and accurately characterise their interacting proteins. Crucially, this is achieved from significantly less amount of starting material, compared to previously published pulldown based methods, meaning that TREX can be readily scaled and used in diverse experimental settings.

[0093] There are several iterations of TREX. In some iterations, the approach uses whole cell or whole tissue lysis to capture all proteins bound to an RNA of interest. In other iterations, subcellular fractionation may be first applied to focus on complexes in a given subcellular compartment. Different crosslinking methods can also be used to capture either direct RNA-protein interactions (using UV-C crosslinking), or indirect interactions (using formaldehyde crosslinking). Finally, in some iterations, RNase-H based degradation can be done for the full length of a target RNA, while in other iterations only a specific region of a target RNA can be targeted, allowing region-specific mapping of interactions. Here we describe a detailed step-by-step protocol for one iteration of TREX, which was used to generate the example data provided for 18S and U1 RNAs.

[0094] Materials - biological materials

[0095] The biological material for our TREX test studies was human HCT116 cells, but any other cell-line can be used in the same manner. It is important to regularly check cell lines to ensure that they are authentic and are not infected with mycoplasma. Handle cell lines according to the supplier’s instructions. Work in a biosafety hood, use sterile equipment, and wear gloves to minimize the risk of contamination.

[0096] Materials - reagents

[0097] • RNaseZAP™ Surface Decontaminant (Sigma, R2020)

[0098] • Nuclease-free water (Thermo, AM9932)

[0099] • Cytosolic Lysis buffer (1% Triton X-100, lOmM Tris-HCl pH 7.5 150 mM NaCl, cOmplete™ protease inhibitor cocktail (Sigma), SUPERase»In™ 1 / 500 v / v (Invitrogen , AM2694)) (Use RNase-free reagents and Nuclease-free water to prepare this buffer)

[0100] ! CAUTION Triton is harmful, and it is an irritant. Triton is hazardous to the environment. Handle solutions containing Triton with care and dispose of waste according to institutional regulations.

[0101] ! CAUTION cOmplete™ protease inhibitor cocktail is an irritant. Handle solutions containing the protease inhibitor mix with care and dispose of waste according to institutional regulations.

[0102] • TRIzol™ LS Reagent (Thermo, 10296028)

[0103] • TRIzol™ Reagent (Fisher Scientific Uk Ltd, 12034977)

[0104] ! CAUTION TRIzol™ reagents are toxic if inhaled and a potential carcinogen. Always handle solutions containing chloroform with care while wearing personal protective equipment in a chemical fume hood and dispose of waste according to institutional regulations.

[0105] • Chloroform

[0106] ! CAUTION Chloroform is volatile and toxic. Chloroform is an irritant. Handle solutions containing chloroform with care and dispose of chloroform waste according to institutional regulations.

[0107] • TE (Tris-Cl pH7,10 mM, EDTA 1 mM) (Use RNase-free reagents and Nuclease-free water to prepare this buffer) • TE+SDS 0.1% (Tris-Ci pH7, 10 mM, EDTA 1 mM, SDS 0.1 %) (Use RNase- free reagents and Nuclease-free water to prepare these buffers)

[0108] • TE+SDS 0.5% (Tris-Cl pH7, 10 mM, EDTA 1 mM, SDS 0.5 %) (Use RNase- free reagents and Nuclease-free water to prepare these buffers)

[0109] • IX Hybridization buffer: (50mM NaCl, ImM EDTA, lOOmM TrisHCl pH 7.0) (Use RNase-free reagents and Nuclease-free water to prepare this buffer)

[0110] • Thermostable RNAse H (NEB, M0523 S)

[0111] • Thermostable RNAse H buffer (NEB, M0523 S)

[0112] • Isopropanol

[0113] • Proteinase K (Thermo, AM2546)

[0114] • Proteinase K buffer (0.1 M NaCl, 10 mM Tris HC1 pH 8.0, 1 mM EDTA, 0.5% SDS) (Use RNase-free reagents and Nuclease-free water to prepare this buffer)

[0115] • Acetone

[0116] • Ethanol

[0117] • Dithiothreitol (DTT) >99.4%, Proteomics Grade (VWR , M109)

[0118] • lodoacetamide (IAA) >98%, Proteomics Grade (Sigma, 11149)

[0119] • Urea >99.5% (Sigma , U1250)

[0120] • ABC buffer (50 mM Ammonium Bicarbonate, pH 7.8)

[0121] • Trypsin, LC-MS / MS grade (Sigma, 650279)

[0122] • Stop buffer (4% acetonitrile, 1% TFA)

[0123] • 5M NaCl (Invitrogen REF AM9760G)

[0124] • Invitrogen GlycoBlue Coprecipitant 15mg / mL (Fisher Scientific Uk Ltd, 10301575)

[0125] • IM Tris HC1 pH 7.0 (Invitrogen, AM9850G)

[0126] • 0.5M EDTA pH 8.0 (Invitrogen, AM9260G)

[0127] • IM MgCL (Invitrogen, AM9530G)

[0128] • 10% SDS (Invitrogen, 15553-035)

[0129] • TURBO DNase 10X buffer (Fisher Scientific, 10646175) • TURBO DNase (Fisher Scientific, 10646175)

[0130] • RNasin® Plus RNase inhibitor (Promega, N2611)

[0131] Materials - equipment

[0132] • Ultraviolet crosslinker (Hoefer, UVC500)

[0133] • Cell scrapers (Coming, cat.no 353087)

[0134] • 150 mm TC-treated Culture Dish

[0135] • RNase-free Microfuge 2 ml tubes (Thermo, AMI 2425)

[0136] • RNase-free Microfuge 1.5 ml tubes

[0137] • 0.2 mL 8-Strip PCR tube (Thermo, AM12425)

[0138] • Hood for Trizol

[0139] • Veriti™ Thermocycler (Thermo, 4375786)

[0140] • Thermomixer comfort ( Eppendorf, No. 5355)

[0141] • Barrier pipette tips

[0142] • Benchtop centrifuge (Eppendord, 5415R)

[0143] • C18 Stage tips

[0144] • Mass spectrometer

[0145] • Vortexer (Scientific industries, SI-0236)

[0146] • Qubit 2.0 fluorometer (Invitrogen, Q32866)

[0147] • QuantStudio 7 Flex Real-Time PCR (Thermo, 4485701)

[0148] • RTqPCR plates

[0149] Materials - software

[0150] • Q7 RTqPCR software

[0151] • Maxquant

[0152] • Perseus

[0153] Protocol steps: 1) In vivo cross-linking - timing 2 min Seed HCT116 cells on 150 mm TC-treated Culture Dish and allow attachment for 16 hours in a 37 °C incubator with 5% CO2 prior to crosslinking.

[0154] 1. Aspirate the media and add 15 ml cold PBS.

[0155] 2. Aspirate the PBS and add another 15 ml of cold PBS.

[0156] 3. Aspirate the PBS.

[0157] 4. Transfer the dish onto ice or an ice receptacle that fits into the cross-linker device (cross-linker device sensor should not be covered to continually measure the UV energy)

[0158] 5. Cross-link cells at 200-400 mJ / cm2on ice (dish lid should be removed from plate to ensure crosslinking efficiency)

[0159] Note: To perform TREX on total cell RNA proceed with 2A. If prior subcellular fractionation is required, proceed with 2B.

[0160] Protocol steps: 2 A) Whole cell lysis - timing 10 min

[0161] 6. Add 1ml of TRIzol™ for every 20 million cells on the plate

[0162] 7. Scrape the cells thoroughly. Make sure the TRIzol™ has covered the whole plate while scraping.

[0163] 8. Homogenize the lysate through repeated pipetting.

[0164] PAUSE POINT Store the lysate for up to 3 months at -80 °C if needed.

[0165] Protocol steps: 2B) Subcellular fractionation - timing 20 min

[0166] 9. Scrape cells in 1 mL of RNase-free ice-cold PBS per 20 million cells. Transfer each 1ml of lysate to an 2 ml RNase-free Eppendorf tube.

[0167] 10. Spin down for 1 minute at 500 g at 4°C to pellet the cells. Discard the PBS supernatant.

[0168] 11. Lyse the cell pellet in 150 pl Cytosolic Lysis buffer.

[0169] 12. Spin for 15 min at 10,000 g at 4°C and collect the supernatant, which contains the cytosolic fraction.

[0170] Note: Alternatively keep the remaining pellet if the nuclear fraction is need. PAUSE POINT Store the fractions for up to 3 months at -80 °C if needed. Protocol steps: 3 A) TRIzol™ phase separation for whole cell lysate (following on from 2 A) Timing Ih

[0171] 13. Incubate the homogenized lysate for 5 min at room temperature (20-

[0172] 25 °C) to dissociate non-crosslinked RNA-protein interactions.

[0173] 14. Add 200 pl of chloroform / 1ml of TRIzol™ and vortex sample until the content is homogenous.

[0174] 15. Centrifuge for 15 min at 12,000 g at 4°C Three phases should now be visible.

[0175] 16. Remove the upper aqueous phase until the interface collapses. Remove as much of the lower organic phase as possible without disrupting the interface.

[0176] 17. Resolubilise the interface in 1ml of TRIzol™ / 20 million cells

[0177] 18. Add 200 pl of chloroform / 1ml of TRIzol™ and vortex sample until the content is homogenous

[0178] 19. Centrifuge for 15 min at 12,000 g at 4°C

[0179] 20. Remove the upper aqueous phase until the interface collapses. Remove as much of the lower organic phase as possible without touching the interface.

[0180] 21. Resolubilise the interface in 1 ml of TRIzol™ / 20 million cells

[0181] PAUSE POINT Store the lysate at -80 °C if needed.

[0182] 22. Add 200 pl of chloroform / 1ml of TRIzol™ and vortex sample until the content is homogenous

[0183] 23. Centrifuge for 15 min at 12,000 g at 4°C

[0184] 24. Remove the upper aqueous phase until the interface collapses. Remove as much of the lower organic phase as possible without touching the interface.

[0185] Note: The following steps are required if you need to work with more than 20 million cells per replicate. 25. Combine the interfaces from up to a 100 million cells in 1ml of TRIzol™ and transfer into a 2 ml tube.

[0186] 26. Add 200 pl of chloroform and vortex sample until the content is homogenous.

[0187] 27. Centrifuge for 15 min at 12,000 g at 4°C.

[0188] 28. Remove the upper aqueous phase until the interface collapses. Remove as much of the lower organic phase as possible without touching the interface.

[0189] Protocol steps: 3B) TRIzol™ LS phase separation for subcellular TREX (following on from 2B) Timing Ih

[0190] 29. Add 1 ml of TRIzol™ LS to the sample and homogenize through repeated pipetting.

[0191] Note: Use standard TRIzol™ if continuing with the nuclear fraction.

[0192] 30. Incubate homogenized lysate for 5 min at room temperature (20-25 °C) to dissociate non-crosslinked RNA-protein interactions.

[0193] 31. Add 200 pl of chloroform and vortex sample until the content is homogenous.

[0194] 32. Centrifuge for 15 min at 12,000 g at 4°C.

[0195] Three phases should now be visible.

[0196] 33. Remove the upper aqueous phase until the interface collapses. Remove as much of the lower organic phase as possible without touching the interface.

[0197] 34. Resolubilise the interface in 1 ml of TRIzol™.

[0198] 35. Add 200 pl of chloroform and vortex.

[0199] 36. Centrifuge for 15 min at 12,000 g at 4°C. Remove upper and lower phases as before.

[0200] 37. Resolubilise the interface in 1 ml of TRIzol™.

[0201] PAUSE POINT Store the lysate at -80 °C if needed.

[0202] 38. Add 200 pl of chloroform and vortex. 39. Centrifuge for 15 min at 12,000 g at 4°C. Remove upper and lower phases as before, leaving a maximum of 50 pl.

[0203] Note: The following steps are required if you work with more than 20 million cells per replicate.

[0204] 40. Combine the interface from up to a 100 million cells in 1ml of TRIzol™ and transfer into a 2 ml tube.

[0205] 41. Add 200 pl of chloroform and vortex sample until the content is homogenous.

[0206] 42. Centrifuge for 15 min at 12,000 g at 4°C.

[0207] 43. Remove the upper aqueous phase until the interface collapses. Remove as much of the lower organic phase as possible without touching the interface.

[0208] Protocols steps: 4) Interface solubilisation andDNA removal Timing 4 h Interface wash:

[0209] 44. Without disturbing the interface, add 1ml of TE buffer by pipetting it slowly along the wall into the tube.

[0210] Do not resuspend or disturb the interface.

[0211] 45. Slowly invert the tube 3 times, taking extra care not to disturb the interface much.

[0212] 46. Centrifuge for 1 min at 5000 g at 4°C.

[0213] 47. Remove the supernatant.

[0214] 48. If needed, repeat step 44 to 47 1 or 2 more times till no more red dye from TRIzol™ is visible.

[0215] Note: After centrifugation a small amount of TRIzol™ might appear trapped under the interface pellet. Repeating the wash (step 44-47) will remove this.

[0216] Breaking down the interface:

[0217] 49. Resuspend the interface through pipetting in 1ml of TE+SDS 0.1% for up 30-60 times. 50. When the interface stops to break down any further, centrifuge for 2 min at 5000 g at room temperature.

[0218] 51. Transfer the supernatant into a new 2ml tube (Collection tube 1)

[0219] Note: From this point onwards, the interface should become increasingly translucent.

[0220] Do not transfer any of the interface into the collection tube.

[0221] 52. Resuspend the remaining interface through repeated pipetting in 1ml of TE+SDS 0.1% for up 30-60 times.

[0222] 53. Centrifuge for 2 min at 5000 g at room temperature.

[0223] 54. Transfer the supernatant into a new 2ml tube (Collection tube 2)

[0224] 55. Resuspend the remaining interface through repeated pipetting in 1ml of TE+SDS 0.5% for up 30-60 times.

[0225] 56. Centrifuge for 2 min at 5000 g at room temperature.

[0226] 57. Transfer the supernatant into a new 2ml tube (Collection tube 3)

[0227] 58. Resuspend the remaining interface through pipetting in 1ml of TE+SDS 0.5% for up 30-60 times.

[0228] 59. Centrifuge for 2 min at 5000 g at room temperature.

[0229] 60. Transfer the supernatant into a new 2ml tube (Collection tube 4)

[0230] Precipitation 1 :

[0231] 61. To each collection tube add 60 pl of NaCl 5M, 1 pl Glycoblue and 1ml of isopropanol.

[0232] 62. Invert the tube multiple times.

[0233] PAUSE POINT Samples can be stored at -20 °C overnight if needed.

[0234] 63. Centrifuge for 15 min at 18000 g (or faster) at 4°C (or colder).

[0235] Ethanol wash:

[0236] 64. Carefully remove the supernatant from all collection tubes.

[0237] 65. Add 1ml of 70% Ethanol to the first collection tube and break down the pellet through repeated pipetting. Then transfer the content to the next collection tube and repeat the process for all 4 tubes until all the pellets are combined in the same Ethanol solution.

[0238] 66. Use another 1 ml of 70% Ethanol to rinse any residual pellets from the now empty collection tubes. Combine this solution with the collection from step 35, leaving you with 2ml of ethanol suspension containing all the pellets.

[0239] 67. Centrifuge for 1 min at 18000 g at room temperature.

[0240] 68. Discard the supernatant. Remove any remaining ethanol using an additional short spin.

[0241] Hydration:

[0242] Note: Skip this step if TREX is being performed on the cytosolic fraction. Note: The following volumes are for up to 100 million cells.

[0243] 69. Add 1.8ml of nuclease-free water to the pellet.

[0244] 70. Invert the tube multiple times or vortex briefly to detach the pellet from the bottom the tube.

[0245] 71. Let the pellet hydrate for 1 hour on ice with occasional inversions. Note: Pellet will become translucent.

[0246] 72. Using a 1 ml pipette, resuspend the pellet repeatedly until it is completely dissolved.

[0247] DNase treatment:

[0248] Note: Skip this step if TREX is being performed on the cytosolic fraction.

[0249] 73. Add 200 pl of IX TURBO DNase buffer and mix the solution through pipetting.

[0250] 74. Add 2 pl of RNase inhibitor and 100 pl of TURBO DNase.

[0251] 75. Invert tube 4-5 times and incubate for 50 min at 37 °C, 700 rpm shaking.

[0252] Precipitation 2:

[0253] Note: Skip this step if TREX is being performed on the cytosolic fraction. 76. Split sample into two 2ml tubes (1ml each)

[0254] 77. To each collection tube add 60 pl of NaCl 5M and 1ml of isopropanol

[0255] 78. Invert the tube multiple times.

[0256] PAUSE POINT Samples can be stored at -20 °C overnight if needed.

[0257] 79. Centrifuge for 15 min at 18000 g (or faster) at 4°C (or colder).

[0258] Ethanol wash 2:

[0259] Note: Skip this step if TREX is being performed on the cytosolic fraction.

[0260] 80. Carefully remove the supernatant from each tube.

[0261] 81. Add 1ml of 70% Ethanol to the first tube and break down the pellet through pipetting. Transfer the content to the second tube and repeat the process until the pellets are combined in the same Ethanol solution.

[0262] 82. Use another 1 ml of 70% Ethanol to rinse any residual pellets from the now empty tubes. Combine this solution with the collection from step 81, leaving you with 2ml of ethanol containing the combined pellet.

[0263] 83. Centrifuge for 1 min at 18000 g at room temperature

[0264] 84. Discard the supernatant and remove any remaining ethanol using an additional short spin.

[0265] Protocol steps: 5) Probe Hybridization to RNA (annealing). Timing Ih

[0266] 85. Dissolve the pellet through repeated pipetting in 110 pl of Hybridization buffer per 20 million cells (e.g. if you started with 100 million cells per replicate, dissolve in 550 pl of hybridization buffer).

[0267] Pipette up and down until the pellet does not dissolve any further.

[0268] 86. Add 6 pl of 1 - 200 pM of tiling oligos* per 20 million cells**

[0269] (e.g.: for U1 RNA, we used 6 pl of 10 pM oligos per 20 million cells. Note that oligos are added to both control and experimental samples)

[0270] 87. Invert samples 4-5 times and give them a brief spin.

[0271] Summary table for the samples:

[0272] *We use antisense tiling DNA oligos designed against a target RNA of interest. Use our oligo design software to generate a list of ~60nt long tiling oligos against any input sequence.

[0273] TREX antisense DNA oligonucleotides can be designed by converting the target RNA sequence to its reverse-complement DNA counterpart, and dividing the resulting sequence into non-overlapping tiling sections of any chosen lengths. In our examples, the length of the oligos were designed to be ~60 nucleotides, with the last oligo at the 3’ end being flexible between 30 and 90, allowing the complete coverage of any target RNA length.

[0274] ** The amount of cells and oligos to use will depend on each target RNA and its copy number and must be empirically determined through qPCR for each set of oligos.

[0275] 88. The annealing will be performed in a Thermomixer. It is important to perform the following steps precisely to guarantee efficient annealing. a) Heat to 95 °C for 2 min with shaking at 1100RPM. b) While shaking, manually decrease the temperature by 2°C every minute, until you reach 50°C. c) Stop the shaking but do not take the samples out or turn off the heating. Move to step 89.

[0276] Do not put a cover over the samples during Annealing since this would not allow an even drop of temperature.

[0277] Protocol steps: 6) Thermostable RNase H Digestion. Timing 1 h 89. Without taking the samples out of the thermomixer, add MgCh to them to the final concentration of 0.55 mM. This is to neutralise the EDTA in the hybridisation buffer.

[0278] 90. Without taking the samples out of the thermomixer, add the following components to the experimental or control samples:

[0279] 91. Close the caps and turn on the shaking (1100RPM)

[0280] 92. Incubate the tubes in the thermomixer for 1 h at 50°C. 93. Centrifuge samples at 12000g for 3 min.

[0281] 94. Transfer the supernatant into a new tube.

[0282] Make sure to not transfer any insoluble particles collected at the bottom.

[0283] 95. Take 1 / 10 of the sample for analysis of RNA degradation by RTqPCR or RNA-sequencing (See Appendix). You can snap freeze and store this sample for up to 3 months at -80°C if needed.

[0284] 96. Proceed with the remaining 9 / 10thof the sample to step 97. Do not stop at this step with the protein preparation.

[0285] Protocol steps: 7) RBPs extraction. Timing 20 h

[0286] 97. Add 900 pl of TRIzol™ LS to each sample and vortex.

[0287] 98. Add 200 pl of chloroform and vortex the sample until the content becomes homogenous.

[0288] 99. Centrifuge for 15 min at 12,000 g at 4°C. Sample should phase separate into the three aqueous, interface, and organic phases.

[0289] It is important to check that an interface still forms after the phase separation. Lack of a visible interface would indicate non-specific degradation of RNA.

[0290] 100. Recover the releases RBPs by pipetting the organic phase into a new tube.

[0291] Only recover the organic phase, free of any interface. If pipetting all of the organic phase while excluding the interface is difficult, some of the organic phase may be left behind to insure no interface carry over.

[0292] 101. Add 100% cold acetone for a final concentration of 80% acetone, in order to precipitate the released RBPs. For example, if 300 pl of the organic phase is collected, add 1200 pl of 100% acetone. Vortex the mix thoroughly. PAUSE POINT Precipitate the released RBPs in acetone by incubating at - 20°C overnight.

[0293] 102. Centrifuge for 20 min at 16,000 g at 4°C.

[0294] 103. Discard the supernatant. RBPs should have precipitated and therefore be present as a pellet after centrifugation (as shown in the image).

[0295] 104. Add 1 ml of 80% acetone to wash the pellet.

[0296] 105. Vortex and centrifuge for 15 min at 16,000 g at 4°C.

[0297] 106. Repeat step 35 to 37.

[0298] 107. Discard supernatant and air-dry the pellet for 5 min.

[0299] Protocol steps: 8) Sample preparation for shotgun proteomics. Timing 2h 108. Recover RBP pellet in 100 pl 8M Urea in ABC buffer.

[0300] Prepare a fresh urea solution on the day.

[0301] Pipette the sample up and down thoroughly to resuspend pellet as much as possible. Any remaining insoluble material should dissolve over the course of the trypsin digestion (step 112).

[0302] 109. Add DTT from a IM stock to reach a final concentration of 10 mM

[0303] DTT. Incubate the samples for 30 min at room temperature (20-25 °C) to reduce the proteins.

[0304] 110. Add IAA from a 0.55M stock to reach a final concentration of 55 mM.

[0305] Incubate the samples for 30 min at room temperature (20-25 °C) in the dark to alkylate the reduced residues.

[0306] IAA is unstable and light-sensitive. Prepare the solution immediately before use and ensure alkylation is performed in the dark.

[0307] 111. Dilute the Urea concentration to 2M by adding ABC buffer to the samples.

[0308] 112. Add I g Trypsin per sample and incubate samples at room temperature (20-25 °C) for 16 h.

[0309] 113. Add 1 volume of Stop 4 buffer to the digested samples.

[0310] 114. Proceed with C18-based clean-up of the samples (e.g. stage tips) to remove salts.

[0311] 115. Inject the purified peptides into LC-MS / MS (procedure is variable depending on the MS instrument used).

[0312] 116. Search and quantify the identified RBPs using Label-free

[0313] Quantification (LFQ). Multiple search and quantification computational pipelines can be used. We routinely use Maxquant and Perseus software packages for the data analysis steps.

[0314] Protocol steps: Appendix) RTqPCR and RNA-seq analysis of RNase H digestion efficiency and specificity To retrospectively analyse the efficiency and specificity of RNase H mediated target degradation, the 1 / 10thaliquots taken after RNase H treatment can be used to purify RNA and analyse this in a targeted (RT-qPCR) or untargeted (RNA-seq) manner. For this purpose, follow the below steps:

[0315] - Add 110 pl of proteinase K buffer to each taken aliquot.

[0316] - Digest the proteins by adding 10 pl Proteinase K and incubating for 1 h at 25 °C.

[0317] - Perform an additional Phase separation by adding TRIzol™ LS to the Proteinase K reaction. Collect the free RNA (upper phase). Purify the RNA according to the TRIzol™ LS manual.

[0318] - Perform RNA quantification using Qubit.

[0319] PAUSE POINT Store extracted RNA for up to 3 months at -80 °C if needed.

[0320] - Perform RT-qPCR quantification of the target as well as independent control transcripts to assess degradation efficiency and specificity. Alternatively, the purified RNA can be subjected to whole-transcriptome RNA-seq to assess RNAse H-mediated degradation efficiency and specificity in a global unbiased manner.

[0321] Results: Analysis of 18s rRNA protein interactors by TREX

[0322] To assess the performance of TREX, we first applied it to analyse the direct interactome of 18S ribosomal RNA (rRNA), the RNA component of the small (40S) ribosomal subunit. As the interactome of 18S is well characterised, including by other Transcript-targeted capture methods, this enables benchmarking the performance of TREX against existing art.

[0323] A total of 4 experimental (+RNase-H) and 4 control samples (-RNase-H), each from 10 million HCT116 human colorectal carcinoma cell, were processed and subjected to TREX, as described in the detailed step-by-step protocol above. Using RT- qPCR, we first confirmed efficient degradation of the target transcript in the RNase-H treated samples (Figure 5). 18S was indeed found to be efficiently degraded in all +RNase-H but not -RNase-H treated samples. Next, the digested protein extracts were analysed by LC-MS / MS, using a Thermo Orbitrap Q Exactive Plus MS Instrument. We identified a total of 480 proteins that were significantly enriched in the +RNase-H samples (Figure 6). Crucially, most of the core 40S small ribosomal proteins were amongst the identified significant interactors, as well as several other known 40S associated proteins (Figure 6).

[0324] Next, we assessed the types of proteins that were overrepresented in our 18S interactome, by performing Category Enrichment Analysis (Tyanova et al., 2016). As expected, the top enriched categories belonged to small ribosomal protein annotations, but other protein categories involved in 18S biogenesis and maturation, as well as categories of proteins involved in protein targeting to ER, which are known to associate with the 40S ribosomes, were also amongst the top enriched (Figure 7). Analysis of existing proteinprotein interactions using the STRING protein-protein interaction database (https: / / string- db.org / ) revealed that the vast majority of the identified 18S interactors are known to associate with each other, further confirming the specificity of the TREX results (Figure 8).

[0325] Finally, we compared the results of our TREX with a previous 18S RAP-MS analysis (McHugh et al, 2015). Over 50% of the hits from the 18S RAP -MS analysis were also identified by TREX (Figure 9). These included the majority of 40S ribosomal proteins. However, TREX identified significantly more proteins, including several 40S ribosomal proteins that were missed by RAP -MS, many other known 40S associated proteins such as translation initiation factors, small ribosomal subunit biogenesis factors, as well as the ER targeting machinery which is known to associate with 40S (Figure 9). In contrast, 52 proteins were identified by RAP -MS that were not in TREX. These were mainly factors involved in the large ribosomal subunit biogenesis, as well as many other nucleolar proteins (Figure 9). Such discrepancy maybe due to the fact that TREX is region-specific, with the probes only acting on the 18S-specific interactors, whereas RAP-MS could be potentially capturing pre- processed rRNA that contains both large as well as the small ribosomal subunit RNAs.

[0326] In conclusion, this example demonstrates that TREX can accurately and efficiently reveal the interactome of 18S rRNA from as little as 10 million cells per sample. This means the total number cells (both - and + RNase-H) used by TREX was only 80 million. As a comparison, 200 or 800 million cells in total were used in the RAP-MS analysis (McHugh et al. 2015). Results: Analysis of U1 RNA protein interactors by TREX

[0327] Next, we applied TREX to analyse the direct interactome of U1 RNA, the well-characterised RNA component of the U1 small nuclear Ribonucleoprotein (U1 snRNP) complex, a key core splicosome complex which recognises 5’ exon-intron junction sites (Kondo et al. 2015). Similar to 18S, the interactome ofUl is well characterised, including by other targeted capture methods, enabling the benchmarking of TREX against the existing art. Moreover, U1 is much smaller than 18S (<200nt vs. >3000nt), thus providing a contrasting test as to whether TREX can also be successfully applied to small RNAs.

[0328] A total of 5 experimental (+RNase-H) and 5 control samples (-RNase-H), each from 20 million HCT116 human colorectal carcinoma cell, were processed and subjected to TREX, as described in the detailed step-by-step protocol above. Using RT-qPCR, we first confirmed efficient degradation of U1 in the RNase-H treated samples (Figure 10). U1 was indeed found to be efficiently degraded in all +RNase-H but not -RNase-H treated samples. Next, the digested protein extracts were analysed by LC-MS / MS, using a Thermo Orbitrap Q Exactive Plus MS instrument. A total of 30 proteins were found to be significantly enriched in the +RNase-H samples (Figure 11). This number is much lower than the number of interacting proteins identified for 18S (Figure 6), but as U1 is a much smaller and less abundant RNA, this is be expected. Crucially, 8 out of the 10 known U1 snRNP proteins were amongst the identified significant interactors, as well as a number of other known splicing factors which are known to associate with the U1 snRNP complex (Figure 11).

[0329] Next, we analysed the existing protein-protein interactions amongst the identified U1 binding proteins, using the STRING protein-protein interaction database (https: / / string-db.org / ). This analysis revealed one major network of U1 snRNPs and splicing associated proteins that was expected, as well as a complex of Thioredoxin (TXN) and Peroxiredoxin-5 (PRDX5). These two proteins are involved in redox control, suggesting a potentially novel interplay between U1 and redox control components in these cells (Figure 12).

[0330] Finally, we compared the results of our TREX with a previous U1 RAP-MS analysis (McHugh et al, 2015). Over 50% (8 out of 14) of the hits in the U1 RAP-MS analysis were also identified by our TREX (Figure 13). Crucially, these were all known U1 snRNP protein components and splicing factors. However, TREX further identified another splicing factor, as well as another U1 snRNP protein (SNRPE) that were missing from RAP- MS. In addition a few metabolic enzymes, which are known to act as non-conventional RBPs (Castello et al 2012), were also identified as U1 interactors by TREX.

[0331] Results: Analysis of region-specific protein interactions of pre-rRNA by TREX

[0332] Next, we applied TREX to region-specific mapping of RNA-protein interactions. For this purpose, we chose 45S pre-rRNA, the precursor RNA that is synthesized as a >13kb long transcript in the nucleolus, before undergoing extensive processing involving cleavage and removal of several spacer regions, ultimately resulting in three fully processed rRNA transcripts (18S, 5.8S, and 28S) that form the major RNA component of the eukaryotic ribosome (Moraleva et al. 2022). The regions of 45S that undergo cleavage and removal are comprised of 5 ’External Transcribed Region (5’ETS), Internal Transcribed Spacers 1 and 2 (ITS1 and ITS2), and 3’ External Transcribed Region (3’ETS) (Figure 14). Little has been known about the endogenous protein interactions of these specific spacer regions up to now.

[0333] As a proof of concept for the region-specific mapping capability of TREX, we applied it to specifically map the protein interactome of the 5’ETS region, the first spacer region that is processed during ribosome biogenesis. A total of 5 experimental (+RNase-H) and 5 control samples (-RNase-H), each from 20 million HCT116 human colorectal carcinoma cell, were processed and subjected to TREX, as described in the detailed step-by-step protocol above. As before, we first confirmed efficient degradation of 5’ETS in the RNase-H treated samples by RT-qPCR (data not shown). Next, the trypsin digested protein extracts were analysed by LC-MS / MS, using a Thermo Orbitrap Q Exactive Plus MS instrument. A total of 16 proteins were found to be significantly enriched in the +RNase-H samples. Crucially, half of these proteins are known to be involved in ribosome biogenesis, with several of the other unknown factors also having been shown to localise to the nucleolus, which is the exclusive site for 5’ETS (Figure 15).

[0334] We also analysed the existing protein-protein interactions amongst the identified ‘5ETS binding proteins, using the STRING protein-protein interaction database (http s : / / strin g-db . org / ). This analysis revealed one major network of ribosome biogenesis factors (Figure 16). Crucially, although most of these factors are known to be important for ribosome biogenesis, particularly the early stages which involves 5’ETS processing, their direct interaction with the RNA within the 5’ETS region is largely a novel and potentially important finding.

[0335] In conclusion, this example demonstrates that from as little as 20 million cells per sample, TREX can accurately and efficiently reveal the interactome of a small RNA (U1 RNA), revealing most of its known but also some novel interactors.

[0336] References

[0337] Castello, A., Fischer, B., Eichelbaum, K., Horos, R., Beckmann, B.M., Strein, C., Davey, N.E., Humphreys, D.T., Preiss, T., Steinmetz, L.M., et al. (2012). Insights into RNA biology from an atlas of mammalian mRNA-binding proteins. Cell 149, 1393-1406.

[0338] 10.1016 / j .cell.2012.04.031.

[0339] Chu, C., Zhang, Q.C., da Rocha, S.T., Flynn, R.A., Bharadwaj, M., Calabrese, J.M., Magnuson, T., Heard, E., and Chang, H.Y. (2015). Systematic discovery of Xist RNA binding proteins. Cell 161, 404-416. 10.1016 / j .cell.2015.03.025.

[0340] Desideri, F., Cipriano, A., Petrezselyova, S., Buonaiuto, G., Santini, T., Kasparek, P., Prochazka, J., Janson, G., Paiardini, A., Calicchio, A., et al. (2020). Intronic Determinants Coordinate Charme IncRNA Nuclear Activity through the Interaction with MATR3 and PTBP1. Cell reports 33, 108548. https: / / doi.Org / 10.1016 / j.celrep.2020.108548.

[0341] Hafner, M., Katsantoni, M., Koster, T., Marks, J., Mukherjee, J., Staiger, D., Ule, J., and Zavolan, M. (2021). CLIP and complementary methods. Nature Reviews Methods Primers 1, 20. 10.1038 / s43586-021-00018-l. Kondo, Y., Oubridge, C., van Roon, A.-M.M., and Nagai, K. (2015). Crystal structure of human U1 snRNP, a small nuclear ribonucleoprotein particle, reveals the mechanism of 5' splice site recognition. eLife 4, e04986. 10.7554 / eLife.04986.

[0342] McHugh, C.A., Chen, C.K., Chow, A., Surka, C.F., Tran, C., McDonel, P., Pandya-Jones, A., Blanco, M., Burghard, C., Moradian, A., et al. (2015). The Xist IncRNA interacts directly with SHARP to silence transcription through HDAC3. Nature 527, 232-236. 10.1038 / naturel4443.

[0343] Moraleva, A. A., Deryabin, A.S., Rubtsov, Y.P., Rubtsova, M.P., and Dontsova, O.A. (2022). Eukaryotic Ribosome Biogenesis: The 40S Subunit. ActaNaturae 74, 14-30. 10.32607 / actanaturae.l 1540.

[0344] Queiroz, R.M.L., Smith, T., Villanueva, E., Marti-Solano, M., Monti, M., Pizzinga, M., Mirea, D.M., Ramakrishna, M., Harvey, R.F., Dezi, V., et al. (2019). Comprehensive identification of RNA-protein interactions in any organism using orthogonal organic phase separation (OOPS). Nat Biotechnol 37, 169-178. 10.1038 / s41587-018-0001-2.

[0345] Trendel, J., Schwarzl, T., Horos, R., Prakash, A., Bateman, A., Hentze, M.W., and Krijgsveld, J. (2019). The Human RNA-Binding Proteome and Its Dynamics during Translational Arrest. Cell 176, 391-403 e319. 10.1016 / j cell.2018.11.004.

[0346] Tyanova, S., Temu, T., Sinitcyn, P., Carlson, A., Hein, M.Y., Geiger, T., Mann, M., and Cox, J. (2016). The Perseus computational platform for comprehensive analysis of (prote)omics data. Nature Methods 13, 731-740. 10.1038 / nmeth.3901.

[0347] Urdaneta, E.C., Vieira- Vieira, C.H., Hick, T., Wessels, H.-H., Figini, D., Moschall, R., Medenbach, J., Ohler, U., Granneman, S., Selbach, M., and Beckmann, B.M. (2019). Purification of cross-linked RNA-protein complexes by phenol-toluol extraction. Nature Communications 10, 990. 10.1038 / s41467-019-08942-3.

Claims

CLAIMS1. A method comprising:(a) contacting a sample with a cross-linking agent thereby providing RNA- protein adducts;(b) performing phase separation on the sample containing the RNA-protein adducts to produce a distinct phase that comprises RNA-protein adducts but is at least substantially free from free protein and free RNA;(c) isolating the RNA-protein adducts from the distinct phase of part (b);(d) contacting the isolated RNA-protein adducts with a plurality of DNA oligonucleotides that comprise complementary nucleotide sequence to a target RNA molecule of interest to thereby form one or more DNA-RNA hybrids;(e) contacting the sample from part (d) with an enzyme that cleaves DNA- RNA hybrids thereby releasing proteins from the RNA molecule of interest.

2. The method according to claim 1, wherein the method further comprises isolating an RNA-binding protein from the sample, the method comprising:(f) performing a further step of phase separation on the sample to produce a distinct phase that comprises released RNA-binding proteins but is at least substantially free from free RNA; and(g) isolating the released RNA-binding proteins from the distinct phase of part (f).

3. The method according to claim 1 or claim 2, wherein the RNA binding protein binds to any one or more of an mRNA, rRNA, 7 SL RNA, tRNA, snRNA, snoRNA, IncRNA, miRNA, and eRNA.

4. The method according to any one of claims 1 to 3, wherein sample comprises material derived from cells that have been cultured in vitro or derived from humanor animal tissue, optionally wherein the material comprises a lysate of cells or human or animal tissue.

5. The method according to claim 4, wherein the material comprises a lysate of cells or human or animal tissue, and wherein the lysate is further subject to subcellular fractionation.

6. The method according to claim 5, wherein the lysate was prepared under conditions free from active ribonuclease enzymes.

7. The method according to any one of claims 1 to 6, wherein the phase separation of part (b) leads to the production of an organic phase, an interphase and an inorganic phase, and wherein the distinct phase that comprises RNA-protein adducts but is at least substantially free from free protein and free RNA corresponds to the interphase.

8. The method according to any one of claims 1 to 7 wherein the enzyme that cleaves DNA-RNA hybrids is RNase H, or a functionally active fragment or analog thereof.

9. The method according to any one of claims 2 to 8, wherein the further step of phase separation of part (f) leads to the production of an organic phase, an interphase and an inorganic phase, and wherein the distinct phase that comprises released RNA- binding proteins but is at least substantially from free RNA corresponds to the organic phase.

10. The method according to any one of claims 1 to 9, wherein the cross-linking agent induces covalent bonds between protein and RNA molecules that are interacting within the sample, optionally wherein the cross-linking agent is UV radiation.

11. The method according to any one of claims 1 to 10, wherein the phase separation comprises the use of one or more of acidic phenol, guanidine isothiocyanate, and chloroform.

12. A method for identifying RNA-binding proteins, the method comprising the steps of the method for isolating an RNA-binding protein from a sample according to any one of claims 2 to 11, and further comprising subjecting the released RNA- binding protein to a technique for protein identification.

13. The method according to claim 12, wherein the technique for protein identification is a biased technique, optionally wherein the technique comprises the use of an antibody for targeting a protein of interest.

14. The method according to claim 12, wherein the technique for protein identification is an unbiased technique, optionally wherein the technique comprises quantitative mass spectrometry.

15. A method for identifying a binding site of an RNA-binding protein in an RNA transcript sequence, the method comprising the steps of the method for isolating an RNA-binding protein from a sample according to any one of claims 2 to 11 and the steps of any one of claims 12 to 14, further wherein the method comprises mapping the plurality of DNA oligonucleotides that comprise complementary nucleotide sequence to a target RNA molecule of interest to an RNA transcriptome, and inferring the binding site of the released RNA-binding protein in the RNA transcript sequence on the basis of the DNA oligonucleotides that are mapped to the RNA transcriptome.