Methods for identifying disease-associated RNA and RNA binding protein (RBP) interactions for diagnosis and treatment selection

A method for identifying RBP-RNA interactions with single nucleotide resolution addresses the limitations of current approaches by stabilizing and sequencing RBP-RNA complexes, facilitating efficient diagnostic and therapeutic target discovery.

AU2024395871A1Pending Publication Date: 2026-07-16MEMORIAL SLOAN KETTERING CANCER CENT +2

Patent Information

Authority / Receiving Office
AU · AU
Patent Type
Applications
Current Assignee / Owner
MEMORIAL SLOAN KETTERING CANCER CENT
Filing Date
2024-12-04
Publication Date
2026-07-16

AI Technical Summary

Technical Problem

Current methodologies for identifying RNA binding protein (RBP)-RNA interactions are limited by their protein-centric or RNA-centric approaches, which hinder the efficient identification of cancer-associated interactions at the scale required for diagnostic and therapeutic target discovery, leading to a massive underestimation of druggable mutations.

Method used

A method involving cross-linking cells with short wavelength ultraviolet light to stabilize RBP-RNA complexes, followed by protease treatment, chemical labeling, RNA fragmentation, and sequencing to identify RBP-RNA interaction sites with single nucleotide resolution across the transcriptome.

Benefits of technology

Enables high-throughput, unbiased identification of dynamic and disease-associated RBP-RNA interactions, allowing for precise detection of enriched sites and RBP identities, which can inform diagnostic and therapeutic strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present disclosure provides methods for simultaneously identifying dynamic and disease-associated RNA binding protein (RBP)-RNA interaction (PRI) sites at single nucleotide resolution across the entire transcriptome. The methods disclosed herein recapitulates PRI profiles obtained with several distinct conventional PRI methods in a single assay at efficiencies that are improved by orders of magnitude, and permit the identification of the specific RBP bound at a PRI.
Need to check novelty before this filing date? Find Prior Art

Description

CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of and priority to U.S. Provisional Patent Application No. 63 / 606,483, filed December 5, 2023, the entire contents of which are incorporated herein by reference. TECHNICAL FIELD

[0002] The present disclosure relates to methods for simultaneously identifying dynamic and disease-associated RNA binding protein (RBP)-RNA interaction (PRI) sites with single nucleotide resolution across the entire transcriptome. The methods disclosed herein recapitulates PRI profiles obtained with several distinct conventional PRI methods in a single assay at efficiencies that are improved by orders of magnitude, and permit the identification of the specific RBP bound at a PRI. GOVERNMENT SUPPORT

[0003] This invention was made with government support under GM124909 awarded by the National Institutes of Health. The government has certain rights in the invention. BACKGROUND

[0004] The following description of the background of the present technology is provided simply as an aid in understanding the present technology and is not admitted to describe or constitute prior art to the present technology.

[0005] Site-specific interactions between RNA binding proteins (RBPs) and regulatory RNA sequences govern RNA stability, processing, and translation - processes essential to the function of many genes necessary for cancer cell survival and proliferation as well as pathogenic processes in many benign diseases1'4. However, these interactions are poorly understood due to major methodologic limitations in identifying these interactions. Current methodologies for identification of pathogenic interactions are largely restricted to proteincentric methods that identify the interactions of a single RBP with RNAs or conversely RNA centric methods that map multiple RBP binding sites on a single RNA5'12. This presents a massive bottleneck that prohibits the scale necessary to annotate cancer- associated interactions at the pace required for efficient diagnostic and therapeutic target discovery. Since RBP-RNA interactions (PRIs) are druggable in principle, the substantial limitations of current methods likely results in a massive underestimation of the druggable mutations in cancer and treatment targets in benign diseases.

[0006] Accordingly, there is an urgent need for high throughput, unbiased and efficient methods for identifying dynamic and disease-associated PRIs at single nucleotide resolution across the entire transcriptome. SUMMARY OF THE PRESENT TECHNOLOGY

[0007] In one aspect, the present disclosure provides a method for identifying one or more RNA binding protein (RBP)-RNA interaction (PRI) sites in at least one transcript derived from a biological sample including a plurality of cells, the method comprising: (a) cross-linking the plurality of cells present in the biological sample with short wavelength ultraviolet light to stabilize a plurality of RBP-RNA complexes comprising RNA transcripts, wherein each RBP-RNA complex corresponds to a PRI; (b) isolating RNA molecules from the cross-linked plurality of cells, wherein the RNA molecules comprise unbound RNA molecules and the RNA transcripts of the plurality of RBP-RNA complexes; (c) contacting the isolated RNA molecules with a protease under conditions that lyse the RBPs of the plurality of RBP-RNA complexes to yield RNA transcripts bound by residual peptides, wherein each bound residual peptide (i) corresponds to a PRI site and (ii) comprises an alpha amino group at its N terminus and a carboxyl group at its C terminus; (d) labeling the RNA transcripts bound by the residual peptides by coupling the alpha amino group or the carboxyl group of each bound residual peptide with a chemical moiety conjugate comprising an affinity ligand; (e) fragmenting the RNA molecules in the presence of heat and divalent cations to generate RNA fragments having a 5’ end and a 3’ end, wherein the RNA fragments comprise unlabeled RNA fragments and RNA fragments labeled with the chemical moiety conjugate comprising the affinity ligand; (f) ligating a sequencing adapter to the 3 ’ end of the RNA fragments to generate adapter tagged RNA fragments; (g) capturing adapter tagged RNA fragments labeled with the chemical moiety conjugate comprising the affinity ligand using affinity purification; (h) reverse transcribing the adapter tagged labeled RNA fragments to generate a plurality of cDNA molecules, optionally wherein the adapter tagged labeled RNA fragments are attached to a solid surface or are in solution; (i) ligating a sequencing adapter to the 3’ end of the cDNA molecules to generate a plurality of adapter tagged cDNA molecules; (j) amplifying the plurality of adapter tagged cDNA molecules to generate a nucleic acid library of transcripts having RNA-RBP interactions; (k) sequencing and mapping the plurality of adapter tagged cDNA molecules of the nucleic acid library to identify individual transcripts; (1) identifying the frequency of cDNA 3’ ends at each nucleotide across the individual transcripts; and (m) detecting enriched RBP bound sites in the individual transcripts when the frequency of cDNA 3’ ends at specific nucleotides within the individual transcripts satisfies a predetermined threshold. In some embodiments, the plurality of cDNA molecules are reverse transcribed using a truncating reverse transcriptase (RT) or a read-through reverse transcriptase (RT). A truncating RT enzyme that truncates at the site of the PRI (i.e., the cross-link site) or a read-through RT enzyme that generates SNVs or indels at the site of the PRI (i.e., the cross-link site). Additionally or alternatively, in some embodiments, the methods disclosed herein further comprise identifying in the individual transcripts (i) short indels (e.g., 1-3 bps) that result from PRIs isolated by affinity purification, (ii) single nucleotide variants (SNVs) that result from PRIs isolated by affinity purification, or (iii) reverse transcriptase (RT) stop sites that result from PRIs isolated by affinity purification.

[0008] The biological sample may be derived from breast tissue, renal tissue, uterine cervical tissue, endometrium tissue, head or neck tissue, gallbladder tissue, parotid tissue, prostate tissue, brain tissue, pituitary gland tissue, kidney tissue, muscle tissue, esophageal tissue, stomach tissue, small intestine tissue, colon tissue, liver tissue, spleen tissue, pancreatic tissue, thyroid tissue, heart tissue, lung tissue, bladder tissue, adipose tissue, lymph node tissue, uterine tissue, ovarian tissue, adrenal tissue, testis tissue, tonsils, thymus, blood, hair, buccal, skin, serum, plasma, CSF, semen, prostate fluid, seminal fluid, urine, feces, sweat, saliva, sputum, mucus, bone marrow, lymph, or tears. In certain embodiments, the biological sample is a fresh or frozen sample.

[0009] In some embodiments, the cross-linking is performed with a UV box at a dose of about 50-600 mJ / cm2 for about 30s-10 minutes. In certain embodiments, the cross-linking is performed with a UV box at a dose of about 50 mJ / cm2, about 100 mJ / cm2, about 150 mJ / cm2, about 200 mJ / cm2, about 250 mJ / cm2, about 300 mJ / cm2, about 350 mJ / cm2, about 400 mJ / cm2, about 450 mJ / cm2, about 500 mJ / cm2, about 550 mJ / cm2, or about 600 mJ / cm2 for about 30s, 45s, 1 minute, 1.5 minutes, 2 minutes, 2.5 minutes, 3 minutes, 3.5 minutes, 4 minutes, 4.5 minutes, 5 minutes, 5.5 minutes, 6 minutes, 6.5 minutes, 7 minutes, 7.5 minutes, 8 minutes, 8.5 minutes, 9 minutes, 9.5 minutes, or about 10 minutes. In some embodiments, the cross-linking is performed with a UV box at a dose of 300-500 mJ / cm2 for about 3 minutes or 100-400 mJ / cm2 for about 10 minutes. In other embodiments, the cross-linking is performed with a UV laser for about 10s-20s.

[0010] Additionally or alternatively, in some embodiments, the isolated RNA molecules are treated with the protease at about 25°C-65°C for about 5 minutes-60 minutes. In certain embodiments, the isolated RNA molecules are treated with the protease at about 25°C, about 27.5°C, about 30°C, about 32.5°C, about 35°C, about 37.5°C, about 40°C, about 42.5°C, 45°C, about 47.5°C, about 50°C, about 52.5°C, at about 55°C, about 57.5°C, about 60°C, about 62.5°C, or about 65°C for about 5 minutes, about 7.5 minutes, about 10 minutes, about 12.5 minutes, about 15 minutes, about 17.5 minutes, about 20 minutes, about 22.5 minutes, about 25 minutes, about 27.5 minutes, about 30 minutes, about 32.5 minutes, about 35 minutes, about 37.5 minutes, about 40 minutes, about 42.5 minutes, about 45 minutes, about 47.5 minutes, about 50 minutes, about 52.5 minutes, about 55 minutes, about 57.5 minutes, or about 60 minutes. In some embodiments, the isolated RNA molecules are treated with the protease at 50°C for about 30 minutes. In certain embodiments, the protease is proteinase K. Additionally or alternatively, in some embodiments, the length of each bound residual peptide is no more than 10 amino acids in length, no more than 9 amino acids in length, no more than 8 amino acids in length, no more than 7 amino acids in length, no more than 6 amino acids in length, no more than 5 amino acids in length, no more than 4 amino acids in length, no more than 3 amino acids in length, no more than 2 amino acids in length, or is a single amino acid.

[0011] Additionally or alternatively, in some embodiments, the methods of the present technology further comprise removing genomic DNA after step (c) to obtained purified RNA molecules. In certain embodiments, the methods disclosed herein further comprise contacting the purified RNA molecules with a protease at 50°C for about 30 minutes prior to step (d), optionally wherein the protease is proteinase K.

[0012] In any of the preceding embodiments, the methods of the present technology further comprise enriching the isolated or purified RNA molecules with poly-dT oligonucleotides or targeted bait capture reagents after step (c), but prior to step (d). In other embodiments, the methods of the present technology further comprise enriching the adapter tagged cDNA molecules with targeted bait capture reagents after step (j), but prior to step (k). In certain embodiments, the targeted bait capture reagents hybridize to one or more disease associated genes, such as genes associated with cancer, neurodegenerative disease, autoimmune disease, metabolic disease etc. Examples of disease associated genes include, but are not limited to, AKT1, ALK, APC, AR, ARAF, ARID 1 A, ARID2, ATM, B2M, BCL2, BCOR, BRAF, BRCA1, BRCA2, CARD11, CBFB, CCND1, CDH1, CDK4, CDKN2A, CIC, CREBBP, CTCF, CTNNB1, DICERI, DIS3, DNMT3A, EGFR, EIF1AX, EP300, ERBB2, ERBB3, ERCC2, ESRI, ETV6, EZH2, FBXW7, FGFR1, FGFR2, FGFR3, FGFR4, FLT3, F0XA1, FOXL2, FOXO1, FUBP1, GATA3, GNA11, GNAQ, GNAS, H3F3A, HIST1H3B, HRAS, IDH1, IDH2, IKZF1, INPPL1, JAKI, KDM6A, KEAP1, KIT, KNSTRN, KRAS, MAP2K1, MAPK1, MAX, MED12, MET, MLH1, MSH2, MSH3, MSH6, MTOR, MYC, MYCN, MYD88, MYODI, NF1, NFE2L2, NOTCH1, NRAS, NTRK1, NTRK2, NTRK3, NUP93, PAK7, PDGFRA, PIK3CA, PIK3CB, PIK3R1, PIK3R2, PMS2, POLE, PPP2R1 A, PPP6C, PRKCI, PTCHI, PTEN, PTPN11, RAC1, RAFI, RBI, RET, RHOA, RIT1, R0S1, RRAS2, RXRA, SETD2, SF3B1, SMAD3, SMAD4, SMARCA4, SMARCB1, S0S1, SPOP, STAT3, STK11, STK19, TCF7L2, TGFBR1, TGFBR2, TP53, TP63, TSC1, TSC2, U2AF1, VHL, XPO1, and TERT.

[0013] In any of the above embodiments, the methods of the present technology further comprise depleting ribosomal RNA from the isolated or purified RNA molecules after step (c), but prior to step (d).

[0014] In any and all embodiments of the methods disclosed herein, labeling the RNA transcripts bound by the residual peptides comprises heat denaturing the isolated or purified RNA molecules, and incubating the isolated or purified RNA molecules with the chemical moiety conjugate comprising the affinity ligand in a reaction buffer supplemented with about 5%-70% DMSO (e.g., 50% DMSO). Examples of suitable reaction buffers include, but are not limited to, phosphate buffers (e.g., PBS buffer), carbonate buffers (e.g., carbonate-bicarbonate buffer), HEPES, borate, 2-[morpholino]ethanesulfonic acid (MES), or any buffer that does not contain free amine or free carboxylate groups. Additionally or alternatively, in some embodiments, heat denaturing comprises heating the isolated or purified RNA molecules at about 50°C-100°C for 1-5 minutes. In some embodiments, heat denaturing comprises heating the isolated or purified RNA molecules at 70°C for 2 min.

[0015] In some embodiments of the methods disclosed herein, the chemical moiety conjugate is coupled to the alpha amino group of each bound residual peptide. In certain embodiments, the chemical moiety conjugate comprises a succinimidyl ester, a carboxylic ester, a tetrafluorophenyl ester, a sulfodichlorophenol ester, a carbonyl azide, an aldehyde, or an isothiocyanate. In other embodiments of the methods disclosed herein, the chemical moiety conjugate is coupled to the carboxyl group of each bound residual peptide. In some embodiments, the chemical moiety conjugate comprises a primary amine, a carboiimide, or an isocyanate. Additionally or alternatively, in some embodiments, the chemical moiety conjugate comprising the affinity ligand is added to a final concentration of about 0.1 mM to about 10 mM and incubated for about 3 minutes to about 5 hours.

[0016] In any and all embodiments of the methods disclosed herein, the affinity ligand comprises biotin, a biotin derivative, digoxin, dinitrophenyl, a click chemistry reagent, sugars or peptides. In some embodiments, the click chemistry reagent comprises an alkyne group, an azide group, a dibenzocyclooctyne group (DBCO), or a bicyclononyne (BCN) group.

[0017] Additionally or alternatively, in some embodiments, fragmenting the RNA molecules comprises heating the RNA molecules at 90°C-95°C for about 3-12 minutes in the presence of about 10mM-30mM divalent cations (e.g., Mg2+, Zn2+). In certain embodiments, fragmenting the RNA molecules comprises heating the RNA molecules at 90°C, 91°C, 92°C, 93°C, 94°C or 95°C for about 3 minutes, about 4 minutes, about 5 minutes, about 6 minutes, about 7 minutes, about 8 minutes, about 9 minutes, about 10 minutes, about 11 minutes, or about 12 minutes in the presence of about lOmM, about 15mM, about 20mM, about 25mM, or about 30mM divalent cations (e.g., Mg2+, Zn2+). The RNA fragments may be about 40-90 nucleotides in length or about 80-250 nucleotides in length.

[0018] In any and all embodiments of the methods disclosed herein, the biological sample is obtained from a subject diagnosed with a disease. In some embodiments, the disease is cancer, autoimmune disease, metabolic disease or neurodegenerative disease. Examples of cancer include, but are not limited to, adrenal cancers, bladder cancers, blood cancers, bone cancers, brain cancers, breast cancers, carcinoma, cervical cancers, colon cancers, colorectal cancers, corpus uterine cancers, ear, nose and throat (ENT) cancers, endometrial cancers, esophageal cancers, gastrointestinal cancers, head and neck cancers, Hodgkin's disease, intestinal cancers, kidney cancers, larynx cancers, leukemias, liver cancers, lymph node cancers, lymphomas, Diffuse large B-cell lymphoma (DLBCL), lung cancers, melanomas, mesothelioma, myelomas, nasopharynx cancers, neuroblastomas, nonHodgkin's lymphoma, oral cancers, ovarian cancers, pancreatic cancers, penile cancers, pharynx cancers, prostate cancers, rectal cancers, sarcoma, seminomas, skin cancers, stomach cancers, teratomas, testicular cancers, thyroid cancers, uterine cancers, vaginal cancers, vascular tumors, and metastases thereof.

[0019] Examples of neurodegenerative disease include, but are not limited to, age-associated memory impairment (AAMI), mild cognitive impairment (MCI), Alzheimer's disease, Down's syndrome, dementia pugilistica, cognitive dysfunction syndrome, multiple system atrophy, inclusion body myositosis, hereditary cerebral hemorrhage with amyloidosis of the Dutch type, Nieman-Pick disease type C, cerebral P-amyloid angiopathy, dementia associated with cortical basal degeneration, the amyloidosis, Creutzfeldt-Jakob disease, Gerstmann-Straussler syndrome, kuru, scrapie, Huntington’s disease, Parkinson’s disease, ataxia, Motor neuron disease, Progressive supranuclear palsy, and Amyotrophic lateral sclerosis (ALS).

[0020] Examples of autoimmune disease include, but are not limited to, vasculitis, e.g., Anti- neutrophil cytoplasm antibodies (ANCA), ANCA-associated vasculitis (AAV) or giant cell arteritis (GCA) vasculitis, Sjogren's syndrome, inflammatory bowel disease (IBD), Pemphigus vulgaris, lupus nephritis, psoriasis, thyroiditis, Type I Diabetes, Idiopathic thrombocytopenic purpura (ITP), Ankylosing spondylitis, Multiple sclerosis, systemic lupus erythematosus (SLE), rheumatoid arthritis, Crohn's disease, Myasthenia Gravis, neuromyelitis optica (NMO), IgG4-related disease, systemic sclerosis, insulindependent diabetes mellitus (IDDM), akylosing spondylitis, atopic dermatitis, uveitis, and Graft-versus- host disease (GVHD).

[0021] Examples of metabolic disease include, but are not limited to, Familial hypercholesterolemia, Gaucher disease, Hunter syndrome, Krabbe disease, Maple syrup urine disease, Metachromatic leukodystrophy, Mitochondrial encephalopathy lactic acidosis stroke-like episodes (MELAS), Niemann-Pick, Phenylketonuria (PKU), Porphyria, Tay-Sachs disease and Wilson's disease.

[0022] In any and all embodiments of the methods disclosed herein, one or more of the enriched RBP bound sites detected in the individual transcripts are allele-specific and / or associated with a disease (e.g., cancer, neurodegenerative disease, autoimmune disease, metabolic disease etc^.

[0023] In any of the foregoing embodiments, the methods of the present technology further comprise determining the identity of a RBP bound to one or more of the enriched RBP bound sites detected in the individual transcripts. In some embodiments, the identity of a RBP bound to one or more of the detected enriched RBP bound sites in the individual transcripts is determined by identifying RBP-specific crosslinking patterns within a sequence motif in the individual transcripts (e.g., via computational analysis). In some embodiments, the crosslinking patterns of RBPs are identified using eCLIP defined consensus motifs, mCross crosslinking patterns, positionally enriched k-mer analysis (PEKA), or sequence motifs defined by RBNS.

[0024] In one aspect, the present disclosure provides a method for evaluating the efficacy of a drug in inhibiting interactions between a specific RNA binding protein (RBP) and RNA transcripts comprising (a) contacting a first cell population with at least one inhibitory oligonucleotide that specifically inhibits expression of a RBP, wherein the first cell population is obtained from a biological sample; (b) contacting a second cell population with a test drug, wherein the second cell population is obtained from the biological sample and is not treated with the at least one inhibitory oligonucleotide; (c) performing the ARORA methods described herein to identify one or more RBP-RNA interaction (PRI) sites in at least one transcript derived from the first cell population, the second cell population and a control cell population, wherein the control cell population is obtained from the biological sample and is not treated with the at least one inhibitory oligonucleotide and the test drug; (d) defining RBP-specific PRIs by identifying crosslink patterns within one or more sequence motifs in the individual transcripts that (i) are present in the control cell population and (ii) absent in the first cell population; and determining that the test drug is effective when the RBP-specific PRIs identified in step (d) are absent in the second cell population. In another aspect, the present disclosure provides a method for evaluating the specificity of a drug in inhibiting interactions between a specific RNA binding protein (RBP) and RNA transcripts comprising (a) contacting a first cell population with at least one inhibitory oligonucleotide that specifically inhibits expression of a RBP, wherein the first cell population is obtained from a biological sample; (b) contacting a second cell population with a test drug, wherein the second cell population is obtained from the biological sample and is not treated with the at least one inhibitory oligonucleotide; (c) performing the ARORA methods described herein to identify one or more RBP-RNA interaction (PRI) sites in at least one transcript derived from the first cell population, the second cell population and a control cell population, wherein the control cell population is obtained from the biological sample and is not treated with the at least one inhibitory oligonucleotide and the test drug; (d) identifying the non-specific effects of the test drug by identifying crosslink patterns within one or more sequence motifs in the individual transcripts that (a) are present in the control cell population and the first cell population and (b) absent in the second cell population. The at least one inhibitory oligonucleotide may be a siRNA, a shRNA, an antisense oligonucleotide, or a sgRNA. In some embodiments, the test drug is a small molecule, a nucleic acid or a peptide. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] FIGs. 1A-H: A universal orthogonal labeling strategy for proteiswmdek acid interactions. FIG. 1A: Schematic for universal labeling of amino acid-nucleic acid crosslinks. PRIs stabilized by UVC crosslinking are subjected to extensive proteolysis, generating a free alpha amino at the N terminus and free carboxyl group at the C terminus regardless of peptide sequence. These are orthogonal chemical moi eties in nucleic acid sequences that can be selective labeled with existing chemistries, including N-hydroxysuccinimide labeling of alpha amino groups. FIG. IB: Dot blot of Streptavidin-1R800 binding to NHSAnotinylated RNA isolated from UVC crosslinked cells, RNA crosslinked with UVC er-ww in the absence of protein, and RNA from noncrosslinked cells. FIG. IC: Denaturing agarose gel electrophoresis of total RNA after NHS-biotinylation. Right, Streptavidin-IR800 binding to NHS-biotinylated total RNA from UVC crosslinked cells. FIG. ID: Quantitative fluorescent measurements of NHS-Alexa488 labeling of RNA isolated from cells after crosslinking with the indicated dose of UVC. FIG. IE: qPCR of enrichment of selected transcripts by NHS biotinylation and streptavidin pulldown with increasing doses of UVC. Lack of recovery in no RT controls (-RT) confirms that RNA, not DNA, is recovered. FIG. IF: Left. Infrared imaging of primer extension assay preformed on a lOObp oligonucleotide with a simulated crosslink (XL) at position 70 with indicated RT enzyme. Specific RT enzymes used are listed above. (FL, full length sequence, Fl.-del, full length sequence with deletion; XL, crosslink-truncated sequence). Right: quantification of the relative abundance of FL, FL-del, and XL fragments obtained with each RT enzyme. FIG. 1G: Rates of deletions (left) and single nucleotide substitutions (right) under the four conditions of ARORA (NCLNN, non-crossiinked RNA input; x400 IN, crosslinked RNA input; NCL PD, non-crosslinked RNA after pulldown with streptavidin beads, x400_PD, crosslinked RNA after pulldown with streptavidin beads). FIG. 1H: Schematic for identification of PRIs in ARORA by identification of both significantly enriched crosslink induced RI' stop sites and crosslinked induced mutation sites. FIG. II: Secondary structure map with PRIs identified by RNF-Alak (left) and ARORA (right). lied annotations indicate significant RI' stop sites, blue annotations indicate significant crosslink induced mutations at residues with significant RT stop sites.

[0026] FIGs. 2A-2D: ARORA identifies dynamic, stress inducible PRIs and variability in ribosome architecture associated with distinct biologic states. FIG. 2A: Diagram of helix 13 and 14 of the small ribosome subunit (18S) and PRIs identified by cryo-EM (left), ARORA (middle), and CLIP-seq (right). ARORA performed in normal, unstress conditions (untreated control) and after ribosomal stress (cycloheximide or anisomycin). Black annotations indicate non-variable PRIs the correlate to RP binding sites on 18S. Red annotations indicated dynamic, stress induced PRIs on h!4, which map precisely to the ZAK a bound site identified by CLIP-seq (red line, right panel). FIG. 2B: Correlation matrix for PRIs in a '- lOOOnt segment of the 18S rRNA performed in 14 breast cancer cell lines. Red ::: 1, White :::: 0, Blue ::: -1). Cohesive domains of high PRI correlation (red nodes) are noted and are delineated by dense regions of rRNA nucleotides bound by RPs (grey lines). Blue: Variation in PRIs between 14 breast cancer cell lines. Below: annotations of 18S helices of interest and RBPs that interact with those sites. FIG. 2C: ARORA PRI signal within the ZAKa binding site on hl-4 of 18S (inset: secondary structure of region identify the nucleotides graphed). Average PRI signal at each nucleotide in this region in the 7 BRCA cell lines surveyed by the CCLE and ARORA, five of which have high ZAKa PRIs and 2 with low ZAKa PRIs. *P<0.01. FIG. 2D: Omacetaxine Area under the curve (AUG) drug sensitivity for each of the 7 cell lines. P<0.02.

[0027] FIGs. 3A-3F: ARORA identifies allele-specific PRIs in regulatory RNA. FIG. 3A: Schematic for interpretation of allele-specific sequencing results in ARORA and how the position of favored crosslinking sites relative to the single nucleotide variant influence the interpretation of allele-specific data when using RT enzymes that preferentially truncate at crosslinked sites. In circumstances where crosslinking preferentially occurs 5’ to the crosslink site, the RT enzyme can still capture the identify of variant nucleotide before the RT is truncated by the crosslink site. However, if crosslinking is favored 3’ to the variant nucleotide, crosslink-truncated reads will not be able capture the identity of the variant nucleotide. Use of read-through RT enzymes can overcome this limitation by reading through the crosslink and capturing the identity of variant nucleotide. FIG. 3B: Allele frequency of normal variant GAPDH rs 1065691 and HepG2 specific MYC 5’UTR variant C184T across four conditions of ARORA performed with truncating RT enzyme Maxima H-. Crosslinked and pulldown conditions significantly depleted C184T, suggesting allele specific binding, but have no impact on GAPDH rsl065691. FKL 3C: ARORA with qPCR of the region immediately 3’ to Cl 84 in three cell lines. Significant enrichment is only noted in HepG2 cells. FIG. 3D: ARORA performed with AffinityScript enriches a region centered on C184T corresponding to a DDX3X eCLIP site in the 5’UTR of MYC. FIGs. 3E-3F: Allele fractions (FIG. 3E) in input and pulldown samples and allele enrichment (FIG. 3F) for ARORA performed with AffinityScript in HepG2 cells. C184T is significantly enriched by ARORA with AffinityScript, indicating C184T has a gain of PRI compared to wildtype.

[0028] FIGs. 4A-4E: Technical points of ARORA sequencing analyses. FIG. 4A: Dot blot of Streptavidin-IR800 binding to EDC / Amino-biotin labeled RNA. isolated from UVC crosslinked cells, RNA crosslinked with UVC ex-vivo in the absence of protein, and RNA from noncrosslinked cells. FIG. 4B: Schematic of the workflow for ARORA sequencing. Protein-crosslinked sites are biotinylated (pink circle) and RNA fragmented. After streptavidin enrichment, reverse transcriptase is performed priming on a 3’ linker with a unique barcode to identify the sample. NGS sequencing is peformed and the sites of RT stops, or crosslink-induced mutations, are used to identify PRIs. FIG. 4C: Nucleotide identify within + / -10 nucleotides of the site of significantly enriched marks: RT stops, Deletions and substitutions. U>OG>A is enriched at positions + / -1 of significantly enriched mark. FIG. 4D: Median relative ARORA signal (reactivity) across all nucleotides of specific identity in NCLPD and x400_PD samples, indicating a slight nonspecific reactivity with A residues. FIG. 4E: Normalization of x400_PD samples to NCL_PD samples corrects this modest artifact and allows comparison between nucleotides. Of note, significantly enriched nucleotides at PRIs typically have reactivity scores >10X to 1000X on this scale, indicating the 1,4x bias towards A in NCL PD and x400_PD samples and the 1.6X bias towards U in x400_PD samples is minor compared to the scale of reactivity towards significant PRIs.

[0029] FIGs. 5A-5C: Optimization of RNA fragmentation in ARORA labeled RNA. FIG. 5A: Denaturing PAGE of RNA ARORA labeled RNAs after increasing times of fragmentation with divalent cation Mg2+ at 94C. FIG. 5B: Bioanalyzer of RNA fragmented for 4 minutes (long fragments) or 9 minutes (short fragments) used in library construction, grey bar indicates fragments in the range of 50-150 nucleotides. FIG. 5C: Input (above) and pulldown (below) dsDNA ARORA libraries. Grey bar indicates substantial truncation in the cDNAs in the pulldown libraries compared to libraries generated with the input material, as expected.

[0030] FIGs. 6A-6C: Effects of RT enzyme and RNA fragment size on RT stops identified by ARORA. FIG. 6A: Correlation matrix for RT stop counts at mapping to 45 S (>95% of reads) in ARORA libraries generate using total RNA. 3 conditions were examined (AffinityScript RT using long RNA fragments “4min”, Maxima H- RT using long RNA fragments “4min”, and Maxima H- RT using short RNA fragments “9min”). NCL_IN, x400_IN, NCLPD, and x400_PD libraries are color coded. Bottom right: four comparisons made between the three conditions. FIG. 6B: Correlation of biologic replicates of x400_PD libraries in each condition is very high and does not vary significantly with RT or fragment size (Cl, ie, reproducibility). FIG. 6C: Correlation of RT stop counts between x400_PD and negative controls significantly decrease in libraries generated by Maxima H- compared to AffinityScript, and is further decreased by using short fragments rather than long fragments.

[0031] FIGs. 7A-7C: Detailed data on RT stop counts mapping to RMRP across conditions and robustness of signal and modest sequencing depth. FIG. 7A: Normalized RT stop counts across mapping to RMRP in four libraries of ARORA. Vertical grey lines indicate the location of statistically significantly enriched RT stops on RMRP corresponding to FIG. II. FIG. 7B: Raw RT stop counts at RMRP in two biologic replicates of each of the 4 ARORA libraries. FIG. 7C: x400_PD RT stop counts in libraries sequenced to depths that differed by nearly 2 orders of magnitude (sequencing depth in mapped reads per library listed on right).

[0032] FIGs. 8A-8E: Low-depth transcriptome-wide ARORA recapitulates PRIs identified by reference methods. FIG. 8A: Schematic for data interpretation. X400 PD samples generated with truncating RT enzymes in ARORA generate peaks in read pileups (Read Depth Peak, grey), but the PRI is identified by the narrow window of highly enriched RT stops (red) at the 5’end of the Read Depth Peak. FIG. 8B: Read Depth Peaks significantly enriched by x400_PD compared to three negative control libraries are readily apparent even in moderate / low abundance transcripts like G6PC. Some RT Stop Peaks can be resolved at this low depth if the RT stops occur with a short sequence space (grey bars), but more dispersed RT stops are difficult to resolve at low sequencing depth. FIG. 8C: Genome browser view of the full length of the abundant housekeeping gene ACTB reveals that most enrichment occurs in the 3’UTR (expanded above). Grey bars indicate the sequences that were enriched >16X in by 1 or more of 8 RBPs in eCLIP that significantly enriched the ACTB 3’UTR in HepG2 cells (equivalent to a Read Depth Peak in ARORA). Multiple significantly RT stop peaks enriched in x400_PD samples are appreciated within these eCLIP enriched regions and at the 5’ end of each of them. FIG. 8D: RT Stop counts in the highly abundant transcript RPL30, encoding ribosomal protein L30. Highly enriched PRIs exist at through the CDS and UTRs of this transcript. Notably, high enrichment is noted in the 5’TOP domain of the 5’UTR which is an essential regulator of RP translation that regulates ribosome biogenesis and RP stochiometry. Enrichment at 5’TOP domains are noted in nearly all transcripts in encoding RPs and a few other regulators of translation, but few other mRNAs. FIG. 8E: RT Stop Peaks and Read Depth Peaks in the 3’UTR of TFRC, where interactions with RBPs have been well characterized by two existing methods, eCLIP and SHAPE Footprinting. Above, identification of significantly enriched RT stop peaks near the 5’ end to significantly enriched eCLIP sequences. Peaks were also identified at the precise location of iron response elements (IRE) denoted by SHAPE Footprinting of Iron Response Protein.

[0033] FIG. 9: RT Stop counts in relation to RP binding sites in the small ribosomal subunit. ARORA performed using total RNA identified significantly enriched RT stop peaks that highly correlated with nucleotides that interact with RPs, as defined by cryo-EM (P<10‘6). Graph depicting the nucleotides of 18S that are in contact with R-Proteins defined by Cryo-EM in a region of the small subunit not visualized in FIG. 2.

[0034] FIG. 10: ARORA in frozen tissue samples recapitulates results from cell culture. Correlation matrix of biological replicate of ARORA obtained from ARORA x400_IN (input) libraries and x400_PD (pulldown) libraries from HepG2 cells UVC crosslinked in culture (original method) and modified protocol that UVC crosslinked material cells recovered from a frozen pellet of HepG2. These results demonstrate concordance of ARORA in samples performed in live and frozen cells, indicating that this method has excellent potential for performing ARORA in frozen clinical tumor samples.

[0035] FIGs. 11A-11C: ARORA identifies RBP interaction on RNA with single nucleotide resolution. FIG. 11 A: A model of a RNA interacting with 3 unique RBPs at distinct sites on the transcript. FIG. 11B: The ARORA workflow and data output: (1) Cells are crosslinked and sites of RBP-RNA interactions are covalently labeled and protein removed. (2) Covalently modified RNA is fragmented reproducibly with heat and divalent cations, followed by tagging RBP interaction sites with an affinity ligand for purification, (3) Labeled fragments are isolated by affinity purification, (4) a sequencing linker is ligated to the 3’ end of purified RNAs followed by first strand synthesis and library amplification, (5) strand specific sequencing and mapping of the first strand orientation delineates sites of RBP binding on the RNA, (6) computational annotation of multiple significantly enriched RBP bound sites on individual RNAs. FIG. 11C: eCLIP workflow: (1) cellular lysates with crosslinked RBP-RNA complexes are treated with RNases to fragment RNA, (2) RNA crosslinked to individual specific RBPs are immunoprecipitated with antibodies, (3) a sequencing linker is ligated to the 3’ end of purified RNAs followed by first strand synthesis and library amplification, (4) strand specific sequencing and mapping of the first strand orientation delineates sites at which an individual RBP bindings RNAs, (5) computational annotation of interactions between a single RBP and RNAs.

[0036] FIGs. 12A-12C: ARORA defines the regional landscape of RBP interactions with mRNAs. FIGs. 12A-12C, Genome browser tracks of the reverse transcription truncation site (RBP-mRNA crosslink site) of the four conditions sequence by ARORA to define RBP-RNA interactions on transcripts for ribosomal protein RPL30 (FIG. 12A), GAPDH (FIG. 12B), and ACTB (FIG. 12C) in HepG2 cells. “Pulldown_x400” indicates samples in which RBP-mRNAs interactions were crosslinked, biochemical modified, and enriched by affinity purification before sequencing. “Pulldown NCL” indicates samples that underwent biochemical modification and enrichment without prior crosslinking (negative control). “Input_x400” and “Input NCL” represent libraries generated from the RNA inputs for these two conditions, but without enrichment by pulldown. Across all tracks, highly specific, unique sites in the “Pulldown_x400” samples represent sites of highly significant enrichment for RBP-mRNA interactions relative to the “Pulldown NCL” and Input samples. x400 and NCL RNA samples are separately indexed and then pooled together before enrichment, thus, the massive enrichment for RBP-mRNA interactions in “Pulldown_x400” samples compared to “Pulldown NCL” samples is a direct measure of the relative retrieval of mRNAs from each condition. RBP interactions with RPL30 are enriched in the 5’UTR especially at the 5’TOP motif that controls translation of ribosomal proteins and additionally throughout the CDS. RBP interactions with GAPDH mRNA are most prominent in exons 3 and 5 of the CDS. RBP interactions with ACTB are most enriched in the 3’UTR. Importantly, for all three transcripts, RBP-mRNA interactions are observed in the 5’UTR, CDS, and 3’UTR. These tracks as scaled demonstrate that certain interactions within certain regions of the mRNA are more enriched relative to other interactions. See also FIG. 15A to view track details scaled to view RBP-mRNA interactions within the 3’UTR of GAPDH mRNA.

[0037] FIGs. 13A-13F: Analyses of crosslink patterns within sequence motifs determines the identity of the specific RBP bound at an RBP-mRNA interaction site. FIGs. 13A-13B: Comparison of ARORA crosslink sites to eCLIP’s annotation of crosslink sites for YBX3 and PUM2 in GAPDH exon 5 (FIG. 13A) and the 3’UTR of T0MM6 (FIG. 13B). Grey box indicates the consensus sequence motif for each of the RBPs. FIGs. 13C-13D, The mCross results for specific nucleotide positions that are crosslinked in YBX3 eCLIP (FIG. 13C) and PUM2 eCLIP (FIG. 13D). Analogous to a fingerprint, YBX3 stereotypically crosslinks at nucleotide 2 of the GUCANC sequence motif while PUM2 crosslinks at position 0 more that position 1 of the PUM2 motif UGUANANA. FIGs. 13E-13F: Heatmap of the crosslinking patterns within the top 200 sites with RBP-RNA interactions identified by ARORA that contain a sequence motif for YBX3 (FIG. 13E) or PUM2 (FIG. 13F), scale bar indicates the crosslinks at each nucleotide as a fraction of all crosslinks within the 11 nucleotide window. A subset of these sites of RBP interactions contain crosslink patterns that are consistent with YBX3 or PUM2. The presence of a crosslinking pattern consistent with the RBP’s stereotypic crosslinking pattern was associated with a -75% probability of an annotated eCLIP peak for the RBP at that precise site, whereas -25% of sites without a crosslinking pattern consistent with the RBP had an eCLIP peak, but usually in a background of many other RBPs with peaks at the site. Therefore, the presence of a crosslinking pattern consistent with a specific RBP within its consensus sequence motif are strongly predictive of whether the RBP is bound at that specific site (P<10‘5). Crucially, RBP-consistent crosslink patterns are only observed in the Pulldown_x400 conditions, but not in the negative control or input conditions.

[0038] FIGs. 14A-14D: ARORA identifies the relative enrichment of specific RBPs on an mRNA, identifying proteins that are critical regulators of the transcript. FIG. 14A: YBX3-consistent RBP interactions with the GAPDH mRNA are by far the most enriched RBP-mRNA interaction, >10X enriched compared to other RBPs. Predominance of a YBX3-consistent RBP interaction is a feature that is common in mRNAs for most proteins involved in glycolysis and a subset of mRNAs for proteins involved in translation, consistent with recent reports that YBX proteins are essential for regulating mRNA stability and translation of glycolysis proteins and some components of the translation apparatus. FIG. 14B: The mRNA for cell cycle regulator cyclin B2 (CCNB2) has a single dominant RBP mRNA interaction in the 3’UTR of the mRNA, which analyses of crosslinking patterns in sequence motifs identifies at PUM2, consistent with eCLIP. FIG. 14C: Interactions at two YBX3-consistent sites and two PUM2-consi stent sites on the RPL14 mRNA show similar relative stoichiometry compared to eCLIP results for each RBP (FIG. 14D), but unlike eCLIP ARORA is able directly compare the relative proportion of interactions of the two RBPs on the same transcript.

[0039] FIGs. 15A-15C: Differential RBP-mRNA interactions between cell lines occur at sites of single nucleotide variants. FIG. 15A: A site of a RBP-mRNA interaction in the GAPDH 3’UTR that differs between HepG2 cells and K562 cells (grey box). FIG. 15B: A RBP-mRNA interaction in the 3’UTR of proto-oncogene CTNNB1 is unique to HepG2 cells and is associated with a single nucleotide variant at the site (*). FIG. 15C: Unique RBP interactions with the 5’UTR of MYC in HepG2 relative to K562 cells is also associate with a 5’UTR single nucleotide variant in the cell line (*). The single nucleotide variant allele in the HepG2 5’UTR was highly enriched relative to the wildtype allele, indicating that this allele likely gains RBP interactions in the MYC 5’UTR.

[0040] FIG. 16A: PUM2 Crosslinking Patterns Identified by ARORA on IncRNA NORAD. FIG. 16B: YBX3 Crosslinking Patterns Identified by ARORA on GAPDH. FIG. 16C: TARDBP Crosslinking Patterns Identified by ARORA on SMAP2. DETAILED DESCRIPTION

[0041] It is to be appreciated that certain aspects, modes, embodiments, variations and features of the present methods are described below in various levels of detail in order to provide a substantial understanding of the present technology.

[0042] In practicing the present methods, many conventional techniques in molecular biology, protein biochemistry, cell biology, immunology, microbiology and recombinant DNA are used. See, e.g., Sambrook and Russell eds. (2001) Molecular Cloning: A Laboratory Manual, 3rd edition; the series Ausubel et al. eds. (2007) Current Protocols in Molecular Biology, the series Methods in Enzymology (Academic Press, Inc., N.Y.); MacPherson et al. (1991) PCR 1: A Practical Approach (IRL Press at Oxford University Press); MacPherson et al. (1995) PCR 2: A Practical Approach', Harlow and Lane eds. (\999) Antibodies, A Laboratory Manual', Freshney (2005) Culture of Animal Cells: A Manual of Basic Technique, 5th edition; Gait ed. (1984) Oligonucleotide Synthesis', U.S. Patent No. 4,683,195; Hames and Higgins eds. (1984) Nucleic Acid Hybridization', Anderson (1999) Nucleic Acid Hybridization', Hames and Higgins eds. (1984) Transcription and Translation; Immobilized Cells and Enzymes (IRL Press (1986)); Perbal (1984) A Practical Guide to Molecular Cloning; Miller and Calos eds. (1987) Gene Transfer Vectors for Mammalian Cells (Cold Spring Harbor Laboratory); Makrides ed. (2003) Gene Transfer and Expression in Mammalian Cells; Mayer and Walker eds. (1987) Immunochemical Methods in Cell and Molecular Biology (Academic Press, London); and Herzenberg et al. eds (1996) Weir’s Handbook of Experimental Immunology. Methods to detect and measure levels of polypeptide gene expression products (i.e., gene translation level) are well-known in the art and include the use of polypeptide detection methods such as antibody detection and quantification techniques. (See also, Strachan & Read, Human Molecular Genetics, Second Edition. (John Wiley and Sons, Inc., NY, 1999)).

[0043] Disclosed herein are methods for simultaneously identifying RBP-RNA interaction (PRI) sites with single nucleotide resolution for hundreds of RBPs across the entire transcriptome (a.k.a., Annotation of RNA Binding Protein Occupancy on RNA (ARORA)).

[0044] Several key technical advances were required for ARORA to reach scalable and sensitive detection of PRIs with high degrees of signal to noise, even at modest read depth, including: (1) highly efficient and unbiased labeling of RNA by maintaining RNA in an unstructured state in high DMSO concentrations during covalent labeling reactions, (2) reduction of molecular complexity of RBP-RNA crosslinked biologic samples using extensive proteolysis while preserving PRI information makes RNA handling identical to standard non-crosslinked RNA and compatible with high throughput handling and automation, (3) utilization of highly reproducible, unbiased fragmentation of RNA with heat and divalent cations to achieve consistent short fragments which are optimal for efficient annotation of PRIs, (4) enriching crosslinked RNA fragments by streptavidin affinity purification and performing cDNA synthesis while RNA templates are immobilized on streptavidin beads, which substantially increases the efficiency of identifying PRIs, (5) ligating unique indexed 3’ linkers to RNA fragments, prior to pooling RNA from UVC crosslinked and non-crosslinked samples into a single sample prior to pulldown and cDNA synthesis, permitting precise, quantitative comparisons of enrichment between UVC crosslinked and non-crosslinked control samples.

[0045] The methods of the present technology recapitulate PRI profiles obtained with three orthogonal conventional methods (eCLIP (enhanced Crosslinking and Immunoprecipitation), RNP-MaP (Ribonucleoprotein networks analyzed by mutational profiling), and SHAPE Footprinting) in a single assay at efficiencies that are improved by orders of magnitude. ARORA is also capable of identifying dynamic and stress responsive PRI interactions and distinguishes biologic states that may be associated with differential drug sensitivity. ARORA also identifies allele-specific sequence variants (e.g., variant in MYC 5’UTR promoting MYC protein translation and cancer cell growth) in regulatory RNA that affect the molecular interactions of RBPs. The Examples described herein further demonstrate that the methods of the present technology can be successfully implemented in frozen tissue samples. Therefore, the methods disclosed herein are useful as a platform technology for diagnostic and drug target discovery. Unlike conventional methods that identify PRIs, ARORA is also amenable to automation, high throughput, unbiased discovery. Definitions

[0046] Unless defined otherwise, all technical and scientific terms used herein generally have the same meaning as commonly understood by one of ordinary skill in the art to which this technology belongs. As used in this specification and the appended claims, the singular forms “a”, “an” and “the” include plural referents unless the content clearly dictates otherwise. For example, reference to “a cell” includes a combination of two or more cells, and the like. Generally, the nomenclature used herein and the laboratory procedures in cell culture, molecular genetics, organic chemistry, analytical chemistry and nucleic acid chemistry and hybridization described below are those well-known and commonly employed in the art.

[0047] Generally, reference to a certain element such as hydrogen or H is meant to include all isotopes of that element. For example, if an R group is defined to include hydrogen or H, it also includes deuterium and tritium. Compounds comprising radioisotopes such as tritium, C14, P32 and S35 are thus within the scope of the present technology. Procedures for inserting such labels into the compounds of the present technology will be readily apparent to those skilled in the art based on the disclosure herein.

[0048] In general, “substituted” refers to an organic group as defined below (e.g., an alkyl group) in which one or more bonds to a hydrogen atom contained therein are replaced by a bond to non-hydrogen or non-carbon atoms. Substituted groups also include groups in which one or more bonds to a carbon(s) or hydrogen(s) atom are replaced by one or more bonds, including double or triple bonds, to a heteroatom. Thus, a substituted group is substituted with one or more substituents, unless otherwise specified. In some embodiments, a substituted group is substituted with 1, 2, 3, 4, 5, or 6 substituents. Examples of substituent groups include: halogens (i.e., F, Cl, Br, and I); hydroxyls; alkoxy, alkenoxy, aryloxy, aralkyloxy, heterocyclyl, heterocyclylalkyl, heterocyclyloxy, and heterocyclylalkoxy groups; carbonyls (oxo); carboxylates; esters; urethanes; oximes; hydroxylamines; alkoxyamines; aralkoxyamines; thiols; sulfides; sulfoxides; sulfones; sulfonyls; pentafluorosulfanyl (i.e., SFs), sulfonamides; amines; N-oxides; hydrazines; hydrazides; hydrazones; azides; amides; ureas; amidines; guanidines; enamines; imides; isocyanates; isothiocyanates; cyanates; thiocyanates; imines; nitro groups; and nitriles (i.e., CN).

[0049] Substituted ring groups such as substituted cycloalkyl, aryl, heterocyclyl and heteroaryl groups also include rings and ring systems in which a bond to a hydrogen atom is replaced with a bond to a carbon atom. Therefore, substituted cycloalkyl, aryl, heterocyclyl and heteroaryl groups may also be substituted with substituted or unsubstituted alkyl, alkenyl, and alkynyl groups as defined below.

[0050] Alkyl groups include straight chain and branched chain alkyl groups having from 1 to 12 carbon atoms, and typically from 1 to 10 carbons or, in some embodiments, from 1 to 8, 1 to 6, or 1 to 4 carbon atoms. Alkyl groups may be substituted or unsubstituted. Examples of straight chain alkyl groups include groups such as methyl, ethyl, n-propyl, n-butyl, n-pentyl, n-hexyl, n-heptyl, and n-octyl groups. Examples of branched alkyl groups include, but are not limited to, isopropyl, iso-butyl, sec-butyl, tert-butyl, neopentyl, isopentyl, and 2,2-dimethylpropyl groups. Representative substituted alkyl groups may be substituted one or more times with substituents such as those listed above, and include without limitation haloalkyl (e.g., trifluoromethyl), hydroxyalkyl, thioalkyl, aminoalkyl, alkylaminoalkyl, dialkylaminoalkyl, alkoxyalkyl, carboxyalkyl, and the like.

[0051] Cycloalkyl groups include mono-, bi- or tricyclic alkyl groups having from 3 to 12 carbon atoms in the ring(s), or, in some embodiments, 3 to 10, 3 to 8, or 3 to 4, 5, or 6 carbon atoms. Cycloalkyl groups may be substituted or unsubstituted. Exemplary monocyclic cycloalkyl groups include, but not limited to, cyclopropyl, cyclobutyl, cyclopentyl, cyclohexyl, cycloheptyl, and cyclooctyl groups. In some embodiments, the cycloalkyl group has 3 to 8 ring members, whereas in other embodiments the number of ring carbon atoms range from 3 to 5, 3 to 6, or 3 to 7. Bi- and tricyclic ring systems include both bridged cycloalkyl groups and fused rings, such as, but not limited to, bicyclo[2.1.1]hexane, adamantyl, decalinyl, and the like. Substituted cycloalkyl groups may be substituted one or more times with, non-hydrogen and non-carbon groups as defined above. However, substituted cycloalkyl groups also include rings that are substituted with straight or branched chain alkyl groups as defined above. Representative substituted cycloalkyl groups may be mono-substituted or substituted more than once, such as, but not limited to, 2,2-, 2,3-, 2,4- 2,5- or 2,6-disubstituted cyclohexyl groups, which may be substituted with substituents such as those listed above.

[0052] Cycloalkylalkyl groups are alkyl groups as defined above in which a hydrogen or carbon bond of an alkyl group is replaced with a bond to a cycloalkyl group as defined above. Cycloalkylalkyl groups may be substituted or unsubstituted. In some embodiments, cycloalkylalkyl groups have from 4 to 16 carbon atoms, 4 to 12 carbon atoms, and typically 4 to 10 carbon atoms. Substituted cycloalkylalkyl groups may be substituted at the alkyl, the cycloalkyl or both the alkyl and cycloalkyl portions of the group. Representative substituted cycloalkylalkyl groups may be mono-substituted or substituted more than once, such as, but not limited to, mono-, di- or tri-substituted with substituents such as those listed above.

[0053] Alkenyl groups include straight and branched chain alkyl groups as defined above, except that at least one double bond exists between two carbon atoms. Alkenyl groups may be substituted or unsubstituted. Alkenyl groups have from 2 to 12 carbon atoms, and typically from 2 to 10 carbons or, in some embodiments, from 2 to 8, 2 to 6, or 2 to 4 carbon atoms. In some embodiments, the alkenyl group has one, two, or three carboncarbon double bonds. Examples include, but are not limited to vinyl, allyl, -CH=CH(CH3), -CH=C(CH3)2, -C(CH3)=CH2, -C(CH3)=CH(CH3), -C(CH2CH3)=CH2, among others. Representative substituted alkenyl groups may be mono-substituted or substituted more than once, such as, but not limited to, mono-, di- or tri-substituted with substituents such as those listed above.

[0054] Cycloalkenyl groups include cycloalkyl groups as defined above, having at least one double bond between two carbon atoms. Cycloalkenyl groups may be substituted or unsubstituted. In some embodiments the cycloalkenyl group may have one, two or three double bonds but does not include aromatic compounds. Cycloalkenyl groups have from 4 to 14 carbon atoms, or, in some embodiments, 5 to 14 carbon atoms, 5 to 10 carbon atoms, or even 5, 6, 7, or 8 carbon atoms. Examples of cycloalkenyl groups include cyclohexenyl, cyclopentenyl, cyclohexadienyl, cyclobutadienyl, and cyclopentadienyl.

[0055] Cycloalkenylalkyl groups are alkyl groups as defined above in which a hydrogen or carbon bond of the alkyl group is replaced with a bond to a cycloalkenyl group as defined above. Cycloalkenylalkyl groups may be substituted or unsubstituted. Substituted cycloalkenylalkyl groups may be substituted at the alkyl, the cycloalkenyl or both the alkyl and cycloalkenyl portions of the group. Representative substituted cycloalkenylalkyl groups may be substituted one or more times with substituents such as those listed above.

[0056] Alkynyl groups include straight and branched chain alkyl groups as defined above, except that at least one triple bond exists between two carbon atoms. Alkynyl groups may be substituted or unsubstituted. Alkynyl groups have from 2 to 12 carbon atoms, and typically from 2 to 10 carbons or, in some embodiments, from 2 to 8, 2 to 6, or 2 to 4 carbon atoms. In some embodiments, the alkynyl group has one, two, or three carboncarbon triple bonds. Examples include, but are not limited to -C=CH, -C=CCH3, -CH2C=CCH3, and -C=CCH2CH(CH2CH3)2, among others. Representative substituted alkynyl groups may be mono-substituted or substituted more than once, such as, but not limited to, mono-, di- or tri-substituted with substituents such as those listed above.

[0057] Aryl groups are cyclic aromatic hydrocarbons that do not contain heteroatoms. Aryl groups herein include monocyclic, bicyclic and tricyclic ring systems. Aryl groups may be substituted or unsubstituted. Thus, aryl groups include, but are not limited to, phenyl, azulenyl, heptalenyl, biphenyl, fluorenyl, phenanthrenyl, anthracenyl, indenyl, indanyl, pentalenyl, and naphthyl groups. In some embodiments, aryl groups contain 6-14 carbons, and in others from 6 to 12 or even 6-10 carbon atoms in the ring portions of the groups. In some embodiments, the aryl groups are phenyl or naphthyl. The phrase “aryl groups” includes groups containing fused rings, such as fused aromatic-aliphatic ring systems (e.g., indanyl, tetrahydronaphthyl, and the like). Representative substituted aryl groups may be mono-substituted (e.g., tolyl) or substituted more than once. For example, monosubstituted aryl groups include, but are not limited to, 2-, 3-, 4-, 5-, or 6-substituted phenyl or naphthyl groups, which may be substituted with substituents such as those listed above.

[0058] Aralkyl groups are alkyl groups as defined above in which a hydrogen or carbon bond of an alkyl group is replaced with a bond to an aryl group as defined above. Aralkyl groups may be substituted or unsubstituted. In some embodiments, aralkyl groups contain 7 to 16 carbon atoms, 7 to 14 carbon atoms, or 7 to 10 carbon atoms. Substituted aralkyl groups may be substituted at the alkyl, the aryl or both the alkyl and aryl portions of the group. Representative aralkyl groups include but are not limited to benzyl and phenethyl groups and fused (cycloalkylaryl)alkyl groups such as 4-indanylethyl. Representative substituted aralkyl groups may be substituted one or more times with substituents such as those listed above.

[0059] Heterocyclyl groups include aromatic (also referred to as heteroaryl) and nonaromatic ring compounds containing 3 or more ring members, of which one or more is a heteroatom such as, but not limited to, N, O, and S. Heterocyclyl groups may be substituted or unsubstituted. In some embodiments, the heterocyclyl group contains 1, 2, 3 or 4 heteroatoms. In some embodiments, heterocyclyl groups include mono-, bi- and tricyclic rings having 3 to 16 ring members, whereas other such groups have 3 to 6, 3 to 10, 3 to 12, or 3 to 14 ring members. Heterocyclyl groups encompass aromatic, partially unsaturated and saturated ring systems, such as, for example, imidazolyl, imidazolinyl and imidazolidinyl groups. The phrase “heterocyclyl group” includes fused ring species including those comprising fused aromatic and non-aromatic groups, such as, for example, benzotriazolyl, 2,3-dihydrobenzo[l,4]dioxinyl, and benzo[l,3]dioxolyl. The phrase also includes bridged polycyclic ring systems containing a heteroatom such as, but not limited to, quinuclidyl. The phrase includes heterocyclyl groups that have other groups, such as alkyl, oxo or halo groups, bonded to one of the ring members, referred to as “substituted heterocyclyl groups”. Heterocyclyl groups include, but are not limited to, aziridinyl, azetidinyl, pyrrolidinyl, imidazolidinyl, pyrazolidinyl, thiazolidinyl, tetrahydrothiophenyl, tetrahydrofuranyl, dioxolyl, furanyl, thiophenyl, pyrrolyl, pyrrolinyl, imidazolyl, imidazolinyl, pyrazolyl, pyrazolinyl, triazolyl, tetrazolyl, oxazolyl, isoxazolyl, thiazolyl, thiazolinyl, isothiazolyl, thiadiazolyl, oxadiazolyl, piperidyl, piperazinyl, morpholinyl, thiomorpholinyl, tetrahydropyranyl, tetrahydrothiopyranyl, oxathiane, dioxyl, dithianyl, pyranyl, pyridyl, pyrimidinyl, pyridazinyl, pyrazinyl, triazinyl, dihydropyridyl, dihydrodithiinyl, dihydrodithionyl, homopiperazinyl, quinuclidyl, indolyl, indolinyl, isoindolyl,azaindolyl (pyrrolopyridyl), indazolyl, indolizinyl, benzotriazolyl, benzimidazolyl, benzofuranyl, benzothiophenyl, benzthiazolyl, benzoxadiazolyl, benzoxazinyl, benzodithiinyl, benzoxathiinyl, benzothiazinyl, benzoxazolyl, benzothiazolyl, benzothiadiazolyl, benzo[1,3]dioxolyl, pyrazolopyridyl, imidazopyridyl (azabenzimidazolyl), triazolopyridyl, isoxazolopyridyl, purinyl, xanthinyl, adeninyl, guaninyl, quinolinyl, isoquinolinyl, quinolizinyl, quinoxalinyl, quinazolinyl, cinnolinyl, phthalazinyl, naphthyridinyl, pteridinyl, thianaphthyl, dihydrobenzothiazinyl, dihydrobenzofuranyl, dihydroindolyl, dihydrobenzodioxinyl, tetrahydroindolyl, tetrahydroindazolyl, tetrahydrobenzimidazolyl, tetrahydrobenzotriazolyl, tetrahydropyrrolopyridy 1, tetrahydropyrazolopyridy 1, tetrahydroimidazopyri dy 1, tetrahydrotriazolopyridyl, and tetrahydroquinolinyl groups. Representative substituted heterocyclyl groups may be mono-substituted or substituted more than once, such as, but not limited to, pyridyl or morpholinyl groups, which are 2-, 3-, 4-, 5-, or 6-substituted, or disubstituted with various substituents such as those listed above.

[0060] Heteroaryl groups are aromatic ring compounds containing 5 or more ring members, of which, one or more is a heteroatom such as, but not limited to, N, O, and S. Heteroaryl groups may be substituted or unsubstituted. Heteroaryl groups include, but are not limited to, groups such as pyrrolyl, pyrazolyl, triazolyl, tetrazolyl, oxazolyl, isoxazolyl, thiazolyl, pyridinyl, pyridazinyl, pyrimidinyl, pyrazinyl, thiophenyl, benzothiophenyl, furanyl, benzofuranyl, indolyl, azaindolyl (pyrrolopyridinyl), indazolyl, benzimidazolyl, imidazopyridinyl (azabenzimidazolyl), pyrazolopyridinyl, triazolopyridinyl, benzotriazolyl, benzoxazolyl, benzothiazolyl, benzothiadiazolyl, imidazopyridinyl, isoxazolopyridinyl, thianaphthyl, purinyl, xanthinyl, adeninyl, guaninyl, quinolinyl, isoquinolinyl, tetrahydroquinolinyl, quinoxalinyl, and quinazolinyl groups. Heteroaryl groups include fused ring compounds in which all rings are aromatic such as indolyl groups and include fused ring compounds in which only one of the rings is aromatic, such as 2,3-dihydro indolyl groups. Representative substituted heteroaryl groups may be substituted one or more times with various substituents such as those listed above.

[0061] Heterocyclylalkyl groups are alkyl groups as defined above in which a hydrogen or carbon bond of an alkyl group is replaced with a bond to a heterocyclyl group as defined above. Heterocyclylalkyl groups may be substituted or unsubstituted. Substituted heterocyclylalkyl groups may be substituted at the alkyl, the heterocyclyl or both the alkyl and heterocyclyl portions of the group. Representative heterocyclyl alkyl groups include, but are not limited to, morpholin-4-yl-ethyl, furan-2-yl-methyl, imidazol-4-yl-methyl, pyridin-3-yl-methyl, tetrahydrofuran-2-yl-ethyl, and indol-2-yl-propyl. Representative substituted heterocyclylalkyl groups may be substituted one or more times with substituents such as those listed above.

[0062] Heteroaralkyl groups are alkyl groups as defined above in which a hydrogen or carbon bond of an alkyl group is replaced with a bond to a heteroaryl group as defined above. Heteroaralkyl groups may be substituted or unsubstituted. Substituted heteroaralkyl groups may be substituted at the alkyl, the heteroaryl or both the alkyl and heteroaryl portions of the group. Representative substituted heteroaralkyl groups may be substituted one or more times with substituents such as those listed above.

[0063] Groups described herein having two or more points of attachment (i.e., divalent, trivalent, or polyvalent) within the compound of the present technology are designated by use of the suffix, “ene.” For example, divalent alkyl groups are alkylene groups, divalent aryl groups are arylene groups, divalent heteroaryl groups are divalent heteroarylene groups, and so forth. Substituted groups having a single point of attachment to the compound of the present technology are not referred to using the “ene” designation. Thus, e.g., chloroethyl is not referred to herein as chloroethylene.

[0064] Alkoxy groups are hydroxyl groups (-OH) in which the bond to the hydrogen atom is replaced by a bond to a carbon atom of a substituted or unsubstituted alkyl group as defined above. Alkoxy groups may be substituted or unsubstituted. Examples of linear alkoxy groups include but are not limited to methoxy, ethoxy, propoxy, butoxy, pentoxy, hexoxy, and the like. Examples of branched alkoxy groups include but are not limited to isopropoxy, sec-butoxy, tert-butoxy, isopentoxy, isohexoxy, and the like. Examples of cycloalkoxy groups include but are not limited to cyclopropyloxy, cyclobutyloxy, cyclopentyloxy, cyclohexyloxy, and the like. Representative substituted alkoxy groups may be substituted one or more times with substituents such as those listed above.

[0065] The terms “alkyloyl” and “alkyloyloxy” as used herein can refer, respectively, to -C(O)-alkyl groups and -O-C(O)-alkyl groups. Similarly, “aryloyl” and “aryloyloxy” refer to -C(O)-aryl groups and -O-C(O)-aryl groups.

[0066] The terms "aryloxy" and “arylalkoxy” refer to, respectively, a substituted or unsubstituted aryl group bonded to an oxygen atom and a substituted or unsubstituted aralkyl group bonded to the oxygen atom at the alkyl. Examples include but are not limited to phenoxy, naphthyloxy, and benzyloxy. Representative substituted aryloxy and arylalkoxy groups may be substituted one or more times with substituents such as those listed above.

[0067] The term “carboxylate” as used herein refers to a -COOH group.

[0068] The term “ester” as used herein refers to -COOR70 and -C(O)O-G groups. R70 is a substituted or unsubstituted alkyl, cycloalkyl, alkenyl, alkynyl, aryl, aralkyl, heterocyclylalkyl or heterocyclyl group as defined herein. G is a carboxylate protecting group. Carboxylate protecting groups are well known to one of ordinary skill in the art. An extensive list of protecting groups for the carboxylate group functionality may be found in Protective Groups in Organic Synthesis, Greene, T.W.; Wuts, P. G. M., John Wiley & Sons, New York, NY, (3rd Edition, 1999) which can be added or removed using the procedures set forth therein and which is hereby incorporated by reference in its entirety and for any and all purposes as if fully set forth herein.

[0069] The term “amide” (or “amido”) includes C- and N-amide groups, i.e., -C(O)NR71R72, and -NR71C(O)R72 groups, respectively. R71 and R72 are independently hydrogen, or a substituted or unsubstituted alkyl, alkenyl, alkynyl, cycloalkyl, aryl, aralkyl, heterocyclylalkyl or heterocyclyl group as defined herein. Amido groups therefore include but are not limited to carbamoyl groups (-C(O)NH2) and formamide groups (-NHC(O)H). In some embodiments, the amide is -NR71C(O)-(Ci-5 alkyl) and the group is termed "carbonylamino," and in others the amide is -NHC(O)-alkyl and the group is termed "alkanoylamino."

[0070] The term “nitrile” or “cyano” as used herein refers to the -CN group.

[0071] Urethane groups include N- and O-urethane groups, i.e., -NR73C(O)OR74 and -OC(O)NR73R74 groups, respectively. R73 and R74 are independently a substituted or unsubstituted alkyl, alkenyl, alkynyl, cycloalkyl, aryl, aralkyl, heterocyclylalkyl, or heterocyclyl group as defined herein. R73 may also be H.

[0072] The term “amine” (or “amino”) as used herein refers to -NR75R76 groups, wherein R75 and R76 are independently hydrogen, or a substituted or unsubstituted alkyl, alkenyl, alkynyl, cycloalkyl, aryl, aralkyl, heterocyclylalkyl or heterocyclyl group as defined herein. In some embodiments, the amine is alkylamino, dialkylamino, arylamino, or alkylarylamino. In other embodiments, the amine is NH2, methylamino, dimethylamino, ethylamino, diethylamino, propylamino, isopropylamino, phenylamino, or benzylamino.

[0073] The term “sulfonamido” includes S- and N-sulfonamide groups, i.e., -SO2NR78R79 and -NR78SO2R79 groups, respectively. R78 and R79 are independently hydrogen, or a substituted or unsubstituted alkyl, alkenyl, alkynyl, cycloalkyl, aryl, aralkyl, heterocyclylalkyl, or heterocyclyl group as defined herein. Sulfonamido groups therefore include but are not limited to sulfamoyl groups (-SO2NH2). In some embodiments herein, the sulfonamido is -NHSO2-alkyl and is referred to as the "alkylsulfonylamino" group.

[0074] The term “thiol” refers to -SH groups, while “sulfides” include -SR80 groups, “sulfoxides” include -S(O)R81 groups, “sulfones” include -SO2R82 groups, “sulfonyls” include -SO2OR83, and “sulfonates” include -SO3 . R80, R81, R82, and R83 are each independently a substituted or unsubstituted alkyl, cycloalkyl, alkenyl, alkynyl, aryl aralkyl, heterocyclyl or heterocyclylalkyl group as defined herein. In some embodiments the sulfide is an alkylthio group, -S-alkyl.

[0075] The term “urea” refers to -NR84-C(O)-NR85R86 groups. R84, R85, and R86 groups are independently hydrogen, or a substituted or unsubstituted alkyl, alkenyl, alkynyl, cycloalkyl, aryl, aralkyl, heterocyclyl, or heterocyclylalkyl group as defined herein.

[0076] The term “amidine” refers to -C(NR87)NR88R89 and -NR87C(NR88)R89, wherein r87 r88, anc[ ^89 are each in(iepen(ientiy hydrogen, or a substituted or unsubstituted alkyl, cycloalkyl, alkenyl, alkynyl, aryl aralkyl, heterocyclyl or heterocyclylalkyl group as defined herein.

[0077] The term “guanidine” refers to -NR90C(NR91)NR92R93, wherein R90, R91, R92 and R93 are each independently hydrogen, or a substituted or unsubstituted alkyl, cycloalkyl, alkenyl, alkynyl, aryl aralkyl, heterocyclyl or heterocyclylalkyl group as defined herein.

[0078] The term “enamine” refers to -C(R94)=C(R95)NR96R97 and -NR94C(R95)=C(R96)R97, wherein R94, R95, R96 and R97 are each independently hydrogen, a substituted or unsubstituted alkyl, cycloalkyl, alkenyl, alkynyl, aryl aralkyl, heterocyclyl or heterocyclylalkyl group as defined herein.

[0079] The term “halogen” or “halo” as used herein refers to bromine, chlorine, fluorine, or iodine. In some embodiments, the halogen is fluorine. In other embodiments, the halogen is chlorine or bromine.

[0080] The term “hydroxyl” as used herein can refer to -OH or its ionized form, -O . A “hydroxyalkyl” group is a hydroxyl-substituted alkyl group, such as HO-CH2-.

[0081] The term “imide” refers to -C(O)NR98C(O)R99, wherein R98 and R99 are each independently hydrogen, or a substituted or unsubstituted alkyl, cycloalkyl, alkenyl, alkynyl, aryl aralkyl, heterocyclyl or heterocyclylalkyl group as defined herein.

[0082] The term “imine” refers to -CR100(NR101) and -N(CR100R101) groups, wherein R100 and R101 are each independently hydrogen or a substituted or unsubstituted alkyl, cycloalkyl, alkenyl, alkynyl, aryl aralkyl, heterocyclyl or heterocyclylalkyl group as defined herein, with the proviso that R100 and R101 are not both simultaneously hydrogen.

[0083] The term “nitro” as used herein refers to an -NO2 group.

[0084] The term “trifluoromethyl” as used herein refers to -CF3.

[0085] The term “trifluoromethoxy” as used herein refers to -OCF3.

[0086] The term “azido” refers to -N3.

[0087] The term “trialkyl ammonium” refers to a -N(alkyl)3 group. A trialkylammonium group is positively charged and thus typically has an associated anion, such as halogen anion.

[0088] The term “isocyano” refers to -NC.

[0089] The term “isothiocyano” refers to -NCS.

[0090] The term “pentafluorosulfanyl” refers to -SF5.

[0091] As understood by one of ordinary skill in the art, “molecular weight” (also known as “relative molar mass”) is a dimensionless quantity but is converted to molar mass by multiplying by 1 gram / mole or by multiplying by 1 Da - for example, a compound with a weight-average molecular weight of 5,000 has a weight-average molar mass of 5,000 g / mol and a weight-average molar mass of 5,000 Da.

[0092] As used herein, the term “about” in reference to a number is generally taken to include numbers that fall within a range of 1%, 5%, or 10% in either direction (greater than or less than) of the number unless otherwise stated or otherwise evident from the context (except where such number would be less than 0% or exceed 100% of a possible value).

[0093] The term “adapter” refers to a short, chemically synthesized, nucleic acid which can be used to ligate to the 3' or 5' end of a nucleic acid sequence in order to facilitate attachment to another molecule. The adapter can be single-stranded or double-stranded. An adapter may incorporate a short (e.g., less than 55 base pairs) sequence useful for PCR amplification or sequencing. The adapter can comprise known sequences, degenerate sequences (a sequence not having a precise definition), or both. A double-stranded adapter may comprise two hybridizable strands. Alternatively, a double-stranded adapter can comprise a hybridizable portion and a non-hybridizable portion. The non-hybridizable portion of a double-stranded adapter comprises two single-stranded regions that are not hybridizable to each other. Within the non hybridizable portion, the strand containing an unhybridized 5'-end is referred to as the 5'-strand and the strand containing an unhybridized 3'-end is referred to as the 3'-strand. In some embodiments, the double-stranded adapter has a hybridizable portion at one end of the adapter and a non-hybridizable portion at the opposite end of the adapter. In some embodiments, the non-hybridizable portion of the double-stranded adapter may be open (Y-shaped adapter). In some embodiments, the adapter may be a U-shaped adapter.

[0094] The term “amino acid” refers to naturally occurring and non-naturally occurring amino acids, as well as amino acid analogs and amino acid mimetics that function in a manner similar to the naturally occurring amino acids. Naturally encoded amino acids are the 20 common amino acids (alanine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, and valine) and pyrolysine and selenocysteine. Amino acid analogs refer to agents that have the same basic chemical structure as a naturally occurring amino acid, i.e., an a carbon that is bound to a hydrogen, a carboxyl group, an amino group, and an R group, such as, homoserine, norleucine, methionine sulfoxide, methionine methyl sulfonium. Such analogs have modified R groups (such as, norleucine) or modified peptide backbones, but retain the same basic chemical structure as a naturally occurring amino acid. In some embodiments, amino acids forming a polypeptide are in the D form. In some embodiments, the amino acids forming a polypeptide are in the L form. In some embodiments, a first plurality of amino acids forming a polypeptide are in the D form, and a second plurality of amino acids are in the L form.

[0095] Amino acids are referred to herein by either their commonly known three letter symbols or by the one-letter symbols recommended by the IUPAC-IUB Biochemical Nomenclature Commission. Nucleotides, likewise, are referred to by their commonly accepted single-letter code.

[0096] As used herein, the terms “amplify” or “amplification” with respect to nucleic acid sequences, refer to methods that increase the representation of a population of nucleic acid sequences in a sample. Nucleic acid amplification methods, such as PCR, isothermal methods, rolling circle methods, etc., are well known to the skilled artisan. Copies of a particular nucleic acid sequence generated in vitro in an amplification reaction are called “amplicons” or “amplification products”.

[0097] “Bait”, as used herein, is a type of hybrid capture reagent that retrieves target nucleic acid sequences for sequencing. A bait can be a nucleic acid molecule, e.g., a DNA or RNA molecule, which can hybridize to (e.g., be complementary to), and thereby allow capture of a target nucleic acid. In one embodiment, a bait is an RNA molecule (e.g., a naturally-occurring or modified RNA molecule); a DNA molecule (e.g., a naturally-occurring or modified DNA molecule), or a combination thereof. In other embodiments, a bait includes a binding entity, e.g., an affinity tag, that allows capture and separation, e.g., by binding to a binding entity, of a hybrid formed by a bait and a nucleic acid hybridized to the bait. In one embodiment, a bait is suitable for solution phase hybridization.

[0098] The term “barcode” refers to a sequence of nucleotides within a polynucleotide that is used to identify a nucleic acid molecule. For example, a barcode can be used to identify molecules when the molecules from several groups are combined for processing or sequencing in a multiplexed fashion. A barcode can be located at a certain position within a polynucleotide (e.g., at the 3'-end, 5'-end, or middle of the polynucleotide) and can comprise sequences of any length (e.g., 1-100 or more nucleotides). Additionally, a barcode can comprise one or more pre-defined sequences. The term “pre-defined” means that sequence of a barcode is predetermined or known prior to identifying or without the need to identify the entire sequence of the nucleic acid comprising the barcode. In some cases, pre-defined barcodes can be attached to nucleic acids for sorting the nucleic acids into groups. In some embodiments, a barcode can comprise artificial sequences, e.g., designed or engineered sequences that are not present in the unaltered (wild-type) genome of a subject. In other embodiments, a barcode can comprise an endogenous sequence, e.g., sequences that are present in the unaltered (wildtype) genome of a subject. In certain embodiments, a barcode can be an endogenous barcode. An endogenous barcode can be a sequence of a genomic nucleic acid, where the sequence is used as a barcode or identifier for the genomic nucleic acid. One or more sequences of the genomic DNA fragment can be an endogenous barcode. Different types of barcodes can be used in combination. For example, an endogenous genomic nucleic acid fragment can be attached to an artificial sequence, which can be used as a unique identifier of the genomic nucleic acid fragment. A "sample-specific barcode" or "patient barcode" refers to a polynucleotide sequence that is used to identify the origin or source of a nucleic acid molecule. For example, a sequence of “AAAA” can be attached to identify nucleic acids isolated from Patient A.

[0099] As used herein, the term “biological sample” means sample material derived from living cells. Biological samples may include tissues, cells, protein or membrane extracts of cells, and biological fluids (e.g., ascites fluid or cerebrospinal fluid (CSF)) isolated from a subject, as well as tissues, cells and fluids present within a subject. Biological samples of the present technology include, but are not limited to, samples taken from breast tissue, renal tissue, the uterine cervix, the endometrium, the head or neck, the gallbladder, parotid tissue, the prostate, the brain, the pituitary gland, kidney tissue, muscle, the esophagus, the stomach, the small intestine, the colon, the liver, the spleen, the pancreas, thyroid tissue, heart tissue, lung tissue, the bladder, adipose tissue, lymph node tissue, the uterus, ovarian tissue, adrenal tissue, testis tissue, the tonsils, thymus, blood, hair, buccal, skin, serum, plasma, CSF, semen, prostate fluid, seminal fluid, urine, feces, sweat, saliva, sputum, mucus, bone marrow, lymph, and tears. Biological samples can also be obtained from biopsies of internal organs or from cancers. Biological samples can be obtained from subjects for diagnosis or research or can be obtained from non-diseased individuals, as controls or for basic research. Samples may be obtained by standard methods including, e.g., venous puncture and surgical biopsy. In certain embodiments, the biological sample is a tissue sample obtained by needle biopsy.

[00100] The terms “cancer” or “tumor” are used interchangeably and refer to the presence of cells possessing characteristics typical of cancer-causing cells, such as uncontrolled proliferation, immortality, metastatic potential, rapid growth and proliferation rate, and certain characteristic morphological features. Cancer cells are often in the form of a tumor, but such cells can exist alone within an animal, or can be a non-tumorigenic cancer cell. As used herein, the term “cancer cells” includes precancerous (e.g., benign), malignant, pre-metastatic, metastatic, and non-metastatic cells. Cancers of virtually every tissue are known to those of skill in the art, including solid tumors such as carcinomas, sarcomas, glioblastomas, melanomas, lymphomas, myelomas, etc., and circulating cancers such as leukemias. Examples of cancer include, but are not limited to, ovarian cancer, breast cancer, colon cancer, lung cancer, prostate cancer, gastric cancer, pancreatic cancer, cervical cancer, ovarian cancer, liver cancer, bladder cancer, cancer of the urinary tract, thyroid cancer, renal cancer, carcinoma, melanoma, head and neck cancer, and brain cancer. The term “cancer cell” refers to a cell that exhibits cancer-like properties, e.g., uncontrollable reproduction, resistance to anti- growth signals, ability to metastasize, and loss of ability to undergo programmed cell death (e.g., apoptosis) or a cell that is derived from a cancer cell, e.g., clone of a cancer cell.

[00101] The terms “complementary” or “complementarity” as used herein with reference to polynucleotides (i. e., a sequence of nucleotides such as an oligonucleotide or a target nucleic acid) refer to the base-pairing rules. The complement of a nucleic acid sequence as used herein refers to an oligonucleotide which, when aligned with the nucleic acid sequence such that the 5' end of one sequence is paired with the 3’ end of the other, is in “antiparallel association.” For example, the sequence “5'-A-G-T-3”’ is complementary to the sequence “3’-T-C-A-5.” Certain bases not commonly found in naturally-occurring nucleic acids may be included in the nucleic acids described herein. These include, for example, inosine, 7-deazaguanine, Locked Nucleic Acids (LNA), and Peptide Nucleic Acids (PNA). Complementarity need not be perfect; stable duplexes may contain mismatched base pairs, degenerative, or unmatched bases. Those skilled in the art of nucleic acid technology can determine duplex stability empirically considering a number of variables including, for example, the length of the oligonucleotide, base composition and sequence of the oligonucleotide, ionic strength and incidence of mismatched base pairs. A complement sequence can also be an RNA sequence complementary to the DNA sequence or its complement sequence, and can also be a cDNA.

[00102] As used herein, the term “conjugated” refers to the association of two molecules by any method known to those in the art. Suitable types of associations include chemical bonds and physical bonds. Chemical bonds include, for example, covalent bonds and coordinate bonds. Physical bonds include, for instance, hydrogen bonds, dipolar interactions, van der Waal forces, electrostatic interactions, hydrophobic interactions and aromatic stacking.

[00103] As used herein, a "control" is an alternative sample used in an experiment for comparison purpose. A control can be "positive" or "negative." A “control nucleic acid sample” or “reference nucleic acid sample” as used herein, refers to nucleic acid molecules from a control or reference sample. In certain embodiments, the reference or control nucleic acid sample is a wild type or a non-mutated DNA or RNA sequence. In certain embodiments, the reference nucleic acid sample is purified or isolated (e.g., it is removed from its natural state).

[00104] “Detecting” as used herein refers to determining the presence of a RBP-RNA interaction (PRI) site in a nucleic acid of interest in a sample. Detection does not require the method to provide 100% sensitivity.

[00105] “Detectable label” as used herein refers to a molecule or a compound or a group of molecules or a group of compounds used to identify a nucleic acid or protein of interest. In some embodiments, the detectable label may be detected directly. In other embodiments, the detectable label may be a part of a binding pair, which can then be subsequently detected. Signals from the detectable label may be detected by various means and will depend on the nature of the detectable label. Detectable labels may be isotopes, fluorescent moieties, colored substances, and the like. Examples of means to detect detectable labels include but are not limited to spectroscopic, photochemical, biochemical, immunochemical, electromagnetic, radiochemical, or chemical means, such as fluorescence, chemifluorescence, or chemiluminescence, or any other appropriate means.

[00106] As used herein, “expression” includes one or more of the following: transcription of the gene into precursor mRNA; splicing and other processing of the precursor mRNA to produce mature mRNA; mRNA stability; translation of the mature mRNA into protein (including codon usage and tRNA availability); and glycosylation and / or other modifications of the translation product, if required for proper expression and function.

[00107] “Gene” as used herein refers to a DNA sequence that comprises regulatory and coding sequences necessary for the production of an RNA, which may have a non-coding function (e.g., a ribosomal or transfer RNA) or which may include a polypeptide or a polypeptide precursor. The RNA or polypeptide may be encoded by a full length coding sequence or by any portion of the coding sequence so long as the desired activity or function is retained. Although a sequence of the nucleic acids may be shown in the form of DNA, a person of ordinary skill in the art recognizes that the corresponding RNA sequence will have a similar sequence with the thymine being replaced by uracil, i.e., "T" is replaced with "U."

[00108] The term “gene region” can refer to a range of sequences within a gene or surrounding a gene, e.g., an intron, an exon, a promoter, a 3’ untranslated region etc.

[00109] The term “hybridize” as used herein refers to a process where two substantially complementary nucleic acid strands (at least about 65% complementary over a stretch of at least 14 to 25 nucleotides, at least about 75%, or at least about 90% complementary) anneal to each other under appropriately stringent conditions to form a duplex or heteroduplex through formation of hydrogen bonds between complementary base pairs. Hybridizations are typically and preferably conducted with probe-length nucleic acid molecules, preferably 15-100 nucleotides in length, more preferably 18-50 nucleotides in length. Nucleic acid hybridization techniques are well known in the art. See, e.g., Sambrook, et al., 1989, Molecular Cloning: A Laboratory Manual, Second Edition, Cold Spring Harbor Press, Plainview, N.Y. Hybridization and the strength of hybridization (i.e., the strength of the association between the nucleic acids) is influenced by such factors as the degree of complementarity between the nucleic acids, stringency of the conditions involved, and the thermal melting point (Tm) of the formed hybrid. Those skilled in the art understand how to estimate and adjust the stringency of hybridization conditions such that sequences having at least a desired level of complementarity will stably hybridize, while those having lower complementarity will not. For examples of hybridization conditions and parameters, see, e.g., Sambrook, etal., 1989, Molecular Cloning: A Laboratory Manual, Second Edition, Cold Spring Harbor Press, Plainview, N.Y.; Ausubel, F. M. et al. 1994, Current Protocols in Molecular Biology, John Wiley & Sons, Secaucus, N.J. In some embodiments, specific hybridization occurs under stringent hybridization conditions. An oligonucleotide or polynucleotide (e.g., a probe or a primer) that is specific for a target nucleic acid will “hybridize” to the target nucleic acid under suitable conditions.

[00110] The term “hybridizable” means that two polynucleotide strands of a nucleic acid are complementary at one or more nucleotide positions, e.g., the nitrogenous bases of the two polynucleotide strands can form two or more Crick-Watson hydrogen bonds. For example, if a polynucleotide comprises 5’ ATGC 3’, it is hybridizable to the sequence 5' GCAT 3'. Under some experimental conditions, if a polynucleotide comprises 5' GGGG 3', it can also be hybridizable to the sequences 5'CCAC 3' and 5' CCCA 3', which are not perfectly complementary.

[00111] The term "non-hybridizable" means that two polynucleotide strands of a nucleic acid are non-complementary, e.g., nitrogenous bases of the two separate polynucleotide strands do not form two or more Crick-Watson hydrogen bonds under stringent hybridization conditions.

[00112] As used herein, the terms “individual”, “patient”, or “subject” can be an individual organism, a vertebrate, a mammal, or a human. In some embodiments, the individual, patient or subject is a human.

[00113] As used herein, the term “library” refers to a collection of nucleic acid sequences, e.g., a collection of nucleic acids derived from whole genomic, subgenomic fragments, cDNA, cDNA fragments, cfDNA, RNA, RNA fragments, or a combination thereof. In one embodiment, a portion or all of the library nucleic acid sequences comprises an adapter sequence. The adapter sequence can be located at one or both ends. The adapter sequence can be useful, e.g., for a sequencing method (e.g., an NGS method), for amplification, for reverse transcription, for sequencing, or for cloning into a vector.

[00114] The library can comprise a collection of nucleic acid sequences, e.g., a target nucleic acid sequence (e.g., a tumor nucleic acid sequence), a reference nucleic acid sequence, or a combination thereof. In some embodiments, the nucleic acid sequences of the library can be derived from a single subject. In other embodiments, a library can comprise nucleic acid sequences from more than one subject (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30 or more subjects). In some embodiments, two or more libraries from different subjects can be combined to form a library having nucleic acid sequences from more than one subject. In one embodiment, the subject has, or is at risk of having, a cancer or tumor.

[00115] A “library nucleic acid sequence” refers to a nucleic acid molecule, e.g., DNA, RNA, or a combination thereof, that is a member of a library. In some embodiments, a library nucleic acid sequence is a DNA molecule, e.g., genomic DNA, cfDNA, or cDNA. In some embodiments, a library nucleic acid sequence is fragmented, e.g., sheared or enzymatically prepared, genomic DNA. In certain embodiments, the library nucleic acid sequences comprise a sequence from a subject and a sequence not derived from the subject, e.g., adapter sequence, a primer sequence, or other sequences that allow for identification, e.g., “barcode” sequences.

[00116] The term "ligating" refers to connecting two molecules by chemical bonds to generate a new molecule. For example, ligating an adapter polynucleotide to another polynucleotide can refer to forming chemical bonds between the adapter and the polynucleotide (e.g., using a ligase or any other method) to generate a single new molecule comprising the adapter and the polynucleotide.

[00117] “Next-generation sequencing or NGS” as used herein, refers to any sequencing method that determines the nucleotide sequence of either individual nucleic acid molecules (e.g., in single molecule sequencing) or clonally expanded proxies for individual nucleic acid molecules in a high throughput parallel fashion (e.g., greater than 103, 104, 105 or more molecules are sequenced simultaneously). In one embodiment, the relative abundance of the nucleic acid species in the library can be estimated by counting the relative number of occurrences of their cognate sequences in the data generated by the sequencing experiment. Next generation sequencing methods are known in the art, and are described, e.g., in Metzker, M. Nature Biotechnology Reviews 11:31-46 (2010).

[00118] As used herein, “oligonucleotide” refers to a molecule that has a sequence of nucleic acid bases on a backbone comprised mainly of identical monomer units at defined intervals. The bases are arranged on the backbone in such a way that they can bind with a nucleic acid having a sequence of bases that are complementary to the bases of the oligonucleotide. The most common oligonucleotides have a backbone of sugar phosphate units. A distinction may be made between oligodeoxyribonucleotides that do not have a hydroxyl group at the 2' position and oligoribonucleotides that have a hydroxyl group at the 2' position. Oligonucleotides may also include derivatives, in which the hydrogen of the hydroxyl group is replaced with organic groups, e.g., an allyl group. One or more bases of the oligonucleotide may also be modified to include a phosphorothioate bond (e.g., one of the two oxygen atoms in the phosphate backbone which is not involved in the intemucleotide bridge, is replaced by a sulfur atom) to increase resistance to nuclease degradation. Oligonucleotides of the method which function as primers or probes are generally at least about 10-15 nucleotides long and more preferably at least about 15 to 55 nucleotides long, although shorter or longer oligonucleotides may be used in the method. The exact size will depend on many factors, which in turn depend on the ultimate function or use of the oligonucleotide. The oligonucleotide may be generated in any manner, including, for example, chemical synthesis, DNA replication, restriction endonuclease digestion of plasmids or phage DNA, reverse transcription, PCR, or a combination thereof. The oligonucleotide may be modified e.g., by addition of a methyl group, a biotin or digoxigenin moiety, a fluorescent tag or by using radioactive nucleotides.

[00119] As used herein, the term “peptide” means a short polymer of 2-50 amino acids joined to each other by peptide bonds or modified peptide bonds, i.e., peptide isosteres. In some embodiments, the peptide is no more than 10 amino acids in length. Peptides may contain amino acids other than the 20 gene-encoded amino acids. Peptides include amino acid sequences modified either by natural processes, such as post-translational processing, or by chemical modification techniques that are well known in the art.

[00120] As used herein, the term “polynucleotide” or “nucleic acid” means any RNA or DNA, which may be unmodified or modified RNA or DNA. Polynucleotides include, without limitation, single- and double-stranded DNA, DNA that is a mixture of single- and double-stranded regions, single- and double-stranded RNA, RNA that is mixture of single-and double-stranded regions, and hybrid molecules comprising DNA and RNA that may be single-stranded or, more typically, double-stranded or a mixture of single- and doublestranded regions. In addition, polynucleotide refers to triple-stranded regions comprising RNA or DNA or both RNA and DNA. The term polynucleotide also includes DNAs or RNAs containing one or more modified bases and DNAs or RNAs with backbones modified for stability or for other reasons.

[00121] The term “stringent hybridization conditions” as used herein refers to hybridization conditions at least as stringent as the following: hybridization in 50% formamide, 5xSSC, 50 mMNaH2PO4, pH 6.8, 0.5% SDS, 0.1 mg / mL sonicated salmon sperm DNA, and 5x Denhart's solution at 42° C. overnight; washing with 2x SSC, 0.1% SDS at 45° C; and washing with 0.2x SSC, 0.1% SDS at 45° C. In another example, stringent hybridization conditions should not allow for hybridization of two nucleic acids which differ over a stretch of 20 contiguous nucleotides by more than two bases.

[00122] As used herein, the terms “target sequence” and “target nucleic acid sequence” refer to a specific nucleic acid sequence in a biological sample to be analyzed. Limitations of Conventional Methods for Identifying PRIs

[00123] A variety of approached have been developed to RNA Binding Protein interactions with RNA molecules, most of which require crosslinking of RBP to RNA to stabilize the otherwise weak intermolecular interaction prior to isolation and in vitro manipulation or chemicals that covalently modify RNA residues that are not bound by RBPs to generate RBP ‘footprints’. Short wavelength ultraviolet light (UVC) crosslinks amino acids with direct contact nucleic acids at sub-nanometer distances and can identify the precise site of PRIs at single nucleotide resolution6. UVC-crosslinking of PRIs form irreversible, covalent linkages between the two macromolecules, and is the primary approach used to stabilize PRIs prior to downstream biochemical steps in most current methods to identify PRIs, including Crosslinking and Immunoprecipitation (CLIP) and its derived methods to identify protein bound sites on RNA, orthogonal organic phase separation (OOPS) and related approaches to identify all proteins capable of direct RNA biding, or mass spectrometry approaches to identify the specific amino acid residues and nucleotides involved in the PRI7

[00124] Utilization of chemical crosslinker such as formaldehyde, glutaraldehyde, and heterobifunctional crosslinkers, or RNA modifying SHAPE reagents have also been used to study PRIs8,10, but are limited in sensitivity including due to several constraints: (1) limited solvent accessibility of many PRIs to chemical crosslinkers, (2) fixation of indirect PRIs due to extensive protein-protein crosslinking, (3) long spacer distances (7nm and longer) capturing PRIs that do not physically interact in vivo, and are thus incapable of single nucleotide resolution of PRIs. Thus, CLIP and derived approaches have emerged as current standard for identifying the PRIs of specific RNAs across the transcriptome. A consortium of investigators recently published their results of nearly a decade of work to map the PRIs of 150 unique RBPs in one or two common cancer cell lines using enhanced CLIP (eCLIP) for the ENCODE project13,14. Yet, OOPS and other related approaches have recently identified more than 1000 human proteins robustly interact with RNA, demonstrating the enormous shortfall of eCLIP and related technologies to map the complete spectrum of PRIs in the human transcriptome7,15. Additionally, eCLIP studies have reported initially reported failure rates of 1 in 3 to 1 in 4 for their initial data deposits14 and a final tally showing a 1 in 2 failure rate.13

[00125] While UVC or chemical crosslinking is a required for identification of PRIs, these generate significant challenges for material handling and storage and preclude high-throughput PRI annotation and the use of automation. Crosslinking of PRIs generates exceedingly large and chemically diverse macromolecular complexes with large hydrophilic nucleic acid polymers attached to bulky protein complexes with hydrophobic properties. This leads to formation of insoluble aggregates with significant loss of material and require the use of strong surfactants that can interfere with enzymatic reactions needed for library construction. Most importantly, all currently available technologies to our knowledge are incompatible to standard liquid handling procedures used in automated library preparation, thus drastically limiting throughput, increasing cost, and escalating the potential for error or contamination. As a result of these limitations, currently approaches fundamentally severely limited in their ability to define the spectrum of dynamic or disease-associated PRIs that will only be identified through systematic, transcriptome-wide annotation of PRIs with all RBPs under multiple biologic conditions. Annotation of RNA Binding Protein Occupancy on RNA (ARORA) Methods of the Present Technology

[00126] In one aspect, the present disclosure provides a method for identifying one or more RNA binding protein (RBP)-RNA interaction (PRI) sites in at least one transcript derived from a biological sample including a plurality of cells, the method comprising: (a) cross-linking the plurality of cells present in the biological sample with short wavelength ultraviolet light to stabilize a plurality of RBP-RNA complexes comprising RNA transcripts, wherein each RBP-RNA complex corresponds to a PRI; (b) isolating RNA molecules from the cross-linked plurality of cells, wherein the RNA molecules comprise unbound RNA molecules and the RNA transcripts of the plurality of RBP-RNA complexes; (c) contacting the isolated RNA molecules with a protease under conditions that lyse the RBPs of the plurality of RBP-RNA complexes to yield RNA transcripts bound by residual peptides (e.g., a single amino acid), wherein each bound residual peptide (e.g., a single amino acid) (i) corresponds to a PRI site and (ii) comprises an alpha amino group at its N terminus and a carboxyl group at its C terminus; (d) labeling the RNA transcripts bound by the residual peptides (e.g., a single amino acid), by coupling the alpha amino group or the carboxyl group of each bound residual peptide with a chemical moiety conjugate comprising an affinity ligand; (e) fragmenting the RNA molecules in the presence of heat and divalent cations to generate RNA fragments having a 5’ end and a 3’ end, wherein the RNA fragments comprise unlabeled RNA fragments and RNA fragments labeled with the chemical moiety conjugate comprising the affinity ligand; (f) ligating a sequencing adapter to the 3’ end of the RNA fragments to generate adapter tagged RNA fragments; (g) capturing adapter tagged RNA fragments labeled with the chemical moiety conjugate comprising the affinity ligand using affinity purification; (h) reverse transcribing the adapter tagged labeled RNA fragments to generate a plurality of cDNA molecules, optionally wherein the adapter tagged labeled RNA fragments are attached to a solid surface or are in solution; (i) ligating a sequencing adapter to the 3’ end of the cDNA molecules to generate a plurality of adapter tagged cDNA molecules; (j) amplifying the plurality of adapter tagged cDNA molecules to generate a nucleic acid library of transcripts having RNA-RBP interactions; (k) sequencing and mapping the plurality of adapter tagged cDNA molecules of the nucleic acid library to identify individual transcripts; (1) identifying the frequency of cDNA 3’ ends at each nucleotide across the individual transcripts; and (m) detecting enriched RBP bound sites in the individual transcripts when the frequency of cDNA 3’ ends at specific nucleotides within the individual transcripts satisfies a predetermined threshold. In some embodiments, the plurality of cDNA molecules are reverse transcribed using a truncating reverse transcriptase (RT) or a read-through reverse transcriptase (RT). A truncating RT enzyme that truncates at the site of the PRI (i.e., the cross-link site) or a read-through RT enzyme that generates SNVs or indels at the site of the PRI (i.e., the cross-link site). Additionally or alternatively, in some embodiments, the methods disclosed herein further comprise identifying in the individual transcripts (i) short indels (e.g., 1-3 bps) that result from PRIs isolated by affinity purification, (ii) single nucleotide variants (SNVs) that result from PRIs isolated by affinity purification, or (iii) reverse transcriptase (RT) stop sites that result from PRIs isolated by affinity purification.

[00127] The biological sample may be derived from breast tissue, renal tissue, uterine cervical tissue, endometrium tissue, head or neck tissue, gallbladder tissue, parotid tissue, prostate tissue, brain tissue, pituitary gland tissue, kidney tissue, muscle tissue, esophageal tissue, stomach tissue, small intestine tissue, colon tissue, liver tissue, spleen tissue, pancreatic tissue, thyroid tissue, heart tissue, lung tissue, bladder tissue, adipose tissue, lymph node tissue, uterine tissue, ovarian tissue, adrenal tissue, testis tissue, tonsils, thymus, blood, hair, buccal, skin, serum, plasma, CSF, semen, prostate fluid, seminal fluid, urine, feces, sweat, saliva, sputum, mucus, bone marrow, lymph, or tears. In certain embodiments, the biological sample is a fresh or frozen sample.

[00128] In some embodiments, the cross-linking is performed with a UV box at a dose of about 50-600 mJ / cm2 for about 30s-10 minutes. In certain embodiments, the cross-linking is performed with a UV box at a dose of about 50 mJ / cm2, about 100 mJ / cm2, about 150 mJ / cm2, about 200 mJ / cm2, about 250 mJ / cm2, about 300 mJ / cm2, about 350 mJ / cm2, about 400 mJ / cm2, about 450 mJ / cm2, about 500 mJ / cm2, about 550 mJ / cm2, or about 600 mJ / cm2 for about 30s, 45s, 1 minute, 1.5 minutes, 2 minutes, 2.5 minutes, 3 minutes, 3.5 minutes, 4 minutes, 4.5 minutes, 5 minutes, 5.5 minutes, 6 minutes, 6.5 minutes, 7 minutes, 7.5 minutes, 8 minutes, 8.5 minutes, 9 minutes, 9.5 minutes, or about 10 minutes. In some embodiments, the cross-linking is performed with a UV box at a dose of 300-500 mJ / cm2 for about 3 minutes or 100-400 mJ / cm2 for about 10 minutes. In other embodiments, the cross-linking is performed with a UV laser for about 10s-20s.

[00129] Additionally or alternatively, in some embodiments, the isolated RNA molecules are treated with the protease at about 25°C-65°C for about 5 minutes-60 minutes. In certain -39- embodiments, the isolated RNA molecules are treated with the protease at about 25°C, about 27.5°C, about 30°C, about 32.5°C, about 35°C, about 37.5°C, about 40°C, about 42.5°C, 45°C, about 47.5°C, about 50°C, about 52.5°C, at about 55°C, about 57.5°C, about 60°C, about 62.5°C, or about 65°C for about 5 minutes, about 7.5 minutes, about 10 minutes, about 12.5 minutes, about 15 minutes, about 17.5 minutes, about 20 minutes, about 22.5 minutes, about 25 minutes, about 27.5 minutes, about 30 minutes, about 32.5 minutes, about 35 minutes, about 37.5 minutes, about 40 minutes, about 42.5 minutes, about 45 minutes, about 47.5 minutes, about 50 minutes, about 52.5 minutes, about 55 minutes, about 57.5 minutes, or about 60 minutes. In some embodiments, the isolated RNA molecules are treated with the protease at 50°C for about 30 minutes. In certain embodiments, the protease is proteinase K. Additionally or alternatively, in some embodiments, the length of each bound residual peptide is no more than 10 amino acids in length, no more than 9 amino acids in length, no more than 8 amino acids in length, no more than 7 amino acids in length, no more than 6 amino acids in length, no more than 5 amino acids in length, no more than 4 amino acids in length, no more than 3 amino acids in length, no more than 2 amino acids in length, or is a single amino acid.

[00130] Additionally or alternatively, in some embodiments, the methods of the present technology further comprise removing genomic DNA after step (c) to obtained purified RNA molecules. In certain embodiments, the methods disclosed herein further comprise contacting the purified RNA molecules with a protease at 50°C for about 30 minutes prior to step (d), optionally wherein the protease is proteinase K.

[00131] In any of the preceding embodiments, the methods of the present technology further comprise enriching the isolated or purified RNA molecules with poly-dT oligonucleotides or targeted bait capture reagents after step (c), but prior to step (d). In other embodiments, the methods of the present technology further comprise enriching the adapter tagged cDNA molecules with targeted bait capture reagents after step (j), but prior to step (k). In certain embodiments, the targeted bait capture reagents hybridize to one or more disease associated genes, such as genes associated with cancer, neurodegenerative disease, autoimmune disease, metabolic disease etc. Examples of disease associated genes include, but are not limited to, AKT1, ALK, APC, AR, ARAF, ARID 1 A, ARID2, ATM, B2M, BCL2, BCOR, BRAF, BRCA1, BRCA2, CARD11, CBFB, CCND1, CDH1, CDK4, CDKN2A, CIC, CREBBP, CTCF, CTNNB1, DICERI, DIS3, DNMT3A, EGFR, EIF1AX, EP300, ERBB2, ERBB3, ERCC2, ESRI, ETV6, EZH2, FBXW7, FGFR1, FGFR2, FGFR3, -40- FGFR4, FLT3, F0XA1, F0XL2, F0X01, FUBP1, GATA3, GNA11, GNAQ, GNAS, H3F3A, HIST1H3B, HRAS, IDH1, IDH2, IKZF1, INPPL1, JAKI, KDM6A, KEAP1, KIT, KNSTRN, KRAS, MAP2K1, MAPK1, MAX, MED12, MET, MLH1, MSH2, MSH3, MSH6, MTOR, MYC, MYCN, MYD88, MYODI, NF1, NFE2L2, NOTCH1, NRAS, NTRK1, NTRK2, NTRK3, NUP93, PAK7, PDGFRA, PIK3CA, PIK3CB, PIK3R1, PIK3R2, PMS2, POLE, PPP2R1 A, PPP6C, PRKCI, PTCHI, PTEN, PTPN11, RAC1, RAFI, RBI, RET, RHOA, RIT1, R0S1, RRAS2, RXRA, SETD2, SF3B1, SMAD3, SMAD4, SMARCA4, SMARCB1, S0S1, SPOP, STAT3, STK11, STK19, TCF7L2, TGFBR1, TGFBR2, TP53, TP63, TSC1, TSC2, U2AF1, VHL, XPO1, and TERT.

[00132] In any of the above embodiments, the methods of the present technology further comprise depleting ribosomal RNA from the isolated or purified RNA molecules after step (c), but prior to step (d).

[00133] In any and all embodiments of the methods disclosed herein, labeling the RNA transcripts bound by the residual peptides comprises heat denaturing the isolated or purified RNA molecules, and incubating the isolated or purified RNA molecules with the chemical moiety conjugate comprising the affinity ligand in a reaction buffer supplemented with about 5%-70% DMSO. In some embodiments, the reaction buffer is supplemented with about 5%, about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, or about 70% DMSO. Examples of suitable reaction buffers include, but are not limited to, phosphate buffers (e.g., PBS buffer), carbonate buffers (e.g., carbonate-bicarbonate buffer), HEPES, borate, 2-[morpholino]ethanesulfonic acid (MES), or any buffer that does not contain free amine or free carboxylate groups.

[00134] Additionally or alternatively, in some embodiments, heat denaturing comprises heating the isolated or purified RNA molecules at about 50°C-100°C for 1-5 minutes. In certain embodiments, heat denaturing comprises heating the isolated or purified RNA molecules at about 50°C, about 55°C, about 60°C, about 65°C, about 70°C, about 75°C, about 80°C, about 85°C, about 90°C, about 95°C, or about 100°C for 1, 2, 3, 4, or 5 minutes.

[00135] In some embodiments of the methods disclosed herein, the chemical moiety conjugate is coupled to the alpha amino group of each bound residual peptide. In certain embodiments, the chemical moiety conjugate comprises a succinimidyl ester, a carboxylic ester, a tetrafluorophenyl ester, a sulfodichlorophenol ester, a carbonyl azide, an aldehyde, or an isothiocyanate. In other embodiments of the methods disclosed herein, the chemical moiety conjugate is coupled to the carboxyl group of each bound residual peptide. In some embodiments, the chemical moiety conjugate comprises a primary amine, a carboiimide, or an isocyanate. In any and all embodiments of the methods disclosed herein, the chemical moiety conjugate comprising the affinity ligand is added to a final concentration of about 0.1 mM to about 10 mM and incubated for about 3 minutes to about 5 hours. In certain embodiments, the chemical moiety conjugate comprising the affinity ligand is added to a final concentration of about 0.1 mM, about 0.2 mM, about 0.3 mM, about 0.4 mM, about 0.5 mM, about 0.6 mM, about 0.7 mM, about 0.8 mM, about 0.9 mM, about 1 mM, about 1.1 mM, about 1.2 mM, about 1.3 mM, about 1.4 mM, about 1.5 mM, about 1.6 mM, about 1.7 mM, about 1.8 mM, about 1.9 mM, about 2 mM, about 2.1 mM, about 2.2 mM, about 2.3 mM, about 2.4 mM, about 2.5 mM, about 2.6 mM, about 2.7 mM, about 2.8 mM, about 2.9 mM, about 3 mM, about 3.1 mM, about 3.2 mM, about 3.3 mM, about 3.4 mM, about 3.5 mM, about 3.6 mM, about 3.7 mM, about 3.8 mM, about 3.9 mM, about 4 mM, about 4.1 mM, about 4.2 mM, about 4.3 mM, about 4.4 mM, about 4.5 mM, about 4.6 mM, about 4.7 mM, about 4.8 mM, about 4.9 mM, about 5 mM, about 5.1 mM, about 5.2 mM, about 5.3 mM, about 5.4 mM, about 5.5 mM, about 5.6 mM, about 5.7 mM, about 5.8 mM, about 5.9 mM, about 6 mM, about 6.1 mM, about 6.2 mM, about 6.3 mM, about 6.4 mM, about 6.5 mM, about 6.6 mM, about 6.7 mM, about 6.8 mM, about 6.9 mM, about 7 mM, about 7.1 mM, about 7.2 mM, about 7.3 mM, about 7.4 mM, about 7.5 mM, about 7.6 mM, about 7.7 mM, about 7.8 mM, about 7.9 mM, about 8 mM, about 8.1 mM, about 8.2 mM, about 8.3 mM, about 8.4 mM, about 8.5 mM, about 8.6 mM, about 8.7 mM, about 8.8 mM, about 8.9 mM, about 9 mM, about 9.1 mM, about 9.2 mM, about 9.3 mM, about 9.4 mM, about 9.5 mM, about 9.6 mM, about 9.7 mM, about 9.8 mM, about 9.9 mM, or about 10 mM and incubated for about 3 minutes, about 4 minutes, about 5 minutes, about 6 minutes, about 7 minutes, about 8 minutes, about 9 minutes, about 10 minutes, about 15 minutes, about 20 minutes, about 25 minutes, about 30 minutes, about 35 minutes, about 40 minutes, about 45 minutes, about 50 minutes, about 55 minutes, about 1 hour, about 1.1 hours, about 1.2 hours, about 1.3 hours, about 1.4 hours, about 1.5 hours, about 1.6 hours, about 1.7 hours, about 1.8 hours, about 1.9 hours, about 2 hours, about 2.1 hours, about 2.2 hours, about 2.3 hours, about 2.4 hours, about 2.5 hours, about 2.6 hours, about 2.7 hours, about 2.8 hours, about 2.9 hours, about 3 hours, about 3.1 hours, about 3.2 hours, about 3.3 hours, about 3.4 hours, about 3.5 hours, about 3.6 hours, about 3.7 hours, about 3.8 hours, about 3.9 hours, about 4 hours, about 4.1 hours, about 4.2 hours, about 4.3 hours, about 4.4 hours, about 4.5 hours, about 4.6 hours, about 4.7 hours, about 4.8 hours, about 4.9 hours, or about 5 hours.

[00136] In any and all embodiments of the methods disclosed herein, the affinity ligand comprises biotin, a biotin derivative, digoxin, dinitrophenyl, a click chemistry reagent, sugars or peptides. In some embodiments, the click chemistry reagent comprises an alkyne group, an azide group, a dibenzocyclooctyne group (DBCO), or a bicyclononyne (BCN) group.

[00137] Additionally or alternatively, in some embodiments, fragmenting the RNA molecules comprises heating the RNA molecules at 90°C-95°C for about 3-12 minutes in the presence of about 10mM-30mM divalent cations (e.g., Mg2+, Zn2+). In certain embodiments, fragmenting the RNA molecules comprises heating the RNA molecules at 90°C, 91°C, 92°C, 93°C, 94°C or 95°C for about 3 minutes, about 4 minutes, about 5 minutes, about 6 minutes, about 7 minutes, about 8 minutes, about 9 minutes, about 10 minutes, about 11 minutes, or about 12 minutes in the presence of about lOmM, about 15mM, about 20mM, about 25mM, or about 30mM divalent cations (e.g., Mg2+, Zn2+). The RNA fragments may be about 40-90 nucleotides in length or about 80-250 nucleotides in length.

[00138] In any and all embodiments of the methods disclosed herein, the biological sample is obtained from a subject diagnosed with a disease. In some embodiments, the disease is cancer, autoimmune disease, metabolic disease or neurodegenerative disease. Examples of cancer include, but are not limited to, adrenal cancers, bladder cancers, blood cancers, bone cancers, brain cancers, breast cancers, carcinoma, cervical cancers, colon cancers, colorectal cancers, corpus uterine cancers, ear, nose and throat (ENT) cancers, endometrial cancers, esophageal cancers, gastrointestinal cancers, head and neck cancers, Hodgkin's disease, intestinal cancers, kidney cancers, larynx cancers, leukemias, liver cancers, lymph node cancers, lymphomas, Diffuse large B-cell lymphoma (DLBCL), lung cancers, melanomas, mesothelioma, myelomas, nasopharynx cancers, neuroblastomas, nonHodgkin's lymphoma, oral cancers, ovarian cancers, pancreatic cancers, penile cancers, pharynx cancers, prostate cancers, rectal cancers, sarcoma, seminomas, skin cancers, stomach cancers, teratomas, testicular cancers, thyroid cancers, uterine cancers, vaginal cancers, vascular tumors, and metastases thereof.

[00139] Examples of neurodegenerative disease include, but are not limited to, age-associated memory impairment (AAMI), mild cognitive impairment (MCI), Alzheimer's disease, Down's syndrome, dementia pugilistica, cognitive dysfunction syndrome, multiple system atrophy, inclusion body myositosis, hereditary cerebral hemorrhage with amyloidosis of the Dutch type, Nieman-Pick disease type C, cerebral P-amyloid angiopathy, dementia associated with cortical basal degeneration, the amyloidosis, Creutzfeldt-Jakob disease, Gerstmann-Straussler syndrome, kuru, scrapie, Huntington’s disease, Parkinson’s disease, ataxia, Motor neuron disease, Progressive supranuclear palsy, and Amyotrophic lateral sclerosis (ALS).

[00140] Examples of autoimmune disease include, but are not limited to, vasculitis, e.g., Anti- neutrophil cytoplasm antibodies (ANCA), ANCA-associated vasculitis (AAV) or giant cell arteritis (GCA) vasculitis, Sjogren's syndrome, inflammatory bowel disease (IBD), Pemphigus vulgaris, lupus nephritis, psoriasis, thyroiditis, Type I Diabetes, Idiopathic thrombocytopenic purpura (ITP), Ankylosing spondylitis, Multiple sclerosis, systemic lupus erythematosus (SLE), rheumatoid arthritis, Crohn's disease, Myasthenia Gravis, neuromyelitis optica (NMO), IgG4-related disease, systemic sclerosis, insulindependent diabetes mellitus (IDDM), akylosing spondylitis, atopic dermatitis, uveitis, and Graft-versus- host disease (GVHD).

[00141] Examples of metabolic disease include, but are not limited to, Familial hypercholesterolemia, Gaucher disease, Hunter syndrome, Krabbe disease, Maple syrup urine disease, Metachromatic leukodystrophy, Mitochondrial encephalopathy lactic acidosis stroke-like episodes (MELAS), Niemann-Pick, Phenylketonuria (PKU), Porphyria, Tay-Sachs disease and Wilson's disease.

[00142] In any and all embodiments of the methods disclosed herein, one or more of the enriched RBP bound sites detected in the individual transcripts are allele-specific and / or associated with a disease (e.g., cancer, neurodegenerative disease, autoimmune disease, metabolic disease etc^. In any of the foregoing embodiments, the methods of the present technology further comprise determining the identity of a RBP bound to one or more of the enriched RBP bound sites detected in the individual transcripts. In some embodiments, the identity of a RBP bound to one or more of the detected enriched RBP bound sites in the individual transcripts is determined by identifying RBP-specific crosslinking patterns within a sequence motif in the individual transcripts (e.g, via computational analysis). In some embodiments, the crosslinking patterns of RBPs are identified using eCLIP defined consensus motifs, mCross crosslinking patterns, positionally enriched k-mer analysis (PEKA), or sequence motifs defined by RBNS.

[00143] In one aspect, the present disclosure provides a method for evaluating the efficacy of a drug in inhibiting interactions between a specific RNA binding protein (RBP) and RNA transcripts comprising (a) contacting a first cell population with at least one inhibitory oligonucleotide that specifically inhibits expression of a RBP, wherein the first cell population is obtained from a biological sample; (b) contacting a second cell population with a test drug, wherein the second cell population is obtained from the biological sample and is not treated with the at least one inhibitory oligonucleotide; (c) performing the ARORA methods described herein to identify one or more RBP-RNA interaction (PRI) sites in at least one transcript derived from the first cell population, the second cell population and a control cell population, wherein the control cell population is obtained from the biological sample and is not treated with the at least one inhibitory oligonucleotide and the test drug; (d) defining RBP-specific PRIs by identifying crosslink patterns within one or more sequence motifs in the individual transcripts that (i) are present in the control cell population and (ii) absent in the first cell population; and determining that the test drug is effective when the RBP-specific PRIs identified in step (d) are absent in the second cell population. In another aspect, the present disclosure provides a method for evaluating the specificity of a drug in inhibiting interactions between a specific RNA binding protein (RBP) and RNA transcripts comprising (a) contacting a first cell population with at least one inhibitory oligonucleotide that specifically inhibits expression of a RBP, wherein the first cell population is obtained from a biological sample; (b) contacting a second cell population with a test drug, wherein the second cell population is obtained from the biological sample and is not treated with the at least one inhibitory oligonucleotide; (c) performing the ARORA methods described herein to identify one or more RBP-RNA interaction (PRI) sites in at least one transcript derived from the first cell population, the second cell population and a control cell population, wherein the control cell population is obtained from the biological sample and is not treated with the at least one inhibitory oligonucleotide and the test drug; (d) identifying the non-specific effects of the test drug by identifying crosslink patterns within one or more sequence motifs in the individual transcripts that (a) are present in the control cell population and the first cell population and (b) absent in the second cell population. The at least one inhibitory oligonucleotide may be a siRNA, a shRNA, an antisense oligonucleotide, or a sgRNA. In some embodiments, the test drug is a small molecule, a nucleic acid or a peptide. EXAMPLES

[00144] The present technology is further illustrated by the following Examples, which should not be construed as limiting in any way. Example 1: Materials and Methods

[00145] Cell lines and treatments

[00146] HepG2, K562, 293T, and BJ Fibroblasts were obtained from ATCC and grown in recommended base media with 10% FBS. Cycloheximide and Anisomycin were obtained from Sigma Aldrich and used at lOOpM concentration in cell treatments for 1 hour to induced translation stress.

[00147] RNA preparation

[00148] Mammalian cells, HEK293T, HepG2, or K562 cells, were grown in 15-cm dishes until a maximum of 80-90% confluency was reached. Cells were washed twice with PBS and cross-linked by UV box (Spectrolinker #XL-1000) at 254 nm with a dose of 100400 mJ / cm2. Immediately after cross-linking, cells were scraped and pelleted. In non-cross-linked controls, cells were directly pelleted and stored at -80°C before use.

[00149] RNA was extracted according to Monarch Total RNA Miniprep Kit (NEB #T2010) with some modifications. Briefly, 600 pL of IX Protection Reagent, 60 pl of Proteinase K reaction buffer, and 30 pl of Proteinase K were added to the samples and incubated at 50°C for 30 min. Samples were then processed according to the protocol to remove genomic DNA with a DNA removal column followed by RNA cleanup and elution. The eluted RNA was then treated once again with Proteinase K at 50°C for 30 minutes, followed by RNA cleanup using the Zymo RNA Clean and Concentrator-5 kit, retaining transcripts >200nt in length. For transcriptome experiments, poly-A tailed RNA was enriched with two sequential enrichments using poly-A purist (Thermo Fisher #AM1922) and followed by ribosomal RNA depletion using the Ribo-depletion kit (Galen Molecular #dp-K024-000054).

[00150] Labeling RNA with N-hydroxysuccinimide (NHS) conjugates

[00151] 1 pg of total RNA or ribo-depleted RNA was heated at 70°C for 2 min and snapped cooled on ice. RNA was diluted in labeling buffer consisting of IX PBS pH 7.4 supplemented with 50% DMSO to limit RNA secondary structure formation which could impair labeling of some crosslinked sites. NHS-dPEG-biotin (Sigma-Aldrich #QBD 10200) added to a ImM final concentration and the reaction was incubated at room temperature for -46- 30 min. Reactions were stopped by Tris and clean-up of RNA was performed by RNeasy Mini Kit (Qiagen #74104).

[00152] RNA Dot blot

[00153] 50 ng of RNA was blotted on the Hybond-N+ membrane (Thermo Fisher #45 000-763). Membranes were then cross-linked by UV at 254 nm with a dose of 120 mJ / cm2 and blotted by IRDye 800CW Streptavidin for Ihr before imaging on LI-COR Odyssey CLx.

[00154] RNA fragmentation, size selection, and dephosphorylation

[00155] 1 pg of Biotinylated RNA per sample in 20 pL 2X 10X Fast AP Buffer (ThermoFisher) is incubated on a thermocycler block for at 94C for 9 minutes to obtain the majority of RNA fragments in the range of 40-100 nucleotides length. RNA fragments then undergo 2 step size selection using AMPure RNAClean XP bead. In the first step, 22 pL of RNAClean XP beads and 12 pL of isopropanol are added. To the flow-through (enriched in fragments < -lOOnucleotides), 25 pL of RNAClean XP beads and 41 pL of isopropanol are added to bind fragments great than 40 nucleotides in length to the beads. Elute, size selected RNA fragments are then treated with FastAP alkaline phosphatase and T4 PNK to dephosphorylate the fragments. Dephosphorylated fragments once again undergo cleanup with RNAClean XP beads and are eluted in 10 pL H2O, RNA content quantified by Qubit, and 1 pL of Superasin added prior to storage at -80C.

[00156] RNA 3 ’ Ligation and Streptavidin Pulldown

[00157] 2 biological replicates of non-crosslinked samples and 2 biological replicates of 400J / m2 crosslinked samples are used for each condition. 200ng of RNA for each sample and lOOpmol each of 2 distinct indexed 3’oligonucleotide are denatured on a heat block at 70C for 2 minutes, followed by snap cooling on ice. T4 RNA ligase 1 mixture is added and the ligation reaction incubated at room temperature for 90 minutes. Ligation reactions are stopped by the addition of RLT buffer and EDTA, then individual, uniquely indexed samples from all four samples are pooled and then undergo cleanup with RNAClean XP beads and elution in 20 pL of H2O. 2 pL of each sample was reserved as input sample and stored on ice until cDNA synthesis.

[00158] MyOne-Cl Streptavidin beads (ThermoFisher, 5 pL beads / 200ng RNA) are washed multiple times with Bind and Wash (B&W) buffer and resuspended in 600 pL of B&W buffer with 3 pL Superasin on ice. 18 pL of 3’ Ligated RNA is denatured at 70C for -47- 2 minutes on a heat block, followed by snap cooling on ice. The denatured RNA is added the MyOne-Cl beads and incubated with end-over-end rotation at 4C for 1 hour. RNA-bound paramagnetic beads are sedimented in a magnetic field and the supernatant discarded. The beads are washed 3 times with ice cold B&W buffer, then 1 time with 5mM Tris 7.5.

[00159] cDNA Synthesis

[00160] Reverse transcription (RT) was performed using Maxima H- reverse transcriptase (ThermoFisher) or AffinityScript (Agilent). RNA-bound paramagnetic beads were resuspended in 20 pL Maxima H- reverse transcriptase reaction buffer or AffinityScript buffer, AR2 primer, and 0.3 pL of Murine RNase inhibitor. Maxima H-reverse transcriptase reaction mix was separated added to input samples to a final volume of 20 pL. Beads were vigorously mixed to ensure they were suspended, then incubated at 55C on a ThermoMixer block (Eppendorf) for 40 minutes, with a 10 second pulse of mixing at 1200rpm every 10 minutes to keep beads in suspension. After 40 minutes of reverse transcription, in pulldown samples with RNA-bound to beads, beads were placed in the magnet, the RT reaction mix was removed, and 200 pL of 5mm Tris 7.5 was added to wash the beads once, then beads were resuspended in 20 pL 5mM Tris 7.5. Next, to all samples (pulldown and inputs), IN NaOH was added and samples incubated on the ThermoMixer at 70C for 12 minutes with constant mixing at 1200rpm to degrade RNA. Paramagnetic beads were subsequently removed in the magnet. cDNA samples underwent cleanup with AMPure beads.

[00161] Ligation of second sequencing adapter to cDNA 3 ’ end

[00162] To each purified cDNA sample, 40pmol of the 3Tr3-10N adapter was added and the mix denatured at 75C for 3 minutes on a heat block, followed by snap cooling. To this, the T4 RNA ligase 1 mixture is added and the ligation reaction incubated at room temperature overnight. The next day, the cDNA undergoes two sequential clean-ups with AMPure beads and then is eluted in 20 pL of 5mM Tris 7.5.

[00163] Library quantification, amplification, and size selection

[00164] 1 pL of each library was diluted 1:10 with nuclease free water and qPCR was performed with 1 pL of diluted cDNA each well in triplicate to obtain the cycle threshold (Ct), a estimate of fragments of library obtained. We typically obtain more that >20X enrichment of crosslinked samples upon pulldown compared to enrichment of noncrosslinked samples. 10 pL of the cDNA libraries were then amplified with barcode Illumine adapters to total of Ct - 3 cycles. AMPure clean-up of libraries was performed to remove excess primer and then size selected in 4% agarose gel to collect fragments from 200-350bp. Libraries were analyzed by Bioanalyzer for molarity and fragment length analyses before sequencing.

[00165] RNA sequencing

[00166] ARORA libraries were sequenced with Illumina PE 100 to a total of 5 million total reads per library for analyses of Ribosomal RNA interactions. ARORA libraries for transcriptome-wide analyses of RBP-RNA interactions were sequenced to a total of 400 million reads per library.

[00167] Computational Analyses

[00168] Paired-end sequencing reads contained two types of molecular barcodes. First, readl contained specific 9bp nucleotide sequences for multiplexing purposes. Secondly, readl and read2 both contained unique molecular identifiers (UMIs) (random oligonucleotide sequences) which were used to uniquely identify input RNA molecules. We used UMI-tools to extract 9bp multiplexing barcodes as well as 4bp (readl) and Ibp (read2) deduplication barcode sequences and appended them to read names. Additionally,^ / ^ / ? was used for adapter and polyG tail trimming. Subsequent to extraction and trimming, paired end reads were aligned to the human reference genome (hgl9) using STAR aligner.

[00169] Each of the initial 4 experiments (2 cell lines, pulldown and input each) was then split into 8 different crosslinking conditions (4 conditions, 2 replicates each) using bamtools filter resulting in a total of 32 BAM files, one for each cell line and condition. Each of the 32 BAM files was deduplicated based on the 14bp UMI (readl + read2) using UMI-tools.

[00170] Finally, we identified single nucleotide deletions, SNVs, and reverse transcriptase (RT) stop sites resulting from PRIs isolated by affinity purification based on CIGAR information contained in the BAM files. Example 2: Crosslinked Oligopeptides as a Specific Attachment Chemistry for Affinity Purification of RBP-RNA Interaction Sites

[00171] UVC crosslinking of RBP-RNA complexes generates irreversible crosslinks at the precise nucleotides that directly contact amino acids6. While intact Protein-RNA complexes are recovered by immunoprecipitation in approaches such as CLIP, we reasoned that the orthogonal chemistry of amino acids could be leveraged to identify PRIs that is universal and specific to amino acids, regardless of the identity of the specific amino acid interacting with the RNA (FIG. 1A). Here, UVC crosslinked protein-RNA complexes are treated with extensive proteolysis to remove the vast majority of the amino acid content of the crosslinked RNA, leaving only minimal residual peptide (or single amino acid) but exposing one alpha amino group at the N-terminus of the peptide and one carboxyl groups at the C terminus of each peptide, regardless of the amino acid identity. Endogenous RNA, including known post-transcriptional chemical modifications, contain no reactive, free amino or carboxyl groups. As a result, this method generates orthogonal chemical moieties on RNA at the precise site of UVC crosslinked PRIs. Furthermore, since the vast majority of protein content is liberated from RNA by proteolysis, the resulting RNA has identical properties as standard non-crosslinked RNA. We have observed that UVC crosslinked, NHS-biotinylated RNA can undergo standard handling for RNA-seq library preparation, including isolation and cleanup with paramagnetic beads, automated liquid handling, consistent and tunable fragmentation with divalent cations and heat, enrichment with poly-dT oligos or targeted hybridization capture, or ribosomal RNA depletion. As a result, this approach generates orthogonal chemical moieties at sites of any PRI and is predicted to be compatible with any commercial pipeline for RNA-seq library preparation.

[00172] ARORA exploits the orthogonal chemistry generated by free alpha amino or carboxyl groups on oligopeptides crosslinked to RNA for site-specific attachment of chemical labels or affinity ligands. N-hyroxysuccinimide (NHS) esters are commonly used for orthogonal attachment chemistry in proteins and react selectively with free alpha or amino groups atN-termini or epsilon amino groups in lysine side chains, but react poorly with RNA or DNA. Optimal labeling was achieved after heat denaturation of RNA followed by incubation of RNA with NHS-biotin at room temperature in PBS buffer supplemented with 50% DMSO to reduce RNA secondary structure during labeling and by enhancing the nucleophilic attack of the NH2 of the amino acid or peptide towards NHS-biotin. NHS-biotin extensively labeled RNA isolated from UVC crosslinked cells, but did not label RNA isolated from mock-irradiated cells (FIG. IB). Importantly, we observed no labeling of RNA crosslinked in vitro with UVC in the absence of proteins, confirming that labeling of free amino groups by NHS-biotin on UVC crosslinked RNA requires amino groups derived from proteins and is not due to generation of free amino groups by UVC-induced damage of RNA nucleotides. RNA from UVC irradiated cells labeled with NHS-biotin is not degraded compared to non-irradiated RNA, unlabeled RNA (RIN scores of 9-10) and RNA gel electrophoresis revealed that biotin labeled RNA existed in a wide range of molecular weights that resembled the typical pattern of total RNA, with high signal corresponding to abundant 28S and 18S ribosomal subunits and lower intensity labeling at molecular weights corresponding to mRNAs (FIG. IC). The degree of RNA labeling by NHS linkers was directly proportion to UV-C dose delivered (FIG. ID). Streptavidin paramagnetic beads retrieve biotinylated RNA from various classes of cellular RNA including the large subunit of ribosome (28S), GAPDH mRNA, and the long noncoding RNA MALAT1 in a UVC dose-dependent manner, where increasing doses of UVC crosslinking result in sequentially increasing yields of each RNA, consistent with increased efficiency of PRI crosslinking with increasing UVC doses (FIG. IE). Thus, NHS attachment chemistries efficiently and specifically label amino acid moieties covalently crosslinked to RNA by UVC.

[00173] We have also found that attachment chemistry targeting the carboxyl group of the C-Terminus also permits highly specific labeling of peptides crosslinked to RNA (FIG. 4A). The chemical l-ethyl-3-(3-dimethylaminorpropyl)carbodiimide (EDC) can activate free carboxyl groups and facilitate attachment of amino-labeled chemicals. RNA isolated from cells after crosslinking with UVC was readily labeled with biotin in a reaction containing EDC and amino-PEG-biotin, whereas minimal labeling of RNA isolated from non-crosslinked cells or RNA that was crosslinked by UVC in vitro in the absence of protein. Therefore, these data demonstrate that both the primary amine at the N-terminus of peptides and the carboxyl group at the C-terminus of the peptide provide unique, orthogonal chemical moieties for labeling of amino acids crosslinked to RNA. Example 3: ARORA Identifies Protein-RNA Interactions at Single Nucleotide Resolution Across Transcripts

[00174] Enrichment of NHS-biotin labeled RNA by streptavidin beads could identify regions of cellular RNAs with high densities of RBP interactions, however identification of PRIs with single nucleotide resolution would offer far greater information about the position, identity, and RNA motifs that determine RBP interactions with RNA. CLIP approaches have previously demonstrated that different reverse transcriptase’s have characteristic behaviors around the site of PRIs which can distinguish the PRI with single nucleotide resolution6,16. Some RTs have a high propensity for crosslink-induced truncations, where the final nucleotide incorporated by the RT enzyme typically corresponds to the nucleotide crosslinked in the PRI. With other RT enzymes, crosslink-induced mutations such as deletions or substitutions occur with high frequency near PRIs. Since ARORA utilized extensive protease treatment followed by covalent addition of a bulky adduct, PRIs marked by ARORA have a chemically distinct amino acid-nucleic acid structures compared to RNA fragments typically sequenced in CLIP. We therefore examined the impact of ARORA labeling of nucleic acids on crosslink-induced effects on various RTs. We generated an in vitro synthesized oligonucleotide template with a single Amino-C2-deoxythymidine residue (amino-C2dT), which is chemically similar to a single lysine amino acid crosslinked to thymidine or uracil, providing a single internal site for NHS-labeling of the oligonucleotide. After ARORA labeling of the oligonucleotide and streptavidin purification, a primer extension assay with four unique RT enzymes revealed two distinct patterns of RT behavior around the amino-C2dT residue. Three RT enzymes (Maxima H-, SuperScript IV, and TGIRT) each results in -50% RT-generated cDNAs with a molecular weight of-70 nucleotides, consistent with truncation of cDNA synthesis at the amino-C2dT residue, with 10-20% of fragments with molecular weight of ~95-98nt, consistent with a 2-5 nucleotide deletion, while 30% of cDNAs were full-length (lOOnt) (FIG. IF). Importantly, highly efficient truncations and mutations were only appreciated when cDNA was synthesized from RNA fragments bound to beads, while fragments in free solution had substantially lower rates of both types of crosslink-induced effect. Without wishing to be bound by theory, it is believed that immobilization of the crosslinked nucleic acid increases steric hinderance for the RT enzymes and prevents processivity through the crosslinked site. In contrast to the three enzymes with high efficiency RT stops and deletions, one enzyme, AffinityScript, had no appreciable truncated cDNAs or cDNAs with deletions, either with cDNA synthesis in solution or bound to beads, consistent with a read-through pattern of reverse transcription.

[00175] Sequencing of ARORA libraries (workflow in FIG. 4B) generated by Maxima H- RT using total cellular RNA confirmed that RNA isolated from cells crosslinked at 400mJ / cm2 and enriched by streptavidin beads resulted in 0.08 deletions per read and 0.0025 substitutions per read on average, significantly more that RNA from non-crosslinked cells and RNA from crosslinked cells that was not enriched by streptavidin beads (FIG. 1G). No appreciable deletions or substitutions were observed in libraries generated using AffinityScript RT. As a result, ARORA libraries generated using Maxima H- RT are suitable for analyses of PRIs identified by crosslink-induced truncations and mutations while ARORA libraries generated with AffinityScript can be utilized in scenarios where read-through of the crosslink site is desirable (for example, in the identification of allelespecific PRIs).

[00176] UVC crosslinks nearly all amino acids with RNA but a bias favoring crosslinking to pyrimidines over purines, U>OG / A. Consistent with this, we observed an enrichment of U>C>G>A at the site of significantly enriched RT stops, deletions, or substitutions (FIG. 4C). Interestingly, we observed a slight enrichment for RT stops at A residues in both NCLPD and x400_PD samples compared to their inputs, suggesting that NHS-labels under these conditions may have modest reactivity to C6 amino group of A in the absence of PRIs, but this enrichment is compensated for when calculating per nucleotide enrichment of X400 PD samples compared to NCL PD samples (FIGs. 4D-4E).

[00177] Optimal fragmentation of RNA is critical for elucidation of PRIs at single nucleotide resolution at is observed with RNA fragment sizes of 40-90 nucleotides in length17. cDNA generated from larger RNA fragments has substantially higher rates of cDNA truncations at non-crosslinked sites and thus significantly reduces that ability to achieve single nucleotide resolution of PRIs. However, transcriptome-wide approaches such as CLIP utilized enzymatic fragmentation of RNA in order maintain protein integrity which is critical for efficient IP with antibodies. This approach unfortunately introduces significant bias due to RNases favoring single or double stranded secondary structures, selectivity towards cutting as specific bases, and limited accessibility of certain regions of RNA at biologic temperatures due to secondary and tertiary RNA structure14. In contrast, RNA generated by ARORA consists of a minimal amino acid adduct on RNA that is essentially irreversible and highly stable to heat and denaturants. ARORA labels on RNA remain intact in chaotropic agents, high salt concentrations, or after boiling for >20 minutes. As a result, ARORA labeled RNA can be fragmented using heat and divalent cations, a method that is free of structure or base biases and can generate RNA fragment size distributions in predictable and reproducible manner (FIG. 5A). We performed ARORA on total RNA from HepG2 cells using Affinity Script RT or Maxima H- RT under conditions of long RNA fragments (4 min fragmentation) and short RNA fragments (9 min fragmentation) (FIG. 5B). Even in libraries generated with short fragments, libraries generated after cDNA synthesis on beads in pulldown samples were significantly truncated compared to libraries generated freely in solution from the same material, confirming significant enrichment of crosslink-induced truncated cDNAs in the pulldown libraries (FIG. 5C).

[00178] In libraries generated from total cellular RNA, >95% of all reads mapped to ribosomal RNA under all conditions, as expected. Analyzing the distribution of RT truncation sites (RT stops), we observed significant differences between libraries generated using truncating RT versus those with read-through RTs, as well as a significant impact of fragmentation times in the ability to distinguish library conditions (FIG. 6A). Libraries generated using RNA isolated from UVC crosslinked cells and enriched by streptavidin bead pulldown (x400_PD) were highly correlated within conditions (R2 ~0.9), indicating that ARORA under all conditions was highly reproducible (FIG. 6B). However, the three different library conditions had substantially distinct ability to discern x400_PD samples from the three control libraries: non-crosslinked, streptavidin enriched (NCLPD) and input (NCL_IN and x400_IN) (FIG. 6C). Libraries generated with Maxima H- significantly reduced the correlation of RT stops between x400_PD and control libraries compared to libraries generated with AffinityScript. These results are consistent with in vitro observations that cDNA generated with AffinityScript are not significantly truncated by crosslinked sites. Furthermore, libraries generated using shorter RNA fragments (40-90nt size after 9 minutes fragmentation) significantly reduced correlations between x400_PD and control libraries compared to libraries generated from larger RNA fragments (80-250nt size after 4 minutes of fragmentation). Therefore, specific RT conditions and small, consistent RNA fragment sizes optimize the ability of ARORA to discriminate the conditions under which crosslinked PRIs can be discerned by streptavidin enrichment of RNA from crosslinked cells (x400_PD condition).

[00179] We next sought to benchmark ARORA identification of PRIs (see model in FIG. 1H) to existing methods. RNP-MaP was previously developed in the Weeks lab to identify the sites of PRIs on single RNAs8. RNP-MaP utilizes a heterobifunctional crosslinker in which an amino reactive linker covalently binds protein while a diazirine moiety crosslinks to nucleic acids upon exposure to UVA light. In RNP-MaP, PRIs are identified by adduct-induced mutations, that occur at a rate of approximately which occur at a background rate of-1 per 10'3 in control samples and increase to 2-5 per 10'3 with in vivo crosslinking of protein and RNA. We compared PRIs identified by ARORA with PRIs identified by RNP-MaP in the RNA component of the ribonucleoprotein complex RNase MRP (RMRP). ARORA agrees significantly at the precise nucleotides predicted to be sites of RNA-protein interactions defined by RNP-MaP (FIGs. II, FIGs. 7A-7B, P=0.0035, Fisher Exact). We additionally analyzed these data relaxing the predicted binding site by 1 nucleotide in either direction to account for the fact that each approach has a distinct nucleotide preferences for crosslinking and that RNP-MaP can crosslink more distant nucleotides due to a 7nm spacer arms in the heterobifunctional crosslinker. After relaxing the PRI annotations by a single nucleotide, we obtained substantially increase concordance between results (P=10‘5), confirming that ARORA largely recapitulates data obtained by RNP-MaP. Nearly all PRIs identified by RNP-MaP were also identified by ARORA, yet ARORA identified additional PRIs beyond those annotated by RNP-MaP. RNP-MaP requires that PRIs are solvent accessible or within 7nm of a lysine residue in order for crosslinking to occur. We hypothesize that sites uniquely identified by ARORA represent PRIs that either occur deep in the ribonucleoprotein complex, in solvent-inaccessible sites that are none-the-less accessible to crosslinking by UVC, or are PRIs that do not contain lysine amino acids.

[00180] RNP-MaP requires 104 unique reads per nucleotide of analyzed transcript to identify statistically significant increases in adduct-induced mutations that identify PRIs. For example, -2,500,000 unique reads are required in RNP-MaP to identify PRIs in the 268 nucleotide long RMRP transcript. In contrast, ARORA utilizes crosslinked-induced truncations, which occur in approximately 50% of reads and deletions occur in approximately 7-8% of reads, resulting in much higher information density that RNP-MaP. With this, ARORA achieves highly reproducible annotation of PRIs in libraries with as few as -1200 total reads of RMRP (FIG. 7C). This translates into an efficiency gain of approximately 2,000-fold compared to RNP-MaP, a crucial gain that makes transcriptomewide annotation of PRIs feasible.

[00181] We have performed low-depth transcriptome-wide sequencing of ARORA libraries in HepG2 and K562 cells to benchmark ARORA against additional reference datasets, including eCLIP annotations from ENCODE and targeted PRI annotations identified by SHAPE-footprinting10’13’14 We have observed at least one region of significant enrichment by ARORA (>30X enrichment) in most mRNAs, indicating that a subset of sequences within a mRNAs, usually in the 3’ or 5’ UTRs, are highly enriched for regulatory PRIs. Individual regions that were significantly enriched usually measured 80-150 nucleotides in length, consistent with a size of 1-2 X the read length for most reads. More interestingly, most regions of significant enrichment are associated with very narrow regions at which the reverse transcriptase truncates (RT stop peaks), which can be clearly visualized in transcripts even with moderately low read depth if the RT stop peaks occur in a narrow widow (FIGs. 8A-8B). RT stop peaks in x400_PD samples were highly enriched compared to inputs and NCL PD samples, and are usually discreet and well-defined, ranging in size from 3 nucleotides - 15 nucleotides in length. This is strikingly similar to the size of RNA motifs that determine sequencing specific PRIs, which are generally in the range of 5-7 nucleotides in length11. We are currently undertaking large-scale computational analyses of these motifs, but a survey of 1000s of individual RT stop peaks has demonstrated that each one conforms to a known RBP motif.

[00182] Due to several technical limitations, eCLIP’s inability to achieve single nucleotide resolution of RBP binding sites based on RT stops is limited. As a result, the eCLIP analysis identified and reported only significantly enriched sequences, the regions of RNAs heavily enriched by pulldown eCLIP of the specific protein. From this, they defined a window within 100 bases of the 5’ of the read to search for RNA motifs. Binding motifs were defined as the motifs that were significantly enriched within the window with lOObp of the RT stop site for each RBP. We have compared the eCLIP defined significantly enriched sequences to data obtained from ARORA. ARORA broadly agrees with eCLIP defined enriched domains. Here, dense RT stop peaks in ARORA are appreciated at the 5’ end of most RNA sequences highly enriched by eCLIP (>16X). Most mRNAs had a moderate number of RT stop peaks (<15) in their 3’UTR and while a significant minority of mRNAs had significant RT stop peaks in the 5’UTR. The ACTB mRNA has been well characterized by eCLIP, which identified 8 RBPs that highly enriched a region of the 3 ’ UTR that occur within a total of 4 distinct 100-150 nucleotide regions. ARORA identified 12 RT stop peaks in the 3’UTR of the ACTB mRNA, all but one of which were contained within or at the 5’ edge of an eCLIP significantly enriched region (FIG. 8C). Generally, significant RT stop peaks were less frequent in the 5’UTRs of transcripts. A notable exception are select mRNAs with known IRES elements within the 5’UTR (ie VEGFA mRNA)18 and mRNAs with 5’TOP domains (ie RPL30 mRNA)19, which are especially frequent in the mRNAs for most proteins required for translation including nearly all RPs and elongation factors (FIG. 8D). Finally, the Transferrin receptor is a moderately abundant mRNA that is the subject of extensive regulation at its 3’UTR by professional RNA binding proteins (profile by eCLIP) as well as Iron Response Protein, which was not included in the eCLIP dataset but was examined by SHAPE footprinting. ARORA identified numerous RT stop peaks within the long 3,500 nucleotide 3’UTR of TFRC, the majority of which correspond to canonical RBP motifs at the 5’ end of regions highly enriched by eCLIP (FIG. 8E). Additionally, two IRP footprints in the TFRC 3’UTR that were identified by SHAPE footprinting were also reproduced by ARORA. In summary, a single experiment with ARORA broadly recapitulates the PRIs that have be previously annotated by three orthogonal technologies: RNP-MaP, eCLIP, and SHAPE footprinting. Example 4: ARORA Reveals Constitutive and Dynamic Protein-RNA Interactions, Elucidating Unique Biologic States

[00183] The ribosome is the single largest ribonucleoprotein in cells, the central apparatus involved in protein synthesis, and overwhelming the most energy intensive process accounting for approximately 1 / 3 of cellular energy consumption in the mammalian cell. The ribosome has generally been considered a homogenous macromoleular machine with little variation between biologic states. However, a more complex picture has emerged in recent years revealing significant context-dependent differences in ribosome composition, regulation, and dynamics, with implications for the pathogenesis of benign diseases including ribosomopathies, neurodegeneration, and cancer20'23. The structure of the ribosome has been extensively studied using cryo-electron microscopy (cryo-EM), revealing the architecture of the ribosome at baseline and dynamically during the process of protein translation24. Identification of ribosomal determinants of disease could reveal new insights in disease pathogenesis and potentially lead to novel diagnostic and therapeutic approaches. However, the technical complexity and cost of cryo-EM has largely limited its application and prevented broad surveys of ribosome architecture across biologic conditions.

[00184] Our initial development of ARORA using total cellular RNA revealed robust, reproducible PRIs across the length of the three ribosomal RNAs encoded in the 45 S pre-rRNA transcript (18S, 5.8S and 28S). We compared the location of PRIs identified by ARORA with the PRIs identified by cryo-EM. We observed a strong similarity in the location of PRIs identified by ARORA and cryo-EM annotated sites of PRIs between ribosomal proteins (RPs) and ribosomal RNA (rRNA) (FIG. 2A, FIG. 9, Fisher Exact p<10'6)25. PRIs identified by cryo-EM are largely limited to the core constituents of the ribosome such as RPs. However, cryo-EM is unable to identify the PRIs of regulatory proteins that dynamically interact with the ribosome due to the transient PRIs and nonuniformity of these PRIs across ribosome particles examined. By contrast, CLIP approaches have identified the interactions of numerous regulatory proteins with rRNA that are not elucidated by cryo-EM26. For example, ribosome collisions induced by protein synthesis inhibitors such as cycloheximide, anisomycin, or homoharringtonine induce p38 cell stress signaling27. This signaling pathway is activated by binding of ZAKa to helix 14 of the ribosomes that have collided. We performed ARORA in normal diploid human fibroblasts which identified PRIs at the sites of RP-rRNA interactions defined by cryo-EM, but no evidence of PRIs at the ZAKa binding site on 18S helix 14. However, ARORA performed on the same cells after 1 hour exposure to either cycloheximide or anisomycin recruited significant a PRI at the precise site of ZAKa binding identified by CLIP, consistent with ARORA detection of a dynamic, stress-inducible interaction between ZAKa and 18S helix 14 (FIG. 2A).

[00185] To evaluate the ability of ARORA to identify variability in ribosome architecture and regulation between differing biologic condition, we performed ARORA in a panel of 14 widely studied, distinct breast cancer cell lines including multiple representative cell lines from each of the three principle types of breast cancer: ER+, HER2+, and triple negative breast cancer. ARORA demonstrated high degrees of correlation and low variation across all samples at sites of cryo-EM defined RP interactions with rRNA (FIG. 2B). Low variability of ARORA reactivity was strongly associated with known sites of RP interaction with rRNA (Fisher exact test, P < 10'5), indicating that PRIs at regions of RP interactions with rRNA are less variable between cell lines that sites of non-RP interactions with rRNA, as expected. Furthermore, PRIs within specific domains of the rRNA demonstrated very high correlation with each other, analogous to topologically associated domains identified by 3C, Hi-C and related technologies which identify functional units within chromatin28. These highly correlated rRNA domains tend to correspond with some secondary structures of rRNA, especially with expansion segments which are evolutionarily the most unique rRNA sequences in human cells and the sites of numerous regulatory PRI that tune translation.

[00186] rRNA Regions with low density of RP contacts, especially those in the expansion segments of 18S and 28S rRNA, demonstrated high degrees of variation between breast cancer cell lines. This suggests that PRIs in these regions between regulatory or dynamic rRNA binding proteins are highly variable between cell lines. Helixes 14, 16, and 18 on the 18S rRNA are enriched for regulatory function (FIG. 2B). Helixes 16 and 18 forming at gate at the mRNA entry channel that regulates translation initiation, elongation and termination and are the sites of binding of numerous regulatory factors including elFs eEFs, and eRFs, as well as mRNA interacting proteins such as DDX3. Helix 14 is involved in signaling ribosome stress is the binding site of ZAKa in response to ribosome collisions27. At each of these sites, high degrees of variable PRIs was noted between breast cancer cell lines. We found it especially interesting that a subset of BRCA cell lines had variability in PRIs within the ribosomal stress responsive site on helix 14 bound by ZAKa. The Cancer Cell Line Encyclopedia examined the sensitivity of cancer cell lines to a variety of clinical and pre-clinical small molecules, including sensitivity to one drug, omacetaxine, the drug name for homoharringtonine, which inhibits translation via a mechanism similar to cycloheximide or anisomycin and induces ribosome collisions and ribosome stress29. We observed that the subset of BRCA cell lines with high PRIs detected by ARORA in the ZAKa-bound region on helix 14 of the 18S were resistant to omacetaxine while BRCA cell lines with low PRIs in this region were significantly more sensitive to the drug (FIGs. 2C-2D). We postulate that high basal PRIs within the ZAKa-bound region on helix 14 identify cell lines with high basal ribosome stress and which have adapted to and overcome this stress, resulting in resistance to drugs that induce ribosome stress. Therefore, ARORA identifies biologically relevant differences in ribosome regulation and architecture with clear potential for direct clinical relevance in identifying cancer cell therapeutic vulnerabilities. Example 5: ARORA Identi fies a Cancer Driver Mutation in Noncoding, Regulatory RNA

[00187] The non-protein coding genome undergoes frequent mutation and alteration in human cancers, contributing to disease pathogenesis30'32. However, we currently lack tools to readily determine whether noncoding mutations are drivers of cancer phenotypes or merely passenger mutations. eCLIP has previously demonstrated an ability to identify allele-specific PRIs for individual RBPs, albeit under conditions not well-suited to single nucleotide resolution of PRIs. We examined whether ARORA could identify mutations in regulatory RNA that altered the interaction of RBPs. In considering how sequencing of ARORA libraries with single nucleotide resolution of PRIs may identify mutations that alter binding, it became clear that the location of PRIs relative to the location of the mutation would have considerable influence on the interpretation of sequencing data (FIG. 3A). When using a RT enzyme that efficiently truncates at the crosslink site and identifies the PRI with single nucleotide resolution, such as Maxima H-, if the predominant crosslinked nucleotides within a PRI are 5’ in the RNA sequence to the single nucleotide variant that alters binding, then the variant nucleotide can be readily quantified and enrichment or depletion of an allele in the x400_PD condition clearly calculated to determine whether the allele is favored in the PRI or disfavored, respectively. However, if the predominant crosslinked nucleotides within a PRI are 3’ to the single nucleotide variant that alters binding, then the RT enzyme will not read through to variant nucleotide since the cDNA is truncated before reaching the variant site. In this case the variant that is favored in binding -59- is depleted in representation in the ARORA sequencing library generated with a truncating RT enzyme. In contrast, ARORA sequencing libraries generated with a read-through RT enzyme, such as Affinity Script, single nucleotide resolution of PRIs are lost, but the capability to read-through the crosslink site can fully capture the variant sequence and quantify allele enrichment or depletion at a PRI.

[00188] We examined the low-depth transcriptome-wide ARORA datasets in HepG2 and K562 to examine whether ARORA identifies allele-specific PRIs. A variety of non-pathogenic, normal variants, such as GAPDH rsl065691, resulted in no change in the allele fractions between the four conditions, confirming no selection by ARORA and consistent with no impact of these normal variants on RBP interactions with mRNAs (FIG. 3B). In contrast, one single nucleotide variant unique to HepG2 cells was noted in the 5’UTR of the oncogene MYC. The frequency of this variant transcript, a C^T transition at residue 184 (C184T), relative to the wildtype transcript was significantly altered by crosslinking and pulldown, suggesting a potential allele-specific PRI (FIG. 3B). The C184T variant was diminished amongst in the total reads in the UVC crosslinked and pulldown samples. While the sequencing depth was insufficient in the MYC 5’UTR to achieve single nucleotide resolution, the ARORA sequencing none-the-less demonstrated significant enrichment (~7X) of RT stops in the nucleotides just 3’ C184T, a feature noted HepG2 samples but not in K562 samples. We performed ARORA with qPCR quantification of this region to precisely quantify the enrichment of this area, confirming that HepG2 cells have a unique, substantial PRI that is not present in K562 or the immortalized, highly proliferative cell line 293T cells (FIG. 3C). Together, these data indicate that HepG2 cells contain a unique MYC 5’UTR variant that is associated with an allele-specific binding and a gain of a PRI that is not present in other cells.

[00189] While the enrichment of RT stops 3’ to C184T and drop-out of the C184T in the total reads suggest that C184T has a gain of RBP binding compared to the wildtype transcript, we performed ARORA in HepG2 with the read-through RT AffinityScript to quantify the relative proportion of MYC alleles recovered by ARORA. Under these conditions, ARORA of HepG2 enriches a sequence overlapping the C184T mutation, again confirming that a significant PRI occurs within this region. In the ENCODE eCLIP database, just two RBPs (DDX3X and DDX6) out of 103 surveyed in HepG2 cells, both involved in translation initiation, strongly enrich the MYC 5’UTR at the precise location identified by ARORA (FIG. 3D), indicating that ARORA likely is detecting the binding of a translation regulator. Furthermore, ARORA with read-through of crosslinked sites reveals that the PRI in the MYC 5’UTR of these cells is indeed allele-specific, since 85% of MYC 5’UTR fragments recovered by ARORA contain the C184T, despite representing 40% of all MYC transcripts in the input samples. This suggests that a potential cis-regulatory factor in the 5’UTR of MYC is predominantly interacting with the MYC transcripts bearing the 5’UTR C184T (FIGs. 3E-3F). In summary, ARORA identified a variant in the 5’UTR of the oncogene MYC that alters the interaction of the MYC 5’UTR with a RBP, likely DDX3X or DDX6, a translation regulators that increases the translation efficiency of a subset of mRNAs through interactions with 5’UTRs.

[00190] Finally, these data demonstrate that ARORA has outstanding potential for development to clinical diagnostic. We have observed that ARORA performance is preserved in samples derived from frozen cell-line “blocks” (simulating frozen tissue), with the ability to annotate PRIs that correlate highly with results obtained by ARORA in the same cell line in cell culture (FIG. 10). Thus, ARORA can be performed on frozen clinical samples standardly used for molecular and genetic diagnostics tests. Example 6: Identi fication of RBPs Bound to Speci fic Sites on RNAs via ARORA

[00191] As described herein, ARORA comprehensively identifies the specific sites of RBP binding on RNA molecules at single nucleotide resolution. However, the specific identity of the RBP bound at any given site was previously unknown. While eCLIP, RNA bind-n-seq, and other technologies have mapped nucleotide sequence motifs that are preferentially bound by specific RBPs, many sequence motifs are bound by different RBPs, limiting the ability of sequence motifs alone to predict the precise RBP bound at a given site on a RNA. Analytical approaches, including mCross and PEKA, have recently provided single nucleotide resolution of RBP crosslinking patterns within sequence motifs, derived from eCLIP. These analyses identify that many individual RBPs have stereotypic patterns of crosslinks with the preferred sequence motif. The mCross analysis demonstrated reproducible, unique crosslinking patterns for individual RBPs within a given sequence motifs, many of which are specific to the RBP and differ from other RBPs bound to the same or similar sequence motif. Therefore, the crosslinking patterns within sequence motifs could identify the exact RBP bound to a specific site on an RNA even in the absence of the prior knowledge of the protein / peptide bound to the RNA site.

[00192] Applying crosslinking patterns identified by mCross to sequencing data from ARORA, we found that crosslinking patterns within sequencing motifs in the ARORA allow for determination of the identity of the RBP bound to a specific site on RNA. Analyzing the mCross crosslinking pattern for three RBPs (PUM2, TARDBP, and YBX3), we found that a subset RNA transcripts containing a sequence motifs specific for each RBP have a stereotypical crosslink pattern within the sequence motif (FIGs. 16A-16C). Sites with RBP-specific crosslinking patterns within the sequence motif correspond to sites of high enrichment by the RBP in eCLIP data and distinguish these from other sites that contain the RBPs preferred motif but which are not bound efficient by the RBP. mCross has identified high confidence crosslinking patterns for 39 RBPs, and lower confidence crosslinking patterns for many more RBPs. Therefore, ARORA’s annotation of crosslinking patterns with a sequence motif and applying experimentally defined RBP-specific crosslinking patterns allows the annotation of sites specifically bound by numerous RBPs simultaneously, in a single assay. Additionally, crosslinking patterns for RBPs not defined by mCross can be discovered de novo using ARORA by identify crosslinking patterns within sequence motifs that are lost after induced loss of the RBP in cell lines. As this does not require the availability of high quality immunoprecipitation antibodies, this presents a substantial advantage over CLIP technologies which are limited by antibody specificity and quality.

[00193] The ability to identify the precise identify of the RBP bound to specific sites on an RNA furthermore facilitates novel insights into RNA regulation. Current methods such as CLIP technologies profile single RBPs in isolation, preventing direct comparison between the strength or stoichiometry of binding between distinct RBPs interacting with an RNA. In contrast, ARORA identifies the relative strength and stoichiometry of RBP-RNA interactions on an individual transcript. For example, eCLIP identified that dozens of RBPs bind to the GAPDH mRNA at various locations throughout the transcript. ARORA recapitulates these bindings sites defined by eCLIP, while also annotating many additional bindings sites (FIG. 16B). Moreover, ARORA also demonstrates that RBP bound sites on GAPDH mRNA are overwhelmingly enriched for YBX3 interactions, while other RBP interactions are orders of magnitude less enriched than YBX3 sites (FIG. 16B). As cold shock proteins, YBX family proteins play a crucial role in maintaining glucose metabolism during cell stress and loss of individual YBX family members results in downregulation of mRNAs associated with glucose metabolism, such as GAPDH and other mRNAs for components of the glycolytic pathway. Highly enriched YBX3 sites were the strongest RBP bound sites on the mRNAs for the majority of proteins involved in glycolysis, demonstrating YBX is a dominant regulator of this critical pathway and identifying potential sites relevant for therapeutic targeting in cancer and other diseases. In contrast to glycolytic genes, most mRNAs for ribosomal proteins have minimal enrichment of YBX3 bound sites. Furthermore, ARORA identified that nearly all mRNAs for ribosomal proteins have dominant enrichment of RBP binding at the 5’ terminal oligopyrimidine tract (5’TOP) site, a conserved feature in the 5’UTR of these mRNAs which is critical for regulation of translation of ribosomal proteins. Therefore, ARORA’s identification of the most strongly bounds sites in RNAs annotates the RBP-RNA interactions that are most crucial for the regulation of the mRNAs function, and therefore identifying sites most relevant to cellular mechanisms, sites with relevant information for disease diagnosis and discovery of disease mechanisms, and for discovery of therapeutic targets.

[00194] We developed ARORA to address the major unmet need to identify new druggable targets for cancer therapy, we developed a novel, rapid, cost-effective technology to comprehensively identify sites protein interactions on RNA. This method utilizes a biochemical process to convert the specific nucleotides bound by RBPs to an orthogonal chemical moiety that is conjugated with a ligand for affinity purification. This results in sequencing libraries in which the precise nucleotide contacting RBPs are unambiguously encoded in each fragment of the library and resolved by next generation sequencing. Critically, this method allows for purified RNA with solubility and handling properties identical to unmodified cellular RNA and thus the process is highly amenable to automation and high throughput processes.

[00195] ARORA represents a major technical advance over existing methods for identifying sites of RBP-RNA interactions in several ways: (1) utilization of a universal chemical moiety in amino acids (alpha amino or carboxyl) makes possible the unbiased labeling of any protein that crosslinks to RNA with UVC, regardless of amino acid identity, allowing for transcriptome-wide analyses of all proteins bounds to RNA in a single assay, (2) the chemistry is fully compatible with standard high-throughput and automated library platforms and RNA in ARORA can be processed with hybridization capture for selective depletion or selective enrichment of transcripts to target specific transcripts for ARORA analyses, (3) permits comprehensive, transcriptome-wide surveys to identifying dynamic, stress-responsive, and allele-specific PRIs that contribute to disease, increasing the potential to identify RBP-RNA interactions exploitable for diagnostic or therapeutic development, (4) ARORA provides direct, quantitative, stochiometric comparisons of PRIs on transcripts and between biologic conditions in contrast to binary (bound vs unbound) results provided by current methods, (5) orders of magnitude improved efficiency in annotation of PRIs compared to existing methods. Example 7: ARORA-based Drug Screening Approaches

[00196] Several small molecule and oligonucleotide drugs that aim to specifically inhibit a single RBP or family of RBPs involved in a specific disease process are currently in development. However, prior reports have shown that a drug believed to specifically inhibit a single RBP in fact has a wide impact on RBP interactions with RNA and numerous off target effects (Walters etal., RNA. 2023 Oct; 29(10): 1458-1470). ARORA, as an assay that comprehensively annotates interactions between all RBPs and all RNA transcripts can in principle identify the efficacy and specificity of specific RBP inhibitor drug. We will use ARORA to determine the efficacy of and specificity of RBP inhibitors (e.g., oligonucleotide inhibitors and small molecule inhibitors).

[00197] Select individual RBPs will be inhibited with antisense oligonucleotides (ASOs) that reduce the mRNA of the targeted RBP and thus aim to reduce RBP abundance in the cell, or small molecule inhibitors that purportedly inhibit the activity of the same RBP. ARORA will be performed in cells after treatment with inhibitory ASOs and compared to untreated controls. ARORA will also be performed in an additional condition, where the efficacy of a small molecule drug with a purported mechanism of action of inhibiting the same RBP will be examined. The efficacy will be defined as the fraction of RBP-specific PRIs, as identified by crosslink patterns and sequence motifs, that are lost by treatment with the small molecule inhibitor. Specificity will be determined by defining the fraction of PRIs specific for non-targeted RBPs that are altered by the treatment with the same small molecule inhibitor. As such, a global view of the effect of a RBP inhibitory drug will be surveyed in a single ARORA assay. Example 8: Identification of RBPs Bound to Specific Sites on RNAs Related to Anauxetic dysplasia (AD)

[00198] AD is a congenital osteochondrodysplasia. Individuals affected by this disorder have bone and cartilage dysplasias resulting in dwarfism and joint hypermobility amongst other traits38. Mutations in POP1, the gene encoding an essential RNA binding protein component of two distinct but related human RNases, RNAse MRP and RNase P, are associated with AD39. POP1 directly binds RMRP, the RNA component of the RNase MRP ribonucleoprotein complex, to form functional RNase MRP40. Separately POP1 binds RPPH1 RNA to form the function ribonucleoprotein complex of RNase P. ARORA identifies strong, discreet protein binding sites in both RMRP and RPPH1 RNAs that correspond to the well-described binding site of POP 1 on the RNA components of RNase MRP and RNase P. Genetic loss of functional POP1 results in failure to incorporate POP1 into these ribonucleoprotein complexes. It is anticipated that cell models of AD with deficiency of POP 1 may be detectable by ARORA as a loss of the protein-RNA interaction at the POP1 binding sites on the RMRP and RPPH1 RNAs.

[00199] ARORA will be performed as described in Example 1 in two cells lines: (1) POP1 wildtype human cells and (2) POP1 mutant cells as a model of AD. The stoichiometry of protein-RNA interactions at the POP1 bound sites on RMRP and RPPH1 RNAs will be quantified and compared to non-POPl protein-RNA interactions on these RNAs, in (1) POP1 wildtype human cells and (2) POP1 mutant cells as a model of AD. It is expected that the stoichiometry of protein-RNA interactions at the POP1 binding may be reduced in POP1 mutant cell lines compared to POP1 wildtype cells, while the protein-RNA interactions of non-POPl protein-RNA interaction sites may be preserved. Thus, ARORA may serve as a sensitive and specific measure of the loss of the POP1 protein in the assembly of RNase MRP and RNaseP associated with the pathogenesis of AD. Example 9: Identi fication of RBPs Bound to Speci fic Sites on RNAs Related to DiamondBlackfan Anemia (DBA)

[00200] DBA is one form of ribosomopathy, rare, autosomal dominant genetic disorders resulting from pathogenic loss of function mutations in one of the many ribosomal proteins (RPs) that are critical structural components of the ribosome41. DBA is classically associated with bone marrow dysfunction with severe disruption of red blood cell production resulting in potentially fatal anemia. While mutations in several genes for ribosomal proteins have been associated with DBA, mutations in RPL5 and RPL11 are amongst most frequently causes of DBA. Loss of function mutations in RP genes lead to a stochiometric imbalance in RPs resulting in impaired ribosome biogenesis or assembly of incomplete ribosomes missing a RP component. As a result, ribosome biogenesis stress induces cellular stress responses, especially p53 signaling, triggering cell cycle arrest or inducing programmed cell death. These cellular responses result deficient hematopoiesis and other tissue disturbances. ARORA demonstrated an ability to identify RP binding sites on the ribosomal RNA (rRNA), with conserved protein binding sites identified in the contact sites of RPL5 and RPL11 with the large subunit (28S) of the rRNA. It is expected that ARORA may identify loss of protein-rRNA contacts at the site of RPL5-28S rRNA contacts or RPL11-28S rRNA contacts in cellular models of DBA, in which RPL5 or RPL11 is genetically deficient.

[00201] ARORA will be performed as described in Example 1 in three cells lines: (1) wildtype human cells, (2) RPL5 mutant cells as a model of DBA, and (3) RPL11 mutant cells as a model of DBA. The stoichiometry of protein-RNA interactions at the RPL5 or RPL11 bound sites on the 28S rRNA will be quantified and compared to protein-RNA interaction of other, non-RPL5 / RPLl 1 proteins bound to 28S rRNA in three cells lines: (1) wildtype human cells, (2) RPL5 mutant cells as a model of DBA, and (3) RPL11 mutant cells as a model of DBA. It is expected that the stoichiometry of protein-RNA interactions at the RPL5 or RPL11 binding sites may be reduced in RPL5 or RPL11 mutant cell lines, respectively, compared to wildtype cells, while the protein-RNA interactions of protein-RNA interactions and other, non-RPL5 / RPLl 1 bound sites may be preserved. Thus, ARORA may serve as a sensitive and specific measure of the loss of the RPL5 or RPL11 protein in the assembly of the large subunit of the ribosome, the pathogenic mechanism of DBA and other ribosomopathies. Example 10: Use of ARORA to Evaluate Small Molecule Efficacy

[00202] Initiation of translation in eukaryotic cells requires the assembly of the elongation initiation factor complex 4F (eIF4F), which recruits the small ribosome subunit to the 5’ cap of mRNAs prior to formation of the preassembled preinitiation complex42. eIF4F consists of 3 subunits: eIF4E which recognized the 5’ cap, the eIF4G scaffold protein, and eIF4A, an RNA helicase. eIF4A helicase activity is necessary for removal of secondary structure in the 5’UTR, a step that facilitates the preinitiation complex’s translocation along the 5’UTR to the start codon. Multiple oncogenes contain 5’UTRs with high secondary structure, and thus are dependent on eIF4A for efficiency translation. As a result, cancer cell growth is frequently dependent on eIF4A helicase activity to maintain oncogenic signaling, making eIF4A an attractive therapeutic target for drug development in oncology. Rocaglates, such as silvestrol, are natural compounds that inhibit translation by modulating the binding of eIF4A with mRNAs and represent of class of molecules of high interest as therapeutics for oncology and viral infections43,44. The rocaglate derivative, zotatifin, recently demonstrated promising results in an initial phase VII clinical trial in breast cancer45. Some rocaglates, such as zotatifin, stabilize eIF4A interaction with purine (AG) rich sequences in the 5’UTR of select mRNAs, creating a potent steric block that prevents the translocation of the preinitiation complex along the 5’UTR, inhibiting translation46. It is expected that that ARORA may be able to determine the efficacy and selectivity of certain rocaglates like zotatifin on eIF4A-mRNA stabilization and translation blocking in 5’UTRs.

[00203] We will perform ARORA as described in Example 1 in two cell conditions: (1) wildtype human cells treated with DMSO vehicle alone and (2) the same human cell line 26 hours after treatment with rocaglate (zotatifin or related molecule) in DMSO. The stoichiometry of protein-RNA interactions at eIF4A bound sites in the 5’UTR of mRNAs will be quantified and compared to protein-RNA interaction of sites within the 3’UTRs which are largely independent of eIF4A, in the two conditions described above. It is expected that the stoichiometry of protein-RNA interactions at the eIF4A binding sites at purine (AG) rich sequences of the 5’UTR may be increased in cells 2-6 hours following treatment with rocaglate (zotatifin or related molecule) in comparison to cells treated with DMSO only. In contrast, the protein-RNA interactions in the 3’UTRs of mRNAs, which are largely eIF4A- independent may be preserved. Thus, ARORA may serve as a sensitive measure of eIF4A-mRNA interaction stabilization and translation inhibition in select mRNAs with highly structure 5’UTRs. EQUIVALENTS

[00204] The present technology is not to be limited in terms of the particular embodiments described in this application, which are intended as single illustrations of individual aspects of the present technology. Many modifications and variations of this present technology can be made without departing from its spirit and scope, as will be apparent to those skilled in the art. Functionally equivalent methods and apparatuses within the scope of the present technology, in addition to those enumerated herein, will be apparent to those skilled in the art from the foregoing descriptions. Such modifications and variations are intended to fall within the scope of the present technology. It is to be understood that this present technology is not limited to particular methods, reagents, compounds compositions or biological systems, which can, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting.

[00205] In addition, where features or aspects of the disclosure are described in terms of Markush groups, those skilled in the art will recognize that the disclosure is also thereby described in terms of any individual member or subgroup of members of the Markush group.

[00206] As will be understood by one skilled in the art, for any and all purposes, particularly in terms of providing a written description, all ranges disclosed herein also encompass any and all possible subranges and combinations of subranges thereof. Any listed range can be easily recognized as sufficiently describing and enabling the same range being broken down into at least equal halves, thirds, quarters, fifths, tenths, etc. As a nonlimiting example, each range discussed herein can be readily broken down into a lower third, middle third and upper third, etc. As will also be understood by one skilled in the art all language such as “up to,” “at least,” “greater than,” “less than,” and the like, include the number recited and refer to ranges which can be subsequently broken down into subranges as discussed above. Finally, as will be understood by one skilled in the art, a range includes each individual member. Thus, for example, a group having 1-3 cells refers to groups having 1, 2, or 3 cells. Similarly, a group having 1-5 cells refers to groups having 1, 2, 3, 4, or 5 cells, and so forth.

[00207] All patents, patent applications, provisional applications, and publications referred to or cited herein are incorporated by reference in their entirety, including all figures and tables, to the extent they are not inconsistent with the explicit teachings of this specification. REFERENCES 1         Investigators, G. P. P. et al. 100,000 Genomes Pilot on Rare-Disease Diagnosis in Health Care - Preliminary Report. The New Englandjournal of medicine 385, 1868-1880, doi: 10.1056 / NEJMoa2035790 (2021). 2         Anastasiadou, E., Jacob, L. S. & Slack, F. J. Non-coding RNA networks in cancer. Nature reviews. Cancer 18, 5-18, doi:10.1038 / nrc.2017.99 (2018). 3         Gebauer, F., Schwarzl, T., Valcarcel, J. & Hentze, M. W. RNA-binding proteins in human genetic disease. Nature reviews. Genetics 22, 185-198, doi:10.1038 / s41576-020-00302-y (2021). 4         Mitschka, S. & Mayr, C. Context-specific regulation and function of mRNA alternative polyadenylation. Nature reviews. Molecular cell biology 23, 779-796, doi: 10.1038 / s41580-022-00507-5 (2022). 5         Ramanathan, M., Porter, D. F. & Khavari, P. A. Methods to study RNA-protein interactions. Nature methods 16, 225-234, doi:10.1038 / s41592-019-0330-l (2019). 6        Hafner, M. etal. CLIP and complementary methods. Nat Rev Method Prime 1, doi:ARTN20 10.103 8 / s43 586-021 -00018-1 (2021). 7         Queiroz, R. M. L. etal. Comprehensive identification of RNA-protein interactions in any organism using orthogonal organic phase separation (OOPS). Nature biotechnology 37, 169-178, doi: 10.103 8 / s415 87-018-0001-2 (2019). 8 Weidmann, C. A., Mustoe, A. M., Jariwala, P. B., Calabrese, J. M. & Weeks, K. M. Analysis of RNA-protein networks with RNP-MaP defines functional hubs on RNA. Nature biotechnology 39, 347-356, doi:10.1038 / s41587-020-0709-7 (2021). 9         Grawe, C., Stelloo, S., van Hout, F. A. H. & Vermeulen, M. RNA-Centric Methods: Toward the Interactome of Specific RNA Transcripts. Trends Biotechnol 39, 890900, doi: 10.1016 / j.tibtech.2020.11.011 (2021). 10 Corley, M. et al. Footprinting SHAPE-eCLIP Reveals Transcriptome-wide Hydrogen Bonds at RNA-Protein Interfaces. Molecular cell 80, 903-914 e908, doi: 10.1016 / j.molcel.2020.11.014 (2020). 11 Dominguez, D. etal. Sequence, Structure, and Context Preferences of Human RNA Binding Proteins. Molecular cell IQ, 854-867 e859, doi:10.1016 / j.molcel.2018.05.001 (2018). 12 Ramanathan, M. et al. RNA-protein interaction detection in living cells. Nature methods 15, 207-212, doi:10.1038 / nmeth.4601 (2018). 13 Van Nostrand, E. L. et al. A large-scale binding and functional map of human RNA-binding proteins. Nature 583, 711-719, doi:10.1038 / s41586-020-2077-3 (2020). 14 Van Nostrand, E. L. et al. Robust transcriptome-wide discovery of RNA-binding protein binding sites with enhanced CLIP (eCLIP). Nature methods 13, 508-514, doi: 10.1038 / nmeth.3810 (2016). 15 Urdaneta, E. C. etal. Purification of cross-linked RNA-protein complexes by phenol-toluol extraction. Nature communications 10, 990, doi:10.1038 / s41467-019-08942-3 (2019). 16 Van Nostrand, E. L., Shishkin, A. A., Pratt, G. A., Nguyen, T. B. & Yeo, G. W. Variation in single-nucleotide sensitivity of eCLIP derived from reverse transcription conditions. Methods 126, 29-37, doi:10.1016 / j.ymeth.2017.08.002 (2017). 17 Zarnegar, B. J. et al. irCLIP platform for efficient characterization of protein-RNA interactions. Nature methods 13, 489-492, doi:10.1038 / nmeth.3840 (2016). 18 Leppek, K., Das, R. & Barna, M. Functional 5' UTR mRNA structures in eukaryotic translation regulation and how to find them. Nature reviews. Molecular cell biology 19, 158-174, doi:10.1038 / nrm.2017.103 (2018). 19 Hsieh, A. C. et al. The translational landscape of mTOR signalling steers cancer initiation and metastasis. Nature 485, 55-61, doi:10.1038 / naturel0912 (2012). 20 Leppek, K. et al. Gene- and Species-Specific Hox mRNA Translation by Ribosome Expansion Segments. Molecular cell 80, 980-995 e913, doi: 10.1016 / j.molcel.2020.10.023 (2020). 21 Simsek, D. et al. The Mammalian Ribo-interactome Reveals Ribosome Functional Diversity and Heterogeneity. Cell 169, 1051-1065 el018, doi:10.1016 / j.cell.2017.05.022 (2017). 22 Ebright, R. Y. et al. Deregulation of ribosomal protein expression and translation promotes breast cancer metastasis. Science 367, 1468-1473, doi:10.1126 / science.aay0939 (2020). 23 Parks, M. M. et al. Variant ribosomal RNA alleles are conserved and exhibit tissue-specific expression. Sci Adv 4, eaao0665, doi:10.1126 / sciadv.aao0665 (2018). 24 Behrmann, E. et al. Structural snapshots of actively translating human ribosomes. Cell 161, 845-857, doi:10.1016 / j.cell.2015.03.052 (2015). 25 Bernier, C. R. et al. RiboVision suite for visualization and analysis of ribosomes. Faraday Discuss 169, 195-207, doi:10.1039 / c3fd00126a (2014). 26 Van Nostrand, E. L. et al. Principles of RNA processing from analysis of enhanced CLIP maps for 150 RNA binding proteins. Genome biology 21, 90, doi: 10.1186 / s 13059-020-01982-9 (2020). 27 Vind, A. C. et al. ZAKalpha Recognizes Stalled Ribosomes through Partially Redundant Sensor Domains. Molecular cell 78, 700-713 e707, doi: 10.1016 / j.molcel.2020.03.021 (2020). 28 Akgol Oksuz, B. et al. Systematic evaluation of chromosome conformation capture assays. Nature methods 18, 1046-1055, doi:10.1038 / s41592-021-01248-7 (2021). 29 Seashore-Ludlow, B. et al. Harnessing Connectivity in a Large-Scale SmallMolecule Sensitivity Dataset. Cancer discovery 5, 1210-1223, doi:10.1158 / 2159-8290.CD-15-0235 (2015). 30 Weinhold, N., Jacobsen, A., Schultz, N., Sander, C. & Lee, W. Genome-wide analysis of noncoding regulatory mutations in cancer. Nature genetics 46, 1160-1165, doi:10.1038 / ng.3101 (2014). 31 Li, K. et al. Noncoding Variants Connect Enhancer Dysregulation with Nuclear Receptor Signaling in Hematopoietic Malignancies. Cancer discovery 10, 724-745, doi: 10.1158 / 2159-8290.CD-19-1128 (2020). 32 Schmitt, A. M. & Chang, H. Y. Long Noncoding RNAs in Cancer Pathways. Cancer cell29, 452-463, doi:10.1016 / j.ccell.2016.03.010 (2016). 33 Guo, Z. et al. A Functional 5'-UTR Polymorphism of MYC Contributes to Nasopharyngeal Carcinoma Susceptibility and Chemoradiotherapy Induced Toxicities. J Cancer 10, 147-155, doi:10.7150 / jca.28534 (2019). 34 Xu-Monette, Z. Y. etal. Clinical and Biologic Significance of MYC Genetic Mutations in De Novo Diffuse Large B-cell Lymphoma. Clin Cancer Res 22, 3593-3605, doi: 10.1158 / 1078-0432.CCR-15-2296 (2016). 35 Chappell, S. A. et al. A mutation in the c-myc-IRES leads to enhanced internal ribosome entry in multiple myeloma: a novel mechanism of oncogene de-regulation. Oncogene 19, 4437-4440, doi: 10.1038 / sj.one. 1203791 (2000). 36 Cobbold, L. C. et al. Upregulated c-myc expression in multiple myeloma by internal ribosome entry results from increased interactions with and expression of PTB-1 and YB-1. Oncogene 29, 2884-2891, doi:10.1038 / onc.2010.31 (2010). 37. Makitie O, Vakkilainen S. Cartilage-Hair Hypoplasia - Anauxetic Dysplasia Spectrum Disorders. In: Adam MP, Feldman J, Mirzaa GM, Pagon RA, Wallace SE, Amemiya A, editors. GeneReviews((R)). Seattle (WA)1993. 38. Barraza-Garcia J, Rivera-Pedroza CI, Hisado-Oliva A, Belinchon-Martinez A, Sentchordi-Montane L, Duncan EL, Clark GR, Del Pozo A, Ibanez-Garikano K, Offiah A, Prieto-Matos P, Cormier-Daire V, Heath KE. Broadening the phenotypic spectrum of POP 1-skeletal dysplasias: identification of POP 1 mutations in a mild and severe skeletal dysplasia. Clin Genet. 2017;92(l):91-8. Epub 20170222. doi: 10.1111 / cge. 12964. PubMedPMID: 28067412. 39. Weidmann CA, Mustoe AM, Jariwala PB, Calabrese JM, Weeks KM. Analysis of RNA-protein networks with RNP-MaP defines functional hubs on RNA. Nature biotechnology. 2021;39(3):347-56. Epub 2020 / 10 / 21. doi: 10.1038 / s41587-020-0709-7. PubMed PMID: 33077962; PMCID: PMC7956044. 40. Kampen KR, Sulima SO, Vereecke S, De Keersmaecker K. Hallmarks of ribosomopathies. Nucleic acids research. 2020;48(3): 1013-28. doi: 10.1093 / nar / gkz637. PubMed PMID: 31350888; PMCID: PMC7026650. 41. Brito Querido J, Diaz-Lopez I, Ramakrishnan V. The molecular basis of translation initiation and its regulation in eukaryotes. Nature reviews Molecular cell biology. 2024;25(3): 168-86. Epub 20231205. doi: 10.1038 / s41580-023-00624-9. PubMedPMID: 38052923. 42. Chu J, Zhang W, Cencic R, O'Connor PBF, Robert F, Devine WG, Selznick A, Henkel T, Merrick WC, Brown LE, Baranov PV, Porco JA, Jr., Pelletier J. Rocaglates Induce Gain-of-Function Alterations to eIF4Aand eIF4F. Cell reports. 2020;30(8):2481-8 e5. doi: 10.1016 / j.celrep.2020.02.002. PubMedPMID: 32101697; PMCID: PMC7077502. 43. Schmidt T, Dabrowska A, Waldron JA, Hodge K, Koulouras G, Gabrielsen M, Munro J, Tack DC, Harris G, McGhee E, Scott D, Carlin LM, Huang D, Le Quesne J, Zanivan S, Wilczynska A, Bushell M. eIF4Al-dependent mRNAs employ purine-rich 5'UTR sequences to activate localised eIF4 Al-unwinding through eIF4Al-multimerisation to facilitate translation. Nucleic acids research. 2023;51(4):1859-79. doi: 10.1093 / nar / gkad030. PubMed PMID: 36727461; PMCID: PMC9976904. 44. Ezra Rosen MS, David Berz, Jennifer Lee Caswell-Jin, Alexander I. Spira, Georgina A. Fulgar, Mark Densel, Nawaid Rana, Samuel Sperry, Douglas Warner, and Funda Meric-Bernstam ASCO Authors' Group. Phase 1 / 2 dose expansion study evaluating first-in-class eIF4A inhibitor zotatifin in patients with ER+ metastatic breast cancer 2023. 45. Gerson-Gurwitz A, Young NP, Goel VK, Earn B, Stumpf CR, Chen J, Fish S, Barrera M, Sung E, Staunton J, Chiang GG, Webster KR, Thompson PA. Zotatifin, an eIF4A-Selective Inhibitor, Blocks Tumor Growth in Receptor Tyrosine Kinase Driven Tumors. Front Oncol. 2021; 11:766298. Epub 20211124. doi: 10.3389 / fonc.2021.766298. PubMed PMID: 34900714; PMCID: PMC8663026.

Claims

WHAT IS CLAIMED IS1. A method for identifying one or more RNA binding protein (RBP)-RNA interaction (PRI) sites in at least one transcript derived from a biological sample including a plurality of cells, the method comprising:(a) cross-linking the plurality of cells present in the biological sample with short wavelength ultraviolet light to stabilize a plurality of RBP-RNA complexes comprising RNA transcripts, wherein each RBP-RNA complex corresponds to a PRI;(b) isolating RNA molecules from the cross-linked plurality of cells, wherein the RNA molecules comprise unbound RNA molecules and the RNA transcripts of the plurality of RBP-RNA complexes;(c) contacting the isolated RNA molecules with a protease under conditions that lyse the RBPs of the plurality of RBP-RNA complexes to yield RNA transcripts bound by residual peptides, wherein each bound residual peptide (i) corresponds to a PRI site and (ii) comprises an alpha amino group at its N terminus and a carboxyl group at its C terminus;(d) labeling the RNA transcripts bound by the residual peptides by coupling the alpha amino group or the carboxyl group of each bound residual peptide with a chemical moiety conjugate comprising an affinity ligand;(e) fragmenting the RNA molecules in the presence of heat and divalent cations to generate RNA fragments having a 5’ end and a 3’ end, wherein the RNA fragments comprise unlabeled RNA fragments and RNA fragments labeled with the chemical moiety conjugate comprising the affinity ligand;(f) ligating a sequencing adapter to the 3 ’ end of the RNA fragments to generate adapter tagged RNA fragments;(g) capturing adapter tagged RNA fragments labeled with the chemical moiety conjugate comprising the affinity ligand using affinity purification;(h) reverse transcribing the adapter tagged labeled RNA fragments to generate a plurality of cDNA molecules, optionally wherein the adapter tagged labeled RNA fragments are attached to a solid surface or are in solution;(i) ligating a sequencing adapter to the 3’ end of the cDNA molecules to generate a plurality of adapter tagged cDNA molecules;(j) amplifying the plurality of adapter tagged cDNA molecules to generate a nucleic acid library of transcripts having RNA-RBP interactions;(k) sequencing and mapping the plurality of adapter tagged cDNA molecules of the nucleic acid library to identify individual transcripts;(1) identifying the frequency of cDNA 3 ’ ends at each nucleotide across the individual transcripts; and(m)detecting enriched RBP bound sites in the individual transcripts when the frequency of cDNA 3’ ends at specific nucleotides within the individual transcripts satisfies a predetermined threshold.

2. The method of claim 1, wherein the cross-linking is performed with a UV box at a dose of 50-600 mJ / cm2 for about 30s-10 minutes.

3. The method of claim 2, wherein the cross-linking is performed with a UV box at a dose of 300-500 mJ / cm2 for about 3 minutes or 100-400 mJ / cm2 for about 10 minutes.

4. The method of claim 1, wherein the cross-linking is performed with a UV laser for about 10s-20s.

5. The method of any one of claims 1-4, wherein the isolated RNA molecules are treated with the protease at 50°C for about 30 minutes, optionally wherein the protease is proteinase K.

6. The method of any one of claims 1-5, wherein each bound residual peptide is a single amino acid or is no more than 10 amino acids in length.

7. The method of any one of claims 1-6, further comprising removing genomic DNA after step (c) to obtained purified RNA molecules.

8. The method of claim 7, further comprising contacting the purified RNA molecules with a protease at 50°C for about 30 minutes prior to step (d), optionally wherein the protease is proteinase K.

9. The method of any one of claims 1-8, further comprising enriching the isolated or purified RNA molecules with poly-dT oligonucleotides or targeted bait capture reagents after step (c), but prior to step (d).

10. The method of any one of claims 1-9, further comprising depleting ribosomal RNA from the isolated or purified RNA molecules after step (c), but prior to step (d).

11. The method of any one of claims 1-10, wherein labeling the RNA transcripts bound by the residual peptides comprises heat denaturing the isolated or purified RNA molecules and incubating the isolated or purified RNA molecules with the chemical moiety conjugate comprising the affinity ligand in a reaction buffer supplemented with 5%-70% DMSO.

12. The method of claim 11, wherein heat denaturing comprises heating the isolated or purified RNA molecules at 50°C-100°C for 1-5 minutes.

13. The method of any one of claims 1-12, wherein the chemical moiety conjugate comprises a succinimidyl ester, a carboxylic ester, a tetrafluorophenyl ester, a sulfodichlorophenol ester, a carbonyl azide, an aldehyde, or an isothiocyanate.

14. The method of claim 13, wherein the chemical moiety conjugate is coupled to the alpha amino group of each bound residual peptide.

15. The method of any one of claims 1-12, wherein the chemical moiety conjugate comprises a primary amine, a carboiimide, or an isocyanate.

16. The method of claim 15, wherein the chemical moiety conjugate is coupled to the carboxyl group of each bound residual peptide.

17. The method of any one of claims 1-16, wherein the affinity ligand comprises biotin, a biotin derivative, digoxin, dinitrophenyl, a click chemistry reagent, sugars or peptides.

18. The method of claim 17, wherein the click chemistry reagent comprises an alkyne group, an azide group, a dibenzocyclooctyne group (DBCO), or a bicyclononyne (BCN) group.

19. The method of any one of claims 1-18, wherein fragmenting the RNA molecules comprises heating the RNA molecules at 90°C-95°C for 3-12 minutes in the presence of 10mM-30mM divalent cations.

20. The method of any one of claims 1-19, wherein the RNA fragments are about 40-90 nucleotides in length or about 80-250 nucleotides in length.

21. The method of any one of claims 1-20, further comprising identifying in the individual transcripts (i) 1-3 base pair indels that result from PRIs isolated by affinity purification, (ii) single nucleotide variants (SNVs) that result from PRIs isolated by affinity purification, or (iii) reverse transcriptase (RT) stop sites that result from PRIs isolated by affinity purification.

22. The method of any one of claims 1-21, wherein the biological sample is obtained from a subject diagnosed with a disease.

23. The method of claim 22, wherein the disease is cancer, an autoimmune disease, a metabolic disease or a neurodegenerative disease.

24. The method of claim 23, wherein the cancer is selected from among adrenal cancers, bladder cancers, blood cancers, bone cancers, brain cancers, breast cancers, carcinoma, cervical cancers, colon cancers, colorectal cancers, corpus uterine cancers, ear, nose and throat (ENT) cancers, endometrial cancers, esophageal cancers, gastrointestinal cancers, head and neck cancers, Hodgkin's disease, intestinal cancers, kidney cancers, larynx cancers, leukemias, liver cancers, lymph node cancers, lymphomas, Diffuse large B-cell lymphoma (DLBCL), lung cancers, melanomas, mesothelioma, myelomas, nasopharynx cancers, neuroblastomas, nonHodgkin's lymphoma, oral cancers, ovarian cancers, pancreatic cancers, penile cancers, pharynx cancers, prostate cancers, rectal cancers, sarcoma, seminomas, skin cancers, stomach cancers, teratomas, testicular cancers, thyroid cancers, uterine cancers, vaginal cancers, vascular tumors, and metastases thereof.

25. The method of any one of claims 1-24, wherein the biological sample is derived from breast tissue, renal tissue, uterine cervical tissue, endometrium tissue, head or neck tissue, gallbladder tissue, parotid tissue, prostate tissue, brain tissue, pituitary gland tissue, kidney tissue, muscle tissue, esophageal tissue, stomach tissue, small intestine tissue, colon tissue, liver tissue, spleen tissue, pancreatic tissue, thyroid tissue, heart tissue, lung tissue, bladder tissue, adipose tissue, lymph node tissue, uterine tissue, ovarian tissue, adrenal tissue, testis tissue, tonsils, thymus, blood, hair, buccal, skin, serum, plasma, CSF, semen, prostate fluid, seminal fluid, urine, feces, sweat, saliva, sputum, mucus, bone marrow, lymph, or tears.

26. The method of any one of claims 1-25, wherein the biological sample is a fresh or frozen sample.

27. The method of any one of claims 1-26, wherein the plurality of cDNA molecules are reverse transcribed using a truncating reverse transcriptase or a read-through reverse transcriptase.

28. The method of any one of claims 9-27, wherein the targeted bait capture reagents hybridize to one or more disease associated genes, optionally wherein the disease is cancer, autoimmune disease, neurodegenerative disease, or metabolic disease.

29. The method of any one of claims 1-28, wherein one or more of the enriched RBP bound sites detected in the individual transcripts are allele-specific and / or associated with a disease.

30. The method of any one of claims 1-29, further comprising determining the identity of a RBP bound to one or more of the enriched RBP bound sites detected in the individual transcripts.

31. The method of claim 30, comprising identifying RBP-specific crosslinking patterns within a sequence motif in the individual transcripts.

32. The method of claim 31, wherein RBP-specific crosslinking patterns are identified using eCLIP defined consensus motifs, mCross crosslinking patterns, positionally enriched k-mer analysis (PEKA), or sequence motifs defined by RBNS.

33. A method for evaluating the efficacy of a drug in inhibiting interactions between a specific RNA binding protein (RBP) and RNA transcripts comprisinga. contacting a first cell population with at least one inhibitory oligonucleotide that specifically inhibits expression of a RBP, wherein the first cell population is obtained from a biological sample;b. contacting a second cell population with a test drug, wherein the second cell population is obtained from the biological sample and is not treated with the at least one inhibitory oligonucleotide;c. performing the method of claim 1 to identify one or more RBP-RNA interaction (PRI) sites in at least one transcript derived from the first cell population, the second cell population and a control cell population, whereinthe control cell population is obtained from the biological sample and is not treated with the at least one inhibitory oligonucleotide and the test drug;d. defining RBP-specific PRIs by identifying crosslink patterns within one or more sequence motifs in the individual transcripts that (i) are present in the control cell population and (ii) absent in the first cell population; ande. determining that the test drug is effective when the RBP-specific PRIs identified in step (d) are absent in the second cell population.

34. A method for evaluating the specificity of a drug in inhibiting interactions between a specific RNA binding protein (RBP) and RNA transcripts comprisinga. contacting a first cell population with at least one inhibitory oligonucleotide that specifically inhibits expression of a RBP, wherein the first cell population is obtained from a biological sample;b. contacting a second cell population with a test drug, wherein the second cell population is obtained from the biological sample and is not treated with the at least one inhibitory oligonucleotide;c. performing the method of claim 1 to identify one or more RBP-RNA interaction (PRI) sites in at least one transcript derived from the first cell population, the second cell population and a control cell population, wherein the control cell population is obtained from the biological sample and is not treated with the at least one inhibitory oligonucleotide and the test drug;d. identifying the non-specific effects of the test drug by identifying crosslink patterns within one or more sequence motifs in the individual transcripts that (a) are present in the control cell population and the first cell population and (b) absent in the second cell population.

35. The method of claim 33 or claim 34, wherein the at least one inhibitory oligonucleotide is a siRNA, a shRNA, an antisense oligonucleotide, or a sgRNA.

36. The method of any one of claims 33-35, wherein the test drug is a small molecule, a nucleic acid or a peptide.