Artificial intelligence (AI)-informed noncoding RNA targeting

The use of BERT-RBP and BERT2-RBP models predicts RNA binding protein interactions for IncRNAs across species, addressing the challenge of sequence conservation and enabling biomarker identification.

WO2026050523A1PCT designated stage Publication Date: 2026-03-05THE UNIV OF NORTH CAROLINA AT CHAPEL HILL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/043979
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-28
Filing Date
2025-08-28
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

Existing methods struggle to accurately predict interactions between long noncoding RNA (IncRNA) sequences and RNA binding proteins due to the lack of linear sequence conservation across species, hindering the classification and characterization of IncRNAs.

Method used

Utilizing machine learning models, specifically BERT-RBP and BERT2-RBP, to predict binding sites between RNA sequences and RNA binding proteins through sliding-window segmentation, enabling the prediction of conserved interactions across multiple species based solely on sequence data.

Benefits of technology

This approach effectively identifies RNA binding protein interactions for IncRNAs, overcoming characterization roadblocks and providing insights into sequence-function relationships, even in the absence of sequence conservation, thus facilitating the identification of biomarkers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025043979_05032026_PF_FP_ABST
    Figure US2025043979_05032026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed are methods, systems, and computer readable media for predicting interactions for RNA sequences, including long noncoding RNA (IncRNA), which can be used, for example, as biomarkers. In an example method of predicting one or more binding sites between a ribonucleotide (RNA) sequence and a RNA binding protein, the method includes at a computing platform including at least one processor and memory: preparing sequence data for a plurality of overlapping RNA segments derived from a target sequence; inputting the sequence data for the plurality of overlapping RNA segments into a machine learning model trained to predict a probability of an interaction between a RNA sequence and a RNA binding protein occurring in a RNA sequence; outputting from the machine learning model one or more segments in the sequence data for the plurality of overlapping RNA segments falling above a probability threshold indicative of a region of interaction between a RNA sequence and a RNA binding protein; and predicting one or more binding sites between a RNA sequence and a RNA binding protein in the sequence data for the plurality of overlapping RNA segments derived from the target sequence based on the output from the model.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Attorney Docket No.: 421 / 553 PCT

[0002] DESCRIPTION

[0003] ARTIFICIAL INTELLIGENCE (AI)-INFORMED NONCODING RNA TARGETING

[0004] CROSS REFERENCE TO RELATED APPLICATIONS

[0005] This application claims priority to and the benefit of U.S. Provisional Patent Application Serial No. 63 / 688,057, filed August 28, 2024; the disclosure of which is incorporated herein by reference in its entirety.

[0006] STATEMENT OF GOVERNMENT SUPPORT

[0007] This invention was made with government support under Grant Numbers 2243666 and 2243562 awarded by the National Science Foundation. The government has certain rights in the invention.

[0008] REFERENCE TO SEQUENCE LISTING XML

[0009] The Sequence Listing XML associated with the instant disclosure has been electronically submitted to the United States Patent and Trademark Office via the Patent Center as a 2,755 byte UTF-8-encoded XML file created on August 28, 2025 and entitled “421-553-PCT_ST26. xml”. The Sequence Listing submitted via Patent Center is hereby incorporated by reference in its entirety.

[0010] TECHNICAL FIELD

[0011] The presently disclosed subject matter relates to evaluation of RNA sequences, including long noncoding RNA (IncRNA). In some examples, the presently disclosed subject matter relates to methods, systems, and computer readable media for predicting interactions for RNA sequences, including long noncoding RNA (IncRNA), which can be used, for example, as biomarkers.

[0012] BACKGROUND

[0013] Long noncoding RNAs (IncRNAs) have been associated with a plethora of molecular and cellular functions in both normal and disease processes (Delas and Hannon, 2017; Andergassen and Rinn, 2022; Rinn and Chang, 2012, 2020; Mattick et al., 2023). While IncRNAs are conventionally transcribed by RNA Polymerase II and bear commonalities with mRNAs including 5’ capping, similar splicing mechanistics, and a poly-A tail, they generally lack linear sequence conservation across species (Mattick et al., 2023; Rinn and Chang, 2020; Quinn and Chang, 2016). Therefore, sequence contexts that have been used to impute conserved protein domains that inform on functionality, are absent for IncRNAs (Mattick et al., 2023). Based on these limitations, general IncRNA classification approaches are lacking, even for IncRNAs with striking functional conservation and impact on organismal biology (Furlan and Rougeulle, 2016). This has implications for classifying newly identified and / or under-characterized IncRNAs. Attorney Docket No.: 421 / 553 PCT

[0014] SUMMARY

[0015] This summary lists several examples of the presently disclosed subject matter, and in many cases lists variations and permutations of these examples. This summary is merely exemplary of the numerous and varied examples. Mention of one or more representative features of a given example is likewise exemplary. Such an example can typically exist with or without the feature(s) mentioned; likewise, those features can be applied to other examples of the presently disclosed subject matter, whether listed in this summary or not. To avoid excessive repetition, this Summary does not list or suggest all possible combinations of such features.

[0016] In some examples, the presently disclosed subject matter provides a method of predicting one or more binding sites between a ribonucleotide (RNA) sequence and a RNA binding protein. In some examples, the method comprises preparing sequence data for a plurality of overlapping RNA segments derived from a target sequence; inputting the sequence data for the plurality of overlapping RNA segments into a machine learning model trained to predict a probability of an interaction between a RNA sequence and a RNA binding protein occurring in a RNA sequence; outputting from the machine learning model one or more segments in the sequence data for the plurality of overlapping RNA segments falling above a probability threshold indicative of an interaction between a RNA sequence and a RNA binding protein; and predicting one or more binding sites between a RNA sequence and a RNA binding protein in the sequence data for the plurality of overlapping RNA segments derived from the target sequence based on the output from the model.

[0017] In some examples, the method further comprises mapping the predicted binding site back to the target sequence.

[0018] In some examples, the overlapping RNA segments derived from a target sequence overlap by a tunable nucleotide parameter. In some examples, the target sequence comprises a mRNA sequence, a spliced transcript sequence, an unspliced transcript sequence, an artificial sequence, or a long noncoding RNA (IncRNA) sequence. In some examples, the target sequence comprises at least one target sequence from at least two different species. In some examples, the method further comprises inferring binding site conservation across the two or more species. In some examples, the target sequence is associated with a disease state.

[0019] In some examples, the method further comprises designing a modulator of binding at the one or more binding sites between a RNA sequence and a RNA binding protein in the sequence data. In some examples, the modulator of binding is a nucleotide-based modulator or a small molecule modulator.

[0020] In some examples, the machine learning model comprises a BERT-RBP model, a BERT2- RBP model, or a combination thereof. Attorney Docket No.: 421 / 553 PCT

[0021] In some examples, the method further comprises providing an interactive graphical user interface that loads and displays probabilities, allows for tuning of parameters, and / or displays performance characteristics about the model that was used.

[0022] In some examples, the presently disclosed subject matter provides a system for predicting one or more binding sites between a ribonucleotide (RNA) sequence and a RNA binding protein. In some examples, the system comprises: a computing platform including at least one processor and memory; and a trained machine learning model stored in the memory and executable by the at least one processor for receiving, as input, sequence data for a plurality of overlapping RNA segments into a machine learning model trained to predict a probability of an interaction between a RNA sequence and a RNA binding protein occurring in a RNA sequence and generating, as output, one or more of the plurality of overlapping RNA segments falling above a probability threshold indicative of an interaction between a RNA sequence and a RNA binding protein.

[0023] In some examples, the overlapping RNA segments derived from a target sequence overlap by a tunable nucleotide parameter. In some examples, the target sequence comprises a mRNA sequence, a spliced transcript sequence, an unspliced transcript sequence, an artificial sequence, or a long noncoding RNA (IncRNA) sequence. In some examples, the target sequence comprises at least one target sequence from at least two different species. In some examples, the target sequence is associated with a disease state.

[0024] In some examples, the machine learning model comprises a BERT-RBP model, a BERT2- RBP model, or a combination thereof.

[0025] In some examples, a method of predicting one or more binding sites between a ribonucleotide (RNA) sequence and a RNA binding protein is disclosed. In some examples, the method comprises receiving, as input, RNA sequence data; precomputing, using a machine learning model, binding probabilities for a plurality of overlapping RNA segments; outputting the binding probabilities; storing in a data structure, the binding probabilities and corresponding RNA segment identifying data; using the RNA sequence data to access the data structure and locate binding probabilities for overlapping segments in the RNA sequence data; and predicting one or more binding sites between a RNA sequence and a RNA binding protein in the RNA sequence data for the plurality of overlapping RNA segments derived from the target sequence based on the output.

[0026] In some examples, a non-transitory computer readable medium comprising computer executable instructions embodied in a computer readable medium that when executed by at least one processor of a computer cause the computer to perform steps of any of the presently disclosed methods.

[0027] In some examples, the presently disclosed subject matter provides a system and / or algorithm configured to execute the methods disclosed herein. Attorney Docket No.: 421 / 553 PCT

[0028] The systems described herein can be implemented in hardware, software, firmware, or combinations of hardware, software and / or firmware. In some examples, the systems described in this specification may be implemented using a non-transitory computer readable medium storing computer executable instructions that when executed by one or more processors of a computer cause the computer to perform operations. Computer readable media suitable for implementing the systems described in this specification include non-transitory computer-readable media, such as disk memory devices, chip memory devices, programmable logic devices, random access memory (RAM), read only memory (ROM), optical read / write memory, cache memory, magnetic read / write memory, flash memory, and application-specific integrated circuits. In addition, a computer readable medium that implements a system described in this specification may be located on a single device or computing platform or may be distributed across multiple devices or computing platforms.

[0029] As used herein, the term “node” refers to a physical computing platform including one or more processors and memory.

[0030] As used herein, the terms “function” or “module” refer to software in combination with hardware and / or firmware for implementing features described herein. In some examples, a module may include a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), or a processor.

[0031] Accordingly, it is an object of the presently disclosed subject matter to provide methods, systems, and computer readable media for predicting one or more binding sites between a ribonucleotide (RNA) sequence and a RNA binding protein. These and other objects are achieved in whole or in part by the presently disclosed subject matter. Further, other objects and advantages of the presently disclosed subject matter will become apparent to those skilled in the art after a study of the following description, Drawings and Examples.

[0032] BRIEF DESCRIPTION OF THE FIGURES

[0033] The presently disclosed subject matter can be better understood by referring to the following figures. A further understanding of the presently disclosed subject matter can be obtained by reference to the illustrations of the accompanying drawings. Although merely exemplary of systems for carrying out the presently disclosed subject matter, both the organization and method of operation of the presently disclosed subject matter, in general, together with further objectives and advantages thereof, may be more easily understood by reference to the drawings and the following description. The drawings are not intended to limit the scope of this presently disclosed subject matter, which is set forth with particularity in the claims as appended or as subsequently amended, but merely to clarify and exemplify the presently disclosed subject matter. Attorney Docket No.: 421 / 553 PCT

[0034] For a more complete understanding of the presently disclosed subject matter, reference is now made to the following drawings in which:

[0035] Figure 1A is a set of violin plots showing distribution of AUROC scores of BERT-RBP models reported by Yamada and Hamada, 2022 (left), and BERT-RBP models fine-tuned in-house (right). Black dots indicate scores of individual RBP models. Figure IB is a set of plots showing precision and recall across thresholds for in-house BERT-RBP models. Bounded area is mean + / - standard deviation across RBPs. Vertical lines labeled 0.5, 0.7 and 0.9 indicate thresholds of 0.5, 0.7 and 0.9, respectively. Precision, solid line; recall, dashed line. Figure 1C is a schematic illustrating analysis approach for running predictions on full-length RNA transcripts. Lines labeled 0.5 (light gray), 0.7 (medium gray) and 0.9 (dark gray) are as in Figure IB. Thicker bars labeled 0.5 (light gray), 0.7 (medium gray) and 0.9 (dark gray) denote predicted binding regions using the corresponding threshold. Figure ID is an example plot showing the effect of varying the subsegment spacing. Points are model outputs of individual segments; lines are smoothed with a gaussian filter (sigma=20nt). Segment spacing, 1 nucleotide (nt), squares; 5 nucleotides (nt), circles; 10 nt, pentagons; 20 nt, inverted triangles; 40 nt, diamonds.

[0036] Figure 2A is a set of genome browser shots of TARDBP showing eCLIP wiggle tracks and called peaks (top, black) relative to model predictions displayed as a BED track and plot of model output (plot, gray). Points represent the model output for the 101 nt sequence centered on the point. Lines are the result of passing model output through a 1-d gaussian filter with sigma of 20 nt. The lower plot shows predictions zoomed in on the region from ChrEl l, 082, 300-11, 083, 952 (region indicated by the gray bar). Figure 2B schematically presents motifs identified by attention analysis of BERT-RBP model for TARDBP. Figure 2C presents a sequence (SEQ ID NO: 1; corresponds to nucleotides 1636-2587 of Accession No. NM_007375.4 and nucleotides 11,022,943-11,023,894 of Accession No. NC_000001.l l of the GENBANK® biosequence database) of zoomed-in region of the human TARDBP mRNA shown in Figure 2A. Predicted TARDBP binding sites are highlighted in solid line boxes. Dashed line boxes indicate previously identified binding sites. Motifs identified in Figure 2B are shown in bold and italics. Figure 2D is a set of genome browser shots of VIM RNA and the RBP METAP2 showing eCLIP wiggle tracks and called peaks (top, black) relative to model predictions displayed as BED track and model output plot (light gray). Figure 2E is a set of genome browser shots of NDEL1 RNA and the RBP RBFOX2 showing eCLIP wiggle tracks and called peaks (top, black) relative to model predictions displayed as BED track and model output plot (light gray). For Figures 2A, 2D and 2E colors indicate, black: K562 eCLIP, dark gray: HepG2 eCLIP, light gray: BERT-RBP model predictions.

[0037] Figure 3 A is a table showing RBPs with most predicted interactions with the human IncRNA XIST as measured by Aggregated Binding, a count of the total number of nucleotides in predicted Attorney Docket No.: 421 / 553 PCT binding regions. Previously identified XIST-interacting proteins are marked with asterisks. Figure 3B is a set of genome browser shots of the XIST RNA showing eCLIP wiggle tracks and called peaks (top, black) relative to model predictions displayed as BED tracks for RBPs (light gray, HNRNPC, HNRNPK, PTBP1, MATR3), established as binding to defined XIST repeat sequences. Figure 3C is a table showing RBPs with most predicted interactions with the human IncRNA MALAT1, as in Figure 3 A. Figure 3D is a set of genome browser shots of the MALAT1 RNA showing eCLIP wiggle tracks and called peaks (top) relative to model predictions displayed as BED tracks for top RBPs in C and RBPs (TARDBP, TRA2A) shown to bind previously in eCLIP studies. Figure 3E is a table showing top predictions for the human IncRNA Hl 9, as in Figure 3 A. Figure 3F is a set of genome browser shots of the Hl 9 RNA showing eCLIP wiggle tracks and called peaks (top) relative to model predictions displayed as BED tracks for the top four RBPs. For Figures 3B, 3D and 3F colors indicate, black: K562 eCLIP, medium gray: HepG2 eCLIP, light gray: BERT- RBP model predictions.

[0038] Figure 4A is a graph showing K-mer analysis using the SEEKR algorithm on human and mouse XIST, Malatl, and H19 IncRNAs. SEEKR correlation of all human and mouse IncRNAs shown in black / gray (mean + / - stdev). Xist, X; Malatl, inverted triangle; Hl 9, gray circles. Figures 4B through 4D are graphs showing correlation analysis between RBPs predicted to interact with human and mouse Xist (Figure 4B), Malatl (Figure 4C) and Hl 9 (Figure 4D) IncRNAs. Figures 4E though 4G are plots showing predicted binding of HNRNPC with human Xist (Figure 4E), mouse Xist (Figure 4F), and opossum Rsx (Figure 4G). Figures 4H through 4J are plots showing predicted binding of MATR3 with human Xist (Figure 4H), mouse Xist (Figure 41), and opossum Rsx (Figure 4J). Cartoon illustration of repeats are shown with open boxes, hatched boxes, and stippled boxes. Figure 4K is a table showing the RBPs predicted to have the most binding with Opossum Rsx, as measured by Aggregated Binding, a count of the total number of nucleotides in predicted binding regions. Human Xist and mouse Xist columns indicate whether the RBP was predicted to bind with human and mouse Xist, respectively. Figures 4L and 4M are plots showing predicted binding of SFPQ (Figure 4L) and SRSF9 (Figure 4M) with opossum Rsx. For Figure 4E, Figure 4F, Figure 4G, Figure 4H, Figure 41, Figure 4J, Figure 4L, Figure 4M, plots include the predicted binding probability after being passed through a ID Gaussian fdter, and bars (top) indicate called regions where the filtered line passed a threshold of 0.9.

[0039] Figure 5A is a set of genome browser shots of the mouse Xist RNA showing conservation (PhyloP) and model predictions displayed as BED tracks for HNRNPC, HNRNPK, PTBP1, and MATR3, established interactors with Xist. Figure 5B is a set of genome browser shots of the mouse Malatl RNA showing conservation (PhyloP) and model predictions displayed as BED tracks. Attorney Docket No.: 421 / 553 PCT

[0040] Figure 5C is asset of genome browser shots of the mouse Hl 9 RNA showing conservation (PhyloP) and model predictions displayed as BED tracks.

[0041] Figure 6 is a set of genome browser shots illustrating RBPs which are predicted to have differential interaction propensity with the XIST RNA in human or mouse, respectively.

[0042] Figure 7 is a flow chart illustrating an exemplary process predicting for prediction one or more binding sites between a ribonucleotide (RNA) sequence and a RNA binding protein.

[0043] Figure 8 is a block diagram illustrating an exemplary system for predicting one or more binding sites between a ribonucleotide (RNA) sequence and a RNA binding protein.

[0044] Figure 9 is an image of an exemplary graphical user interface (GUI) for use in a process and / or system predicting one or more binding sites between a ribonucleotide (RNA) sequence and a RNA binding protein.

[0045] Figure 10 is a flow chart illustrating an exemplary process predicting for prediction one or more binding sites between a ribonucleotide (RNA) sequence and a RNA binding protein, using a data structure.

[0046] Figure 11A is a set of violin plots showing distribution of AUROC scores of BERT-RBP models reported by Yamada et al, and BERT-RBP and BERT2-RBP models finetuned in-house. Black dots indicate scores of individual RBP models. Figure 1 IB is a set of plots showing precision and recall across thresholds for in-house BERT-RBP and BERT2-RBP models. Bounded is mean + / - standard deviation across RBPs. Vertical lines labeled 0.5, 0.7 and 0.9 indicate thresholds of 0.5, 0.7 and 0.9, respectively. Precision, solid line; recall, dashed line. Figure 11C is a schematic detailing analysis approach for running predictions on RNA transcripts. Lines labeled 0.5 (light gray), 0.7 (medium gray) and 0.9 (dark gray) are as in Figure IB. Thicker bars labeled 0.5 (light gray), 0.7 (medium gray) and 0.9 (dark gray) denote predicted binding regions using the corresponding threshold.

[0047] Figure 12A is a set of genome browser shots of TARDBP showing eCLIP wiggle tracks and called peaks (top) relative to model predictions displayed as BED tracks and plots of model output. Points represent the model output for the 101 nt sequence centered on the point. Lines are the result of passing model output through a 1-d gaussian filter with sigma of 20 nt. The lower plot shows the zoomed in predictions in the region from Chrl : 11,082,300-11,083952 (indicated by the gray bar). Figure 12B schematically presents motifs identified by attention analysis of BERT-RBP model for TARDBP. Figure 12C presents a sequence (SEQ ID NO:1) of zoomed-in region of TARDBP mRNA shown in Figure 12A. Predicted TARDBP binding sites are highlighted in solid line boxes. Dashed line boxes indicate previously identified binding sites. Motifs identified in Figure 12B are shown in bold and italics. Figure 12D is a set of genome browser shots of VIM RNA and the RBP METAP2 showing eCLIP wiggle tracks and called peaks (top) relative to model Attorney Docket No.: 421 / 553 PCT predictions displayed as BED tracks and model output plots. Figure 12E is a set of genome browser shots of NDEL1 RNA and the RBP RBF0X2 showing eCLIP wiggle tracks and called peaks (top) relative to model predictions displayed as BED tracks and model output plots (intronic sequence representation shown as triangles). In Figures 12A, 12D and 12E colors indicate: black: K562 eCLIP, medium gray: HepG2 eCLIP, light gray: BERT-RBP model predictions, dark gray: BERT2-RBP model predictions, darker gray: Intersection of BERT-RBP and BERT2-RBP model predictions.

[0048] Figure 13 A is a table showing RBPs with most predicted interactions with the human IncRNA XIST as measured by Aggregated Region Lengths, a count of the number of nucleotides in regions predicted by both BERT-RBP and BERT2-RBP models. Figure 13B is a set of genome browser shots of the XIST RNA showing eCLIP wiggle tracks and called peaks (top) relative to model predictions displayed as BED tracks for RBPs (HNRNPC, HNRNPK, PTBP1), established as binding to defined XIST repeat sequences and the top predicted interacting RBP, MATR3. Figure 13C is a table showing RBPs with most predicted interactions with the human IncRNA Malatl, as in Figure 13 A. Figure 13B is a set of genome browser shots of the MALAT1 RNA showing eCLIP wiggle tracks and called peaks (top) relative to model predictions displayed as BED tracks for top RBPs in C and RBPs (TARDBP, TRA2A) shown to bind previously in eCLIP studies. Figure 13E is a table showing top predictions for the human IncRNA Hl 9, as in Figure 13 A. Figure 13F is a set of genome browser shots of the Hl 9 RNA showing eCLIP wiggle tracks and called peaks (top) relative to model predictions displayed as BED tracks for the top four RBPs. In Figures 13B, 13D and 13F colors indicate: black: K562 eCLIP, medium gray: HepG2 eCLIP, light gray: BERT-RBP model predictions, dark gray: BERT2-RBP model predictions, darker gray: Intersection of BERT-RBP and BERT2-RBP model predictions.

[0049] Figures 14A to 14C are a set of graphs showing K-mer analysis using the SEEKR algorithm on XIST (Figure 14A), Malatl (Figure 14B), and H19 (Figure 14C) IncRNAs. Raw, solid line and solid circles; normalized, dashed line and open circles. Figure 14D is a graph showing correlation analysis between top RBPs predicted to interact with human and mouse XIST IncRNA. Figure 14E is a set of genome browser shots of the mouse XIST RNA showing conservation (PhyloP) and model predictions displayed as BED tracks for HNRNPC, HNRNPK, PTBP1 as established interactors with XIST, and the top predicted interacting RBP in human, MATR3. Figure 14F is a graph showing correlation analysis between top RBPs predicted to interact with human and mouse Malatl IncRNA. Figure 14G is a set of genome browser shots of the mouse Malatl RNA showing conservation (PhyloP) and model predictions displayed as BED tracks showing top predictions in human for comparison, and top predicted interacting RBPs in mouse. Figure 14H is a graph showing a correlation analysis between top RBPs predicted to interact with human and mouse Hl 9 IncRNA. Attorney Docket No.: 421 / 553 PCT

[0050] Figure 141 is a set of genome browser shots of the mouse H19 RNA showing conservation (PhyloP) and model predictions displayed as BED tracks, showing top predictions in both mouse and human for comparison, and top predicted interacting RBPs in mouse. In Figures 14E, 14G and 141 colors indicate: light gray: BERT-RBP model predictions, medium gray: BERT2-RBP model predictions, dark gray: Intersection of BERT-RBP and BERT2-RBP model predictions.

[0051] DETAILED DESCRIPTION

[0052] The presently disclosed subject matter pertains in some examples to a toolkit and methodology based on computational algorithms, including machine learning, that focus on the outputs from the noncoding genome. This toolkit can be used to identify and predict properties for noncoding RNAs. One use-case of the toolkit is the identification of biomarker candidates and predictions of their properties that could be used to expand the biomarker development pipeline to support precision medicine.

[0053] In some examples, the presently disclosed subject matter relates to prediction of conserved interactions for long noncoding RNAs via natural language processing. Long noncoding RNA (IncRN A) genes outnumber mRNA genes in the human genome and the majority remain uncharacterized. A major difficulty in generalizing understanding of IncRNA function is the absence of gross sequence conservation, both for IncRNAs across species and for IncRNAs that perform similar functions within a species. Machine learning based methods which harness vast amounts of information on RNAs are increasingly used to impute certain biological characteristics. This includes interactions with proteins that are important mediators of RNA function, thus enabling the generation of knowledge in contexts for which experimental data are lacking. In accordance with some examples of the presently disclosed subject matter, a natural language-based machine learning approach was applied that provided for the identification of RNA binding protein interactions in IncRNA transcripts, using only RNA sequence as an input. Moreover, it was found that this predictive method is a powerful approach to infer conserved binding across species as distant as human and opossum, even in the absence of sequence conservation, thus informing on sequence-function relationships for these poorly understood RNAs. In some examples, these sequence-function relationships can serve as biomarkers.

[0054] In accordance with some examples of the presently disclosed subject matter, an approach to use BERT-RBP and / or BERT2-RBP is provided to address the IncRNA-functional prediction problem, applying sliding- window segmentation to facilitate RBP-lncRNA interaction predictions in full transcripts. We found we were able to predict conserved IncRNA interactions in multiple species, thereby overcoming roadblocks in IncRNA characterization. Additionally, since it uses only sequence as input, this is an accessible approach that is applicable across a range of species. Attorney Docket No.: 421 / 553 PCT

[0055] The presently disclosed subject matter now will be described more fully hereinafter, in which some, but not all examples of the presently disclosed subject matter are described. Indeed, the presently disclosed subject matter can be embodied in many different forms and should not be construed as limited to the examples set forth herein; rather, these examples are provided so that this disclosure will satisfy applicable legal requirements.

[0056] I Definitions

[0057] The terminology used herein is for the purpose of describing particular examples only and is not intended to be limiting of the presently disclosed subject matter.

[0058] While the following terms are believed to be well understood by one of ordinary skill in the art, the following definitions are set forth to facilitate explanation of the presently disclosed subject matter.

[0059] All technical and scientific terms used herein, unless otherwise defined below, are intended to have the same meaning as commonly understood by one of ordinary skill in the art. References to techniques employed herein are intended to refer to the techniques as commonly understood in the art, including variations on those techniques or substitutions of equivalent techniques that would be apparent to one of skill in the art. Accordingly, for the sake of clarity, this description will refrain from repeating every possible combination of the individual steps in an unnecessary fashion. Nevertheless, the specification and claims should be read with the understanding that such combinations are entirely within the scope of the invention and the claims.

[0060] Following long-standing patent law convention, the terms “a”, “an”, and “the” refer to “one or more” when used in this application, including the claims. Thus, for example, reference to "a cell" includes a plurality of such cells, and so forth.

[0061] Unless otherwise indicated, all numbers expressing quantities of ingredients, reaction conditions, and so forth used in the specification and claims are to be understood as being modified in all instances by the term “about”. Accordingly, unless indicated to the contrary, the numerical parameters set forth in this specification and attached claims are approximations that can vary depending upon the desired properties sought to be obtained by the presently disclosed subject matter.

[0062] As used herein, the term “about,” when referring to a value or to an amount of a composition, dose, sequence identity (e.g., when comparing two or more nucleotide or amino acid sequences), mass, weight, temperature, time, volume, concentration, percentage, etc. , is meant to encompass variations of in some examples ±20%, in some examples ±10%, in some examples ±5%, in some examples ±1%, in some examples ±0.5%, and in some examples ±0.1% from the specified amount, as such variations are appropriate to perform the disclosed methods or employ the disclosed compositions. Attorney Docket No.: 421 / 553 PCT

[0063] As used herein, ranges can be expressed as from “about” one particular value, and / or to “about” another particular value. It is also understood that there are a number of values disclosed herein, and that each value is also herein disclosed as “about” that particular value in addition to the value itself. For example, if the value “10” is disclosed, then “about 10” is also disclosed. It is also understood that each unit between two particular units are also disclosed. For example, if 10 and 15 are disclosed, then 11, 12, 13, and 14 are also disclosed.

[0064] The term “comprising”, which is synonymous with “including” “containing” or “characterized by” is inclusive or open-ended and does not exclude additional, unrecited elements or method steps. “Comprising” is a term of art used in claim language which means that the named elements are essential, but other elements can be added and still form a construct within the scope of the claim.

[0065] As used herein, the phrase “consisting of’ excludes any element, step, or ingredient not specified in the claim. When the phrase “consists of’ appears in a clause of the body of a claim, rather than immediately following the preamble, it limits only the element set forth in that clause; other elements are not excluded from the claim as a whole.

[0066] As used herein, the phrase “consisting essentially of’ limits the scope of a claim to the specified materials or steps, plus those that do not materially affect the basic and novel characteristic(s) of the claimed subject matter.

[0067] With respect to the terms “comprising”, “consisting of’, and “consisting essentially of’, where one of these three terms is used herein, the presently disclosed and claimed subject matter can include the use of either of the other two terms.

[0068] As used herein, the term “and / or” when used in the context of a listing of entities, refers to the entities being present singly or in combination. Thus, for example, the phrase “A, B, C, and / or D” includes A, B, C, and D individually, but also includes any and all combinations and subcombinations of A, B, C, and D.

[0069] II. Methods. Systems, and Computer Readable Media

[0070] Given the difficulties in assigning IncRNA function from sequence information, there is a need for methodology to characterize the plethora of uncharacterized IncRN As in humans and other species (Mattick et al., 2023). Emerging approaches are based on short sequence motifs (Kirk et al., 2018; Ross et al., 2021), which suggest that IncRNAs may be functionally classified based on related motifs / k-mers (Kirk et al. 2018; Ross et al. 2021). These k-mers are often interaction sites for RNA-binding proteins (RBPs) (Lambert et al., 2014; Kuret et al., 2022), which are regulatory mediators and functional partners of IncRNAs (Briata and Gherzi, 2020; Huang et al., 2021; Ferre et al., 2016; Noh et al., 2018). Attorney Docket No.: 421 / 553 PCT

[0071] High-throughput approaches have been used to experimentally map RBP-RNA interactions (Ule et al., 2018; Kuret et al., 2022; Van Nostrand et al., 2016, 2020b). These results have been used to provide input for machine learning to generate predictions in the absence of direct experimental data (Pan et al., 2019; Moore and ’t Hoen, 2019; Horlacher et al., 2023). Recently self-attention dependent, deep learning methods have shown promise for complex language tasks using massive sequence-based datasets, including entire genomes (luchi et al., 2021; Ji et al., 2021). Bidirectional encoder representations from transformer (BERT) has been pretrained on massive corpora of DNA sequence to create DNABERT (Zhou et al., 2023; Ji et al., 2021; Devlin et al., 2018). While DNABERT predicts regulatory regions such as promoters and transcription factor binding sites (Ji et al. 2021; luchi et al. 2021), further fine-tuning on RBP-RNA binding data to create BERT-RBP (Yamada and Hamada, 2022), attempts prediction of whether short RNA sequences bind specific RBPs.

[0072] In some examples, the presently disclosed subject matter provides an approach to use BERT-RBP (Yamada and Hamada, 2022) to address the IncRN A-functional prediction problem, applying sliding-window segmentation to facilitate RBP-lncRNA interaction predictions in full transcripts. The presently disclosed subject matter provides for the prediction of conserved IncRNA interactions in multiple species, thereby overcoming roadblocks in IncRNA characterization. Additionally, since it uses only sequence as input, this is an accessible approach that is applicable across a range of species.

[0073] Disclosed herein is the use of the BERT-derived natural language processing (NLP) approaches, BERT-RBP and BERT2-RBP, which are based on DNABERT / DNABERT2 respectively (Yamada and Hamada 2022; Ji et al. 2021; Zhou et al. 2023), to predict RBP-lncRNA interactions across mammalian species. In some examples, the presently disclosed subject matter pertains to the use of these predictive models, which run predictions solely based on sequence input, to overcome initial roadblocks pertaining to functional likelihoods for novel and uncharacterized IncRNAs across a range of species, thereby overcoming a major limitation in the study of IncRNAs.

[0074] Figure 7 is a flow chart illustrating an exemplary process for predicting one or more binding sites between a ribonucleotide (RNA) sequence and a RNA binding protein. Referring to Figure 7, step 700 comprises, optionally at a computing platform including at least one processor and memory, preparing sequence data for a plurality of overlapping RNA segments derived from a target sequence. For example, the sequences can be divided in preparation for running the models. Indeed, the presently disclosed subject matter provides division & assembly methods for handling long RNA sequences. In some examples, the overlapping RNA segments derived from a target sequence overlap by a tunable nucleotide parameter. A representative tunable nucleotide parameter is overlap by nucleotides of tunable length, such as but not limited to 1-40 nucleotides, including Attorney Docket No.: 421 / 553 PCT

[0075] 10 nucleotides. See, e.g., Figure ID. The RNA segments themselves typically range in length from about 50 to about 200 nucleotides, with 101 nucleotides being a representative, non-limiting length.

[0076] Continuing with reference to Figure 7, step 702 comprises inputting the sequence data for the plurality of overlapping RNA segments into a machine learning model trained to predict a probability of an interaction between a RNA sequence and a RNA binding protein occurring in a RNA sequence. In some examples, the machine learning module is BERT-RBP, BERT2-RBP, or a combination thereof. For example, BERT-RBP and BERT2-RBP can be run individually. Then, after running individually, an ensemble approach can be implemented, arriving at an intersection between both models.

[0077] Continuing with reference to Figure 7, step 704 comprises outputting from the machine learning model one or more of the plurality of overlapping RNA segments falling above a probability threshold indicative of an interaction between a RNA sequence and a RNA binding protein. For example, then, output from the machine learning model is a probability of a segment (e.g., 101 nt) binding. In some examples, the presently disclosed method looks at probabilities for overlapping segments and use their correlation to call binding regions.

[0078] Continuing with reference to Figure 7, step 706 comprises predicting one or more binding sites between a RNA sequence and a RNA binding protein in the sequence data for the plurality of overlapping RNA segments derived from the target sequence based on the output from the model. For example, the site- or region-related analyses happens before (breaking up into component sequences) and after (correlating the overlapping segment probabilities) the machine learning step. In some examples, the method further comprises mapping the predicted binding site back to the target sequence. For example, coordinate tags, such (chromosomal) coordinate tags, associated with the sequence, that can be used to map back to the genome / transcriptome.

[0079] Figure 8 is a block diagram illustrating an exemplary system for predicting one or more binding sites between a ribonucleotide (RNA) sequence and a RNA binding protein. Referring to Figure 8, a computing platform 800 includes at least one processor 802 and memory 804. Computing platform 800 also includes a trained machine learning model 806. In one example, machine learning model 806 may be a trained machine learning model stored in the memory and executable by the at least one processor for receiving, as input, sequence data for a plurality of overlapping RNA segments into a machine learning model trained to predict a probability of an interaction between a RNA sequence and a RNA binding protein occurring in a RNA sequence and generating, as output, one or of the plurality of overlapping RNA segments falling above a probability threshold indicative of an interaction between a RNA sequence and a RNA binding protein. Machine learning model 806 may be implemented using computer executable instructions Attorney Docket No.: 421 / 553 PCT stored in memory 804 and executed by processor 802. Processor 802 may be a GPU. By way of example, a python script and model are stored on disk and when the script is executed, it loads the model into memory and then runs sequences through the model.

[0080] In some examples, the presently disclosed subject matter provides a visualization approach, such as an interactive graphical user interface (GUI) that loads and displays probabilities 902. Referring now to Figure 9, GUI 900 allows for tuning of parameters 904, e.g. thresholds and also displays performance characteristics 906 about the model that was used. To illustrate, this is how the plots displayed in Figures 2A and 4E for example, are generated.

[0081] In some examples the presently disclosed subject matter provides a method of predicting one or more binding sites between a ribonucleotide (RNA) sequence and a RNA binding protein, employing a data structure for storing probabilities and their associated sequences. Referring now to Figure 10, step 1000 comprises receiving, as input, RNA sequence data. Step 1002 comprises precomputing, using a machine learning model, binding probabilities for a plurality of overlapping RNA segments. Step 1004 comprises outputting the binding probabilities. Step 1006 comprises storing in a data structure, the binding probabilities and corresponding RNA segment identifying data. Step 1008 comprises using the RNA sequence data to access the data structure and locate binding probabilities for overlapping segments in the RNA sequence data. Step 1010 comprises predicting one or more binding sites between a RNA sequence and a RNA binding protein in the RNA sequence data for the plurality of overlapping RNA segments derived from the target sequence based on the output.

[0082] A non-transitory computer readable medium is also disclosed. In some examples, the non- transitory computer readable medium comprises computer executable instructions embodied in a computer readable medium that when executed by at least one processor of a computer cause the computer to perform steps as disclosed herein.

[0083] In some examples, the target sequence comprises a mRNA sequence, a spliced transcript sequence, an unspliced transcript sequence, an artificial sequence, or a long noncoding RNA (IncRNA) sequence. In some examples, the target sequence comprises at least one target sequence from at least two different species. In some examples, this can only be 1 sequence. In some examples, this can be two or more sequences from the same species, which can be compared. In some examples, this can be two or more sequences from different species, i.e. two or more different species, such as but not limited to one sequence each from the two or more different species, which can then be compared. This can grow depending on the use-case. In some examples, then, the method further comprises inferring binding site conservation across the two or more species.

[0084] In some examples, the target sequence is associated both normal and disease processes with a disease state. Any disease state as would be apparent to one of ordinary skill in the art upon a Attorney Docket No.: 421 / 553 PCT review of instant disclosure is provided in accordance with the presently disclosed subject matter. In some examples, the disease state is a cancer, such as prostate cancer and neurological diseases and rare diseases.

[0085] In examples, the presently disclosed methods further comprise designing a modulator of binding at the one or more binding sites between a RNA sequence and a RNA binding protein in the sequence data. In some examples, the modulator of binding is a nucleotide-based modulator or a small molecule modulator. Any suitable approach for designing a modulator can be applied. For example, CRISPR techniques can be employed for nucleotide-based modulators.

[0086] EXAMPLES

[0087] The following examples are included to further illustrate various aspects of the presently disclosed subject matter. However, those of ordinary skill in the art should, in light of the present disclosure, appreciate that many changes can be made in the specific Examples which are disclosed and still obtain a like or similar result without departing from the spirit and scope of the presently disclosed subject matter.

[0088] Materials and Methods Employed in EXAMPLES

[0089] Training data preparation

[0090] The benchmark training dataset was downloaded from RBPSuite (Pan et al., 2020) and comprised sequences around eCLIP peaks from K562 and HepG2 cell lines for 154 RBPs (Van Nostrand et al., 2016; Kagda et al., 2023). The positive dataset comprised up to 60,000 nucleotide (nt) sequences of 101 nt, derived from eCLIP peaks for each RBP. The negative dataset comprised matched regions without any peaks from the same gene. For fine-tuning, we randomly selected 15,000 positive and negative sequences and split them into training, evaluation, and test sets, using scripts from BERT-RBP (Yamada and Hamada, 2022).

[0091] Fine-tuning BERT-RBP models

[0092] BERT-RBP models were fine-tuned from the pretrained 3-mer DNABERT model (Ji et al., 2021) following instructions from Yamada & Hamada 2022 (https: / / github.com / kkyamada / bert- rbp). We first used default hyperparameters of 4 GPUs with a per-gpu-batch-size of 32 for a total batch size of 128. However, this method gave unstable results as also found by Horlacher et al., 2023 (with 16 / 154 models lacking any classification ability). We therefore trained additional model sets with batch sizes of 192 and 256 (using 6 and 8 GPUs, respectively), which showed improved but not perfect stability. Based on training differences, we selected models trained with batch sizes of 192, except when their AUROC negatively deviated from previously reported values by more than 0.005, in which case we used the higher performing of the 128 or 256 batch size model (Table 1). Attorney Docket No.: 421 / 553 PCT

[0093] Evaluating model performance

[0094] BERT-RBP model performance (Yamada and Hamada, 2022) was evaluated using withheld test data, with outputs normalized to a range of 0-1 using Min-Max normalization. Performance metrics including accuracy, precision, recall, Fl score, Matthew’s correlation coefficient, average precision (area under precision-recall curve) and AUROC were calculated using the scikit-leam python package, version 1.3.0 (Table 2).

[0095] Using models to predict binding for RNA transcripts

[0096] Training data were RNA sequences of 101 nucleotides (Pan et al., 2020). We generated segments of overlapping 101 nt sequences, used a lOnt sliding window for segments to overlap neighbors by 90 nt, and ran each segment through the given BERT-RBP model to obtain a binding prediction for that segment. We used Min-Max normalization for model outputs, with minimum and maximum values from model predictions on the test dataset. To facilitate comparisons, we used a 1-D gaussian filter (scipy.ndimage.gaussian_filterld, scipy version 1.10.1) with a sigma of 20 nt and linearly interpolated results. Regions of the resulting curve above a threshold of 0.9 are classified as binding regions.

[0097] Sequence Acquisition

[0098] All sequences were downloaded with exons and introns demarcated (Kent et al., 2002) as detailed in Table 4 using UCSC’s Table Browser tool (Karolchik et al., 2004).

[0099] Motif Analysis

[0100] Motif analysis from BERT-RBP models was done similar to prior work (Ji et al., 2021; Yamada and Hamada, 2022), with modifications to how motif candidates were merged. Briefly, attention vectors from the CLS token to the last layer of the model were extracted for each sequence in the test dataset and summed across attention heads. For positive sequences, we selected contiguous 5-10 nucleotide regions where attention was either greater than the average attention for the sequence or greater than 10 times the minimum attention for the sequence. For these putative motifs, we performed a hypergeometric test to test if it occurs more often in positive sequences than in negative ones, kept motifs where the p-value from the hypergeometric test is <0.005 after correction for multiple tests, and merged similar motifs together using an Unweighted Pair Group Method with Arithmetic mean (UPGMA) algorithm. Similarities between each motif were measured using the alignment score from biopython’s Pairwise Aligner. The similarity matrix was then inverted to create a distance matrix, which was put into scipy ’s UPGMA clustering algorithm (scipy. cl uster.hierarchy. average), with motifs grouped until the distance between groups reached a threshold corresponding to the inversion of a minimum alignment score of 5. For alignment scoring, matching nucleotides scored 2, mismatches scored -1, open and extended gaps scored -0.75 and internal gaps were not allowed. Alignment scores were adjusted for length so that scores for longer Attorney Docket No.: 421 / 553 PCT motifs were similar to scores for shorter motifs. For each motif group, component sequences were collected and extended to 12 nt. Depictions of motifs were generated from these sequences using WebLogo version 3.7.12.

[0101] EXAMPLE 1

[0102] Evaluation of model performance and practical application for full-length transcripts

[0103] We fine-tuned DNABERT-derived (Ji et al., 2021), BERT-RBP models following procedures from Yamada and Hamada (2022), with customized hyperparameters (see Methods). Training data originated from ENCODE eCLIP data for 154 RBPs (Pan et al. 2020; Yamada and Hamada 2022; Van Nostrand et al. 2016), for sequences around eCLIP peak sites in K562 and HepG2 cells as the positive class, and matched genic regions without peaks as the negative class.

[0104] We utilized Area Under the Receiver Operating Characteristic (AUROC) scores to evaluate model performance, and scores were comparable to original reports (Yamada and Hamada, 2022) (Yamada & Hamada: 0.786 + / - 0.041; In-house: 0.791 + / - 0.037, two-sided t-test p-value: 0.250) (Figure 1A).

[0105] BERT-RBP prediction outputs for a 101 nt RNA sequence is a binding probability ranging from 0 to 1. We selected thresholds to generate binary binding / not-binding decisions from binding probabilities and used precision-recall curves to guide threshold selection decisions (Figure IB). Lower threshold values prioritize recall while higher threshold values prioritize precision (Figures 1B-1C), and we used a threshold of 0.9 to prioritize precision of most-likely binding sites vs. capturing all binding sites.

[0106] Given that RNA transcript lengths vary significantly (Quinn and Chang, 2016; Ransohoff et al., 2018; St Laurent et al., 2015; Mattick et al., 2023), we developed a segmentation approach to run predictions on longer sequences using overlapping 101 nt sequence segments and a 10 nt sliding window, based on abalance between output and computational efficiency (Figures 1C-1D). Our primary goal was to develop a methodology to guide assessment of IncRNA function. To call predicted binding sites, we passed the model output through a gaussian filter (sigma: 20 nt) (Figures 1C-1D), set a threshold of 0.9, and defined regions above threshold as probable binding sites.

[0107] EXAMPLE 2

[0108] Models predict known RBP interactions with mRNAs

[0109] RBPs interact with RNAs to regulate cellular processes such as splicing and translation (Gerstberger et al., 2014; Hentze et al., 2018). We tested predictive ability using the established model of autoregulation of TARDBP via its 3’UTR (Ayala et al., 2011; Sun et al., 2014; Bhardwaj et al., 2013) and successfully predicted interactions over a region of approximately 500 nt in the 3’UTR of TARDBP, similar to a region of enrichment identified in prior work (Wolin et al., 2023; Van Nostrand et al., 2016) (Figure 2A). We found that model-derived TARDBP binding motifs Attorney Docket No.: 421 / 553 PCT were similar to those previously identified experimentally (Figure 2B). We queried motif locations within predicted interacting sites and found interactions within a 600 nt region enriched with GAAUG and (UG)n repeat motifs (Bhardwaj et al., 2013). Motif sites were within or proximal to predicted binding regions (Figure 2C; SEQ ID NO:1) and included well-known TARDBP interaction elements including U rich sequences, GU repeats, and the GAAUG motif (Figure 2C), relative to a neighboring region with low probability of predicted interactions (Figure 2C).

[0110] We also examined RBPs with different binding characteristics and functional properties (Van Nostrand et al., 2020b; Briata and Gherzi, 2020) to test the capacity to predict interactions among different RBP / RNAs (Figures 2D-2E). Relative to ENCODE eCLIP peaks, we successfully predicted interactions in similar patterns as ENCODE e-CLIP data for METAP2 and VIM (Figure 2D), and RBFOX2 and NDEL1 (Figure 2E) (Van Nostrand et al., 2016; Kagda et al., 2023).

[0111] EXAMPLE 3

[0112] Models predict biologically meaningful interactions in IncRNAs

[0113] A primary goal of the present study was to determine if this predictive approach could inform on functional and / or regulatory aspects of IncRNA biology. To test this, we used highly studied candidate RNAs, beginning with the IncRNA XIST, which has well-defined roles in X- chromosome inactivation (Loda and Heard, 2019; Sahakyan et al., 2018; Furlan and Rougeulle, 2016). Examination of the top 20 predicted interactions revealed RBPs including MATR3 , KHSRP, PTBP1, and HNRNPC (Figure 3 A) that were previously shown to interact with XIST in various contexts (Minajigi et al., 2015; Chu et al., 2015; Teng et al., 2020; Li et al., 2014). Moreover, model predictions recapitulated enrichment patterns for XIST’s modular repeat domains (Brockdorff, 2002; Brockdorff et al., 2020), with HNRNPC binding within Repeat A, HNRNPK within midregions containing Repeats B, C, and D, and PTBP1 within Repeat E (Fig. 3B). Interaction sites for another top interactor, MATR3, overlapped with PTBPl’s predicted interaction in Repeat E, similar to prior findings (Pandya- Jones et al., 2020; Jacobson et al., 2022).

[0114] MALAT1 differs from XIST in its sequence and structural characteristics, and has roles in splicing, transcription, and as a competing endogenous RNA, with numerous disease implications (Arun et al., 2020; Zhang et al., 2017; Wu et al., 2015). We predicted differentially localized interactions, consistent with different RBP functions (Figures 3C-3D), as seen with WDR3 and HLTF and previously described interactors, TARDBP and TRA2A (Van Nostrand et al., 2016; Kagda et al., 2023). This included relatively more 5’-localized predictions for WDR3, with similar predictions for HLTF (Figure 3D). Predictions for HLTF were intriguing since they were similar to CLIP-seq peaks in K562 cells but different from more widespread binding in HepG2 cells. We predicted more discrete 3’ binding for TARDBP, relative to broader binding to the 5’ of the Attorney Docket No.: 421 / 553 PCT transcript for TRA2A (Figure 3D), similar to experimentally mapped interactions (Van Nostrand et al., 2016; Kagda et al., 2023).

[0115] We examined interactions for Hl 9 (Figure 3E), which differs from XIST and MALAT1 based on its shorter length, cytoplasmic subcellular location, and associated roles including functioning as a miRNA precursor (Jonas et al., 2020; Noh et al., 2018; Liang et al., 2015). Among top predicted interactors, GRSF1 is known to bind to IncRNAs as it has been shown to interact with the IncRNA RMRP for its mitochondrial localization, while EWSR1 has been implicated in cancer biology like H19 (Liang et al., 2015; Matouk et al., 2015; Lee et al., 2019). GRSF1 and EWSR1 were predicted to interact at the 5’ end of the transcript, while LARP7 and EIF4G2 were predicted to show more dispersed binding (Figure 3F), similar to that observed for ENCODE e-CLIP.

[0116] Altogether, these data suggest that our approach is able to predict RBP interactions with IncRNAs of varying lengths, and different cellular / molecular regulatory and functional aspects.

[0117] EXAMPLE 4

[0118] Models predict RBP-lncRNA interactions across species

[0119] A major gap for the IncRNA field is the ability to predict functional properties for IncRNAs a priori, particularly across species as can be done for mRNAs (Necsulea et al., 2014; Johnsson et al., 2014). One approach to deduce functional conservation could be through identifying conserved interactions with proteins. Since many RNA binding proteins bind through short motifs (Lambert et al., 2014; Kuret et al., 2022), we undertook k-mer content comparisons of human and mouse XIST, MALAT1, and Hl 9 RNAs and found similar sequence composition, particularly for MALAT1 and XIST (Figure 4 A), suggesting functional commonalities similar to prior observations (Kirk et al., 2018). However, the sequence content alone did not easily inform on possible functional and / or regulatory interactions, particularly between the species. We postulated that RBP interaction predictions could fdl this gap (Huang et al., 2021; Briata and Gherzi, 2020), given that RBPs are functionally conserved across vertebrates (Gerstberger et al., 2014).

[0120] Even though our approach is based on human RBP CLIP-Seq-derived training data (Ji et al., 2021; Yamada and Hamada, 2022), we sought to determine whether sequence-based machine learning could be used beyond human sequence. We began by examining whether we could predict interactions with mouse sequence and found congruence in predictions of top RBP interactors between mouse and human Xist and Malatl, and less so for Hl 9 (Figures 4B-4D). This suggests models’ capacity to predict protein-lncRNA interactions for species beyond human, and to glean potential cross-species differences.

[0121] We next examined well-characterized binding characteristics using Xist. Despite profound roles in dosage compensation, there is sequence divergence across human and mouse Xist (see PhyloP conservation, Figure 5). Despite this, for mouse Xist, we found overlap with 15 of the top Attorney Docket No.: 421 / 553 PCT

[0122] 20 interacting RBPs identified for human (Figure 4B). The exceptions for mouse were at positions 32 (HNRNPC), 82 (DDX59), 22 (HNRNPK), 72 (CPEB4) and 37 (NSUN2) (Figure 6). Differences could be due to model limitations or could point to species-specific interactions and functional differences. Indeed, further examination revealed general similarities in binding as well as speciesspecific intricacies, including differential binding propensity (Figure 6).

[0123] We also examined interactions for Malatl and found common predictions between human and mouse for 17 of the top RBPs including for HLTF and WDR3 (Figure 4C). Interestingly, prediction patterns showed similarity between human and mouse even in 5’ regions of the RNA that lack sequence conservation between human and mouse (Figure 5B, see PhyloP track). Binding predictions for WDR3 were enriched in the 5’ region of the transcript, which was also where most HLTF predicted binding was observed, similar to what was seen for K562 enrichment in human MALAT1, and suggesting that this region is a highly regulated MALAT1 segment (Figure 5B). Predicted binding patterns for previously shown interactors, TARDBP and TRA2A (Van Nostrand et al., 2016), were also similar to human binding patterns. This confirms the ability of models to identify possible RBP-mediated functional conservation without linear sequence conservation.

[0124] Of the top 20 predicted interactions, we found 10 common between human and mouse Hl 9 (Figure 4C). H19’s expression pattern and functions are conserved, with roles in early embryonic development and cancers (Matouk et al., 2015; Jonas et al., 2020; Noh et al., 2018). Even though there were fewer common predictions for Hl 9, we found similar binding patterns between human and mouse for some RBPs, e.g. EWSR1 (Figure 5C), suggesting this approach can be used to query cross-species similarities and differences in function.

[0125] Using this foundation of human training data facilitating predictions for mouse sequence, we further probed applicability across evolutionary distance. We compared model predictions for human and mouse Xist, particularly repeat region predictions, to Rsx, a IncRNA that mediates X chromosome inactivation in marsupials. Rsx uses mechanisms similar to Xist, including recruitment of silencing factors through repeat domains, despite having no overt linear sequence conservation with Xist (Sahakyan et al., 2018; Loda and Heard, 2019; Furlan and Rougeulle, 2016).

[0126] First, as seen with human (Figure 3B, Figure 4E), we found that models successfully predicted commonalities for repeat regions in mouse Xist (Figure 4F), including HNRNPC being predicted to interact with Repeat A, albeit less so in mouse (Figure 4F, Figure 6), and MATR3 predicted interactions with Repeat E (Figure 4H-I). Relative to the 5’ enrichment of binding to Xist in human and mouse (Figure 4E-F), we also predicted interactions with HNRNPC for opossum Rsx in the 3’ end of the transcript (Figure 4G). These regions of Rsx and Xist were shown to have similar motif binding patterns in prior work (Sprague et al., 2019). Rsx repeat domains have been found to be more expansive than Xist (Grant et al., 2012; Sprague et al., 2019), aligning with our Attorney Docket No.: 421 / 553 PCT interaction predictions (Figure 4G). Similarly, predictions for interactions with MATR3 were also enriched in the 3’ repeat region of Rsx (Figure 4J), suggesting more similarities also with this region and Repeat E of Xist.

[0127] We found common predicted interactions with Rsx and human and mouse Xist, with prominent RBPs including TRA2A, FUBP3, SFPQ, and HNRNPK (Figure 4K, Table 3), which were also recently identified in Rsx interactome mapping (McIntyre et al., 2024). Intriguingly, model predictions also indicated striking differences based on interaction potential between 5’ and 3’ halves (approximately) of Rsx as indicated by SFPQ and SRSF9 interactions (Figures 4L-4M), which suggests the ability to predict functional and / or regulatory differences across the transcript.

[0128] Altogether, these results between human and mouse Xist, and extending further to marsupial Rsx, suggests our approach is useful in determining functional conservation or divergence across species, including for syntenic or non-syntenic IncRNAs.

[0129] Discussion of EXAMPLES 1-4

[0130] Deep learning-based approaches are increasingly used to characterize biological molecules, including for the prediction of intermolecular interactions (Horlacher et al., 2023; Moore and ’t Hoen, 2019; Pan et al. , 2019; Y amada and Hamada, 2022). Using BERT -RBP, we leveraged natural language processing principles applicable to genome-wide nucleic acid sequence contexts (Ji et al., 2021; Zhou et al., 2023; Yamada and Hamada, 2022), to design apractical and applicable machine learning-based approach to study IncRNAs, since they remain poorly understood.

[0131] Foundationally, we tested the predictive utility using well-established RBP-mRNA interactions. We then found that predictions for IncRNAs aligned with prior interactions observed in experimentation, including CLIP-Seq and mass spectrometry-based analyses (Pandya-Jones et al., 2020; Yi et al., 2020; Chu et al., 2015; McHugh et al., 2015; Minajigi et al., 2015; Van Nostrand et al., 2016). Altogether therefore, we successfully predicted both coding- and noncoding RNA- RBP interactions, capturing differing functions, gene lengths, and regulatory qualities. Since IncRNA genes outnumber mRNA genes in the human genome (Quinn and Chang, 2016), this approach can provide a foundation for de novo predictions for uncharacterized IncRNAs to generate candidate interactions to guide experimentation, such as in vivo perturbation analysis via CRISPR. We anticipate that additional training data including from complementary approaches (Wolin et al., 2023) or cell- and / or context-specific binding data would support broader prediction of interactions.

[0132] Our predictive approach can fill experimental gaps by imputing protein binding sites for transcripts showing little to no expression, as has been observed for many IncRNAs (Rinn and Chang, 2020; St Laurent et al., 2015; Mattick et al., 2023). To capture such variation, models may require additional training data, with a limitation of these approaches being differences in training data sets used by different research groups (Horlacher et al., 2023). Inherent data quality and Attorney Docket No.: 421 / 553 PCT predictive issues are limited by experimentally queried RBPs, experimental CLIP protocols and associated limitations including resolution, and computational analysis design including training and evaluation parameters. Deviation between CLIP-Seq and our model predictions may be due to our decision to prioritize precision vs. recall, or other elements influencing model optimization including RNA structure or combinations of RBP interactions (Sun et al. 2021).

[0133] As shown for Rsx, our approach can be used for IncRNAs that may have convergently evolved similar function, or to inform on the function of syntenic IncRNAs, both cases in which sequence conservation across species is lacking (Ulitsky, 2016; Necsulea et al., 2014; Johnsson et al., 2014). While an abundance of experimental data such as CLIP-seq for numerous proteins exists for human samples (Van Nostrand et al. 2020; Gerstberger et al. 2014), such data are often limiting for other species, even for conventional animal models such as mouse, and more so for more diverse or non-model species. Our approach therefore provides a foundation for further study across many organisms since RBP properties are generally evolutionarily conserved.

[0134] Interestingly, we found that some predictions aligned preferentially with CLIP-Seq data from a specific cell line (e.g. in HepG2 but not K562). Such contextual specificity is anticipated given that CLIP-Seq output is dependent on protein target abundance in a given cell type (Van Nostrand et al., 2020a, 2016, 2020b). Our results therefore suggest that models can detect contextual elements for specific RBP-RNA interactions, with only sequence data as an input. These intricacies are considerations for positive and negative labels in classification methods, as a negative sequence in a specific cell type may not be negative in another. Future work may determine whether models effectively leam cell-specific interactions to identify a spectrum of interactions that may be missed in a single study, e.g. due to cell type, experimental timing, or experimental technique limitations.

[0135] Since model fine-tuning produces a set of weights that provide sufficient solutions for the given task, different sets of weights can create models with similar overall performance, but which vary in their response to a particular input. Therefore, one way to improve our approach would be to create an ensemble of models, either with additional instances of BERT-RBP models that have been trained with a different subset of training data or have been initialized with different random seeds, or with additional types of models. Additional improvements might be obtained by converting these models from binary classifiers to multi-class models.

[0136] Analysis parameters can be adjusted to favor increased precision to capture high- confidence interactions vs. recall to identify all potential binding sites within a specific transcript more comprehensively. Future work to experimentally validate predictions will aid in setting thresholds and analysis parameters, and studies characterizing validated interactions will enhance our understanding of how RBP-lncRNA interactions direct IncRNA function. Predictions that Attorney Docket No.: 421 / 553 PCT consider combinatorial RBP interactions would also increase the capacity to predict RBP impact on associated RNAs, to influence the capacity to predict more granular facets of IncRNA biology (Noh et al., 2018; Ang et al., 2020; Rinn and Chang, 2020; Mercer and Mattick, 2013).

[0137] A major advantage of our approach is that predictions are based solely on sequence, lowering the barrier for IncRNA characterization, which is of particular importance since most IncRNAs are uncharacterized with limited / no experimental data. Moreover, our easy-to-use method bridges the gap between computational fields focused on developing machine learning algorithms, and wet lab work, by providing experimental targets for biochemistry and molecular biology studies.

[0138] Table 1. Training Hyperparameters

[0139]

[0140]

[0141]

[0142]

[0143]

[0144] Table 2. Model performance metrics

[0145]

[0146]

[0147]

[0148]

[0149]

[0150]

[0151] Attorney Docket No.: 421 / 553 PCT

[0152] Table 3. Binding predictions for Opossum Rsx Attorney Docket No.: 421 / 553 PCT Attorney Docket No.: 421 / 553 PCT Attorney Docket No.: 421 / 553 PCT

[0153] Table 4. RNA Sequence information Attorney Docket No.: 421 / 553 PCT

[0154] EXAMPLE 5

[0155] BERT2-RBP AND ENSEMBLE APPROACHES

[0156] Example 5, while similar to prior examples, describes among other aspects how BERT2- RBP is used in a complementary manner as well as the implementation of the ensemble approach.

[0157] We finetuned RNA-RBP binding prediction models based on pre-trained DNABERT and DNABERT2 models (Ji et al. 2021; Zhou et al. 2023; Yamada and Hamada 2022). Training data for finetuning both types of models originates from ENCODE eCLIP data (Pan et al. 2020; Yamada and Hamada 2022; Van Nostrand et al. 2016) comprising samples of sequence around eCLIP peak sites in K562 and HepG2 cells for 154 RBPs as the positive class, and matched regions from the same gene that did not contain any peaks as the negative class.

[0158] BERT-RBP models use k-mer encoding and were finetuned from DNABERT (Ji et al. 2021) following procedures for RBP-based targets (Yamada and Hamada 2022). When finetuning BERT-RBP models using default hyperparameters, 136 / 154 models showed Area Under the Receiver Operating Characteristic (AUROC) scores similar to AUROC scores originally reported (Y amada and Hamada 2022), but 16 models lacked classification ability and 2 models had AUROC scores much lower than originally reported.

[0159] A recent study also reported instability when training BERT-RBP models with default parameters (Horlacher et al. 2023). Default hyperparameters used 4 GPUs with a batch size of 32 per GPU for a total batch size of 128. To improve model performance, we also trained all 154 models with larger batch sizes of 192 and 256 (Table 1) . We found that only 6 models failed to train upon training with a batch size of 192, and training with a batch size of 256 resulted in 2 models failing to train and 3 having lower performance than previously reported. We observed that the models that failed to train were different for the different training sets, therefore we opted to use models trained with a batch size of 192 moving forward, except when their AUROC was more than 0.005 less than the AUROC scores previously reported, in which case we used the higher scoring 128 or 256 batch size trained model (Table 1). AUROC scores of the resulting set of models were slightly higher, in comparable range to those originally reported (Yamada and Hamada 2022) (Yamada & Hamada: 0.786 + / - 0.041; In-house: 0.791 + / - 0.037) (Figure 11A).

[0160] We carried out comparative predictive analysis using BERT2-RBP models, which were finetuned on top of DNABERT2 (Zhou et al. 2023), which uses byte-pair encoding instead of kmer- based encoding. BERT2-RBP models performed comparably to both BERT-RBP models trained in-house and those reported for BERT-RBP (Yamada and Hamada 2022) (BERT2-RBP AUROC: 0.800 + / - 0.037) (Figure 11 A).

[0161] For both DNABERT and DNABERT2-derived approaches (Yamada and Hamada 2022; Ji et al. 2021; Zhou et al. 2023), the output of running an RNA sequence through both models is a Attorney Docket No.: 421 / 553 PCT binding probability ranging from 0 to 1. In order to convert binding probability into a binary binding / not-binding decision, we selected thresholds where sequences with binding probabilities above that threshold are classified as binding and those with probabilities below are classified as non-binding. We examined the Precision-Recall curves for each model to guide our threshold selection decision for each model (Figure 11B). Since for any binary model precision and recall are always in balance, we tested effects of favoring precision over recall (Figure 11C), and subsequently opted to use a threshold of 0.9, which prioritized precision, to facilitate the prediction of most-likely binding sites rather than predicting all possible binding sites.

[0162] BERT-RBP and BERT2-RBP models run predictions on RNA sequences 101 nucleotides (nt) long. Given that the length of RNA transcripts vary significantly and can extend to thousands of kilobases (Quinn and Chang 2016; Ransohoff et al. 2018; St Laurent et al. 2015; Mattick et al. 2023), we used a segmentation approach in order to run predictions on longer sequences (Figure 11C). We generated sequence segments of overlapping 101 nt sequences and used a lOnt sliding window which results in each segment overlapping its neighbors by 90 nt. The result of running each segment through BERT-RBP or BERT2-RBP models are binding probability predictions for each segment, with a single binding site being present in multiple adjacent segments as well as correlation in model predictions for adjacent segments (Figure 11C).

[0163] RBPs interact with RNAs to regulate well-studied regulatory processes such as splicing and translation (Gerstberger et al. 2014; Hentze et al. 2018). We used the established model of autoregulation of TARDBP via its 3’UTR (Ayala et al. 2011; Sun et al. 2014; Bhardwaj et al. 2013) to determine whether we could predict established RBP-RNA interactions. Both BERT-RBP and BERT2-RBP predicted interactions over a region of approximately 500 nt in the 3’UTR of TARDBP, similar to a region of enrichment identified in prior work (Wolin et al. 2023; Van Nostrand et al. 2016) (Figure 12A). We sought to explore motifs identified for TARDBP protein binding by BERT-RBP and found top motifs similar to those identified in previous studies as sequence elements that were determinants for interaction predictions (Figure 12B). We queried the location of motifs within predicted interacting sites, given that the binding region for the TARDBP protein has been mapped in biochemical studies, with interactions have been identified across a 600 nt region containing specific sites surrounding a region enriched with the GAAUG and (UG)n repeat motifs (Bhardwaj et al. 2013). We found these sites to be in the vicinity of predicted binding regions (Figure 12C; SEQ ID NO:1). Indeed, in regions of significant overlap, we found sequences containing well-known TARDBP interaction elements including U rich sequences, GU repeats, and the GAAUG motif (Figure 12C), relative to a neighboring region with low probability of predicted interactions (Figure 12C). We also found predicted interacting regions that did not Attorney Docket No.: 421 / 553 PCT overlap with binding motifs, suggesting that models are attentive to other elements within this region.

[0164] We next tested the capacity of models to predict interactions among different RBP / RNAs, focusing on RBPs with different functional properties that regulate various types of RNAs (Figures 12D-12E). For these RBPs, binding to UTRs and intronic sequences are regulatory paradigms that drive RNA metabolic processes including processing and stability (Van Nostrand et al. 2020; Briata and Gherzi 2020). We observed that we were able to predict interactions between RBPs and RNA targets including METAP2 and VIM, and RBFOX2 an^ NDELl, observations which aligned with ENCODE eCLIP peaks (REF: van Nostrand et al). Regarding METAP2 with VIM RNA (Figure 12D), we predicted interactions primarily in an early exon, similar to patterns observed with ENCODE e-CLIP data (Van Nostrand et al. 2016; Kagda et al. 2023). Similarly, both models primarily predicted interactions in the most 3’ intron of NDEL1, which overlapped with a region of prominent enrichment in eCLIP in K562 and HepG2 cells (Van Nostrand et al. 2016; Kagda et al. 2023) (Figure 12E).

[0165] The ability to predict functional properties for IncRNAs a priori is limited due to the absence of long stretches of sequence conservation, as is characteristic for mRNAs (Necsulea et al. 2014; Johnsson et al. 2014). We set out to determine if this deep learning approach could provide an opportunity to generate hypotheses for functional and / or regulatory aspects of IncRNAs based on their interactions with RBPs.

[0166] For IncRNAs, we used the ensemble approach based on the agreement in predictions from both model types to examine predictive capabilities for highly studied candidate RNAs. Beginning with the IncRNA XIST, which has well-defined roles in X-chromosome inactivation (Loda and Heard 2019; Sahakyan et al. 2018; Furlan and Rougeulle 2016), we examined the top 20 predicted interactions, and found that we identified RBPs previously shown to interact with XIST in various contexts (Figure 3A). Indeed, of the top 20 predicted RBPs, this ensemble approach successfully predicted interactions for 17 RBPs identified previously in mass spectrometry -based studies or eCLIP studies (Minajigi et al. 2015; Chu et al. 2015; Teng et al. 2020; Li et al. 2014). Predicted interactions included previously identified RBPs including MATR3, KHSRP, PTBP1, and HNRNPC (Li et al. 2014; Chu et al. 2015) (Figure 13A).

[0167] XIST is organized in modular repeat domains, each of which harbor specific RBP-binding enrichment (Brockdorff 2002; Brockdorff et al. 2020). Both model predictions successfully reproduced these binding enrichment patterns with HNRNPC binding to regions encompassing Repeat A, HNRNPK binding to mid-regions containing Repeats B, C, and D, and PTBP1 showing enriched predictions in Repeat-E, as similarly seen with eCLIP (Fig. 13B). We found that interaction sites for our top predicted RBP interactor, MATR3, primarily overlapped with predicted Attorney Docket No.: 421 / 553 PCT sites for PTBP1 in Repeat-E, similar to that shown in prior work (Jacobson et al, 2022, Pandya- Jones et al, 2020).

[0168] The IncRNA MALAT1 has well-demarcated molecular characteristics (Arun et al. 2020; Zhang et al. 2017) including roles in splicing, transcription, and actions as a competing endogenous RNA, and has been implicated in numerous diseases (Arun et al. 2020; Zhang et al. 2017; Wu et al. 2015), but is vastly different from XIST in its sequence and structural characteristics. Using the prior ensemble approach, we were also able to predict interactions with RBP with various functions (Figure 13C). Prominent interactions were seen with WDR3 and HLTF, as well as previously shown interactors, TARDBP and TRA2A (Van Nostrand et al. 2016) (Figure 13D). Models were able to predict different interacting pattern-types among top-binding RBPs (Figure 13D). These patterns correlated with those seen via CLIP-Seq (Van Nostrand et al. 2016; Kagda et al. 2023). We observed dispersed binding throughout the transcript for WDR3 with enrichment in the 3’ region of the transcript (Figure 3D). Intriguingly, predictions for HLTF were more restricted to the 3’ end of the molecule, similar to CLIP-seq peaks called in the K562 cell line, and differed from the more widespread binding in CLIP-seq seen in HepG2 cells (Figure 13D). Predictions using BERT -RBP extended to these regions seen in HepG2, while BERT2-RBP predictions were more narrow. We also examined congruence between previously identified interactions, including for TARDBP and TRA2A, and predicted more discrete 5’ binding predictions for TARDBP, relative to broader binding to the 3 ’ of the transcript for TRA2A, similar to the region predicted to bind for HLTF in K562 cells (Van Nostrand et al. 2016; Kagda et al. 2023)(Figure 13D).

[0169] We also examined interactions for H19 (Figure 13E), which differs from XIST and MALAT1 based on its shorter length, cytoplasmic subcellular location and associated roles including functioning as a miRNA precursor (Jonas et al. 2020; Noh et al. 2018; Liang et al. 2015). Among the top predicted interactors GRSF1 preferentially binds G-rich sequences, and has been shown previously to interact with the IncRNA RMRP for its mitochondrial localization, while EWRS1 has been implicated in cancer biology similar to H19 (Liang et al. 2015; Matouk et al. 2015; Lee et al. 2019). GRSF1 andEWSRl were predicted to interact at the 5’ end of the transcript, while LARP7 and EIF4G2 were predicted to show more dispersed binding, similar to that observed for ENCODE e-CLIP (Figure 13F).

[0170] Altogether, these data suggest that both model types are able to predict RBP interactions with IncRNAs of varying lengths, as well as the cellular and molecular regulation and functions as mediated by RBPs. The training data for these models originated from human nucleotide sequences (Ji et al. 2021; Yamada and Hamada 2022). Given the lack of conventional sequence conservation for IncRNAs, imputing function and consistency of properties across species is a major limitation in the field. One mechanism for functional conservation in the absence of broad sequence Attorney Docket No.: 421 / 553 PCT conservation is the conservation of interactions with proteins, which is frequently driven by short motif sequences (Lambert et al. 2014; Kuret et al. 2022). Indeed, k-mer content comparisons of XIST, MALAT1, and H19 RNAs from both human and mouse suggest that the overall sequence composition is similar, especially based on 3-mer and 4-mer comparisons, particularly for XIST, confirming prior observations (Kirk et al. 2018) andMALATl (Figure 14A-14C).

[0171] Given that RBPs are highly conserved across vertebrates at the sequence and functional levels (Gerstberger et al. 2014), we examined whether BERT-derived models could predict interactions with mouse sequence, which would suggest possible functional conservation for IncRNAs across species. Intriguingly, using this approach, we found congruence in the predictions of top interactors between mouse and human (Figures 14D,14F, 14H), confirming that protein- IncRNA interactions are central mediators of IncRNA function (Huang et al. 2021; Briata and Gherzi 2020). This was highest for Xist and Malatl, and less so for H19 (Figure 14D, 14F, 14H). We more closely examined common interactions between human and mouse Xist, which maintains its function to silence X chromosomes across species, even though it shows sequence divergence across species including for repeat sequences (see PhyloP conservation, Figure 14E).

[0172] We found that 13 of the top 20 predicted RNAs in human were in the top 20 predictions for mouse Xist based on aggregated binding counts (Figure 14D). The exceptions, HNRNPC, HNRNPK, DDX59, KHDRBS1, TAF15, SBDS, and FTO were at positions 27 (HNRNPC), 31 (HNRNPK), 88 (DDX59), 106 (KHDRBS1), 69 (TAF15), 63 (SBDS) and 48 (FTO) based on this metric. Such interactions are the basis of conserved roles for IncRNAs such as XIST in dosage compensation, despite limited conservation at the sequence level. This approach could therefore be useful in determining whether syntenic IncRNAs (given deviation in sequence conservation) could be functionally conserved across species.

[0173] Similar to human XIST, modular repeat domains are generally functionally conserved through protein-interaction capacity in mouse Xist (Sahakyan et al. 2018; Loda and Heard 2019; Furlan and Rougeulle 2016). We found that both models successfully identified commonalities in repeat-RBP interactions between human and mouse. In mouse, HNRNPC was also predicted to interact with Repeat-A, while HNRNP-K bound downstream in the region corresponding to Repeats B, C, and D (Figure 14E). As with human XIST, PTBP1 was also predicted to interact with Repeat-E (Figure 14E).

[0174] We also examined interactions for Malatl. Here we found predictions common between human and mouse for 15 of the top RBPs including for HLTF and WDR3 (Figure 14F). Interestingly, binding prediction patterns showed similarity between human and mouse even in 5’ regions of the RNA that lack sequence conservation between human and mouse (Figure 14G, see PhyloP track). Enrichment of binding predictions were observed in the 5’ region of the transcript Attorney Docket No.: 421 / 553 PCT for WDR3. This region was also where most HLTF predicted binding was observed, similar to what was seen for K562 enrichment in human MALAT1 (Figure 14G). Predicted binding patterns for previously shown interactors, TARDBP and TRA2A (Van Nostrand et al. 2016), were similar to patterns shown for human with more 3’ and 5’ enrichment respectively relative to each other. We also examined predictions for other top RBP-binders, Z3CH8 and CPEB4, in mouse and found similar 5’ predicted enrichment, suggesting that this region is a highly regulated segment of the transcript. Altogether, this again is suggestive of possible RBP-mediated functional conservation in the absence of sequence conservation.

[0175] We found 13 of the top 20 predicted interactions common between human and mouse Hl 9 (Figure 14H). Despite this RNA being very different from Xist and Malatl in terms of sequence organization, it maintains expression patterns and functions with roles in early embryonic development and cancers (Matouk et al. 2015; Jonas et al. 2020; Noh et al. 2018). Examination of top predictions revealed similar binding patterns generally, with several top RBP interactors binding to the most 5’ region of the RNA, followed by the 3’ end of exon 1 (Figure 141). For GRSF1, predicted binding was similar to that seen for the human RNA, however interactions in mouse covered a smaller region and were also consistently predicted at the 3’ end of exon 1 for both models. EWSR1, LARP7, and EI4G2 predictions were generally similar between human and mouse, with the exception of predicted binding in the 3 ’ exon for human Hl 9, that was not strongly predicted for mouse Hl 9, and was not overtly represented in human e-CLIP data, except for LARP7. We also examined two additional top binders, PCBP2 and SRSF9 (Figure 41), and found different binding patterns that were more 3’ or mid-transcript localized, relative to the other RBPs that preferentially bound to the 5 ’ regions, suggesting differential regulation by these RBPs in those regions of the RNA.

[0176] Discussion

[0177] Deep learning-based approaches are increasingly used to predict intermolecular interactions, with a plethora of methods for development and refinement for RBP-interaction predictions (Horlacher et al. 2023; Moore and ’t Hoen 2019; Pan et al. 2019; Yamada and Hamada 2022). Here, we used BERT-RBP and BERT2-RBP, two BERT-derived architectures with numerous trainable parameters (Ji et al. 2021; Zhou et al. 2023; Yamada and Hamada 2022) that apply natural language processing-based principles to genome-wide nucleic acid sequence contexts. The primary distinction between the foundational models, DNABERT and DNABERT2 and their respective RBP interaction prediction models, is in how DNA and RNA sequences are tokenized to be fed into the model, where DNABERT uses k-mer tokenization, while DNABERT2 uses byte pair encoding (Sennrich et al. 2015; Ji et al. 2021; Zhou et al. 2023). The use of BPE, among other optimizations, allows the DNABERT2 and BERT2-RBP models (Zhou et al. 2023) Attorney Docket No.: 421 / 553 PCT to be more computationally efficient than their predecessors, by eliminating the redundancy inherent in the overlapping kmer encoding used by DNABERT and BERT-RBP (Yamada and Hamada 2022; Ji et al. 2021), which results in shorter tokenized sequences. DNABERT2 also overcomes the input length limitation of DNABERT (Zhou et al. 2023; Ji et al. 2021), a feature we have not yet taken advantage of with BERT2-RBP because the finetuning data that we use is limited to 101 nt sequences.

[0178] We found that predictions of both model types were generally in agreement when they were confident about binding / not-binding (produced probabilities close to 0 or 1) but had more variable results when the output probability was closer to 0.5. In consideration of future methodological applications such experimental validation where a manageable list of candidate sites would be ideal, we used an ensemble approach to examine predictive utility for IncRNA sequences. Combining output from the two model types in a small ensemble increased our confidence in their predictions and aligned with prior interactions observed in experimentation including CLIP-Seq and mass spectrometry-based analyses (Pandya-Jones et al. 2020; Yi et al. 2020; Chu et al. 2015; McHugh et al. 2015; Minajigi et al. 2015; Van Nostrand et al. 2016). Fine-tuning of models produces a set of model weights that provide a sufficient solution to the given task, and there are many sets of weights that create models with similar overall performance, but have variation in their response to a particular input. One way to improve this method would be to expand our ensemble of models either with additional types of models, or with additional instances of BERT- RBP or BERT2-RBP models that have been trained with a different subset of training data or have been initialized with different random seeds. Additionally, it would be interesting to convert these models from binary classifiers to multi-class models.

[0179] Interestingly, we found that some predictions aligned preferentially with CLIP-Seq data from a specific cell line, suggesting contextual elements may come into play to direct specific RBP- RNA interactions in a cell-line specific way or that models leam cell-specific interactions. Contextual specificity is anticipated given that CLIP-Seq output is dependent on protein target abundance in a given cell type (Van Nostrand, Freese, et al. 2020; Van Nostrand et al. 2016; Van Nostrand, Pratt, et al. 2020), and is an important consideration for positive and negative labels in classification methods as a given negative sequence in a specific cell type may not be negative in another.

[0180] Nevertheless, our predictive approach can fill experimental gaps by imputing binding sites in the case of transcripts showing little to no expression, as has been observed for many IncRNAs (Rinn and Chang 2020; St Laurent et al. 2015; Mattick et al. 2023). In order to capture such variation, models may require additional training data, with a major limitation of these approaches being differences in training data sets used by different research groups (Horlacher et al. 2023). Attorney Docket No.: 421 / 553 PCT

[0181] Inherent data quality and predictive issues are therefore limited by RBPs that are experimentally queried, experimental CLIP protocols and associated limitations including resolution, and computational analysis design including parameters of training and evaluation sets. Deviation between CLIP-Seq and our model predictions also suggest additional elements such RNA structure, and possibly combinatorial interactions should be considered for model optimization (Sun et al. 2021). However, one advantage of BERT-RBP and BERT2-RBP models is that they run predictions based solely on sequence, which make them especially useful for studying novel IncRNAs for which additional information about structure or combinatorial interactions is not yet known.

[0182] We found similar capabilities for interaction prediction for coding and noncoding RNAs, and for RNAs with differing functions, gene lengths, and regulatory properties. While we anticipate that additional training data including from complementary approaches (Wolin et al. 2023) and / or containing cell-and / or context-specific binding data would support broader prediction of interactions, further work will be needed to determine what elements drive false positives / negatives in model predictions.

[0183] Here we demonstrate the utility of our approach for interaction comparisons for IncRNAs that may only be syntenically conserved, ie. lacking in primary sequence conservation across species (Ulitsky 2016; Necsulea et al. 2014; Johnsson et al. 2014). With IncRNA genes outnumbering mRNA genes in the human genome (Quinn and Chang 2016), we envision this approach will also provide a foundation for de novo predictions for novel IncRNAs to generate extensive lists of candidate interactions to guide experimentation such as in vivo perturbation analysis via CRISPR or other approaches.

[0184] Depending on the study objective, in addition to the ensemble approach using the intersection between model predictions, model and analysis parameters can be adjusted based on study goals to favor increased precision for capture of high-confidence interactions vs. recall to more comprehensively identify all potential binding sites within a specific transcript for a particular RBP. Future work to experimentally validate predictions will aid in setting thresholds and model parameters, and studies characterizing validated interactions will enhance our understanding of how RBP-lncRNA interactions direct IncRNA function. Moreover, when additional training data is available including on additional sequence and / or structure contexts or combinatorial RBP interactions to improve the design and interpretability of predictive models, we anticipate the capacity to predict RBP impact on associated RNAs generally. We should also then be able to predict more granular facets of IncRNA biology, such as specific characteristics associated with properties that determine subcellular functions (Noh et al. 2018; Ang et al. 2020; Rinn and Chang Attorney Docket No.: 421 / 553 PCT

[0185] 2020; Mercer and Mattick 2013), e.g. elements that drive based on nuclear vs. cytoplasmic location and associated roles.

[0186] We present a use-case for NLP-based predictions as a tool to predict interactions, and when knowledge becomes more readily available, predict functions for IncRNAs. In this way, hypotheses on IncRNA functions and regulation can be inferred in the absence of sequence conservation and without experimental data. While an abundance of experimental data such as CLIP-seq for numerous proteins exists for human samples (Van Nostrand et al. 2020; Gerstberger et al. 2014), such data are often limiting for other species, including conventional animal models such as mouse or more diverse non-model species. Given the high sequence conservation of RBPs (Gerstberger et al. 2014), driven by the sequence and associated functional conservation of constituent RNA binding domains, this predictive approach can be readily transferred to other species where no experimental data is available.

[0187] Methods Employed In Example 5

[0188] Finetuning models

[0189] All models (Yamada and Hamada 2022; Zhou et al. 2023) were finetuned on the Longleaf Computing Cluster at UNC-Chapel Hill, using either GeForce GTX1080 or Tesla V100-SXM2 GPUs. Analyses and statistics were performed using python and R.

[0190] Training data preparation

[0191] The benchmark training dataset was downloaded from RBPSuite (Pan et al. 2020) and included sequences around eCLIP peaks from K562 and HepG2 cell lines for 154 RBPs (Van Nostrand et al. 2016; Kagda et al. 2023). Briefly, the positive dataset was comprised of up to 60,000 nucleotide sequences of 101 nt in length, derived from eCLIP peaks for each RBP. The negative dataset included matched regions without any peaks from the same gene. For finetuning, we randomly selected 15,000 positive and negative sequences and split them into training (9,600), evaluation (2,400) and test (3,000) sets, using scripts from BERT-RBP (Yamada and Hamada 2022). The same sequence sets were used to finetune both BERT-RBP and BERT2-RBP models.

[0192] Finetuning BERT-RBP models

[0193] BERT-RBP models were finetuned on the pretrained 3-mer DNABERT model (Ji et al. 2021 (following instructions from Yamada et al. 2022 (https: / / github.com / kkyamada / bert-rbp). We first used default hyperparameters of 4 GPUs with a per-gpu-batchsize of 32 for a total batchsize of 128. However, when this method gave unstable results (with 16 models lacking any classification ability) we trained model sets with batch sizes of 192 and 256 (using 6 and 8 GPUs, respectively). Models trained with larger batch sizes showed improved, but not perfect stability. For running predictions, we selected models trained with batch sizes of 192, except when the AUROC score of Attorney Docket No.: 421 / 553 PCT the model was more than 0.005 less than that originally reported by Yamada & Hamada (2022), in which case we used the higher performing of the 128 or 256 batch size model (Table 1).

[0194] Finetuning BERT2-RBP models

[0195] BERT2-RBP models were finetuned on top of the pretrained DNABERT2 model (Zhou et al. 2023) using a script developed in-house. We used Optuna to perform a hyperparameter search along the batch-size, learning rate, weight decay and dropout hyperparameters for 3 of 154 RBPs (TIAL1, HNRNPK, KHSRP). Guided by this search, we used a batch size of 128, a learning rate of 5e-5, a dropout of 0.1 and weight decay of 0.01 (Table 1) to finetune models for the rest of the RBPs. We leveraged the Trainer class of the transformers python library (version 4.29.2) to evaluate the model throughout the finetuning process and to save the version that performed best on the evaluation dataset.

[0196] Evaluating model performance

[0197] Performance of both BERT-RBP and BERT2-RBP models (Zhou et al. 2023; Yamada and Hamada 2022) was evaluated by running binding probability predictions on the test dataset withheld from finetuning. Outputs were normalized to a range of 0-1 using Min-Max normalization, where the min and max were the minimum and maximum values of predictions in the test dataset. Performance metrics including accuracy, precision, recall, Fl score, Matthew’s correlation coefficient, average precision (area under precision-recall curve) and AUROC were calculated using functions from the scikit-leam python package, version 1.3.0 (Table 2).

[0198] Using models to predict binding for RNA transcripts

[0199] The training data for both BERT-RBP and BERT2-RBP models was RNA sequences 101 nucleotides (nt) long (Pan et al. 2020), thus the sequences that we run predictions on should also be 101 nt long. Given that the length of RNA transcripts varies significantly and can extend to thousands of kilobases, we used a segmentation approach to run predictions on longer sequences (Figure 11C). We generated sequence segments of overlapping 101 nt sequences and used a lOnt sliding window which results in each segment overlapping its neighbors by 90 nt. We then run each segment through the given BERT-RBP or BERT2-RBP model, to obtain a binding prediction for that segment. We normalize the model output using Min-Max normalization, with minimum and maximum values from the model’s predictions on the test dataset. Because each segment overlaps with its neighbors significantly, we then smooth the resulting prediction curve with a 1-D gaussian filter (scipy.ndimage.gaussian_filterld, scipy version 1.10.1) with a sigma of 20 nt, followed by linearly interpolating the result. Regions of the resulting curve above a threshold of 0.9 are classified as binding regions and are exported in BED format for viewing in the UCSC genome browser (http: / / genome.ucsc.edu). Attorney Docket No.: 421 / 553 PCT

[0200] Sequence Acquisition

[0201] All sequences were downloaded using UCSC’s Table Browser tool (Karolchik et al. 2004). Human sequences are from the hgl9 GENCODE V441ift37 track (Kent et al. 2002). Mouse sequences are from the mmlO All GENCODE VM25 track. Ensembl IDs for each RNA transcript are listed in Table 4. Sequences were downloaded in FASTA format.

[0202] Motif Analysis

[0203] Motif analysis from Bert-RBP models was done as previously developed (Ji et al. 2021; Yamada and Hamada 2022), with modifications to how motif candidates were merged. Briefly, attention vectors from the CLS token to the last layer of the model were extracted for each sequence in the test dataset. Attention vectors were summed across attention heads. For positive sequences, we selected contiguous 5-10 nucleotide regions of the sequence where attention was either greater than the average attention for the sequence or greater than 10 times the minimum attention for the sequence. These regions are putative motifs. We then perform a hypergeometric test for each putative motif to test if it occurs more often in positive sequences than in negative ones. We keep motifs where the p-value from the hypergeometric test is <0.005 after correction for multiple tests.

[0204] We then merge similar motifs together using an Unweighted Pair Group Method with Arithmetic mean (UPGMA) algorithm. Similarities between each motif were measured using the alignment score from biopython’s Pairwise Aligner. The similarity matrix was then inverted to create a distance matrix. This distance matrix was then put into scipy’s UPGMA clustering algorithm (scipy.cluster.hierarchy. average), and motifs were grouped until the distance between groups reached a threshold corresponding to the inversion of a minimum alignment score of 5. For alignment scoring matching nucleotides scored 2, mismatches scored -1, open and extended gaps scored -0.75 and internal gaps were not allowed. Alignment scores were adjusted for length so that scores for longer motifs were similar to scores for shorter motifs. For each motif group, component sequences were collected and extended to 12 nt. Depictions of motifs were generated from these sequences using WebLogo version 3.7.12.

[0205] REFERENCES

[0206] All references listed herein including but not limited to all patents, patent applications and publications thereof, scientific journal articles, and database entries (e.g., GENBANK® database entries and all annotations available therein) are incorporated herein by reference in their entireties to the extent that they supplement, explain, provide a background for, or teach methodology, techniques, and / or compositions employed herein.

[0207] Andergassen, D., and Rinn, J. L. (2022). From genotype to phenotype: genetics of mammalian long non-coding RNAs in vivo. Nat. Rev. Genet. 23, 229-243. doi:10.1038 / s41576-021-00427- 8. Attorney Docket No.: 421 / 553 PCT

[0208] Ang, C. E., Trevino, A. E., and Chang, H. Y. (2020). Diverse IncRNA mechanisms in brain development and disease. Curr. Opin. Genet. Dev. 65, 42-46. doi: 10.1016 / j.gde.2020.05.006.

[0209] Arun, G., Aggarwal, D., and Spector, D. L. (2020). MALAT1 Long Non-Coding RNA: Functional Implications. Noncoding RNA 6. doi:10.3390 / ncma6020022.

[0210] Ayala, Y. M., De Conti, L., Avendano- Vazquez, S. E., Dhir, A., Romano, M., D’Ambrogio, A., Tollervey, J., Ule, J., Baralle, M., Buratti, E., et al. (2011). TDP-43 regulates its mRNA levels through a negative feedback loop. EMBO J. 30, 277-288. doi: 10.1038 / emboj.2010.310.

[0211] Bhardwaj, A., Myers, M. P., Buratti, E., and Baralle, F. E. (2013). Characterizing TDP-43 interaction with its RNA targets. Nucleic Acids Res. 41, 5062-5074. doi: 10.1093 / nar / gktl 89.

[0212] Briata, P., and Gherzi, R. (2020). Long Non-Coding RNA-Ribonucleoprotein Networks in the Post- Transcriptional Control of Gene Expression. Noncoding RNA 6. doi: 10.3390 / ncma6030040.

[0213] Brockdorff, N., Bowness, J. S., and Wei, G. (2020). Progress toward understanding chromosome silencing by Xist RNA. Genes Dev. 34, 733-744. doi:10.1101 / gad.337196.120.

[0214] Brockdorff, N. (2002). X-chromosome inactivation: closing in on proteins that bind Xist RNA. Trends Genet. 18, 352-358. doi: 10.1016 / s0168-9525(02)02717-8.

[0215] Chu, C., Zhang, Q. C., da Rocha, S. T., Flynn, R. A., Bharadwaj, M., Calabrese, J. M., Magnuson, T., Heard, E., and Chang, H. Y. (2015). Systematic discovery of Xist RNA binding proteins. Cell 161, 404-416. doi:10.1016 / j.cell.2015.03.025.

[0216] Delas, M. J., and Hannon, G. J. (2017). IncRNAs in development and disease: from functions to mechanisms. Open Biol. 7. doi: 10.1098 / rsob.170121.

[0217] Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2018). BERT: Pre-training of deep bidirectional transformers for language understanding. arXiv. doi: 10.48550 / arxiv.1810.04805.

[0218] Ferre, F., Colantoni, A., and Helmer-Citterich, M. (2016). Revealing protein-lncRNA interaction. Brief. Bioinformatics 17, 106-116. doi:10.1093 / bib / bbv031.

[0219] Furlan, G., and Rougeulle, C. (2016). Function and evolution of the long noncoding RNA circuitry orchestrating X-chromosome inactivation in mammals. Wiley Interdiscip. Rev. RNA 7, 702-722. doi:10.1002 / wma,1359.

[0220] Gerstberger, S., Hafner, M., and Tuschl, T. (2014). A census of human RNA-binding proteins. Nat. Rev. Genet. 15, 829-845. doi: 10.1038 / nrg3813. Attorney Docket No.: 421 / 553 PCT

[0221] Grant, J., Mahadevaiah, S. K., Khil, P., Sangrithi, M. N., Royo, H., Duckworth, J., McCarrey, J. R., VandeBerg, J. L., Renfree, M. B., Taylor, W., et al. (2012). Rsx is a metatherian RNA with Xist-like properties in X-chromosome inactivation. Nature 487, 254-258. doi: 10.1038 / naturel l l71.

[0222] Hentze, M. W., Castello, A., Schwarzl, T., and Preiss, T. (2018). A brave new world of RNA- binding proteins. Nat. Rev. Mol. Cell Biol. 19, 327-341. doi: 10.1038 / nrm.2017.130.

[0223] Horlacher, M., Cantini, G, Hesse, J., Schinke, P., Goedert, N., Londhe, S., Moyon, L., andMarsico, A. (2023). A systematic benchmark of machine learning methods for protein-RNA interaction prediction. Brief. Bioinformatics 24. doi: 10.1093 / bib / bbad307.

[0224] Huang, Y., Qiao, Y., Zhao, Y., Li, Y., Yuan, J., Zhou, J., Sun, H., and Wang, H. (2021). Large scale RNA-binding proteins / LncRNAs interaction analysis to uncover IncRNA nuclear localization mechanisms. Brief. Bioinformatics 22. doi:10.1093 / bib / bbabl95. luchi, H., Matsutani, T., Yamada, K., Iwano, N., Sumi, S., Hosoda, S., Zhao, S., Fukunaga, T., and Hamada, M. (2021). Representation learning applications in biological sequence analysis. Comput. Struct. Biotechnol. J. 19, 3198-3208. doi: 10.1016 / j.csbj.2021.05.039.

[0225] Jacobson, E. C., Pandya-Jones, A., and Plath, K. (2022). A lifelong duty: how Xist maintains the inactive X chromosome. Curr. Opin. Genet. Dev. 75, 101927. doi: 10.1016 / j.gde.2022.101927.

[0226] Ji, Y., Zhou, Z., Liu, H., andDavuluri, R. V. (2021). DNABERT: pre-trained Bidirectional Encoder Representations from Transformers model for DNA-language in genome. Bioinformatics 37, 2112-2120. doi:10.1093 / bioinformatics / btab083.

[0227] Johnsson, P., Lipovich, L., Grander, D., and Morris, K. V. (2014). Evolutionary conservation of long non-coding RNAs; sequence, structure, function. Biochim. Biophys. Acta 1840, 1063-1071. doi:10.1016 / j.bbagen.2013.10.035.

[0228] Jonas, K., Cahn, G. A., and Pichler, M. (2020). RNA-Binding Proteins as Important Regulators of Long Non-Coding RNAs in Cancer. Int. J. Mol. Sci. 21. doi:10.3390 / ijms21082969.

[0229] Kagda, M. S., Lam, B., Litton, C., Small, C., Sloan, C. A., Spragins, E., Tanaka, F., Whaling, I., Gabdank, I., Youngworth, I., et al. (2023). Data navigation on the ENCODE portal. arXiv. doi: 10.48550 / arxiv.2305.00006.

[0230] Karolchik, D., Hinrichs, A. S., Furey, T. S., Roskin, K. M., Sugnet, C. W., Haussler, D., and Kent, W. J. (2004). The UCSC Table Browser data retrieval tool. Nucleic Acids Res. 32, D493- 6. doi: 10.1093 / nar / gkhl03.

[0231] Kent, W. J., Sugnet, C. W., Furey, T. S., Roskin, K. M., Pringle, T. H., Zahler, A. M., and Haussler, D. (2002). The human genome browser at UCSC. Genome Res. 12, 996-1006. doi: 10.1101 / gr.229102. Attorney Docket No.: 421 / 553 PCT

[0232] Kirk, J. M., Kim, S. O., Inoue, K., Smola, M. J., Lee, D. M., Schertzer, M. D., Wooten, J. S., Baker, A. R., Sprague, D., Collins, D. W., et al. (2018). Functional classification of long noncoding RNAs by k-mer content. Nat. Genet. 50, 1474-1482. doi:10.1038 / s41588-018- 0207-8.

[0233] Kuret, K., Amalietti, A. G., Jones, D. M., Capitanchik, C., and Ule, J. (2022). Positional motif analysis reveals the extent of specificity of protein-RNA interactions observed by CLIP. Genome Biol. 23, 191. doi:10.1186 / sl3059-022-02755-2.

[0234] Lambert, N., Robertson, A., Jangi, M., McGeary, S., Sharp, P. A., and Burge, C. B. (2014). RNA Bind-n-Seq: quantitative assessment of the sequence and structural binding specificity of RNA binding proteins. Mol. Cell 54, 887-900. doi:10.1016 / j.molcel.2014.04.016.

[0235] Lee, J., Nguyen, P. T., Shim, H. S., Hyeon, S. J., Im, H., Choi, M.-H., Chung, S., Ko wall, N. W., Lee, S. B., and Ryu, H. (2019). EWSR1, a multifunctional protein, regulates cellular function and aging via genetic and epigenetic pathways. Biochim. Biophys. Acta Mol. Basis Dis. 1865, 1938-1945. doi:10.1016 / j.bbadis.2018.10.042.

[0236] Liang, W.-C., Fu, W.-M., Wong, C.-W., Wang, Y., Wang, W.-M., Hu, G.-X., Zhang, L., Xiao, L - J., Wan, D. C.-C., Zhang, J.-F., et al. (2015). The IncRNA Hl 9 promotes epithelial to mesenchymal transition by functioning as miRNA sponges in colorectal cancer. Oncotarget 6, 22513-22525. doi:10.18632 / oncotarget.4154.

[0237] Li, J.-H., Liu, S., Zhou, H., Qu, L.-H., and Yang, J.-H. (2014). starBase v2.0: decoding miRNA- ceRNA, miRNA-ncRNA and protein-RNA interaction networks from large-scale CLIP- Seq data. Nucleic Acids Res. 42, D92-7. doi:10.1093 / nar / gktl248.

[0238] Loda, A., and Heard, E. (2019). Xist RNA in action: Past, present, and future. PLoS Genet. 15, el008333. doi: 10.1371 / joumal.pgen.1008333.

[0239] Matouk, I. J., Halle, D., Gilon, M., and Hochberg, A. (2015). The non-coding RNAs of the H19- IGF2 imprinted loci: a focus on biological roles and therapeutic potential in Lung Cancer. J. Transl. Med. 13, 113. doi: 10.1186 / sl2967-015-0467-3.

[0240] Mattick, J. S., Amaral, P. P., Caminci, P., Carpenter, S., Chang, H. Y., Chen, L.-L., Chen, R., Dean, C., Dinger, M. E., Fitzgerald, K. A., et al. (2023). Long non-coding RNAs: definitions, functions, challenges and recommendations. Nat. Rev. Mol. Cell Biol. 24, 430-447. doi: 10.1038 / s41580-022-00566-8.

[0241] McHugh, C. A., Chen, C.-K., Chow, A., Surka, C. F., Tran, C., McDonel, P., Pandya-Jones, A., Blanco, M., Burghard, C., Moradian, A., et al. (2015). The Xist IncRNA interacts directly with SHARP to silence transcription through HDAC3. Nature 521, 232-236. doi : 10.1038 / nature 14443. Attorney Docket No.: 421 / 553 PCT

[0242] McIntyre, K. L., Waters, S. A., Zhong, L., Hart-Smith, G., Raftery, M., Chew, Z. A., Patel, H. R., Graves, J. A. M., and Waters, P. D. (2024). Identification of the RSX interactome in a marsupial shows functional coherence with the Xist interactome during X inactivation. Genome Biol. 25, 134. doi:10.1186 / sl3059-024-03280-0.

[0243] Mercer, T. R., and Mattick, J. S. (2013). Structure and function of long noncoding RNAs in epigenetic regulation. Nat. Struct. Mol. Biol. 20, 300-307. doi: 10.1038 / nsmb.2480.

[0244] Minajigi, A., Froberg, J., Wei, C., Sunwoo, H., Kesner, B., Colognori, D., Lessing, D., Payer, B., Boukhali, M., Haas, W., et al. (2015). Chromosomes. A comprehensive Xist interactome reveals cohesin repulsion and an RNA-directed chromosome conformation. Science 349. doi: 10.1126 / science.aab2276.

[0245] Moore, K. S., and ’t Hoen, P. A. C. (2019). Computational approaches for the analysis of RNA- protein interactions: A primer for biologists. J. Biol. Chem. 294, 1-9. doi: 10.1074 / jbc.REVl 18.004842.

[0246] Necsulea, A., Soumillon, M., Wamefors, M., Liechti, A., Daish, T., Zeller, U., Baker, J. C., Grtitzner, F., and Kaessmann, H. (2014). The evolution of IncRNA repertoires and expression patterns in tetrapods. Nature 505, 635-640. doi: 10.1038 / nature 12943.

[0247] Noh, J. H., Kim, K. M., McClusky, W. G, Abdelmohsen, K., and Gorospe, M. (2018). Cytoplasmic functions of long noncoding RNAs. Wiley Interdiscip. Rev. RNA 9, el471. doi: 10.1002 / wma.1471.

[0248] Pandya-Jones, A., Markaki, Y., Serizay, J., Chitiashvili, T., Mancia Leon, W. R., Damianov, A., Chronis, C., Papp, B., Chen, C.-K., McKee, R., et al. (2020). A protein assembly mediates Xist localization and gene silencing. Nature 587, 145-151. doi:10.1038 / s41586-020-2703- 0.

[0249] Pan, X., Fang, Y., Li, X., Yang, Y., and Shen, H.-B. (2020). RBPsuite: RNA-protein binding sites prediction suite based on deep learning. BMC Genomics 21, 884. doi:10.1186 / sl2864-020- 07291-6.

[0250] Pan, X., Yang, Y., Xia, C.-Q., Mirza, A. H., and Shen, H.-B. (2019). Recent methodology progress of deep learning for RNA-protein interaction prediction. Wiley Interdiscip. Rev. RNA 10, el 544. doi: 10.1002 / wma.1544.

[0251] Quinn, J. J., and Chang, H. Y. (2016). Unique features of long non-coding RNA biogenesis and function. Nat. Rev. Genet. 17, 47-62. doi: 10.1038 / nrg.2015.10.

[0252] Ransohoff, J. D., Wei, Y., and Khavari, P. A. (2018). The functions and unique features of long intergenic non-coding RNA. Nat. Rev. Mol. Cell Biol. 19, 143-157. doi:10.1038 / nrm.2017.104. Attorney Docket No.: 421 / 553 PCT

[0253] Rinn, J. L., and Chang, H. Y. (2012). Genome regulation by long noncoding RNAs. Annu. Rev. Biochem. 81, 145-166. doi:10.1146 / annurev-biochem-051410-092902.

[0254] Rinn, J. L., and Chang, H. Y. (2020). Long noncoding mas: molecular modalities to organismal functions. Annu. Rev. Biochem. 89, 283-308. doi:10.1146 / annurev-biochem-062917- 012708.

[0255] Ross, C. J., Rom, A., Spinrad, A., Gelbard-Solodkin, D., Degani, N., and Ulitsky, I. (2021). Uncovering deeply conserved motif combinations in rapidly evolving noncoding sequences. Genome Biol. 22, 29. doi: 10.1186 / sl3059-020-02247-l.

[0256] Sahakyan, A., Yang, Y., and Plath, K. (2018). The Role of Xist in X-Chromosome Dosage Compensation. Trends Cell Biol. 28, 999-1013. doi: 10.1016 / j.tcb.2018.05.005.

[0257] Sprague, D., Waters, S. A., Kirk, J. M., Wang, J. R., Samollow, P. B., Waters, P. D., and Calabrese, J. M. (2019). Nonlinear sequence similarity between the Xist and Rsx long noncoding RNAs suggests shared functions of tandem repeat domains. RNA 25, 1004-1019. doi: 10.1261 / ma.069815.118.

[0258] St Laurent, G, Wahlestedt, C., and Kapranov, P. (2015). The Landscape of long noncoding RNA classification. Trends Genet. 31, 239-251. doi: 10.1016 / j .tig.2015.03.007.

[0259] Sun, Y., Arslan, P. E., Won, A., Yip, C. M., and Chakrabartty, A. (2014). Binding of TDP-43 to the 3’UTR of its cognate mRNA enhances its solubility. Biochemistry 53, 5885-5894. doi: 10.1021 / bi500617x.

[0260] Teng, X., Chen, X., Xue, H., Tang, Y., Zhang, P., Kang, Q., Hao, Y., Chen, R., Zhao, Y., and He, S. (2020). NPInter v4.0: an integrated database of ncRNA interactions. Nucleic Acids Res. 48, D160-D165. doi: 10.1093 / nar / gkz969.

[0261] Ule, J., Hwang, H.-W., and Darnell, R. B. (2018). The Future of Cross-Linking and Immunoprecipitation (CLIP). Cold Spring Harb. Perspect. Biol. 10. doi: 10.1101 / cshperspect.a032243.

[0262] Ulitsky, I. (2016). Evolution to the rescue: using comparative genomics to understand long noncoding RNAs. Nat. Rev. Genet. 17, 601-614. doi:10.1038 / nrg.2016.85.

[0263] Van Nostrand, E. L., Freese, P., Pratt, G. A., Wang, X., Wei, X., Xiao, R., Blue, S. M., Chen, I- Y., Cody, N. A. L., Dominguez, D., et al. (2020a). A large-scale binding and functional map of human RNA-binding proteins. Nature 583, 711-719. doi:10.1038 / s41586-020- 2077-3.

[0264] Van Nostrand, E. L., Pratt, G. A., Shishkin, A. A., Gelboin-Burkhart, C., Fang, M. Y., Sundararaman, B., Blue, S. M., Nguyen, T. B., Surka, C., Elkins, K., et al. (2016). Robust transcriptome-wide discovery of RNA-binding protein binding sites with enhanced CLIP (eCLIP). Nat. Methods 13, 508-514. doi:10.1038 / nmeth.3810. Attorney Docket No.: 421 / 553 PCT

[0265] Van Nostrand, E. L., Pratt, G. A., Yee, B. A., Wheeler, E. C., Blue, S. M., Mueller, J., Park, S. S., Garcia, K. E., Gelboin-Burkhart, C., Nguyen, T. B., et al. (2020b). Principles of RNA processing from analysis of enhanced CLIP maps for 150 RNA binding proteins. Genome Biol. 21, 90. doi:10.1186 / sl3059-020-01982-9.

[0266] Wolin, E., Guo, J. K., Blanco, M. R., Perez, A. A., Goronzy, I. N., Abdou, A. A., Gorhe, D., Guttman, M., and Jovanovic, M. (2023). SPIDR: a highly multiplexed method for mapping RNA-protein interactions uncovers a potential mechanism for selective translational suppression upon cellular stress. BioRxiv. doi: 10.1101 / 2023.06.05.543769.

[0267] Wu, Y., Huang, C., Meng, X., and Li, J. (2015). Long Noncoding RNA MALAT1 : Insights into its Biogenesis and Implications in Human Disease. Curr. Pharm. Des. 21, 5017-5028. doi: 10.2174 / 1381612821666150724115625.

[0268] Yamada, K., and Hamada, M. (2022). Prediction of RNA-protein interactions using a nucleotide language model. Bioinformatics Advances 2, vbac023. doi:10.1093 / bioadv / vbac023.

[0269] Yi, W., Li, J., Zhu, X., Wang, X., Fan, L., Sun, W., Liao, L., Zhang, J., Li, X., Ye, J., et al. (2020). CRISPR-assisted detection of RNA-protein interactions in living cells. Nat. Methods 17, 685-688. doi:10.1038 / s41592-020-0866-0.

[0270] Zhang, X., Hamblin, M. H., and Yin, K.-J. (2017). The long noncoding RNA Malatl: Its physiological and pathophysiological functions. RNA Biol. 14, 1705-1714. doi: 10.1080 / 15476286.2017.1358347.

[0271] Zhou, Z., Ji, Y., Li, W., Dutta, P., Davuluri, R., and Liu, H. (2023). Dnabert-2: Efficient foundation model and benchmark for multi-species genome. arXiv preprint arXiv 2306.15006.

[0272] No admission is made that any reference, including any non-patent or patent document cited in this specification, constitutes prior art. In particular, it will be understood that, unless otherwise stated, reference to any document herein does not constitute an admission that any of these documents forms part of the common general knowledge in the art in the United States or in any other country. Any discussion of the references states what their authors assert, and the applicant reserves the right to challenge the accuracy and pertinence of any of the documents cited herein. All references cited herein are fully incorporated by reference, unless explicitly indicated otherwise. The present disclosure shall control in the event there are any disparities between any definitions and / or description found in the cited references.

[0273] One skilled in the art will readily appreciate that the present disclosure is well adapted to carry out the objects and obtain the ends and advantages mentioned, as well as those inherent therein. The disclosures described herein are presently representative and / or exemplary, and are not intended as limitations on the scope of the present disclosure. Changes therein and other uses Attorney Docket No.: 421 / 553 PCT will occur to those skilled in the art which are encompassed within the spirit of the present disclosure as defined by the scope of the claims.

[0274] It will be understood that various details of the presently disclosed subject matter may be changed without departing from the scope of the presently disclosed subject matter. Furthermore, the foregoing description is for the purpose of illustration only, and not for the purpose of limitation.

Claims

Attorney Docket No.: 421 / 553 PCTCLAIMSWhat is claimed is:

1. A method of predicting one or more binding sites between a ribonucleotide (RNA) sequence and a RNA binding protein, the method comprising: preparing sequence data for a plurality of overlapping RNA segments derived from a target sequence; inputting the sequence data for the plurality of overlapping RNA segments into a machine learning model trained to predict a probability of an interaction between a RNA sequence and a RNA binding protein occurring in a RNA sequence; outputting from the machine learning model one or more segments in the sequence data for the plurality of overlapping RNA segments falling above a probability threshold indicative of an interaction between a RNA sequence and a RNA binding protein; and predicting one or more binding sites between a RNA sequence and a RNA binding protein in the sequence data for the plurality of overlapping RNA segments derived from the target sequence based on the output from the model.

2. The method of claim 1, further comprising mapping the predicted binding site back to the target sequence.

3. The method of claim 1 or claim 2, wherein the overlapping RNA segments derived from a target sequence overlap by a tunable nucleotide parameter.

4. The method of any one of claims 1-3, wherein the target sequence comprises a mRNA sequence, a spliced transcript sequence, an unspliced transcript sequence, an artificial sequence, or a long noncoding RNA (IncRNA) sequence.

5. The method of any one of claims 1-4, wherein the target sequence comprises at least one target sequence from at least two different species.

6. The method of claim 5, further comprising inferring binding site conservation across the two or more species.

7. The method of any one of claims 1-6, wherein the target sequence is associated with a disease state.

8. The method of claim 7, further comprising designing a modulator of binding at the one or more binding sites between a RNA sequence and a RNA binding protein in the sequence data.

9. The method of claim 8, wherein the modulator of binding is a nucleotide-based modulator or a small molecule modulator.

10. The method of any one of claims 1-9, wherein the machine learning model comprises a BERT-RBP model, a BERT2-RBP model, or a combination thereof.Attorney Docket No.: 421 / 553 PCT11. The method of any one of claims 1-10, further comprising providing an interactive graphical user interface that loads and displays probabilities, allows for tuning of parameters, and / or displays performance characteristics about the model that was used.

12. A system for predicting one or more binding sites between a ribonucleotide (RNA) sequence and a RNA binding protein, the system comprising: a computing platform including at least one processor and memory; and a trained machine learning model stored in the memory and executable by the at least one processor for receiving, as input, sequence data for a plurality of overlapping RNA segments into a machine learning model trained to predict a probability of an interaction between a RNA sequence and a RNA binding protein occurring in a RNA sequence and generating, as output, one or more of the plurality of overlapping RNA segments falling above a probability threshold indicative of an interaction between a RNA sequence and a RNA binding protein.

13. The system of claim 12, wherein the overlapping RNA segments derived from a target sequence overlap by a tunable nucleotide parameter.

14. The system of claim 12 or claim 13, wherein the target sequence comprises a mRNA sequence, a spliced transcript sequence, an unspliced transcript sequence, an artificial sequence, or a long noncoding RNA (IncRNA) sequence.

15. The system of any one of claims 12-14, wherein the target sequence comprises at least one target sequence from at least two different species.

16. The system of any one of claims 12-15, wherein the target sequence is associated with a disease state.

17. The system of any one of claims 12-16, wherein the machine learning model comprises a BERT-RBP model, a BERT2-RBP model, or a combination thereof.

18. A method of predicting one or more binding sites between a ribonucleotide (RNA) sequence and a RNA binding protein, the method comprising: receiving, as input, RNA sequence data; precomputing, using a machine learning model, binding probabilities for a plurality of overlapping RNA segments; outputting the binding probabilities; storing in a data structure, the binding probabilities and corresponding RNA segment identifying data; using the RNA sequence data to access the data structure and locate binding probabilities for overlapping segments in the RNA sequence data; andAttorney Docket No.: 421 / 553 PCT predicting one or more binding sites between a RNA sequence and a RNA binding protein in the RNA sequence data for the plurality of overlapping RNA segments derived from the target sequence based on the output.

19. A non-transitory computer readable medium comprising computer executable instructions embodied in a computer readable medium that when executed by at least one processor of a computer cause the computer to perform steps of any one of claims 1-11.

20. A non-transitory computer readable medium comprising computer executable instructions embodied in a computer readable medium that when executed by at least one processor of a computer cause the computer to perform steps of claim 18.