Method and system for predicting binding site of non-targeted probing RNA-binding proteins on RNA

By integrating mutation and stop signals using the STONE method, the observation dimensions are broadened, solving the accuracy problem of predicting RNA-binding protein binding sites in low-accessibility regions. This improves the precision and breadth of RNA structure analysis and is applicable to the study of both short and long RNAs.

CN120072034BActive Publication Date: 2026-05-05SHANDONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANDONG UNIV
Filing Date
2025-02-08
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing technologies cannot accurately predict the binding sites of RNA-binding proteins in low-accessibility regions, resulting in insufficient accuracy in RNA structure analysis.

Method used

The STONE method is employed to broaden the observation dimensions by integrating mutation and stop signals. The enhanced difference is used to infer RBP binding sites, thereby improving the accuracy of RNA structure analysis in low-accessibility regions.

Benefits of technology

It improves the accuracy of identifying RNA-binding protein binding sites, enhances the depth and breadth of RNA structure analysis, is applicable to the analysis of both short and long RNAs, and provides a more powerful tool for RNA biology research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120072034B_ABST
    Figure CN120072034B_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for predicting RNA binding protein binding sites on RNA without targeting, relating to the field of bioinformatics. The method includes the following steps: acquiring raw RNA structure data; expanding and filtering the preprocessed data; training an RNA analysis model using the filtered data; processing the RNA to be analyzed using the trained RNA analysis model, calculating the probability that each nucleotide is single-stranded, and obtaining a structure score; identifying low-accessibility regions based on the difference between the structure score output by the RNA analysis model and the structure score calculated by the original method; and identifying potential RNA binding protein binding sites by comparing the signal difference with that of the original method in a single RNA structure detection experiment. This invention utilizes the STONE method to improve the accuracy of RNA structure analysis in low-accessibility regions, and uses the improved difference to infer RBP binding sites, thereby improving the accuracy of RBP binding site identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of bioinformatics, and in particular to a method and system for predicting the binding sites of RNA-binding proteins on RNA without targeting. Background Technology

[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.

[0003] RNA-binding proteins (RBPs), as key RNA regulators, contain important clues about their binding sites within RNA structural information. By binding to specific RNA sequences or structures, RBPs participate in a variety of biological processes, including RNA splicing, translation, stability, and transport. These processes play crucial roles in cellular function and development, and abnormal RBP-RNA interactions are often associated with various diseases, such as cancer and neurodegenerative diseases. Therefore, accurately identifying and locating RBP binding sites not only helps elucidate the fundamental mechanisms of RNA biology but also provides potential opportunities for research into disease mechanisms and the discovery of novel therapeutic targets.

[0004] In recent years, with in-depth research on the mechanism of action of RBP and the continuous advancement of related technologies and methods, these binding sites can be revealed more accurately, providing a more powerful tool for RNA biology research.

[0005] Traditionally, methods for detecting RBP binding sites can be broadly categorized into two types: experimental methods and computational prediction methods. Experimental methods are further divided into targeted and non-targeted approaches. Experimental methods for detecting RBP binding sites primarily rely on techniques such as cross-linked immunoprecipitation and sequencing (CLIP-seq). While these experimental methods provide important information for revealing RBP binding sites, they still have limitations due to the limited availability of human data and the complexity of the experiments. Computational prediction methods, combining known RBP binding site data with RNA structural features, train models to identify potential binding sites, significantly improving prediction accuracy. For example, PrismNet, a deep learning-based computational tool, can efficiently and accurately predict RBP binding sites. Smola et al. non-targetedly detected the RBP binding site of mouse Xist lncRNA by comparing the differences in SHAPE scores under in vitro and in vivo conditions, and verified that their results were highly consistent with CLIP-seq data. However, there is currently no method that relies entirely on RNA structural signals obtained from a single experiment to identify RBP binding sites.

[0006] RNA structure is fundamental to understanding RNA function because it dynamically changes under different cellular conditions. Furthermore, the relationship between RNA structure and RNA-binding proteins (RBPs) is close and complex. RBPs regulate RNA's biological functions by interacting with specific sequences or structural regions of RNA, while RNA structure also determines the binding properties and function of RBPs to some extent. Therefore, accurate RNA structure analysis is crucial for revealing these complexities.

[0007] Methods for determining RNA structure are mainly divided into experimental methods and predictive methods. Compared with other methods, chemical probing methods have unique advantages. Weeks et al. proposed that chemical modification data can not only provide information on which regions of RNA are exposed, but also reflect how RNA folds into different structures under different conditions. By comparing the patterns of chemical modification under different experimental conditions, researchers can identify changes in the structural state of RNA molecules. Chemical modification reagents can react with specific sites on RNA, modifying single-stranded RNA regions, leading to mutations in cDNA or early termination of synthesis during reverse transcription, and the extent of these reactions is closely related to the spatial structure of RNA. When the RNA strand folds into secondary structures, some nucleotides are exposed, while others are "hidden" because they are wrapped in double strands or form complex spatial structures. This structural difference causes exposed and hidden nucleotides to exhibit different characteristics in their reactivity with chemical modification reagents. Researchers can use this data to construct RNA folding models or further refine the secondary and three-dimensional structures of RNA. Based on this, a series of methods have been developed, among which DMS and SHAPE reagents are the most commonly used tools.

[0008] RNA-protein interactions limit the accessibility of chemical modification reagents. Many RNA-binding proteins bind to RNA at sites located in structurally compact or complex regions, which may be poorly characterized in standard RNA structure probing due to their low accessibility. While masking inaccessible regions can improve the accuracy of RNA structure analysis, most RNAs lack comprehensive accessibility data. Although RNA structure probing methods (such as chemical modification methods, SHAPE, DMS, etc.) can provide detailed information about RNA folding states, there is no method to directly infer the binding site of RNA-binding proteins (RBPs) from single RNA structure data.

[0009] Therefore, how to improve the accuracy of predicting the binding sites of RNA-binding proteins in low-accessibility regions has become a technical problem that needs to be solved by existing technologies. Summary of the Invention

[0010] To address the shortcomings of existing technologies, the present invention aims to provide a method and system for predicting the binding sites of RNA-binding proteins on RNA without targeting. By utilizing the STONE method and integrating mutation and stop signals, the observation dimensions are broadened, improving the accuracy of RNA structure analysis in low-accessibility regions. The improved difference is then used to infer the RBP binding site, thereby enhancing the accuracy of RBP binding site identification.

[0011] To achieve the above objectives, the present invention is implemented through the following technical solution:

[0012] The first aspect of this invention provides a method for predicting the binding site of RNA-binding proteins on RNA without targeting, comprising the following steps:

[0013] Obtain raw RNA structure data and preprocess the raw data;

[0014] The preprocessed data is expanded and filtered, and the filtered data is used to train the RNA analysis model.

[0015] The trained RNA analysis model is used to process the RNA to be analyzed, calculate the probability that each nucleotide is single-stranded, and obtain the structure score.

[0016] Based on the difference between the structure score output by the RNA analysis model and the structure score calculated by the original method, low-accessibility regions are identified.

[0017] Potential RNA-binding protein binding sites are identified by comparing the signal differences with those of the original method in a single RNA structure detection experiment.

[0018] Further steps in preprocessing the raw data include removing low-quality and duplicate data.

[0019] Furthermore, the specific steps for training the RNA analysis model are as follows:

[0020] The count and depth data of the processed signal are obtained by using sequence alignment and signal calling methods, and the ratio of stop signal to abrupt signal at each position is calculated.

[0021] The preprocessed data is expanded and filtered, and the filtered data is divided into training set and test set;

[0022] An RNA analysis model was constructed, trained using a training set, and tested using a test set.

[0023] Furthermore, the steps for expanding and filtering the preprocessed data include:

[0024] Feature engineering is used to expand the features of the preprocessed data;

[0025] Input features are selected based on the learning and generalization abilities of the RNA analysis model.

[0026] Furthermore, the RNA analysis models include the SHAPE model and the DMS model. The training data for the SHAPE model comes from human 18S rRNA data and mouse data; the training data for the DMS model includes 18S rRNA data from Arabidopsis thaliana, human data, and yeast data.

[0027] Furthermore, the RNA analysis model was trained using a 10-fold cross-validation method, and the optimal model parameters were determined based on the highest AUC value obtained from the dataset.

[0028] Furthermore, when calculating the difference between the structure score output by the RNA analysis model and the structure score calculated by the original method, in order to mitigate the impact of outliers and ensure data comparability within a consistent range, the structure scores output by both the original method and the RNA analysis model were adjusted by quantile normalization. Nonlinear transformations were also applied to the normalized stop and mutation signals to adjust the weight differences within their respective numerical ranges.

[0029] A second aspect of the present invention provides a predictive system for non-targeted detection of RNA-binding protein binding sites on RNA, comprising:

[0030] The data acquisition module is configured to acquire raw RNA structure data and preprocess the raw data.

[0031] The feature expansion module is configured to expand and filter the preprocessed data, and use the filtered data to train the RNA analysis model;

[0032] The RNA analysis module is configured to process the RNA to be analyzed using a trained RNA analysis model, calculate the probability that each nucleotide is single-stranded, and obtain a structure score.

[0033] The region identification module is configured to identify low-accessibility regions based on the difference between the structure score output by the RNA analysis model and the structure score calculated by the original method.

[0034] The binding site prediction module is configured to identify potential RNA-binding protein binding sites by comparing the signal differences with those of the original method in a single RNA structure detection experiment.

[0035] A third aspect of the present invention provides a medium having a program stored thereon, which, when executed by a processor, implements the steps of the method for predicting the binding site of RNA-binding proteins on RNA as described in the first aspect of the present invention.

[0036] A fourth aspect of the present invention provides an apparatus including a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor executes the program to perform the steps of the method for predicting the binding site of an RNA-binding protein without targeting, as described in the first aspect of the present invention.

[0037] The above one or more technical solutions have the following beneficial effects:

[0038] This invention discloses a non-targeted method and system for predicting the binding sites of RNA-binding proteins on RNA. The proposed STONE (Determine RNA Structure with Truncation-mutation Signals within ONE Experiment) method is applied to RNA structure analysis. Through the application of feature engineering and machine learning models, it can effectively integrate multiple signal features and improve the prediction accuracy of RNA binding sites.

[0039] This invention enhances the analytical capabilities for low-accessibility regions. Traditional RNA structure detection methods often face problems such as weak signals and difficulty in parsing low-accessibility regions. The STONE method, by broadening the observation dimensions, can provide more accurate RNA structure analysis in these low-accessibility regions, overcoming the limitations of previous methods.

[0040] This invention enables efficient RBP binding site estimation. The STONE method uses the increased difference to infer RBP binding sites, identifying potential binding sites even in the absence of comprehensive and available data. This characteristic provides new ideas and tools for RNA biology research.

[0041] The STONE method of this invention is applicable not only to short RNAs but also to the effective identification of RBP-binding regions in long RNAs, demonstrating its broad applicability in the analysis of different types of RNA. Furthermore, the STONE method can be combined with existing chemical modification techniques (such as SHAPE and DMS) to further enhance the depth and breadth of RNA structure analysis, providing a more powerful tool for RNA biology research.

[0042] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0043] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0044] Figure 1 This is a flowchart of the STONE method in Embodiment 1 of the present invention;

[0045] Figure 2 This is a schematic diagram illustrating the contribution of quantified features from four datasets to the final STONE AUC value in Embodiment 1 of the present invention.

[0046] Figure 3 This is a visual schematic diagram of IGV in Embodiment 1 of the present invention;

[0047] Figure 4 This is a schematic diagram of the RNA and RBP structure and interaction in the human 18S rRNA region 2 in Embodiment 1 of the present invention.

[0048] Figure 5 This is a schematic diagram illustrating STONE's performance in identifying the location of a single strand of human 18S rRNA in Embodiment 1 of the present invention without providing accessibility information.

[0049] Figure 6 A schematic diagram illustrating the calculation of ΔSHAPE scores using STONE-derived structural scoring in Embodiment 1 of the present invention to identify RBP binding sites in human U1 snRNA and human RNaseP.

[0050] Figure 7 This is a schematic diagram comparing the region level of the ΔSHAPE RBP binding data calculated in Embodiment 1 of the present invention with CLIP-seq data on human XIST long non-coding RNA.

[0051] Figure 8 This is a schematic diagram comparing the RBP binding regions on human XIST long non-coding RNA calculated using ΔSHAPE and RNP-MaP in Embodiment 1 of the present invention. Detailed Implementation

[0052] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0053] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0054] Example 1:

[0055] Currently, existing RNA structure analysis methods suffer from low accuracy in low-accessibility regions. In well-characterized structures such as 18S rRNA, the signal intensity in certain regions is significantly reduced. This is mainly due to RNA-protein interactions limiting the accessibility of chemical modification reagents; therefore, these regions are termed low-accessibility regions. While masking low-accessibility regions can improve the accuracy of RNA structure analysis, comprehensive accessibility data is lacking for most RNAs.

[0056] In RNA analysis, protein-RNA interactions have a significant impact on RNA secondary structure and folding, especially in low-accessibility regions. These regions may have their structural protection further enhanced by protein binding, making it difficult for modifying reagents to enter and bind to the target site.

[0057] Proteins stabilize RNA folding by binding to specific domains, particularly in inaccessible regions. These regions are typically located in the core or relatively enclosed parts of the RNA, and protein binding makes them more stable, thus shielding them from the entry of modifying agents. This means that although modifying agents (such as DMS or SHAPE) are designed to label exposed bases of RNA, they often struggle to penetrate and react effectively with RNA in inaccessible regions due to the protective effect of proteins. Consequently, the chemical modification signals in these regions are weak, affecting the accuracy of subsequent analysis and sequencing.

[0058] Reverse transcriptase plays a crucial role in RNA structure analysis. It transcribes RNA into cDNA, allowing the sequencing process to read the RNA sequence. However, reverse transcriptase typically acts after RNA has undergone chemical modification or thermal denaturation. Chemical modification agents alter the chemical properties of RNA by reacting with exposed bases, revealing specific structural features and providing information for reverse transcription. In low-accessibility regions, due to their high structural stability, they are often difficult for chemical modification agents and reverse transcriptase to effectively interact with. Modification agents label RNA before reverse transcription to reveal its structural features, but the protective structures of low-accessibility regions often prevent modification agents from penetrating these areas, thus affecting the accuracy of transcription in these regions.

[0059] To address the aforementioned deficiencies in existing technologies, Embodiment 1 of this invention provides a method for predicting the binding sites of RNA-binding proteins on RNA without targeting. The STONE method can broaden the observation dimensions and improve the accuracy in low-accessibility regions.

[0060] The STONE method involves performing sequence alignment and signal retrieval on the data, merging and processing the signal data, and then integrating the mutation and stop signals from a single experiment through feature engineering and model training, thereby improving the accuracy of signal analysis.

[0061] like Figure 1 As shown, STONE integrates two signals through feature engineering and a machine learning model. The workflow involves analyzing mutation and stop signals separately, then expanding the features of these signals through feature engineering, and finally applying them to a machine learning model to calculate the probability that each nucleotide is single-stranded, i.e., the STONE score.

[0062] Specifically, the following steps are included:

[0063] Step 1: Obtain raw RNA structure data and preprocess the raw data.

[0064] In one specific implementation, the raw data includes mutation signals and stop signals. In this embodiment, the purpose of the model prediction process is to highlight the characteristics of these two signals in order to perform RNA structure analysis.

[0065] Step 1.1: Raw data were downloaded from the RASP database or other existing databases. The final raw data dataset contained 11 studies involving five species, six small molecule reagents, and four reverse transcriptases. All studies were included in the signal distribution calculations, except for the single-cell dataset. Human, mouse, viral, and yeast RNA were analyzed in particular for evaluating RNA with known structures. High-depth representative datasets were used for genome-wide assessments. The single-cell dataset included two batches of H9 cells, totaling 76 cells. CLIP-seq data and RNP-MaP data from the HEK293T cell line were used for the RNA-binding protein (RBP) binding site identification application. All data were currently available, known, and legitimate.

[0066] Step 1.2: The preprocessing steps for the raw data include removing low-quality data and duplicate data.

[0067] After obtaining the raw sequencing files in FASTQ format, low-quality reads were removed using the FASTP software with default parameter settings of 43. FASTP was used for read filtering, with parameters set to remove duplicates, apply a minimum Phred score of 20, exclude reads with >20% low-quality bases, and retain a minimum read length of 5 bases. This rigorous preprocessing step is crucial for minimizing noise and improving the reliability of subsequent analyses. Adapters were removed either automatically or by specifying details from the original literature.

[0068] Step 2: Expand and filter the preprocessed data, and use the filtered data to train the RNA analysis model.

[0069] Step 2.1: Obtain the count and depth data of the processed signal using sequence alignment and signal calling methods, and calculate the ratio of stop signal to abrupt change signal at each position.

[0070] The ratio of stop signal to mutation signal at each position is used to assess the balance between the two signals generated by different chemical modification strategies. The STONE method performs better when both signals are balanced and both have high signal intensity.

[0071] In one specific implementation, Bowtie2 is used to construct a genome reference index, and reads are aligned to a reference genome. The `--local` option is used to optimize alignment efficiency. The generated BAM file is then used for mutation signal invocation via the RNA Framework, with configured parameters to exclude insertions while retaining default settings for deletion operations. The `--orc` option is implemented to include raw mutation counts summarizing mutation directions, insertions, and deletions. For stop signal invocation, the default settings recommended by icSHAPE-pipe are followed. Finally, the ratio of stop signals to mutation signals at each location is calculated based on the counts and depths of the processed signals generated from icSHAPE-pipe and the RNA Framework.

[0072] Step 2.2: Expand and filter the preprocessed data.

[0073] Step 2.2.1: Use feature engineering to expand the features of the preprocessed data.

[0074] Specifically, 42 features were added based on the merged and processed signal data, such as rf_mutation_Count (mutation count) and pipe_truncation_RT (stop count).

[0075] Step 2.2.2: Select input features based on the learning and generalization abilities of the RNA analysis model.

[0076] After evaluating the model's learning and generalization abilities under various feature combinations, the number of input features was ultimately reduced to 15.

[0077] Specifically, the model's learning and generalization abilities were evaluated by calculating the AUC values ​​under various feature combinations. Feature combinations that resulted in high AUC values ​​were selected, and ultimately 15 features were chosen.

[0078] Step 2.3: Divide the filtered data into training set and test set.

[0079] Step 2.4: Construct an RNA analysis model, train the RNA analysis model using the training set, and test the trained RNA analysis model using the test set.

[0080] Step 2.4.1: Construct an RNA analysis model.

[0081] In this embodiment, the RNA analysis models include the SHAPE model and the DMS model. The training data for the SHAPE model comes from human 18S rRNA data and mouse data; the training data for the DMS model includes 18S rRNA data from Arabidopsis thaliana, human data, and yeast data. Positive samples represent unpaired sites, while negative samples represent paired sites. The SHAPE model has 7,106 positive samples and 6,907 negative samples; the DMS model has 1,814 positive samples and 1,471 negative samples. The data were first merged and then divided into training and testing sets, with 75% of the samples used for training and the remaining 25% used for testing.

[0082] Both the SHAPE and DMS models utilize the STONE (Determine RNA Structure with Truncation-mutation Signals within ONE Experiment) method, differing only in the chemical modification reagents used. The SHAPE model determines RNA structure using stop mutation signals against the SHAPE reagent in a single experiment, while the DMS model determines RNA structure using stop mutation signals against the DMS reagent in a single experiment. In this embodiment, different models are selected based on the type of reagent. Features from the different datasets obtained by the two models are used to train a signal integration model using a random forest model. The AUC value is calculated by inputting the obtained features into the random forest model, and parameters are adjusted by comparing the AUC values ​​to select higher-performing feature combinations.

[0083] In addition to the random forest method, gradient boosting, extreme gradient boosting, and decision tree methods can also be used to train two models, and hyperparameter tuning is performed to optimize model performance.

[0084] Step 2.4.2: Train the RNA analysis model using the training set.

[0085] The training set data was input into the random forest model, and hyperparameters were tuned to optimize model performance. The training process specifically includes the following steps:

[0086] A random forest model consists of multiple decision trees. Each tree is trained by randomly sampling subsets of data and features during training, thereby enhancing the model's robustness. The structure of each tree is based on the CART (Classification and Regression Tree) algorithm, using a binary tree structure. Each node in the tree represents a split of a feature, while leaf nodes represent classification results or regression values. The depth of the tree and the number of leaf nodes are determined by the hyperparameters `max_depth` and `min_samples_leaf`, while the number of features used for each node split is determined by `max_features`.

[0087] To optimize model performance, this embodiment employs a grid search method to optimize hyperparameters. The optimized hyperparameters include the number of trees (n_estimators), the maximum tree depth (max_depth), the minimum number of sample splits (min_samples_split), the minimum number of sample leaf nodes (min_samples_leaf), and the maximum number of features (max_features).

[0088] Grid search selects the optimal hyperparameter combination by iterating through all possible combinations of hyperparameters and using cross-validation to evaluate the performance of each combination.

[0089] During optimization, cross-validation helps reduce errors that may arise from uneven data splitting, and the optimal model is selected by averaging the performance from multiple evaluations. The optimization principle is based on reducing overfitting and underfitting. By adjusting hyperparameters, the complexity and expressive power of the model are optimized to achieve good performance on both the training set and unseen datasets.

[0090] After training, the model was evaluated on 25% of the test dataset and an independent validation dataset. Model performance was evaluated using multiple metrics, including AUC, F1 score, precision, and accuracy. AUC values ​​were calculated using the `roc_curve` and `auc` functions from the scikit-learn library. These evaluation metrics comprehensively reflect the model's performance, ensuring its generalization ability across different datasets.

[0091] Step 2.4.3: Train the RNA analysis model using the 10-fold cross-validation method, and determine the optimal model parameters based on the highest AUC value obtained from the dataset.

[0092] In this embodiment, the performance of the DMS and SHAPE models was validated using independent datasets by comparing structure scores with known RNA structures in the RNA STRAND database. DMS validation used datasets from DMS-seq, DMS-MaPseq, Mod-seq, and Structure-seq; SHAPE validation used datasets from SHAPE-Seq, icSHAPE, SHAPE-MaP, and sc-SPORT. Sites with missing structure scores were excluded, and the remaining sites were evaluated to determine if they corresponded to unpaired nucleotides, thereby calculating AUC values.

[0093] To further evaluate the model's generalization ability and potential overfitting, 10-fold cross-validation was employed, dividing the dataset into ten non-overlapping subsets. In each training iteration, nine subsets were used for training, while the remaining subset was used as the test set. The optimal model was selected based on the highest AUC value obtained on both the validation and test sets.

[0094] Step 3: Use the trained RNA analysis model to process the RNA to be analyzed, calculate the probability that each nucleotide is single-stranded, and obtain the structure score, i.e., the STONE score.

[0095] The STONE score represents the probability that each nucleotide is single-stranded, while the STONE signal is the probability obtained after integrating the mutation signal and the stop signal. The mutation signal is the probability obtained using the original RNA Framework method, and the stop signal is the probability obtained using the icSHAPE-pipe method.

[0096] The final structure scores (including stop, mutation, and stone signals) were visualized using IntegrativeGenomics Viewer and displayed alongside a reference sequence to demonstrate the structural signals at single-nucleotide resolution. To facilitate visualization and assessment of the scoring accuracy of different methods, a custom Python script was used to map these scores onto the secondary structure of human 18S rRNA.

[0097] Accessibility information was obtained from the 60S ribosome crystal structure (PDB ID: 5LKS), and the solvent-accessible surface area of ​​each nucleotide was calculated using PyMOL.

[0098] Step 4: Identify low-accessibility regions based on the difference between the structure score output by the RNA analysis model and the structure score calculated by the original method.

[0099] In this embodiment, the original methods used refer to icSHAPE-pipe and RNA Framework. icSHAPE-pipe is a method that analyzes only the stop signal to obtain the structure score, while RNA Framework is a method that analyzes the mutation signal to obtain the structure score. This embodiment selects these two methods to obtain detailed information about the files, while other existing methods are also applicable to this invention.

[0100] The original method works by generating a reference index and then performing read mapping. Specifically, it relies on Cutadapt to sequence adapters and low-quality base clips, and Bowtie 2 for sequence alignment. SAMtools is used to automatically sort the output alignments and convert them to BAM format. Then, the number of RT stops, mutations, and coverage depth for each base are calculated, known as RNA count (RC).

[0101] RC files store transcript sequences, RT stop / mutation counts for each base, read coverage for each base, and the total number of reads covering the transcripts. RC files from RNA structure detection experiments are reactivity-normalized, and the raw signal for each base is stored. The calculation formula is:

[0102] .

[0103] Here, nTi and cTi represent the mutation count and read coverage at transcript position i, respectively. After fractional calculation, the original reactivity is normalized: original reactivity values ​​above the 95th percentile are set to the 95th percentile, and original reactivity values ​​below the 5th percentile are set to the 5th percentile. Then, each reactivity value is divided by the 95th percentile, producing a reactivity value of 0 to 1, representing a low or high single-stranded tendency of the residues, respectively.

[0104] Step 4.1: In this embodiment, when calculating the difference between the structure score output by the RNA analysis model and the structure score calculated by the original method, in order to mitigate the impact of outliers and ensure the data is comparable within a consistent range, the structure scores output by both the original method and the RNA analysis model are adjusted to the [0, 1] interval by quantile normalization, and nonlinear transformations are applied to the normalized stop signal and mutation signal to adjust the weight differences within their respective numerical ranges.

[0105] Step 4.2: Calculate the ΔSHAPE value and compare the STONE output with the original signal. By comparing the signal difference (ΔSHAPE score) between the conventional method and the STONE method in a single RNA structure detection experiment, potential RBP binding sites can be identified.

[0106] The filtering steps are as follows:

[0107] 1. ΔSHAPE calculation: First, the difference in SHAPE reactivity under intracellular and extracellular conditions (ΔSHAPE) is calculated. This is obtained by subtracting the extracellular reactivity from the intracellular SHAPE reactivity and averaging over a trinucleotide sliding window to reduce local signal fluctuations.

[0108] 2. Z-factor test: The Z-factor test is used to assess the significance of the ΔSHAPE value. This test compares the ΔSHAPE value with associated extracellular and intracellular measurement errors, identifying nucleotide sites where the error magnitude is not significant with respect to the ΔSHAPE value. The Z-factor test requires that the difference in SHAPE reactivity between extracellular and intracellular measurements exceed 1.96 standard deviations (Z-factor > 0), and ensures that the 95% confidence intervals of the two measurements do not overlap.

[0109] 3. Standard Score Calculation: For each nucleotide, calculate the standard score of the ΔSHAPE value relative to the ΔSHAPE values ​​of all other nucleotides. This requires that the ΔSHAPE value of each nucleotide deviates from the average ΔSHAPE value by at least one standard deviation (absolute value ≥ 1), meaning that only the largest ΔSHAPE value is considered for further analysis.

[0110] Combining Z-factor and standard score: When finally determining RNA-protein interaction sites, both the Z-factor and standard score are considered. If at least three nucleotides within a five-nucleotide window have a Z-factor >0 and an absolute standard score ≥1, these nucleotide sites are considered to have undergone significant changes in SHAPE reactivity due to cellular environmental influences.

[0111] To calculate the ΔSHAPE value, the location of different signals (2964 sites) was identified using a sliding window method, and consecutive locations were grouped into regions (628 regions).

[0112] Specifically, within each window (size = 5 nt), the position identified by ΔSHAPE is considered valid only if at least three positions are identified. The calculation steps for ΔSHAPE are shown in formula (1):

[0113] (1).

[0114] In formula (1), depth represents the reading depth of a specific nucleotide, and structural scores represent the STONE score, mutation score, or stop score. The standard error (SEscore) of each score is calculated.

[0115] SEscore is defined as the SHAPE reactivity standard error. The standard error (SEscore) is defined as the reactivity measurement error estimated based on a Poisson distribution. Specifically, the SHAPE reactivity standard error is calculated based on the mutation rate (mutation count) per nucleotide and the read depth at that position. In the formula:

[0116] .

[0117] Here, λ represents the number of observed mutation events, "reads" represents the sequencing depth of the nucleotide, and "rate" is the number of mutation events per read. The variance of the Poisson distribution is equal to the number of events, therefore the standard error reflects the accuracy of each mutation event read.

[0118] (2).

[0119] The ΔSHAPE value for each nucleotide is calculated using formula (2), where S and O represent the structure scores obtained from the STONE and original methods, respectively. The ΔSHAPE value captures the differences between the scores and is averaged over a sliding window of three nucleotides.

[0120] To incorporate smoothing, the standard error values ​​calculated using formula (1) are averaged to ensure consistency with the method established by Smola et al. The smoothing error estimate for each nucleotide position is denoted by formula (3):

[0121] (3).

[0122] In formula (3), the Z factor (Z) is the value of each nucleotide (Z). The Z-value is calculated, where the subscripts S and O represent STONE and the original method, respectively. A Z-value greater than 0 indicates a significant change in the SHAPE reactivity of the nucleotide.

[0123] The standard fraction (Si) of each nucleotide was rigorously calculated, and the identification of potential binding sites was strictly performed according to the method of Smola et al.

[0124] Specifically, Smola et al. used two methods, SHAPE-MaP and icSHAPE, to study the sites of RNA-protein interactions, particularly by comparing changes in RNA SHAPE reactivity (ΔSHAPE) under different environments.

[0125] RNA was modified using 1M7 reagent to measure its SHAPE reactivity under both in vivo and in vitro conditions. These reactivity data can help identify interaction sites between RNA and protein complexes.

[0126] Nucleotides exhibiting significant changes in reactivity can be identified by calculating ΔSHAPE (i.e., the difference in SHAPE reactivity between in vitro and intracellular samples) and the Z factor for each nucleotide.

[0127] The difference between their methods is that their ΔSHAPE is the difference between experiments, while the ΔSHAPE score in this embodiment is the difference between the value obtained when calculating the AUC of the single-cell dataset generated by the SHAPE-MaP method and the AUC value calculated by applying the STONE model to the features selected in this embodiment. It is the subtraction between different calculation methods and can help identify potential RBP binding sites.

[0128] Step 5: Identify potential RNA-binding protein binding sites by comparing the signal differences with those of the original method in a single RNA structure detection experiment.

[0129] Step 6: Evaluate the ΔSHAPE result.

[0130] CLIP-seq data were obtained from the POSTAR3 CLIPdb database in BED format. Six experimental methods were used: PAR-CLIP, PIP-seq, eCLIP, HITS-CLIP, iCLIP, and 4SU-iCLIP. To extract data for specific genes, CLIP-seq RBP binding sites were annotated onto the corresponding transcriptomes using the GAP tool.

[0131] RNP-MaP site extraction was based on the method of Weidmann et al. The processed data files contained mutation frequencies and read depths for both cross-linked and non-cross-linked samples. All calculations were performed using a custom R script, generating candidate RNP-MaP sites on human XIST long non-coding RNAs (lncRNAs).

[0132] To verify the accuracy of the ΔSHAPE calculation, the ΔSHAPE results were compared with CLIP-seq and RNP-MaP results. In this analysis, the continuous nucleotide positions in the ΔSHAPE, CLIP-seq, and RNP-MaP data were grouped to independently represent their RBP binding regions. To minimize the possibility of random overlap, two criteria were used to determine overlapping regions: for CLIP regions less than or equal to 6 nt in length, overlap was considered valid if the overlap length exceeded 3 nt; for CLIP regions longer than 6 nt, an overlap threshold of at least 5 nt was set for validity.

[0133] The statistical significance of observed overlap was assessed by randomly sampling nucleotides. Specifically, 2964 nucleotides were randomly selected, equivalent to the number of binding sites identified by ΔSHAPE. For each RBP, ten sets of randomly selected locations were generated, matching the number of binding sites identified by CLIP, and statistical significance was tested using a one-sample t-test. For visualization, the midpoint of the CLIP-identified region was used to represent each region, clearly showing the spatial distribution of RBP binding regions in the figure.

[0134] To better demonstrate the superiority of the method of this invention, STONE was comprehensively evaluated among five models, using up to 42 features, such as the relative proportion of each base in the total sequencing depth. To achieve a balance between learning and generalization capabilities, while maintaining model simplicity and reducing parameter usage, 15 key features were ultimately selected, and the signal integration model was trained using a random forest model. Due to significant differences in chemical modification mechanisms, training was divided into two models: one for the SHAPE reagent and the other for the DMS reagent. Both models were trained using only 18S rRNA data.

[0135] The training dataset for the SHAPE model comes from the research of Sun et al., Spitale et al., Marinus et al., and Wang et al., while the DMS model uses datasets from Ding et al., Rouskin et al., Zubradt et al., and Talkish et al. During parameter tuning, 25% of the nucleotide sites in each dataset (including positive samples (unpaired sites) and negative samples (paired sites) are retained as a validation set.

[0136] In the DMS model, the main contributing features are the base frequencies of cytosine and adenine, while in the SHAPE model, multiple features contribute evenly. This is entirely consistent with the understanding of chemical modification mechanisms, namely that DMS primarily targets cytosine and adenine, while SHAPE can unbiasedly modify all four bases. Specifically, the contribution of different types of parameters to the final AUC value varies under different strategies, but all features are more effective at improving accuracy than using only traditional observation features, such as... Figure 2 As shown. Through these salient features, STONE can effectively integrate signal information, such as... Figure 3 As shown.

[0137] During testing, it was found that the accuracy improved in four different regions of the DMS dataset. The improvement was most significant in the second region, as this region is located in the core area where ribosomal proteins are densely distributed, making it difficult for small molecules to enter and modify, resulting in lower signal intensity and making it difficult to interpret using traditional calculation methods. Figure 4As shown, the RNA and RBP structures and interactions in region 2 of human 18S rRNA are illustrated (highlighted in red). The structural model of human 18S rRNA is from the Protein Data Bank. RNA is shown in cyan, and proteins in wheat brown. The STONE method significantly improves accuracy in low-accessibility regions.

[0138] The performance of STONE in identifying single-stranded sites of human 18S rRNA without providing accessibility information was assessed. Paired t-tests were used to calculate p-values. The number of sites used was as follows: 175 sites for high accessibility and 769 sites for low accessibility. Figure 5 As shown, in the low-accessibility single-chain region, the structure score of the STONE method was not significantly affected compared to the high-accessibility region (ΔScoreSTONE = 0.028, p = 0.317); while under the same conditions, the structure score of the traditional method decreased significantly (ΔScoreoriginal = 0.093, p = 0.002).

[0139] The STONE method in this embodiment can infer the RBP binding site by using the increased difference.

[0140] In 18S rRNA studies lacking information on small molecule accessibility, the STONE method showed a smaller decrease in AUC compared to the original method (ΔAUCSTONE = 0.049, ΔAUCoriginal = 0.082). This indicates that even in the absence of accessibility data, the STONE method can still identify regions affected by low accessibility by comparing its structure score with that calculated by the original method. Furthermore, compared to the original method, the STONE method consistently exhibited a significantly higher AUC when no accessibility data was available (paired t-test, p = 7.92e-11), demonstrating a lower dependence on accessibility information.

[0141] Next, we analyzed the main factors contributing to the improved accuracy of the STONE method in low-accessibility regions (protein-RNA spatial distance <3 Å). In low-accessibility single-stranded regions, the structure score of the STONE method was not significantly affected compared to high-accessibility regions (ΔScore). STONE = 0.028, p = 0.317); while under the same conditions, the structure score of the traditional method decreased significantly (ΔScore). original= 0.093, p = 0.002. This observation demonstrates the advantage of the STONE method in analyzing RNA regions affected by protein binding. Furthermore, this capability suggests that comparing the signal difference (ΔSHAPE score) between conventional and STONE methods in a single RNA structure detection experiment can help identify potential RBP binding sites.

[0142] To further validate the effectiveness of the STONE method, its performance on short RNAs (such as U1 snRNA and RNase p RNA) and long RNAs (such as XIST lncRNA) was evaluated and compared with antibody-dependent CLIP-seq and the nonspecific RNP-MaP method. RNP-MaP is a live-cell chemical probing strategy that maps RNA-protein interaction networks at nucleotide resolution using a heterogeneous bifunctional cross-linking agent. For U1 snRNA and RNase p RNA, the significantly differentially identified nucleotides (ΔSTONE) by the STONE method were highly consistent with CLIP-seq results, achieving concordance rates of 81.4% and 34.9% with RNP-MaP, respectively. Figure 6 As shown, the calculated sites are visualized using known RNA structures from the RNA Central database. ΔSHAPE results (colored nucleotides) are compared with CLIP-seq (represented by plotted RBPs) and RNP-MaP (represented by blue lines) results, with the percentage of matches for each comparison labeled. For long RNAXIST lncRNAs, in the analysis of 33 CLIP-seq datasets from the HEK293 cell line, the STONE method identified interaction regions for the vast majority of RBPs, with 31 RBPs showing an overlap exceeding 15%, such as... Figure 7 As shown, the calculated ΔSHAPERBP binding data at the regional level on human XIST long non-coding RNAs were compared with CLIP-seq data, comparing the matching level of each RBP with the matching level of randomly sampled regions. Although CLIP-seq experiments are inherently limited by the available antibody RBPs, which restricts their applicability, the STONE method still achieved a concordance rate of 51.5% compared to non-targeted RBP detection techniques (such as RNP-MaP). Figure 8 As shown, the RBP binding regions on human XIST long non-coding RNAs calculated using ΔSHAPE and RNP-MaP are compared. This result is consistent with the common overlap (21.5%–56.0%) between experimental techniques, different cell lines, or between experimental results and AI predictions.

[0143] Traditional RNA chemical modification signal analysis typically relies on dominant signals and a limited feature set. These methods usually calculate signal counts and site depths from alignment results and use simple formulas to calculate structure scores. In contrast, this embodiment proposes the STONE method, which significantly enhances RNA structure signal analysis by integrating mutation and stop signals and using 15 engineered features. The accuracy of STONE's DMS and SHAPE models on known RNAs was validated using a large dataset, demonstrating that the STONE method has higher accuracy than traditional methods.

[0144] A significant advantage of STONE is its ability to learn accessibility information, making it valuable in applications such as RNA-binding protein (RBP) binding site identification. In vivo RNA structure analysis is particularly challenging due to the complexity and dynamics of RNA folding and its extensive interactions with proteins. The effectiveness of chemical modifications varies significantly in protein regions tightly bound to RNA, making single-probe experiments ineffective in distinguishing these regions from fully double-stranded RNA regions.

[0145] In summary, the STONE method developed in this embodiment enhances RNA structure signal analysis based on integration mutations and stop signals, broadens the observation dimensions, and determines RNA structure more accurately than traditional methods. Furthermore, it uses the enhanced difference to infer RBP binding sites, enabling the identification of RBP binding sites based on RNA structure signals obtained in a single experiment.

[0146] Example 2:

[0147] Embodiment 2 of the present invention provides a prediction system for non-targeted detection of RNA-binding protein binding sites on RNA, comprising:

[0148] The data acquisition module is configured to acquire raw RNA structure data and preprocess the raw data.

[0149] The feature expansion module is configured to expand and filter the preprocessed data, and use the filtered data to train the RNA analysis model;

[0150] The RNA analysis module is configured to process the RNA to be analyzed using a trained RNA analysis model, calculate the probability that each nucleotide is single-stranded, and obtain a structure score.

[0151] The region identification module is configured to identify low-accessibility regions based on the difference between the structure score output by the RNA analysis model and the structure score calculated by the original method.

[0152] The binding site prediction module is configured to identify potential RNA-binding protein binding sites by comparing the signal differences with those of the original method in a single RNA structure detection experiment.

[0153] Example 3:

[0154] Embodiment 3 of the present invention provides a medium on which a program is stored. When the program is executed by a processor, it implements the steps in the method for predicting the binding site of RNA-binding proteins on RNA as described in Embodiment 1 of the present invention.

[0155] Example 4:

[0156] Embodiment 4 of the present invention provides a device including a memory, a processor, and a program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps in the method for predicting the binding site of RNA-binding proteins on RNA as described in Embodiment 1 of the present invention.

[0157] The steps and methods involved in Examples 2, 3 and 4 above correspond to those in Example 1. For specific implementation details, please refer to the relevant description section of Example 1.

[0158] Those skilled in the art will understand that the modules or steps of the present invention described above can be implemented using general-purpose computer devices. Optionally, they can be implemented using computer-executable program code, thereby allowing them to be stored in a storage device for execution by a computer device, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. The present invention is not limited to any particular combination of hardware and software.

[0159] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.

Claims

1. A method for predicting the binding site of RNA-binding proteins on RNA without targeting, characterized in that, include: Obtain raw RNA structure data and preprocess the raw data; The preprocessed data is expanded and filtered, and the filtered data is used to train the RNA analysis model. The trained RNA analysis model is used to process the RNA to be analyzed, calculate the probability that each nucleotide is single-stranded, and obtain the structure score. Based on the difference between the structure score output by the RNA analysis model and the structure score calculated by the original method, low-accessibility regions are identified. Potential RNA-binding protein binding sites were identified by comparing the signal differences with those of the original method in a single RNA structure detection experiment. The specific steps for training the RNA analysis model are as follows: The count and depth data of the processed signal are obtained by using sequence alignment and signal calling methods, and the ratio of stop signal to abrupt signal at each position is calculated. The preprocessed data is expanded and filtered, and the filtered data is divided into training set and test set; An RNA analysis model was constructed, trained using a training set, and tested using a test set. To mitigate the impact of outliers and ensure data comparability within a consistent range, the structure scores output by both the original method and the RNA analysis model were adjusted using quantile normalization when calculating the difference between the structure scores output by the RNA analysis model and the structure scores calculated by the original method. Furthermore, nonlinear transformations were applied to the normalized stop and mutation signals to adjust for weight differences within their respective numerical ranges.

2. The method for predicting the binding site of RNA-binding proteins by non-targeted detection as described in claim 1, characterized in that, The steps for preprocessing raw data include removing low-quality data and duplicate data.

3. The method for predicting the binding site of RNA-binding proteins by non-targeted detection as described in claim 1, characterized in that, The steps for expanding and filtering the preprocessed data include: Feature engineering is used to expand the features of the preprocessed data; Input features are selected based on the learning and generalization abilities of the RNA analysis model.

4. The method for predicting the binding site of RNA-binding proteins by non-targeted detection as described in claim 1, characterized in that, RNA analysis models include the SHAPE model and the DMS model. The training data for the SHAPE model comes from human 18S rRNA data and mouse data; the training data for the DMS model includes 18S rRNA data from Arabidopsis thaliana, human data, and yeast data.

5. The method for predicting the binding site of RNA-binding proteins on RNA without targeting, as described in claim 1, is characterized in that... The RNA analysis model was trained using a 10-fold cross-validation method, and the optimal model parameters were determined based on the highest AUC value obtained from the dataset.

6. A predictive system for non-targeted detection of RNA-binding protein binding sites on RNA, characterized in that, include: The data acquisition module is configured to acquire raw RNA structure data and preprocess the raw data. The feature expansion module is configured to expand and filter the preprocessed data, and use the filtered data to train the RNA analysis model; The RNA analysis module is configured to process the RNA to be analyzed using a trained RNA analysis model, calculate the probability that each nucleotide is single-stranded, and obtain a structure score. The region identification module is configured to identify low-accessibility regions based on the difference between the structure score output by the RNA analysis model and the structure score calculated by the original method. The binding site prediction module is configured to identify potential RNA-binding protein binding sites by comparing the signal differences with those of the original method in a single RNA structure detection experiment. The specific steps for training the RNA analysis model are as follows: The count and depth data of the processed signal are obtained by using sequence alignment and signal calling methods, and the ratio of stop signal to abrupt signal at each position is calculated. The preprocessed data is expanded and filtered, and the filtered data is divided into training set and test set; An RNA analysis model was constructed, trained using a training set, and tested using a test set. To mitigate the impact of outliers and ensure data comparability within a consistent range, the structure scores output by both the original method and the RNA analysis model were adjusted using quantile normalization when calculating the difference between the structure scores output by the RNA analysis model and the structure scores calculated by the original method. Furthermore, nonlinear transformations were applied to the normalized stop and mutation signals to adjust for weight differences within their respective numerical ranges.

7. A computer-readable storage medium, characterized in that, It stores multiple instructions, which are adapted to be loaded by the processor of a terminal device and executed by the method for predicting the binding site of non-targeted RNA-binding proteins on RNA as described in any one of claims 1-5.

8. A terminal device, characterized in that, The method includes a processor and a computer-readable storage medium, the processor being used to implement various instructions; the computer-readable storage medium being used to store multiple instructions adapted for loading by the processor and executing the method for predicting the binding site of non-targeted RNA-binding proteins on RNA as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Method and system for constructing models for predicting protein-RNA interaction binding sites

    CN111192631A

  • RNA and protein binding site recognition method based on deep learning

    CN113178229A