Prediction method and system for non-targeted detection of binding site of RNA binding protein on RNA
The STONE method integrates mutation and stop signals in RNA structure analysis to broaden the observation dimension, solving the problem of low accuracy of RNA binding protein binding site prediction in low-accessible regions, and achieving higher RNA structure analysis accuracy and RBP binding site recognition accuracy.
Patent Information
- Application Number
- CN202510140903.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-08
AI Technical Summary
The prior art is difficult to accurately predict the binding sites of RNA-binding proteins in low access areas, resulting in a decrease in the accuracy of RNA structure analysis.
The STONE method is adopted to integrate mutation signals and stop signals, broaden the observation dimension, improve the accuracy of RNA structure analysis in low-accessibility regions, and use the increased difference to infer RBP binding sites.
It improves the accuracy of RBP binding site recognition, enhances the analytical ability of low-accessibility regions, and overcomes the problem that traditional methods have weak signals and are difficult to analyze in these regions.
Smart Images

Figure CN120072034A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of bioinformatics technology, and in particular, to a method and system for predicting the binding sites of RNA-binding proteins on RNA by non-targeted detection. Background Art
[0002] The statements in this section merely provide background technical information related to the present invention and do not necessarily constitute prior art.
[0003] As a key RNA regulatory factor, RNA-binding protein (RBP) contains important clues about RBP binding sites in RNA structure information. By binding to specific RNA sequences or structures, it participates in various biological processes, including RNA splicing, translation, stability, transport, and so on. These processes play a crucial role in cell function and development, and abnormal RBP-RNA interactions are often associated with various diseases (such as cancer, neurodegenerative diseases, etc.). Therefore, accurately identifying and locating RBP binding sites not only helps to analyze the basic mechanisms of RNA biology but also provides potential opportunities for the study of disease mechanisms and the discovery of new therapeutic targets.
[0004] In recent years, with the in-depth study of the mechanism of action of RBP and the continuous progress of related technologies and methods, it has become possible to more accurately reveal these binding sites and provide more powerful tools for RNA biology research.
[0005] Traditionally, the detection methods of RBP binding sites can be roughly divided into two categories: experimental methods and computational prediction methods. Experimental methods are divided into two major categories: targeted and non-targeted. The methods for detecting RBP binding sites in experimental methods mainly rely on technologies such as cross-linking immunoprecipitation and sequencing (CLIP-seq). Although these experimental methods provide important information for revealing RBP binding sites, there are still certain limitations due to the relatively small amount of measurable data in the human body and the complexity of the experiments. Computational prediction methods, combined with known RBP binding site data and the characteristics of RNA structure, train models to identify potential binding sites, significantly improving the prediction accuracy. For example, PrismNet, a computational tool based on deep learning, can efficiently and accurately predict RBP binding sites. Smola et al. non-targetedly detected the RBP binding sites of mouse Xist lncRNA by comparing the differences in SHAPE scores under in vitro and in vivo conditions and verified that the results were highly consistent with CLIP-seq data. However, there is no method that completely relies on the RNA structure signals obtained in a single experiment to identify RBP binding sites.
[0006] RNA structure is the basis for understanding RNA function because it changes dynamically according to different cellular conditions. The relationship between RNA structure and RNA binding protein (RBP) is close and complex. RBP regulates the biological function of RNA by interacting with specific sequences or structural regions of RNA. At the same time, the structure of RNA also determines the binding characteristics and functions of RBP to a certain extent. Therefore, accurate RNA structural analysis is crucial to revealing these complexities.
[0007] The methods used to determine RNA structure are mainly divided into experimental methods and predictive methods. Compared with other methods, chemical detection methods have unique advantages. Weeks et al. proposed that chemical modification data can not only provide which regions of RNA are exposed, but also reflect how RNA folds into different structures under different conditions. By comparing the patterns of chemical modification under different experimental conditions, researchers can identify changes in the structural state of RNA molecules. Chemical modification reagents can react with specific parts of RNA, modify single-stranded RNA regions, and cause mutations in cDNA or early termination of synthesis during reverse transcription, and the extent of these reactions is closely related to the spatial structure of RNA. When the RNA chain folds into a secondary structure, some nucleotides are exposed, while other parts are "hidden" because they are wrapped in double strands or form complex spatial structures. This structural difference causes exposed and hidden nucleotides to show different characteristics in their reactivity with chemical modification reagents. Researchers can use these data to construct RNA folding models or further refine the secondary and three-dimensional structures of RNA. Based on this, a series of methods have been developed, among which DMS and SHAPE reagents are the most commonly used tools.
[0008] RNA-protein interactions limit the accessibility of chemical modification reagents. The binding sites of many RNA-binding proteins and RNA are usually located in regions with tighter or more complex structures. These regions may not be well characterized in standard RNA structure detection due to their low accessibility. Although shielding low-accessibility regions can improve the accuracy of RNA structure analysis, most RNAs lack comprehensive accessibility data. Although RNA structure detection methods (such as chemical modification methods, SHAPE, DMS, etc.) can provide detailed information about the folding state of RNA, there is no method to directly infer the binding sites of RNA-binding proteins (RBPs) from single RNA structure data.
[0009] Therefore, how to improve the accuracy of predicting the binding sites between RNA-binding proteins and RNA in low-accessibility regions has become a technical problem that needs to be urgently solved in existing technologies. Summary of the invention
[0010] Aiming at the deficiencies of the existing technology, the purpose of the present invention is to provide a method and system for predicting the binding sites of RNA-binding proteins on RNA non-targetedly. By using the STONE method, by integrating mutation signals and stop signals, the observation dimension is broadened, and the accuracy of RNA structure analysis in low-accessibility regions is improved. The improved difference is used to infer the RBP binding sites, thereby improving the accuracy of RBP binding site recognition.
[0011] To achieve the above object, the present invention is implemented through the following technical solutions:
[0012] The first aspect of the present invention provides a method for predicting the binding sites of RNA-binding proteins on RNA, including the following steps:
[0013] Obtain the original RNA structure data and preprocess the original data;
[0014] Expand and screen the preprocessed data, and use the screened data to train the RNA analysis model;
[0015] Use the trained RNA analysis model to process the RNA to be analyzed, calculate the probability of each nucleotide being single-stranded, and obtain the structure score;
[0016] Identify low-accessibility regions according to the difference between the structure score output by the RNA analysis model and the structure score calculated by the original method;
[0017] Identify potential RNA-binding protein binding sites by comparing the signal differences with the original method in a single RNA structure detection experiment.
[0018] Further, the step of preprocessing the original data includes removing low-quality data and duplicate data.
[0019] Further, the specific steps of training the RNA analysis model are:
[0020] Use sequence alignment and signal calling methods to obtain the count and depth data of the processed signals, and calculate the ratio of stop signals and mutation signals at each position;
[0021] Expand and screen the preprocessed data, and divide the screened data into a training set and a test set;
[0022] Construct an RNA analysis model, use the training set to train the RNA analysis model, and use the test set to test the trained RNA analysis model.
[0023] Even further, the steps of expanding and screening the preprocessed data include:
[0024] Feature engineering is used to perform feature expansion on the preprocessed data of the data pair.
[0025] According to the learning ability and generalization ability of the RNA analysis model, the input features are screened.
[0026] Furthermore, the RNA analysis model includes the SHAPE model and the DMS model. The training data of the SHAPE model comes from human 18S rRNA data and mouse data; the training data of the DMS model includes 18S rRNA data from Arabidopsis thaliana, human data, and yeast data.
[0027] Further, the 10-fold cross-validation method is used to train the RNA analysis model, and the optimal model parameters are determined according to the highest AUC value obtained in the dataset.
[0028] Further, when calculating the difference between the structure score output by the RNA analysis model and the structure score calculated by the original method, in order to reduce the influence of outliers and ensure the comparability of data within a consistent range, the structure scores output by both the original method and the RNA analysis model are adjusted by quantile normalization, and a non-linear transformation is applied to the normalized stop signal and mutation signal to adjust the weight difference within their respective numerical ranges.
[0029] The second aspect of the present invention provides a prediction system for non-targeted detection of RNA-binding protein binding sites on RNA, including:
[0030] A data acquisition module, configured to acquire the original RNA structure data and preprocess the original data;
[0031] A feature expansion module, configured to expand and screen the preprocessed data, and use the screened data to train the RNA analysis model;
[0032] An RNA analysis module, configured to use the trained RNA analysis model to process the RNA to be analyzed, calculate the probability of each nucleotide being single-stranded, and obtain a structure score;
[0033] A region recognition module, configured to identify low accessibility regions according to the difference between the structure score output by the RNA analysis model and the structure score calculated by the original method;
[0034] The binding point prediction module is configured to identify potential RNA-binding protein binding sites by comparing the signal differences with the original method in a single RNA structure detection experiment.
[0035] The third aspect of the present invention provides a medium, on which a program is stored, and when the program is executed by a processor, it implements the steps in the prediction method for non-targeted detection of RNA-binding protein binding sites on RNA as described in the first aspect of the present invention.
[0036] In the fourth aspect of the present invention, there is provided a device, including a memory, a processor, and a program stored on the memory and executable on the processor. When the processor executes the program, the steps in the method for predicting the binding sites of non-targeted detection of RNA-binding proteins on RNA as described in the first aspect of the present invention are implemented.
[0037] The above one or more technical solutions have the following beneficial effects:
[0038] The present invention discloses a method and a system for predicting the binding sites of non-targeted detection of RNA-binding proteins on RNA. The proposed STONE (Determine RNA Structure with Truncation-mutation Signals within ONE Experiment) method is applied to RNA structure analysis. Through the application of feature engineering and machine learning models, it can effectively integrate various signal features and improve the prediction accuracy of RNA binding sites.
[0039] The method of the present invention can enhance the analysis ability of low-accessibility regions. Traditional RNA structure detection methods often face problems such as weak signals and difficulty in parsing when dealing with low-accessibility regions. The STONE method can provide more accurate RNA structure analysis in these low-accessibility regions by broadening the observation dimension, overcoming the limitations of previous methods.
[0040] The present invention can achieve effective speculation of RBP binding sites. The STONE method uses the enhanced difference to inversely speculate RBP binding sites. Even in the absence of comprehensive accessibility data, it can still identify potential binding sites. This feature provides new ideas and tools for RNA biology research.
[0041] The STONE method of the present invention is not only applicable to short RNAs, but also can effectively identify RBP binding regions in long RNAs, demonstrating its wide applicability in the analysis of different types of RNAs. The STONE method can also be combined with existing chemical modification techniques (such as SHAPE and DMS) to further enhance the depth and breadth of RNA structure analysis, providing a more powerful tool for RNA biology research.
[0042] The advantages of the additional aspects of the present invention will be partly given in the following description, partly will become obvious from the following description, or will be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The specification drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention.
[0044] Figure 1 This is the flowchart of the STONE method in the first embodiment of the present invention;
[0045] Figure 2 This is a schematic diagram showing the contribution of four datasets' quantified features to the final STONE AUC value in the first embodiment of the present invention;
[0046] Figure 3 This is a schematic diagram of IGV visualization in the first embodiment of the present invention;
[0047] Figure 4 This is a schematic diagram of the RNA and RBP structures and their interactions in the human 18S rRNA region 2 in the visualization diagram of the first embodiment of the present invention;
[0048] Figure 5 This is a schematic diagram showing the performance of STONE in identifying the single-stranded positions of human 18S rRNA in the case where contact information is not provided in the first embodiment of the present invention;
[0049] Figure 6 This is a schematic diagram showing the calculation of ΔSHAPE scores using the structure scores derived by STONE to identify the RBP binding sites in human U1snRNA and human RNaseP in the first embodiment of the present invention;
[0050] Figure 7 This is a schematic diagram showing the comparison of the calculated ΔSHAPE RBP binding data with CLIP-seq data at the region level on the human XIST long non-coding RNA in the first embodiment of the present invention;
[0051] Figure 8 This is a schematic diagram showing the comparison of the RBP binding regions on the human XIST long non-coding RNA calculated using ΔSHAPE and RNP-MaP in the first embodiment of the present invention. Detailed implementation manners
[0052] It should be noted that the following detailed description is exemplary and is intended to provide further illustration of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0053] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments of the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof;
[0054] Example 1:
[0055] Currently, existing RNA structure analysis methods have the phenomenon of low accuracy in low-accessibility regions. In well-characterized structures such as 18S rRNA, the signal intensity in certain regions is significantly reduced, mainly because RNA-protein interactions limit the accessibility of chemical modification reagents, so these regions are called low-accessibility regions. Although shielding low-accessibility regions can improve the accuracy of RNA structure analysis, most RNAs lack comprehensive accessibility data.
[0056] In RNA analysis, the interaction between proteins and RNA has an important impact on the secondary structure and folding of RNA, especially in low-accessibility regions, where the structural protection may be further strengthened due to protein binding, resulting in difficulty for modification reagents to enter and bind to the target positions.
[0057] Proteins bind to specific domains of RNA to stabilize the folding of RNA, especially in low-accessibility regions. These regions are usually located in the core or more enclosed parts of RNA, and the binding of proteins can make these regions more stable, thus shielding the entry of modification reagents. This means that although modification reagents (such as DMS or SHAPE) are designed to label the exposed bases of RNA, in low-accessibility regions, due to the protective effect of proteins, the modification reagents often cannot effectively penetrate and react with RNA. Therefore, the chemical modification signals in these regions are weak, affecting the accuracy of subsequent analysis and sequencing.
[0058] Reverse transcriptase plays a crucial role in RNA structure analysis. It reads the sequence of RNA during sequencing by transcribing RNA into cDNA. However, the timing of the action of reverse transcriptase is usually after RNA has been chemically modified or heat-denatured. Chemical modification reagents react with the exposed bases of RNA to change the chemical properties of RNA and expose specific structural features, providing information for reverse transcription. In low-accessibility regions, due to their relatively high structural stability, they are usually difficult to be effectively acted upon by chemical modification reagents and reverse transcriptase. Modification reagents label RNA before reverse transcription to reveal the structural features of RNA, but the protective structures in low-accessibility regions often prevent the modification reagents from entering these regions, thus affecting the accuracy of these parts.
[0059] In view of the above defects of the prior art, Example 1 of the present invention provides a method for predicting the binding sites of non-targeted probing RNA-binding proteins on RNA. The STONE method can broaden the observation dimension and improve the accuracy of low-accessibility regions.
[0060] Among them, the STONE method refers to performing sequence alignment and signal calling on data, merging and processing the resulting signal data, and then integrating the mutations and stop signals in a single experiment through feature engineering and model training, thereby improving the accuracy of signal analysis.
[0061] As Figure 1 shown, STONE integrates two types of signals through feature engineering and machine learning models. The workflow includes separately analyzing the mutation signal and the stop signal, then expanding the features of these signals through feature engineering, and finally applying them to a machine learning model to calculate the probability that each nucleotide is single-stranded, namely the STONE score.
[0062] Specifically, it includes the following steps:
[0063] Step 1: Obtain the original RNA structure data and preprocess the original data.
[0064] In a specific implementation, the original data includes mutation signals and stop signals. The purpose of this embodiment in the model prediction process is to highlight the features of these two signals for RNA structure analysis.
[0065] Step 1.1: Download the original data from the RASP database or other existing databases. The final original data dataset contains 11 studies, involving five species, six small molecule reagents, and four reverse transcriptases. In the signal distribution calculation, all studies are included, except for single-cell datasets. When evaluating RNAs with known structures, the RNAs of humans, mice, viruses, and yeast were mainly analyzed. For genome-wide evaluation, a representative dataset with high depth was used. The single-cell dataset includes two batches of H9 cells, totaling 76 cells. For the application of identifying RNA-binding protein (RBP) binding sites, CLIP-seq data and RNP-MaP data from the HEK293T cell line were used. All data are existing, known, and legal data that can be obtained.
[0066] Step 1.2: The steps for preprocessing the original data include removing low-quality data and duplicate data.
[0067] After obtaining the original sequencing files in FASTQ format, the fastp software is used to remove low-quality reads with the default parameter settings of 43. fastp is used for read filtering, setting parameters to remove duplicates, apply a minimum Phred score of 20, exclude reads with >20% low-quality bases, and retain a minimum read length of 5 bases. This strict preprocessing step is crucial for minimizing noise and improving the reliability of subsequent analysis. Adapters are removed by automatic detection or specified according to the details in the original literature.
[0068] Step 2: Expand and filter the preprocessed data, and use the filtered data to train the RNA analysis model.
[0069] Step 2.1: Obtain the count and depth data of the processed signals using sequence alignment and signal calling methods, and calculate the ratio of stop signals and mutation signals at each position.
[0070] The ratio of stop signals and mutation signals at each position is used to evaluate the balance of the two signals generated by different chemical modification strategies. The STONE method performs better when both signals are balanced and have high signals.
[0071] In a specific implementation, use Bowtie2 to build a genomic reference index and align the reads to the reference genome. Use the --local option to optimize the alignment efficiency. The generated BAM file is then used for mutation signal calling through RNA Framework, configuring the parameters to exclude insertions while retaining the default settings for deletions. Implement the --orc option to include the original mutation counts summarizing mutation directions, insertions, and deletions. For stop signal calling, follow the default settings recommended by icSHAPE-pipe. Finally, calculate the ratio of stop signals and mutation signals at each position based on the count and depth of the processed signals generated from icSHAPE-pipe and RNA Framework.
[0072] Step 2.2: Expand and filter the preprocessed data.
[0073] Step 2.2.1: Use feature engineering to perform feature expansion on the preprocessed data.
[0074] Specifically, 42 features are expanded based on the merged and processed signal data, such as rf_mutation_Count (mutation count number), pipe_truncation_RT (pause number), etc.
[0075] Step 2.2.2: Select the input features according to the learning ability and generalization ability of the RNA analysis model.
[0076] After evaluating the learning ability and generalization ability of the model under various feature combinations, the number of input features is finally reduced to 15.
[0077] Specifically, evaluate the learning ability and generalization ability of the model through the AUC values calculated by the model under various feature combinations, select the feature combination that makes the AUC value of the model high, and finally select 15 features.
[0078] Step 2.3: Divide the filtered data into a training set and a test set.
[0079] Step 2.4: Construct an RNA analysis model, train the RNA analysis model using the training set, and test the trained RNA analysis model using the test set.
[0080] Step 2.4.1: Construct an RNA analysis model.
[0081] In this embodiment, the RNA analysis model includes a SHAPE model and a DMS model. The training data of the SHAPE model comes from human 18S rRNA data and mouse data; the training data of the DMS model includes 18S rRNA data from Arabidopsis thaliana, human data, and yeast data. Positive samples represent unpaired sites, while negative samples represent paired sites. The SHAPE model has 7,106 positive samples and 6,907 negative samples; the DMS model has 1,814 positive samples and 1,471 negative samples. The data is first merged and then divided into a training set and a test set, where 75% of the samples are used for training and the remaining 25% are used for testing.
[0082] Among them, both the SHAPE model and the DMS model apply the STONE (Determine RNA Structure with Truncation-mutation Signals within ONE Experiment) method. The difference is that different chemical modification reagents are used in the experiment. The SHAPE model is a model that determines the RNA structure using stop mutation signals for the SHAPE reagent in one experiment, and the DMS model is a model that determines the RNA structure using stop mutation signals for the DMS reagent in one experiment. In this embodiment, different models are selected according to different reagent types. The features of different data sets obtained by the two models are used to train a signal integration model through a random forest model. By inputting the obtained features into the random forest model to calculate the AUC value, the parameters are adjusted by comparing the magnitudes of the AUC values, and a higher feature combination is selected.
[0083] In addition to the random forest method, gradient boosting, extreme gradient boosting, and decision tree methods can also be used to train the two models, and hyperparameter tuning is performed to optimize the model performance.
[0084] Step 2.4.2: Train the RNA analysis model using the training set.
[0085] The training set data is input into the random forest model and hyperparameter tuning is performed to optimize the model performance. The training process specifically includes the following steps:
[0086] The random forest model consists of multiple decision trees. During the training process, each tree is trained by randomly sampling subsets of data and subsets of features, thereby enhancing the robustness of the model. The structure of each tree is based on the CART (Classification and Regression Tree) algorithm and adopts the form of a binary tree. Each node of the tree represents the split of a feature, while the leaf nodes represent the classification results or regression values. The depth of the tree and the number of leaf nodes are determined by the hyperparameters max_depth and min_samples_leaf, while the number of features used for splitting at each node is determined by max_features.
[0087] To optimize the performance of the model, this embodiment adopts the Grid Search method to optimize the hyperparameters. The optimized hyperparameters include the number of trees (n_estimators), the maximum depth of the tree (max_depth), the minimum number of samples for splitting (min_samples_split), the minimum number of samples for leaf nodes (min_samples_leaf), and the maximum number of features (max_features).
[0088] Grid Search traverses all possible combinations of hyperparameters and uses cross-validation to evaluate the performance of each combination, thereby selecting the best combination of hyperparameters.
[0089] During the optimization process, cross-validation helps reduce the errors that may be caused by uneven data splitting and selects the best model based on the average performance obtained through multiple evaluations. The principle of optimization is based on reducing overfitting and underfitting. By adjusting the hyperparameters, the complexity and expressive power of the model are optimized to achieve better performance on both the training set and unseen datasets.
[0090] After training, the model was evaluated on 25% of the test dataset and an independent validation dataset. The performance of the model was evaluated using multiple metrics, including AUC, F1-score, Precision, and Accuracy. The AUC value was calculated using the roc_curve and auc functions in the scikit-learn library. These evaluation metrics comprehensively reflect the performance of the model and ensure its generalization ability on different datasets.
[0091] Step 2.4.3: Train the RNA analysis model using the 10-fold cross-validation method and determine the optimal model parameters based on the highest AUC value obtained from the dataset.
[0092] In this example, the performance of the DMS and SHAPE models was verified using independent datasets by comparing the structure scores with known RNA structures in the RNA STRAND database. The DMS verification used datasets from DMS-seq, DMS-MaPseq, Mod-seq, and Structure-seq; the SHAPE verification used datasets from SHAPE-Seq, icSHAPE, SHAPE-MaP, and sc-SPORT. Sites with missing structure scores were excluded, and the remaining sites were evaluated to determine if they corresponded to unpaired nucleotides, and the AUC value was calculated.
[0093] To further evaluate the generalization ability of the model and potential overfitting problems, 10-fold cross-validation was adopted, dividing the dataset into ten non-overlapping subsets. In each training iteration, nine subsets were used for training, and the remaining one subset was used as the test set. The optimal model was selected based on the highest AUC value obtained on the validation set and the test set.
[0094] Step 3: Use the trained RNA analysis model to process the RNA to be analyzed, calculate the probability of each nucleotide being single-stranded, and obtain the structure score, i.e., the STONE score.
[0095] Among them, the STONE score is the probability of each nucleotide being single-stranded, and the STONE signal is the probability obtained by integrating the mutation signal and the stop signal. The mutation signal is the probability obtained by the original method RNAFramework, and the stop signal is the probability obtained by the method of icSHAPE-pipe.
[0096] The final structure scores (including the stop signal, mutation signal, and STONE signal) were visualized using IntegrativeGenomics Viewer and displayed together with the reference sequence to show the structure signals at single-nucleotide resolution. To facilitate visual evaluation of the scoring accuracy of different methods, a custom Python script was used to map these scores to the secondary structure of human 18S rRNA.
[0097] The accessibility information was from the 60S ribosome crystal structure (PDB ID: 5LKS), and the solvent-accessible surface area of each nucleotide was calculated using PyMOL.
[0098] Step 4: Identify low-accessibility regions based on the difference between the structure scores output by the RNA analysis model and the structure scores calculated by the original method.
[0099] Among them, the original methods used in this embodiment refer to icSHAPE-pipe and RNA Framework. icSHAPE-pipe is a method that only analyzes stop signals to obtain structure scores, while RNA Framework is a method that analyzes mutation signals to obtain structure scores. The details of the files that can be obtained by selecting these two methods in this embodiment are selected, and other existing methods are also applicable to the present invention.
[0100] The principle of the original method is to generate a reference index and then perform read mapping. Specifically, Cutadapt is relied on to clip adapters and low-quality bases for sequencing, and Bowtie 2 is relied on for sequence alignment. SAMtools is used to automatically sort the output alignment and convert it to the BAM format. Then, the RT stop count, mutation count, and coverage depth of each base are calculated, which is called RNA count (RC).
[0101] The RC file stores the transcript sequence, the RT stop / mutation count of each base, the read coverage of each base, and the total number of reads covering the transcript. The RC file for processing the RNA structure probing experiment is used for reactivity normalization, and the original signal S of each base i The calculation formula is:
[0102]
[0103] where nTi and cTi are the mutation count and read coverage at transcript position i, respectively. After the score calculation, the original reactivity is normalized: the original reactivity values higher than the 95th percentile are set to the 95th percentile, and the original reactivity values lower than the 5th percentile are set to the 5th percentile, and then each reactivity value is divided by the 95th percentile. Reactivity values ranging from 0 to 1 are generated, indicating that the residue has a low or high single-stranded tendency, respectively.
[0104] Step 4.1: In this embodiment, when calculating the difference between the structure score output by the RNA analysis model and the structure score calculated by the original method, in order to reduce the influence of outliers and ensure the comparability of data within a consistent range, the structure scores output by the original method and the RNA analysis model are both adjusted to the interval [0,1] through quantile normalization, and a non-linear transformation is applied to the normalized stop signal and mutation signal to adjust the weight difference within their respective numerical ranges.
[0105] Step 4.2: Calculate the ΔSHAPE value and compare the STONE output with the original signal. By comparing the signal differences (ΔSHAPE scores) between the traditional method and the STONE method in a single RNA structure probing experiment, potential RBP binding sites can be identified.
[0106] The filtering steps are as follows:
[0107] 1. ΔSHAPE calculation: First, calculate the difference in SHAPE reactivity (ΔSHAPE) under intracellular and extracellular conditions. This is obtained by subtracting the intracellular SHAPE reactivity from the extracellular reactivity and averaged over a three-nucleotide sliding window to reduce local signal fluctuations.
[0108] 2. Z-factor test: Use the Z-factor test to evaluate the significance of the ΔSHAPE values. This test compares the ΔSHAPE values with the associated extracellular and intracellular measurement errors, identifying nucleotide positions where the error magnitude is not significant for the ΔSHAPE values. The Z-factor test requires that the difference in SHAPE reactivity between extracellular and intracellular is more than 1.96 standard deviations (Z-factor > 0), ensuring that the 95% confidence intervals of the two measurements do not overlap.
[0109] 3. Standard Score calculation: For each nucleotide, calculate the standard score of the ΔSHAPE value relative to the ΔSHAPE values of all other nucleotides. This requires that the ΔSHAPE value of each nucleotide deviates from the average ΔSHAPE value by at least one standard deviation (absolute value ≥ 1), meaning that only the largest ΔSHAPE values are considered for further analysis.
[0110] Combining Z-factor and standard score: When finally determining RNA-protein interaction sites, both the Z-factor and the standard score are considered. If at least three nucleotides within a pentanucleotide window have a Z-factor > 0 and an absolute standard score ≥ 1, then these nucleotide positions are considered to have undergone significant changes in SHAPE reactivity due to the influence of the cellular environment.
[0111] To calculate the ΔSHAPE values, a sliding window method was used to identify the positions of different signals (2964 positions), and consecutive positions were grouped into regions (628 regions).
[0112] Specifically, within each window (size = 5 nt), the positions identified by ΔSHAPE are considered valid only when at least three positions are identified. The calculation steps of ΔSHAPE are shown in formula (1):
[0113]
[0114] In formula (1), depth represents the read depth of a specific nucleotide, and structural scores represent STONE scores, mutation scores, or stop scores. The standard error (SEscore) of each score is calculated.
[0115] The definition of SEscore is the standard error of SHAPE reactivity. The standard error (SEscore) is defined as the measurement error of reactivity estimated according to the Poisson distribution. Specifically, the standard error of SHAPE reactivity is calculated based on the mutation rate (mutation count) of each nucleotide and the read depth at that position. In the formula:
[0116]
[0117] Among them, λ represents the number of observed mutation events, "reads" represents the sequencing depth of this nucleotide, and "rate" is the number of mutation events per read. The variance of the Poisson distribution is equal to the number of events, so the standard error reflects the precision of the mutation events per read.
[0118]
[0119] The ΔSHAPE value for each nucleotide is calculated by formula (2), where S and O represent the structure scores obtained from STONE and the original method respectively. The ΔSHAPE value captures the difference between the scores and is averaged within a sliding window of three nucleotides.
[0120] To incorporate smoothing, the standard error values calculated using formula (1) are averaged to ensure consistency with the method established by Smola et al. The smoothed error estimate value at each nucleotide position is denoted as formula (3):
[0121]
[0122] In formula (3), the Z factor (Z) is calculated for each nucleotide (Z i ), where the subscripts S and O represent STONE and the original method respectively. When the Z value is greater than 0, it is considered that a significant change has occurred in the SHAPE reactivity of this nucleotide.
[0123] The standard score (Si) of each nucleotide is strictly calculated, and the identification of potential binding sites is strictly carried out according to the method of Smola et al.
[0124] Specifically, Smola et al. used two methods, SHAPE-MaP and icSHAPE, to study the sites of RNA-protein interactions, especially by comparing the changes in SHAPE reactivity (ΔSHAPE) of RNA in different environments.
[0125] The RNA was modified with 1M7 reagent to measure the SHAPE reactivity of RNA in living cells and in vitro conditions. These reactivity data can help determine the interaction sites of RNA-protein complexes.
[0126] By calculating ΔSHAPE (i.e., the difference in SHAPE reactivity between in vitro and intracellular), and the Z-factor for each nucleotide, nucleotides with significant reactivity changes can be identified.
[0127] The difference from their method is that their ΔSHAPE is the difference between experiments, while the ΔSHAPE score in the method of this example is the difference between the value obtained when calculating the AUC using the single-cell dataset generated by the SHAPE-MaP method and the AUC value calculated by applying the STONE model to the features screened in this example. It is the subtraction between different calculation methods and can help identify potential RBP binding sites.
[0128] Step 5: Identify potential RNA-binding protein binding sites by comparing the signal differences with the original method in a single RNA structure probing experiment.
[0129] Step 6: Evaluate the ΔSHAPE results.
[0130] The CLIP-seq data is from the POSTAR3CLIPdb database and obtained in BED format. Six experimental methods were used: PAR-CLIP, PIP-seq, eCLIP, HITS-CLIP, iCLIP, and 4SU-iCLIP. To extract data for specific genes, the CLIP-seq RBP binding sites were annotated to the corresponding transcriptome using the GAP tool.
[0131] The extraction of RNP-MaP sites is based on the method of Weidmann et al. The processed data files contain the mutation frequencies and read depths of cross-linked and non-cross-linked samples. All calculations were completed using custom R scripts, and candidate RNP-MaP sites on the human XIST long non-coding RNA (lncRNA) were generated.
[0132] To verify the accuracy of the ΔSHAPE calculation, the results of ΔSHAPE were compared with the CLIP-seq and RNP-MaP results. In this analysis, the consecutive nucleotide positions in the ΔSHAPE, CLIP-seq, and RNP-MaP data were grouped to independently represent their RBP binding regions. To minimize the possibility of random overlap, two criteria were used to determine the overlapping regions: for CLIP regions with a length less than or equal to 6 nt, if the overlapping length exceeds 3 nt, the overlap is considered valid; for CLIP regions with a length greater than 6 nt, the overlap threshold is set to at least 5 nt to be considered valid.
[0133] The statistical significance of the observed overlaps was evaluated by randomly sampling nucleotides. Specifically, 2964 nucleotides were randomly selected, equivalent to the number of binding sites identified by ΔSHAPE. For each RBP, ten sets of randomly drawn positions were generated that matched the number of binding sites identified by CLIP, and the statistical significance was tested by a one-sample t-test. For visualization, the midpoints of the CLIP-identified regions were used to represent each region, clearly showing the spatial distribution of the RBP binding regions in the figure.
[0134] To better demonstrate the superiority of the method of the present invention, STONE was comprehensively evaluated in five models, using up to 42 features, such as the relative proportion of each base in the total sequencing depth. To balance learning ability and generalization ability, while keeping the model simple and reducing parameter usage, 15 key features were finally selected to train the signal integration model through a random forest model. Due to significant differences in the chemical modification mechanisms, the training was divided into two models: one for SHAPE reagents and the other for DMS reagents. Both models were trained using only 18S rRNA data.
[0135] The training dataset for the SHAPE model was sourced from the studies of Sun et al., Spitale et al., Marinus et al., and Wang et al., while the DMS model used the datasets of Ding et al., Rouskin et al., Zubradt et al., and Talkish et al. During the parameter adjustment process, 25% of the nucleotide sites in each dataset (including positive samples, i.e., unpaired sites, and negative samples, i.e., paired sites) were reserved as the validation set.
[0136] In the DMS model, the main contributing features are the base frequencies of cytosine and adenine, while in the SHAPE model, multiple features contribute evenly. This is completely consistent with the understanding of the chemical modification mechanism, that is, DMS mainly targets cytosine and adenine, while SHAPE can modify all four bases without bias. Specifically, the contributions of various types of parameters to the final AUC value under different strategies are different, but all features can improve the accuracy more effectively than only using traditional observation features, as Figure 2 shown. Through these significant features, STONE can effectively integrate signal information, as Figure 3 shown.
[0137] During the testing process, it was found that in the DMS dataset, the accuracy was improved in four different regions. Among them, the improvement in the second region was the most significant because this region is located in the core area with a dense distribution of ribosomal proteins, where small molecules are difficult to enter and modify, resulting in low signal intensity and difficult to analyze by traditional computational methods, as Figure 4Shown are the RNA and RBP structures and interactions (marked in red) in the human 18S rRNA region 2. The structural model of human 18S rRNA is from the Protein Data Bank. RNA is shown in cyan and proteins are shown in wheat. The STONE method significantly improves the accuracy in low-accessibility regions.
[0138] Performance of STONE in identifying single-stranded positions of human 18S rRNA without providing accessibility information. Paired t-tests were used to calculate p-values. The number of positions used was as follows: 175 positions for high accessibility and 769 positions for low accessibility, as Figure 5 shown. In the single-stranded regions of low accessibility, the structural scores of the STONE method were not significantly affected compared to the high-accessibility regions (ΔScoreSTONE = 0.028, p = 0.317); while under the same conditions, the structural scores of the traditional method decreased significantly (ΔScoreoriginal = 0.093, p = 0.002).
[0139] The STONE method of this example can use the improved difference to infer RBP binding sites in turn.
[0140] In the study of 18S rRNA without including small molecule accessibility information, the decrease in AUC of the STONE method was less than that of the original method (ΔAUCSTONE = 0.049, ΔAUCoriginal = 0.082). This indicates that even in the absence of accessibility data, the STONE method can still identify regions affected by low accessibility by comparing its structural scores with those calculated by the original method. In addition, compared with the original method, the AUC of the STONE method is always significantly higher in the absence of accessibility data (paired t-test, p = 7.92e-11), indicating its lower dependence on accessibility information.
[0141] Next, the main factors contributing to the improved accuracy of the STONE method in low-accessibility regions (protein and ) were analyzed. In the single-stranded regions of low accessibility, the structural scores of the STONE method were not significantly affected compared to the high-accessibility regions (ΔScore STONE = 0.028, p = 0.317); while under the same conditions, the structural scores of the traditional method decreased significantly (ΔScore original = 0.093, p = 0.002). This observation indicates that the STONE method has an advantage in analyzing RNA regions affected by protein binding. In addition, this ability indicates that by comparing the signal differences (ΔSHAPE scores) between the traditional method and the STONE method in a single RNA structure probing experiment, potential RBP binding sites can be identified.
[0142] To further validate the effectiveness of the STONE method, its performance in short RNAs (such as U1 snRNA and RNase P RNA) and long RNAs (such as XIST lncRNA) was evaluated and compared with the antibody-dependent CLIP-seq technique and the non-specific RNP-MaP method. RNP-MaP is a live-cell chemical probing strategy that maps RNA-protein interaction networks at nucleotide resolution using a heterobifunctional crosslinker. For U1 snRNA and RNase P RNA, the significantly different nucleotides (ΔSTONE) identified by the STONE method were highly consistent with the CLIP-seq results, achieving agreement rates of 81.4% and 34.9% with RNP-MaP, respectively, as Figure 6 shown, the calculated sites were visualized through known RNA structures in the RNA Central database. The ΔSHAPE results (colored nucleotides) were compared with the CLIP-seq (represented by drawing RBPs) and RNP-MaP (represented by blue lines) results, and the matching percentages for each comparison were annotated. For the long RNA XIST lncRNA, in the analysis of 33 CLIP-seq datasets in the HEK293 cell line, the STONE method identified interaction regions for the vast majority of RBPs, with an overlap rate of over 15% for 31 RBPs, as Figure 7 shown, the calculated ΔSHAPE RBP-binding data was aligned with the CLIP-seq data at the region level on the human XIST long non-coding RNA, comparing the matching levels of each RBP with those of randomly sampled regions. Although CLIP-seq experiments are inherently limited by the RBPs for which antibodies are available, which restricts their scope of application, the STONE method still achieved an agreement rate of 51.5% compared with non-targeted RBP detection techniques (such as RNP-MaP), as Figure 8 shown, comparing the RBP-binding regions on the human XIST long non-coding RNA calculated using ΔSHAPE and RNP-MaP. This result is consistent with the common overlap rates (21.5% - 56.0%) between experimental techniques, between different cell lines, or between experimental results and AI predictions.
[0143] In traditional RNA chemical modification signal analysis, methods typically rely on primary signals and limited feature sets. These methods usually calculate signal counts and site depths obtained from alignment results and use simple formulas to calculate structure scores. In contrast, this embodiment proposes the STONE method, which significantly enhances RNA structure signal analysis using 15 engineered features by integrating mutation and stop signals. By leveraging a large dataset, the accuracy of the STONE DMS and SHAPE models on known RNAs was verified, and the results showed that the STONE method has higher accuracy compared with traditional methods.
[0144] A significant advantage of STONE is its ability to learn accessibility information, which makes it valuable in applications such as the identification of RNA-binding protein (RBP) binding sites. Analyzing RNA structures in vivo is particularly challenging due to the complexity and dynamics of RNA folding and its extensive interactions with proteins. The effectiveness of chemical modifications varies significantly in regions of proteins that bind tightly to RNA, making it impossible for a single probe experiment to effectively distinguish these regions from fully double-stranded RNA regions.
[0145] In summary, the STONE method developed in this embodiment is based on integrating mutations and stop signals to enhance RNA structure signal analysis, broadening the observation dimension to determine RNA structures more accurately than traditional methods, and using the enhanced difference to inversely infer RBP binding sites, enabling the identification of RBP binding sites relying on the RNA structure signals obtained in a single experiment.
[0146] Example Two:
[0147] Embodiment Two of the present invention provides a prediction system for non-targeted detection of RNA-binding protein binding sites on RNA, including:
[0148] A data acquisition module, configured to acquire raw RNA structure data and preprocess the raw data;
[0149] A feature expansion module, configured to expand and screen the preprocessed data, and use the screened data to train an RNA analysis model;
[0150] An RNA analysis module, configured to process the RNA to be analyzed using the trained RNA analysis model, calculate the probability of each nucleotide being single-stranded, and obtain a structure score;
[0151] A region identification module, configured to identify low-accessibility regions based on the difference between the structure score output by the RNA analysis model and the structure score calculated by the original method;
[0152] A binding point prediction module is configured to identify potential RNA-binding protein binding sites by comparing the signal differences with the original method in a single RNA structure detection experiment.
[0153] Example Three:
[0154] Embodiment Three of the present invention provides a medium on which a program is stored, and when the program is executed by a processor, it implements the steps in the prediction method for non-targeted detection of RNA-binding protein binding sites on RNA as described in Embodiment One of the present invention.
[0155] Example Four:
[0156] Embodiment 4 of the present invention provides a device, including a memory, a processor, and a program stored on the memory and executable on the processor. When the processor executes the program, the steps in the method for predicting the binding sites of non-targeted detection of RNA-binding proteins on RNA as described in Embodiment 1 of the present invention are implemented.
[0157] The steps involved in Embodiments 2, 3, and 4 above correspond to those in Method Embodiment 1. For specific implementation manners, reference may be made to the relevant description part of Embodiment 1.
[0158] Those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computer device. Optionally, they can be implemented by program codes executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.
[0159] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present invention. Those skilled in the art should understand that based on the technical solutions of the present invention, various modifications or deformations that can be made without creative efforts by those skilled in the art are still within the protection scope of the present invention.
Claims
1. A method for predicting the binding sites of RNA binding proteins on RNA by non-targeted detection, characterized in that: include: Obtaining raw RNA structure data and preprocessing the raw data; Expand and filter the preprocessed data, and use the filtered data to train the RNA analysis model; The trained RNA analysis model is used to process the RNA to be analyzed, and the probability of each nucleotide being single-stranded is calculated to obtain the structure score; Identify low-accessibility regions based on the difference between the structure scores output by the RNA analysis model and those calculated by the original method; Potential RNA-binding protein binding sites were identified by comparing the signal differences with the original method in a single RNA structure probing experiment.
2. The method for predicting the binding site of an RNA-binding protein on RNA by non-targeted detection according to claim 1, wherein: The steps of preprocessing the raw data include removing low-quality data and duplicate data.
3. The method for predicting the binding site of an RNA-binding protein on RNA by non-targeted detection according to claim 1, characterized in that: The specific steps for training the RNA analysis model are: Sequence alignment and signal calling methods were used to obtain count and depth data of processed signals, and the ratio of stop signals and mutation signals at each position was calculated; Expand and filter the preprocessed data, and divide the filtered data into training sets and test sets; Construct an RNA analysis model, train the RNA analysis model using the training set, and test the trained RNA analysis model using the test set.
4. The method for predicting the binding site of an RNA-binding protein on RNA by non-targeted detection according to claim 3, wherein: The steps to expand and filter the preprocessed data include: Use feature engineering to expand the features of preprocessed data; Input features are selected based on the learning and generalization capabilities of the RNA analysis model.
5. The method for predicting the binding site of an RNA-binding protein on RNA by non-targeted detection according to claim 1, characterized in that: RNA analysis models include the SHAPE model and the DMS model. The training data of the SHAPE model comes from human 18SrRNA data and mouse data; the training data of the DMS model includes 18SrRNA data from Arabidopsis, human data and yeast data.
6. The method for predicting the binding site of an RNA-binding protein on RNA by non-targeted detection according to claim 1, characterized in that: The RNA analysis model was trained using a 10-fold cross-validation method, and the optimal model parameters were determined based on the highest AUC value obtained in the dataset.
7. The method for predicting the binding site of an RNA-binding protein on RNA by non-targeted detection according to claim 1, characterized in that: In order to mitigate the impact of outliers and ensure that the data are comparable within a consistent range, the structural scores output by the original method and the RNA analysis model were adjusted by quantile normalization based on the difference between the structural scores output by the RNA analysis model and the structural scores calculated by the original method. Nonlinear transformation was applied to the normalized stop signal and mutation signal to adjust the weight differences within their respective numerical ranges.
8. A prediction system for non-targeted detection of RNA binding protein binding sites on RNA, characterized in that: include: A data acquisition module is configured to acquire raw RNA structure data and pre-process the raw data; A feature expansion module is configured to expand and filter the preprocessed data and use the filtered data to train the RNA analysis model; An RNA analysis module is configured to process the RNA to be analyzed using the trained RNA analysis model, calculate the probability of each nucleotide being a single strand, and obtain a structure score; a region identification module configured to identify low accessibility regions based on the difference between the structure scores output by the RNA analysis model and the structure scores calculated by the original method; The binding site prediction module is configured to identify potential RNA binding protein binding sites by comparing the signal differences with the original method in a single RNA structure probing experiment.
9. A computer-readable storage medium, characterized in that: A plurality of instructions are stored therein, and the instructions are suitable for being loaded by a processor of a terminal device and executed by the method for predicting the binding site of an RNA-binding protein on RNA by non-targeted detection according to any one of claims 1 to 7.
10. A terminal device, characterized in that: The method comprises a processor and a computer-readable storage medium, wherein the processor is used to implement each instruction; and the computer-readable storage medium is used to store multiple instructions, wherein the instructions are suitable for being loaded by the processor and executing the method for predicting the binding site of a non-targeted detection RNA-binding protein on RNA according to any one of claims 1 to 7.
Citation Information
Patent Citations
Prediction method and system for RNA binding sites in protein molecules
CN106446602A
Method and system for constructing models for predicting protein-RNA interaction binding sites
CN111192631A
RNA and protein binding site recognition method based on deep learning
CN113178229A
Method and system for predicting protein-polypeptide binding site
CN113593631A
Method and Media of Predicting protein-binding regions in RNA Using Nucleotide Profiles and Compositions
KR1020180017827A