Method and system for predicting RNA-protein interaction based on computer

CN120833849APending Publication Date: 2025-10-24HAINAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510988343.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

Existing technologies for predicting RNA-protein interactions suffer from problems such as high cost, long processing time, strong sample dependence, low throughput, limited coverage, and poor cross-species adaptability. Furthermore, traditional methods are difficult to verify the biological validity of the prediction results.

Method used

We employ a multimodal feature fusion approach, combining sequence and structural features of RNA and proteins. We extract features using pre-trained large language models RNAErnie and ESM2, and train the model using a graph neural network architecture including graph attention network (GAT) and gated graph convolutional network (GGCN) to improve prediction accuracy and stability.

Benefits of technology

It achieves rapid, low-cost, and high-accuracy prediction of RNA-protein interactions, breaking through the limitations of traditional manual feature encoding, improving the model's ability to capture complex interactions, and enhancing cross-species adaptability and predictive performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120833849A_ABST
    Figure CN120833849A_ABST
Patent Text Reader

Abstract

The invention discloses a method and a system for predicting RNA-protein interaction based on a computer. The method comprises the following steps: acquiring RPI data and preprocessing the RPI data to obtain a data set; the RNAErnie model extracts RNA sequence features and fuses the RNA sequence features with the RNA secondary structure features to obtain RNA features; the ESM2 model extracts protein sequence features and fuses the protein sequence features with protein secondary structure features to obtain protein features; fusing the RNA features with the protein features to obtain node features; constructing a binary adjacent matrix; extracting a k-hop closed subgraph taking the RNA node and the protein node as the center, and performing node marking; converting into a line graph structure, and obtaining graph topological characteristics; the node features and the graph topology features are input into a graph neural network architecture for joint processing, and a prediction result of RNA-protein interaction is output; and evaluating the model. According to the invention, more accurate and reliable RPI prediction can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of bioinformatics, and in particular to a method and system for predicting RNA-protein interactions based on a computer. BACKGROUND

[0002] RNA and protein have extensive interactions in most life activities. From the perspective of biological function, the combination of RNA and protein plays a core role in regulating gene expression, participating in signal transduction, and affecting key life processes such as cell metabolism. These interactions not only maintain the normal function of cells, but also participate in the regulation of various physiological and pathological processes. From the perspective of disease research, abnormal RNA-protein interactions are closely related to various diseases. For example, in cancer, abnormal combination of RNA and protein can lead to dysregulation of gene expression, thereby promoting the occurrence and development of tumors; in neurodegenerative diseases, the dysfunction of RNA-binding proteins is closely related to the progression of the disease. Therefore, in-depth study of the interaction between RNA and protein not only helps to reveal the pathogenesis of diseases, but also provides important theoretical basis and potential targets for the development of new treatment strategies. Currently, RNA-protein interaction prediction has become one of the key directions in functional genomics research, with important scientific research value and broad clinical application prospects.

[0003] Existing prediction methods: 1) Traditional biological experimental methods for RNA-protein interaction (RPI) prediction mainly include cross-linking immunoprecipitation (CLIP), RNA immunoprecipitation (RIP), and mass spectrometry (MS) techniques. These methods have high accuracy, can directly capture real RNA-protein binding events, provide structural and functional biological information, and support in-depth research and verification of interaction mechanisms. Therefore, they are still widely regarded as key technical means for RNA-protein interaction research in current studies. However, traditional biological experiments have the following problems: most wet experiments are not only expensive and time-consuming, but also have high detection costs; they are highly dependent on samples, have low throughput and limited coverage; they are limited by experimental conditions and technical maturity; and they lack scalability and flexibility; 2) Computer-based prediction methods mostly use machine learning and deep learning algorithms to extract sequence features, physical and chemical properties, evolutionary information, or structural features of RNA and protein to build prediction models. In recent years, deep learning technologies such as convolutional neural networks (CNN), recurrent neural networks (RNN), and graph neural networks (GNN) have been introduced into this field, further improving the model's ability to model complex sequence relationships. However, these methods are heavily dependent on traditional feature encoding techniques, have poor cross-species or cross-condition adaptability, and are difficult to verify the biological effectiveness of the prediction results. SUMMARY

[0004] In order to solve the above technical problems, the present application provides a method and system for predicting RNA-protein interaction based on computer. In the method and system, (1) a fast, low-cost and high-accuracy computer prediction method is provided; (2) the problem of limited expression ability of traditional manual feature coding is broken through, and the advantages of pre-training large language model are effectively utilized; (3) a method based on multi-modal feature fusion is proposed, which combines the sequence and structure feature information of RNA and protein; (4) the extraction and modeling ability of key features in the RNA-protein interaction relationship is enhanced by introducing a graph neural network architecture, further improving the accuracy and stability of RPI prediction.

[0005] In order to achieve the above purpose, the technical scheme of the present application is as follows:

[0006] The method for predicting RNA-protein interaction based on computer comprises the following steps:

[0007] Obtain RPI data and preprocess to obtain a data set;

[0008] Construct an RPI-PLMGNN model, and the training process of the RPI-PLMGNN model is as follows: 1) input the RNA sequence in the data set into the RNAErnie model to extract the RNA sequence feature and fuse it with the RNA secondary structure feature to obtain the RNA feature; input the protein sequence in the data set into the ESM2 model to extract the protein sequence feature and fuse it with the protein secondary structure feature to obtain the protein feature; fuse the RNA feature and the protein feature to obtain the node feature; 2) construct a binary adjacency matrix based on the RNA-protein interaction relationship pairs verified in the data set; extract the k-hop closed subgraph centered on the RNA node and the protein node based on the binary adjacency matrix and mark the nodes; convert the closed subgraph into a line graph structure to obtain the graph topology feature; 3) input the node feature and the graph topology feature into a graph neural network architecture for joint processing, and output the prediction result of the RNA-protein interaction, wherein the graph neural network architecture comprises a graph attention network module GAT and a gated graph convolution network module GGCN;

[0009] Model evaluation.

[0010] Preferably, the obtaining of RPI data and the preprocessing to obtain a data set comprises the following steps:

[0011] Collect experimentally verified RNA-protein interaction relationship pairs as positive samples from RPI7317, RPI2241, RPI369 and NPInter v2.0 benchmark data sets; remove RNA sequences and their corresponding RPI data that do not meet the preset requirements in terms of length from the positive samples;

[0012] Randomly select RNA and protein without experimental verification as negative samples, the number of which is the same as that of positive samples.

[0013] Preferably, the model evaluation comprises the following steps:

[0014] The RPI-PLMGNN model is trained using four benchmark datasets respectively;

[0015] Determine evaluation indicators, perform K-fold cross-validation on each model, and retain the best model weight. The performance of the RPI-PLMGNN model is verified using several independent test sets across species.

[0016] Preferably, the evaluation indicators include sensitivity, specificity, precision, accuracy, Matthew correlation coefficient, and area under the curve.

[0017] Preferably, it further comprises the following steps:

[0018] The RNAErnie model adopts motif-level masking, subsequence-level masking, and motif-level random masking strategies, and combines RNA types as word labels;

[0019] The RNAErnie model appends word labels to the end of the RNA sequence to enhance the context representation of the sequence.

[0020] Preferably, the ESM2 model comprises 12 layers of stacked Transformer layers, each Transformer layer consisting of multi-head attention mechanism and feedforward neural network, combined with layer normalization and residual connection.

[0021] Preferably, the RNA secondary structure feature acquisition method comprises the following steps:

[0022] Based on the principle of minimum free energy, the free energy of RNA in different folding states is calculated, the prediction results are obtained, and the base pairing in the prediction results is represented by different specific characters respectively;

[0023] Based on the combination frequency of K adjacent specific characters in the RNA sequence fragment, the RNA secondary structure feature is determined.

[0024] Preferably, the protein secondary structure feature acquisition method comprises the following steps:

[0025] The SOPMA program is used for protein secondary structure prediction, and the predicted secondary structure types α-helix, β-sheet, β-turn, and random coil are respectively represented by different specific characters;

[0026] Based on the occurrence frequency of K adjacent specific characters in the protein sequence, the protein secondary structure feature is determined.

[0027] Preferably, the graph attention network module is used to aggregate the neighbor information of nodes in the graph through a multi-head attention mechanism to highlight local important features, update the feature representation of the node and pass it to the gated graph convolutional network module; the gated graph convolutional network module includes a graph convolutional network, a gating mechanism and an attention mechanism, and the graph convolutional network is used to aggregate the feature information of its neighboring nodes for each node to generate an aggregation vector, thereby capturing global structural information; the gating mechanism will filter the output of the graph convolutional network, dynamically control the flow of information, retain important features and suppress irrelevant or noise information; the attention mechanism is used to adaptively adjust the importance of neighboring nodes, thereby enhancing the effectiveness of information dissemination.

[0028] Based on the above, the present invention further discloses a system for predicting RNA-protein interactions based on computers, comprising:

[0029] The acquisition module is used to obtain RPI data and preprocess it to obtain a data set;

[0030] A training module is used to construct an RPI-PLMGNN model. The RPI-PLMGNN model training process includes the following steps: 1) inputting RNA sequences in a data set into an RNAErnie model to extract RNA sequence features and fusing them with RNA secondary structure features to obtain RNA features; inputting protein sequences in a data set into an ESM2 model to extract protein sequence features and fusing them with protein secondary structure features to obtain protein features; fusing the RNA features with the protein features to obtain node features; 2) constructing a binary adjacency matrix based on RNA-protein interaction relationship pairs experimentally verified in the data set; extracting a k-hop closed subgraph centered on RNA nodes and protein nodes based on the binary adjacency matrix and marking the nodes; converting the closed subgraph into a line graph structure to obtain graph topology features; 3) inputting the node features and graph topology features into a graph neural network architecture for joint processing to output the prediction results of RNA-protein interaction, wherein the graph neural network architecture includes a graph attention network module GAT and a gated graph convolutional network module GGCN;

[0031] Evaluation module, used for model evaluation.

[0032] Based on the above technical scheme, the beneficial effects of the present application are: the present application innovatively combines pre-trained large language models and graph neural network architecture, fully utilizes the sequence features and structural features of RNA and proteins, and improves the prediction ability of RPI. In the aspect of node feature extraction, the present application first applies two advanced pre-trained large language models, RNAErnie model and ESM2 model, to sequence feature extraction of RNA and proteins respectively, and fuses them with their respective secondary structure features to generate more rich node features. This innovation can effectively capture the complex interaction between RNA and proteins. In the aspect of graph topology modeling, the present application adopts linear graph topology to depict the RPI network, so as to better capture the deep interaction relationship and ensure the accuracy and comprehensiveness of the topology information. Finally, the multi-modal node features and graph topology features are jointly processed by the graph neural network architecture, which integrates the graph attention network (GAT) module and the gated graph convolution network (GGCN) module to generate the final interaction prediction. GAT can effectively identify key interaction sites by adaptively assigning weights to different fragments, while GGCN can enhance the modeling ability of the model for long-distance dependencies and suppress noise through the gating mechanism to ensure the robustness of the model in complex data. The synergistic effect of the two makes the present application be able to balance local and global features, thereby effectively improving the prediction performance of RPI. In summary, the present application innovatively combines RNAErnie and ESM2 pre-trained large language models, adopts linear graph topology structure, and uses a hybrid graph neural network architecture to propose a more accurate and reliable RPI prediction method. The present application not only provides a new technical means for studying the interaction mechanism of RNA and proteins, but also has important scientific value and clinical application prospect. BRIEF DESCRIPTION OF DRAWINGS

[0033] Figure 1 FIG. 1 is a flow diagram of a method for predicting RNA-protein interactions based on a computer in an embodiment;

[0034] Figure 2 FIG. 2 is a model framework diagram of RPI-PLMGNN in an embodiment;

[0035] Figure 3 FIG. 3 is a schematic diagram of the overall performance comparison of the RPI-PLMGNN model and other five existing models on four benchmark datasets in an embodiment. The performance of NPIGNN, RPITER, IPMiner, GATLGEMF and RPI-CapsuleGAN on datasets RPI2241 (a), RPI7317 (b), RPI369 (c) and NPIter (d) is compared by five-fold cross-validation;

[0036] Figure 4Figure 1 is a schematic diagram of the predicted performance of different combinations of node features in an embodiment, where (a), (b) represent polar coordinate bar charts of performance comparison on RPI7317 and RPI2241 datasets, respectively; (c), (d) represent box plots of performance comparison on NPInter and RPI369 datasets, respectively. “SEQ+STRUCT” represents the combined sequence and structure features. “SEQ” represents the fused sequence features, where RNA features are extracted using RNAErnie pre-trained large language model and protein features are extracted using ESM2 pre-trained large language model. “STRUCT” represents the fused structure features, where RNA structure is predicted by RNAFold and protein structure is predicted by SOPMA. “Non-ESM 2” represents the protein sequence features excluding only ESM2 extracted, while “Non-RNAErnie” represents the RNA sequence features excluding only RNAErnie extracted;

[0037] STRUCT” represents the combined sequence and structure features. “SEQ” represents the fused sequence features, where RNA features are extracted using RNAErnie pre-trained large language model and protein features are extracted using ESM2 pre-trained large language model. “STRUCT” represents the fused structure features, where RNA structure is predicted by RNAFold and protein structure is predicted by SOPMA. “Non-ESM 2” represents the protein sequence features excluding only ESM2 extracted, while “Non-RNAErnie” represents the RNA sequence features excluding only RNAErnie extracted;

[0038] Figure 5 Figure 2 is a schematic diagram of the performance comparison of two RNA pre-trained large language models (RNAErnie and RNAAFM) in an embodiment, where (a) is a radar chart comparison schematic of NPInter dataset; (b) is a bullet chart of five indicators of RPI7317 (top) and RPI2241 (bottom) datasets; (c) is a bar chart analysis of RPI369 dataset;

[0039] Figure 6 Figure 3 is a schematic diagram of box plots illustrating the comparison of the performance of four protein pre-trained large language models (ESM2-35M, ESM2-8M, ESM-C and Protrans) on datasets RPI7317 (a), RPI2241 (b), RPI369 (c) and NPInter (d) in an embodiment;

[0040] Figure 7 Figure 4 is a schematic diagram of the performance comparison of different graph neural network architectures in an embodiment, where the radar chart compares six performance indicators of seven GNN architectures on datasets NPInter (a) and RPI7317 (b). (c) is a two-way bar chart of accuracy (ACC) comparison: top for results on dataset RPI2241, bottom for results on dataset RPI369;

[0041] Figure 8 Figure 5 is a schematic diagram of the prediction results of RPI-PLM GNN model for Drosophila melanogaster species in dataset RPI_D in an embodiment, where the ellipses and rectangles represent RNA and protein, respectively, and the black solid line represents the RPI pairs successfully predicted by the RPI-PLM GNN model, and the red dashed line represents the RPI pairs unsuccessfully predicted by the model. DETAILED DESCRIPTION

[0042] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application.

[0043] As shown in the embodiment, a method for predicting RNA-protein interaction based on a computer is provided, comprising the following steps: Figure 1

[0044] In step 101, RPI data is obtained and preprocessed to obtain a data set.

[0045] In this embodiment, four data sets RPI7317, RPI2241, RPI369 and NPInter v2.0 are selected, which respectively contain 7,301 pairs, 2,241 pairs, 369 pairs and 10,412 pairs of RPI samples, to ensure that the training data of the model has wide representativeness.

[0046] ​RPI2241 and RPI369 were extracted by Muppirala et al. from the Protein-RNA Interaction Database (PRIDB). The RPI2241 dataset contains 2,241 experimentally validated RPI pairs, covering all RNA-protein interaction types, but with a bias towards ribosomal interactions. To eliminate this bias, the RPI369 dataset was constructed, excluding all ribosome-related interactions, leaving 369 non-ribosome-specific interactions. The NPInterv2.0 dataset contains 10,412 experimentally validated RPI pairs, covering 4,636 non-coding RNAs (ncRNAs) and 449 proteins in 6 different organisms. To analyze the secondary structure of ncRNAs, the RNAfold tool in the ViennaRNA package (version 2.7.0) was used. This tool cannot predict the structure of RNA sequences longer than 32,700 nucleotides. However, in the RPI7317 dataset, some ncRNAs exceeded this limit, resulting in their secondary structures not being calculated. Therefore, this embodiment performed data cleaning based on the RPI7317 dataset, including removing ncRNAs with sequences that did not meet the length requirement and their corresponding RPI samples. At the same time, for those protein sequences that lost their pairing relationship due to the absence of ncRNAs, corresponding removal processing was also performed to maintain the integrity and consistency of the data. The final RPI7317 dataset retained 7,30 positive sample pairs, involving 1,872 ncRNAs and 117 proteins. The sequence lengths of ncRNAs in the remaining three datasets all met the requirements. To construct negative samples, the present invention randomly selected ncRNAs and proteins that lacked experimental validation to ensure the same number of positive samples.

[0047] In addition, to comprehensively evaluate the generalization ability of the model, this embodiment used the independent test set screened by Fan et al. This test set covers 6 different species: Caenorhabditis elegans, Drosophila melanogaster, Escherichia coli, Homo sapiens, Mus musculus, and Saccharomyces cerevisiae (RPI_C, RPI_D, RPI_E, RPI_H, RPI_M, and RPI_S). These species correspond to datasets containing 9, 67, 183, 7,317, 1,847, and 670 RPI samples, respectively.

[0048] Table 1 Datasets and test sets

[0049]

[0050] Step 102, constructing an RPI-PLMGNN model, the training process of which is as follows: 1) inputting RNA sequences in the data set into an RNAErnie model to extract RNA sequence features and fuse them with RNA secondary structure features to obtain RNA features; inputting protein sequences in the data set into an ESM2 model to extract protein sequence features and fuse them with protein secondary structure features to obtain protein features; fusing the RNA features and the protein features to obtain node features; 2) constructing a binary adjacency matrix based on the RNA-protein interaction relationships verified in the data set; extracting k-hop closed subgraphs centered on RNA nodes and protein nodes based on the binary adjacency matrix and performing node labeling; converting the closed subgraphs into line graph structures to obtain graph topology features; 3) inputting the node features and the graph topology features into a graph neural network architecture for joint processing to output a prediction result of the RNA-protein interaction, the graph neural network architecture comprising a graph attention network module GAT and a gated graph convolution network module GGCN.

[0051] The RPI-PLMGNN model and the training process provided by the embodiment are described in detail below with reference to Figure 2 :

[0052] 1. Node features:

[0053] RNA sequence features are obtained through an RNAErnie model and fused with RNA secondary structure information; protein sequence features are extracted using an ESM2 model and combined with protein secondary structure features. Finally, the RNA and protein features are fused as node representations. This multi-modal feature fusion method combines sequence features and structure features, which can more comprehensively represent the complexity of RPI and provide more valuable information for subsequent analysis.(1) RNAErnie model:

[0054] RNA sequence features are extracted using an RNAErnie model. The RNAErnie model is a pre-trained large language model designed specifically for RNA sequences, and its core framework is based on Enhanced Representation of Knowledge Integration (ERNIE), which consists of 12 layers of multi-head Transformer blocks, each with 768-dimensional hidden states.

[0055] First, the RNAErnie model is based on the Transformer architecture, one of its core components is the multi-head attention mechanism. This mechanism captures different aspects of the input sequence by parallel computing multiple attention heads, thereby enhancing the ability to understand RNA sequences. For each attention head, the input matrix X is mapped to the query (Q), key (K), and value (V) matrices through h different linear transformations, and the calculation formula involved in this process is as follows:

[0056]

[0057] Each head computes attention scores through self-attention mechanism. Among them is a scaling factor to prevent the gradient from vanishing due to too large dot product. The calculation formula is as follows:

[0058]

[0059] The outputs of multiple attention heads are connected and then linearly transformed to get the final result. Among them W O is the output projection matrix, head i =Attention(Q i , K i , V i ), the calculation formula involved in this process is as follows:

[0060] MultiHead(Q,K,V)=Concat(head1,head2,…,head h )W O

[0061] Secondly, the RNAErnie model takes RNA motifs as biological prior knowledge, introduces a random masking strategy at the motif / subsequence level, and combines it with masked language modeling for pre-training. Specifically, it includes three kinds of primitive awareness pre-training strategies: motif-level masking, subsequence-level masking, and primitive-level random masking. During the pre-training phase, all three masking strategies are applied to comprehensively capture the multi-level features of RNA sequences.

[0062] In addition, the model uses RNA types (such as miRNAs, lncRNAs) as stop word markers and appends them to the sequence to enhance the context representation of the sequence. In addition, for tasks that distribute RNA sequences not seen in the pre-training phase, the RNAErnie model proposes a type-guided fine-tuning strategy. The algorithm first predicts the possible RNA sequence types, and then adds the prediction results to the tail of the sequence to optimize the feature embedding.

[0063] In summary, the advantages of using the RNAErnie model lie in its strong feature representation ability and deep understanding of the complexity of RNA sequences. The first time the RNAErnie model is applied to RNA and protein interaction prediction fully captures the biological characteristics of RNA, improves the effectiveness of RNA sequence feature extraction, and provides more rich feature representation for subsequent research, thereby significantly improving the performance of RPI prediction.

[0064] (2) RNA secondary structure:

[0065] The secondary structure prediction of RNA uses the RNAFold tool in the Vienna RNA package (version 2.7.0). The prediction process is based on the principle of minimum free energy (MFE), and the results are represented in dot bracket format: a dot (".") represents an unpaired base. Base pairing is represented by parentheses "(" and ")". This representation visualizes the secondary structure characteristics of RNA. For a combination of K adjacent secondary structure symbols in a RNA sequence fragment, the frequency R SS , is as follows:

[0066]

[0067] where S j (j = 1, 2,..., 2 K ) represents the frequency of K adjacent symbols within the RNA sequence. In the k-mer, each position can be represented by a dot or a bracket, and a 2 K dimensional R SS characteristic can be generated for any RNA sequence. The value of K ranges from 1 to 4, generating dimensional SS characteristics, which are used to represent the secondary structure information of the RNA sequence. This method can effectively capture the secondary structure characteristics of the RNA sequence, providing support for subsequent analysis.

[0068] (3) ESM2 model:

[0069] The ESM2 model is a pre-trained large language model, and the protein sequence features are extracted using the ESM2 model. The core of the ESM2 model is the Transformer architecture, which is composed of multiple layers of self-attention mechanisms and feedforward neural networks. The ESM2 model used in this embodiment is esm2_t12_35M_UR50D, which has 12 layers of Transformer, 8M parameter size and 480 embedding dimensions. The ESM2 model uses self-attention mechanisms to capture long-range dependencies in the sequence. In addition, evolutionary information is introduced into the ESM2 model, which is beneficial for the model to learn the conservation of protein sequences during evolution, thereby improving the reliability of the prediction. ESM2 encodes the position along the input sequence, which helps to understand the relative relationship of different positions in the sequence. By generating protein sequence embeddings using ESM2, a biologically rich feature representation is obtained, which significantly improves the prediction performance of RPI.

[0070] (4) Protein secondary structure:

[0071] The secondary structure information of proteins is crucial in RNA-protein interaction prediction, which can provide key information about the spatial structure of proteins, thus helping to identify potential interaction sites. In this embodiment, the secondary structure prediction of proteins is performed using the SOPMA program, and the prediction includes four main types: a-helix, b-sheet, b-turn, and random coil. These secondary structure types correspond to the letters H, E, T, and C, which can quickly understand the structural characteristics of proteins. The following formula describes the frequency of occurrence of different protein secondary structure combinations:

[0072]

[0073] where S j S K is the frequency of occurrence of K adjacent secondary structure symbols within the protein sequence. Since each position within a K-length subsequence can be H, E, T, or C, for each protein sequence, a SS feature can be successfully constructed. To more comprehensively represent the secondary structure information of proteins, K is defined in the range of 1-4, and dimensional SS features are generated to represent the secondary structure information of protein sequences.

[0074] 2. Graph topology features:

[0075] To characterize the RNA-protein interaction (RPI) relationship, a binary adjacency matrix A ∈ {0, 1} N×M is constructed based on known RPI data, where N is the total number of RNA sequences, and M is the total number of protein sequences. Specifically, if the i-th RNA interacts with the j-th protein, the value of the corresponding position in the matrix is set to 1; otherwise, it is set to 0.

[0076] Next is the closed subgraph extraction. A 2-hop closed subgraph centered on the RNA node v i and the protein node v j is extracted, which contains the target nodes and all adjacent nodes within the 2-hop range and connecting edges, as shown in the following formula:

[0077]

[0078] where min(d(v, v i ), d(v, v j )) refers to the minimum distance from node v to the two target nodes v i and v j . G(v i , v j ) represents the 2-hop closed subgraph centered on the nodes v i and v j .

[0079] Then, a node labeling method is used to assign labels to each node in the closed subgraph as follows:

[0080]

[0081] where v, v i and v j denote the current node, the target ncRNA node and the target protein node, respectively. d(v, v1) and d(v, v2) denote the path distance between node v and the target ncRNA node v1 and between node v and the target protein node v2. d s = d(v, v1) + d(v, v2) denotes the distance from node v to both target nodes. The target ncRNA node v1 and the target protein node v2 are both labeled as 1, i.e., f1(v1) = 1 and f1(v2) = 1. For other nodes, they are labeled as f l (v) = 0 if there is no path between v and v1 or v2. The node labeling method helps the model to distinguish the target nodes from other nodes and capture the structural importance of each node to the target nodes. This provides important information for subsequent feature learning and model prediction.

[0082] Finally, the closed subgraph is converted into a line graph. Each edge of the original closed subgraph becomes a new node in the linear graph. In addition, if two edges share the same vertex in the original graph, a new edge connecting them is created in the linear graph. Therefore, the nodes of RNA and protein in the original graph represent individuals, while the edges represent their interactions. The linear graph focuses more on interactions rather than individual molecules. In addition, the line graph can learn the similarity between interactions, for example, if RNA fragment A interacts with protein B, and RNA fragment A also interacts with protein C, then these two interactions are connected in the line graph, indicating that they may have similar mechanisms of action. Therefore, the line graph topology can reveal potential RNA-protein interaction patterns and effectively improve RPI prediction.

[0083] 3. Graph neural network architecture:

[0084] (1) Graph Attention Network (GAT)

[0085] Graph Attention Network (GAT) is a graph neural network model for processing graph-structured data. Its core idea is to aggregate the neighbor information of nodes in the graph through attention mechanism, and update the feature representation of nodes. The following is the basic principle of GAT: for each node i in the graph, GAT calculates the attention coefficient e ij between it and each neighbor node j, which represents the importance of node j to node i. Through a shared linear transformation W, the node h iand h j are mapped into a high-dimensional space. Then the mapped features are concatenated together and a scalar value e ij is computed by a single-layer feedforward neural network a

[0086] e ij i j

[0087] Next, the e ij is normalized using a softmax function to obtain attention coefficients a ij :

[0088]

[0089] After obtaining the attention coefficients, GAT uses these coefficients to weight and sum the features of neighboring nodes to update the feature representation of the current node. The formula is as follows:

[0090]

[0091] where σ is an activation function, and h′ i is the updated feature representation of node i.

[0092]

[0093] GAT usually adopts a multi-head attention mechanism. That is, multiple attention heads are independently calculated, and their outputs are spliced together. The formula is as follows:

[0094]

[0095] where K is the number of attention heads, and W k are the attention coefficients and weight matrices of the Kth attention head, respectively. In this way, GAT can adaptively learn the relationships between nodes in the graph and effectively aggregate neighbor information, thereby improving the representation of graph structure data.

[0096] (2) Gated Graph Convolutional Neural Network (GGCN)

[0097] Gated Graph Convolutional Neural Network (GGCN) is a deep learning model that combines Graph Convolutional Network (GCN) and gating mechanism. Its core concept is to represent nodes and edges in the graph as feature vectors and weights, and use convolution operations to learn the feature representations of these nodes and edges. The convolution operation of GGCN mainly consists of two key parts: neighbor feature aggregation and gating mechanism.

[0098] ​​​In GGCN, the aggregation of neighbor features is usually represented as shown in the following formula. GGCN aggregates the feature information of its neighboring nodes for each node to generate an aggregated vector and uses it to update the representation of the central node. This process can be implemented through weighted summation, while more complex mechanisms (such as attention mechanism) can be introduced to adaptively adjust the importance of neighboring nodes, thereby enhancing the effectiveness of information propagation.

[0099]

[0100] where is used to represent the state of node t at layer l+1. N(t) is the neighbor set of the node. W is a trainable weight matrix. σ is a nonlinear activation function. The normalization term prevents nodes with higher heights from having too much influence.

[0101] GGCN dynamically adjusts the flow of information by introducing a gating mechanism. The update gate controls the weight of the neighboring nodes, while the reset gate controls the weight of the node itself. and represent the candidate state and hidden state at the current time, respectively, as follows:

[0102]

[0103]

[0104] where N(v) denotes the neighbor set of node v. ⊙ represents element-by-element multiplication (Hadamard product) for the gating mechanism. σ is the Sigmoid function, which is used to normalize the gating value to the range of (0, 1). The hyperbolic tangent function tanh is used for nonlinear transformation. W u ,W r ,W h ,U u ,U r ,U h are trainable weight matrices with different weight matrices controlling different information flows. In this embodiment, GGCN captures long-term dependencies by effectively extracting global information and can selectively convey key features. Its multi-layer convolutional network enables the model to identify complex long-distance interactions while automatically filtering out information that is crucial to the interaction, thereby improving prediction accuracy.

[0105] Step 103, model evaluation.

[0106] To evaluate the performance of the model, evaluation metrics are used, including sensitivity (SEN), specificity (SPE), precision (PRE), accuracy (ACC), Matthews correlation coefficient (MCC), and area under the curve (AUC). The formulas for these metrics are as follows:

[0107]

[0108] where TP, TN, FN, and FP represent the number of true positives, true negatives, false negatives, and false positives, respectively. SEN represents the proportion of correctly identified positive samples. SPE represents the proportion of correctly identified negative samples. PRE represents the proportion of samples that the model predicts as positive that are actually positive. ACC represents the proportion of all samples that are correctly classified. MCC measures the correlation between the true value and the predicted value, ranging from [-1, 1]. A value close to 1 indicates better model performance. AUC represents the ability of the model to distinguish between positive and negative samples, with a value closer to 1 indicating better performance. In general, the higher the value of the above metrics, the better the performance of the model.

[0109] The experiments provided by the embodiment:

[0110] 1. Comparison of experimental results of existing models on five-fold cross-validation

[0111] To comprehensively evaluate the performance of the RPI-PLMGNN model in predicting RNA-protein interactions (RPI), the present application uses a five-fold cross-validation method to conduct comparative experiments on four benchmark datasets, namely RPI7317, RPI2241, RPI369, and NPInterv2.0. Each dataset is divided into five relatively independent subsets, four of which are used as the training set, and the remaining one is used as the test set to evaluate the prediction performance of the model. This process is performed for five iterations, with a different subset selected as the test set in each iteration. The final performance results are calculated by averaging the performance indicators of the five iterations to obtain more robust and reliable evaluation results.

[0112] To demonstrate the superiority of RPI-PLMGNN, a comprehensive set of evaluation metrics is used, including SEN, SPE, PRE, ACC, MCC, and AUC, to ensure a comprehensive performance comparison. In addition, five representative existing models, namely NPI-GNN, RPITER, IPMiner, GATLGEMF, and RPI-CapsuleGAN, are selected for comparison with the RPI-PLMGNN model. As shown in Table 1, the RPI-PLMGNN model outperforms the existing models in terms of all evaluation metrics, demonstrating its superior performance in predicting RNA-protein interactions. Figure 3As shown, the proposed model outperforms these state-of-the-art models on all indicators. This superiority mainly comes from two aspects. First, pre-training large language models perform well in feature extraction, utilizing a large amount of prior knowledge to capture complex patterns and relationships in RNA-protein interactions. Second, graph neural network architectures such as GAT and GGCN effectively capture complex dependencies when processing graph-structured data, enhancing the learning ability of the model. In summary, RPI-PLMGNN has high accuracy in RNA-protein interaction prediction, providing a more reliable prediction tool for research in this field.

[0113] 2. Performance comparison between different species and other methods

[0114] In addition, the present invention further evaluates the performance of the RPI-PLMGNN model across different species. The RPI-PLMGNN model is trained on the benchmark dataset RPI7317 and evaluated for cross-species performance using six independent test sets RPI_C, RPI_D, RPI_E, RPI_H, RPI_M, and RPI_S. We compare with existing methods RPI-CapsuleGAN, RPI-MDLStack, LPICNNCP, LPI-BLS, RPISeq-RF, and RPISeqSVM, and the results are shown in Table 2. On the RPI_C, RPI_D, RPI_E, and RPI_S datasets, the prediction accuracy of RPI-PLMGNN reaches 94.2%, 92.8%, 94.5%, and 97.1%, respectively, with better performance compared to other algorithms. On the RPI_H dataset, 7,317 human RPI pairs are actually predicted, with an accuracy of 97.5%, second only to RPI MDLStack. In addition, RPI-PLMGNN predicts 1,847 RPI pairs in the dataset RPI_M, with an accuracy of 98.2%. Therefore, the RPI-PLMGNN model has better generalization ability and stable robustness.

[0115] Table 2 Performance comparison of RPI-PLMGNN and other prediction methods on independent test datasets

[0116]

[0117] 3. Ablation experiment

[0118] (1) Performance comparison of different node feature combinations:

[0119] To further explore the impact of different node features on the performance of RPI-PLM GNN, a series of comparative experiments covering various feature combinations are designed. Specifically, the following feature combinations are considered for their impact on performance: "SEQ+STRUCT", "SEQ", "STRUCT", "Non-ESM 2", and "Non-RNA Ernie". Among them, the "SEQ+STRUCT" feature combination is the feature fusion method adopted in this embodiment.

[0120] The results of the performance comparison experiments conducted on the benchmark dataset show that when all features including sequence and structure features are input to the model, the model's prediction performance reaches the best level. This result indicates that the fusion of sequence features and structure features can provide the model with more comprehensive and richer information, thereby enhancing its prediction ability. In addition, when only sequence features are input, the model's performance is second. In contrast, when only structure features are used, the model's performance is relatively poor. From Figure 4 It can be seen that in the RPI2241 dataset, the ACC prediction value of the sequence features extracted using the pre-trained large language model is significantly better than that of the structure features, with a value exceeding 3.85%. This indicates that the pre-trained large language model can capture more rich semantic information and biological characteristics when extracting features, providing strong support for model prediction.

[0121] In addition, in order to further analyze the specific impact of the two pre-trained large language models, RNAErnie and ESM2, on model performance, experiments were conducted to remove RNAErnie and ESM2, respectively. The experimental results show that when protein sequence features are not extracted using ESM2, the model's performance decreases more than when RNA sequence features are not extracted using RNAErnie. As Figure 3 shown, in the dataset RPI7317, the ACC prediction value decreases by 0.27% and the AUC prediction value decreases by 0.6% when protein sequence features are not extracted using ESM2 compared to when RNA sequence features are not extracted using RNAErnie. This result indicates that the ESM2 pre-trained large language model plays a more critical role in the RPI prediction task. The possible reason is that ESM2 can more effectively extract biological features related to RNA interaction when processing protein sequences, thereby providing more accurate and valuable information for model prediction. This shows that in the RPI prediction task, accurately representing protein sequence features is crucial to improving model performance, and the ESM2 pre-trained model has shown significant advantages in this regard.

[0122] (2) Contribution of different pre-trained large language models

[0123] The present application evaluates the impact of embedding features generated by different pre-trained large language models on the RPI prediction task. Based on the limitation of computing resources, two RNA pre-trained large language models, RNAErnie and RNA-FM, are compared in the RNA feature extraction task. See Figure 5 The experimental results show that the performance of RNAErnie is significantly better than that of RNA-FM in all tasks. As shown in the radar chart of Figure 5 (a), RNAErnie shows more comprehensive advantages, with a coverage significantly larger than that of RNA-FM, and shows stronger performance in multiple key indicators. This may be due to the fact that RNAErnie combines biological priors and multi-level masking as well as type-guided fine-tuning strategies, providing accurate and highly generalizable features for downstream tasks, thereby significantly improving the performance of RNA-related tasks.

[0124] In the protein feature extraction task, four protein large language models, ESM2-8M, ESM2-35M, ESM-C(300M) and ProtTrans, are evaluated. As shown in the box plot in Figure 6 ESM2-35M significantly outperforms other protein large language models in all evaluation indicators on the four benchmark datasets, fully demonstrating the superiority of the model in the RPI prediction task. ProtTrans has poorer overall performance compared to the Evolutionary Scale Modeling (ESM) model, which may be attributed to the larger feature dimension (1024 dimensions) extracted by it, which can lead to overfitting. On the other hand, the advantage of ESM models lies in their training with evolutionary information, which can effectively capture complex patterns in protein sequences, thereby providing more biologically meaningful feature representations. ESM2-35M extracts features with a dimension of 480, while ESM2-8M extracts features with a dimension of 320. The performance of two different versions of protein pre-trained large language models is compared, and ESM2-35M is generally superior to ESM2-8M, because ESM2-35M is trained on a larger dataset, which can learn more complex protein sequence features, thereby obtaining more general and efficient feature representations. Although ESM-C is superior to ESM2 in many areas, its performance is not the best. Therefore, it is believed that this may be mainly due to the fact that ESM-C extracts features with a dimension of 960, while a part of the dataset is relatively small, which may lead to overfitting.

[0125] Based on the above experimental results and theoretical analysis, RNAErnie and ESM2-35M are finally selected as the sequence feature extraction tools of the present application.

[0126] (3) Performance comparison of different graph neural network architectures:

[0127] To systematically evaluate the effectiveness of the proposed model, the performance of multiple graph neural network architectures was compared on four benchmark datasets RPI7317, RPI2241, RPI369 and NPInter.

[0128] Figure 7 The comparison results of different models on different datasets are shown, and it can be seen that the performance of GGCN is significantly better than GCN, for example, the ACC is increased by 2.25% on the dataset RPI2241 and 0.87% on the dataset RPI369, which is due to the addition of its gating mechanism. This mechanism enhances the model's ability to handle long dependencies by controlling information flow, thus more effectively capturing the global features of RNA-protein interactions. In addition, the performance of GAT+GGCN is better than single models and other model combinations, for example, on the NPInter dataset, the value of ACC reaches 97.74% and AUC reaches 98.98%. The reason may be that the attention mechanism of GAT can adaptively learn the weights of key interaction sites in RNA and protein sequences to highlight local important features, while the gating mechanism of GGCN complements the ability to model global dependencies. The combination of the two achieves a synergistic optimization of local and global features, and the performance is significantly improved. Experiments show that by fusing the graph attention mechanism and the gating mechanism, the combined structure of GAT+GGCN has obvious advantages in the RNA-protein interaction prediction task.

[0129] (4) Prediction of RNA-protein interaction network:

[0130] The prediction results of RPI-PLMGNN model on 67 RPIs in RPI_D were analyzed. The RNA-protein interaction network graph was drawn with RNA and protein as nodes and interaction relationship as edges. In this network, the oval frame represents RNA, the rectangular frame represents protein, the black solid line represents the RPI pair successfully predicted by the RPI-PLMGNN model, and the red dashed line represents the RPI pair not successfully predicted. This analysis method helps to understand the interaction relationship between RNA and protein, and provides visual support for subsequent functional research. The network structure is shown in Figure 8

[0131] ​Drosophila melanogaster is a classic organism that has been widely used in biological research. It is small in size, short in lifespan, strong in reproductive ability, and easy to genetically manipulate, making it an ideal species for studying genetic mechanisms, gene functions, and signal pathway regulation. For example, P40417 is a mitogen-activated protein kinase ERK-A from Drosophila melanogaster. ERK-A can help cells respond to external growth signals, drive the cell cycle forward, and thus promote cell proliferation. ERK-A regulates the expression of genes related to the cell cycle, proliferation, and apoptosis, thereby affecting cell function and state. In addition, the Drosophila genome contains a large number of RNA-protein interactions (RPIs) related to transcriptional regulation. Systematic analysis of the RPI network of Drosophila melanogaster not only helps to reveal the regulatory mechanisms of gene expression, but also has important significance for the study of life activities such as cell differentiation and development.

[0132] It should be understood that although each step in the above flowchart is displayed in sequence according to the direction of the arrow, these steps are not necessarily executed in sequence according to the direction of the arrow. Unless otherwise specified herein, there is no strict order limitation for the execution of these steps, and these steps can be executed in other orders. Moreover, at least part of the steps in the above flowchart can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of these sub-steps or stages is not necessarily sequential, but can be alternately executed with at least part of other steps or sub-steps or stages of other steps.

[0133] The above are only preferred embodiments of the present application, and are not used to limit the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for in silico prediction of RNA-protein interactions, characterized in that, The method comprises the following steps: Obtaining RPI data and preprocessing to obtain a data set; Building an RPI-PLMGNN model, and the training process of the RPI-PLMGNN model comprises the following steps: 1) inputting RNA sequences in the data set into an RNAErnie model to extract RNA sequence features and fuse the RNA sequence features with RNA secondary structure features to obtain RNA features; inputting protein sequences in the data set into an ESM2 model to extract protein sequence features and fuse the protein sequence features with protein secondary structure features to obtain protein features; Fusing the RNA features and the protein features to obtain node features; 2) constructing a binary adjacency matrix based on RNA-protein interaction relationships verified by experiments in the data set; extracting a k-hop closed subgraph centered on RNA nodes and protein nodes based on the binary adjacency matrix and performing node labeling; Converting the closed subgraph into a line graph structure to obtain graph topology features; 3) inputting the node features and the graph topology features into a graph neural network architecture for joint processing to output a prediction result of the RNA-protein interaction, wherein the graph neural network architecture comprises a graph attention network module GAT and a gated graph convolution network module GGCN; Model evaluation.

2. The method of computer-based prediction of RNA-protein interactions according to claim 1, characterized in that, The method for obtaining RPI data and preprocessing to obtain a data set comprises the following steps: Collecting RNA-protein interaction relationships verified by experiments as positive samples from RPI7317, RPI2241, RPI369 and NPInter v2.0 benchmark data sets; removing RNA sequences and corresponding RPI data whose lengths do not meet preset requirements from the positive samples; Randomly selecting RNA and protein lacking experimental verification as negative samples, wherein the number of the negative samples is the same as that of the positive samples.

3. The method of computer-based prediction of RNA-protein interactions according to claim 2, characterized in that, The model evaluation comprises the following steps: Training the RPI-PLMGNN model using four benchmark data sets respectively; Determining evaluation indexes, performing K-fold cross-validation on each model, and reserving the best model weight to verify the performance of the RPI-PLMGNN model using several independent test sets across species.

4. The method of computer-based prediction of RNA-protein interactions according to claim 3, characterized in that, The evaluation indexes comprise sensitivity, specificity, precision, accuracy, Matthews correlation coefficient and area under the curve.

5. The method of computer-based prediction of RNA-protein interactions according to claim 1, characterized in that, The method further comprises the following steps: The RNAErnie model adopts a motif-level mask, a subsequence-level mask and a motif-level random mask strategy, and combines RNA types as word labels; The RNAErnie model appends word labels to the end of the RNA sequence to enhance the context representation of the sequence.

6. The method of computer-based prediction of RNA-protein interactions according to claim 1, characterized in that, The ESM2 model comprises 12 stacked Transformer layers, each of which is composed of a multi-head attention mechanism and a feedforward neural network, and combines layer normalization and residual connection.

7. The method of computer-based prediction of RNA-protein interactions according to claim 1, characterized in that, The method for obtaining RNA secondary structure features comprises the following steps: Based on the principle of minimum free energy, calculating the free energy of RNA in different folding states to obtain a prediction result, and representing the base pairing in the prediction result with different specific characters respectively; Determining RNA secondary structure features based on the combination frequency of K adjacent specific characters in an RNA sequence fragment.

8. The method for predicting RNA-protein interactions based on a computer according to claim 1, wherein: The protein secondary structure feature acquisition method comprises the following steps: The SOPMA program is used for protein secondary structure prediction, and the predicted secondary structure types alpha-helix, beta-sheet, beta-turn and random coil are respectively represented by different specific characters; Based on the frequency of K adjacent specific characters in the protein sequence, the protein secondary structure feature is determined.

9. The method of computer-based prediction of RNA-protein interactions according to claim 1, characterized in that, The graph attention network module is used to aggregate the neighbor information of nodes in the graph through a multi-head attention mechanism to highlight local important features, update the feature representation of the nodes and pass it to the gated graph convolution network module; the gated graph convolution network module comprises a graph convolution network, a gating mechanism and an attention mechanism, the graph convolution network is used to aggregate the feature information of adjacent nodes for each node to generate an aggregation vector, thereby capturing global structure information; the gating mechanism filters the output of the graph convolution network, dynamically controls the flow of information, retains important features and suppresses irrelevant or noise information; the attention mechanism is used to adaptively adjust the importance of adjacent nodes, thereby enhancing the effectiveness of information propagation.

10. A system for in silico prediction of RNA-protein interactions, characterized in that, It comprises: The acquisition module is used for acquiring RPI data and pre-processing to obtain a data set; The training module is used for constructing an RPI-PLMGNN model, and the RPI-PLMGNN model training process comprises the following steps: 1) inputting the RNA sequence in the data set into the RNAErnie model to extract the RNA sequence feature and fuse it with the RNA secondary structure feature to obtain the RNA feature; inputting the protein sequence in the data set into the ESM2 model to extract the protein sequence feature and fuse it with the protein secondary structure feature to obtain the protein feature; Fuse the RNA feature and the protein feature to obtain the node feature; 2) based on the RNA-protein interaction relationship verified by experiments in the data set, a binary adjacency matrix is constructed; based on the binary adjacency matrix, a k-hop closed subgraph centered on the RNA node and the protein node is extracted and the node is labeled; Convert the closed subgraph into a line graph structure to obtain the graph topology feature; 3) input the node feature and the graph topology feature into a graph neural network architecture for joint processing to output the prediction result of the RNA-protein interaction, the graph neural network architecture comprises a graph attention network module GAT and a gated graph convolution network module GGCN; The evaluation module is used for model evaluation.