Prediction method for identifying RNA methylation sites
Through the combination of multimodal feature fusion and Transformer encoder, the problem of insufficient feature extraction in the existing RNA methylation site prediction methods is solved, and higher prediction accuracy and generalization capabilities are achieved.
Patent Information
- Application Number
- CN202510428519.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-11
AI Technical Summary
The existing RNA methylation site prediction methods rely on artificial feature engineering. The performance is affected by the quality of feature design and model integration methods, making it difficult to fully extract nucleotide features, resulting in insufficient prediction accuracy.
A multimodal feature fusion architecture based on attention mechanism is adopted, through one-heat encoding, nucleotide chemical property encoding, electron-ion interaction potential encoding and local nucleotide composition encoding, feature embedding is performed in combination with the DNA2Vec model, and long-range dependencies are captured using the Transformer encoder to optimize the model to improve prediction accuracy.
It significantly improves the prediction accuracy of RNA methylation sites, improves the sensitivity, specificity and Matthews correlation coefficient of the model, and enhances the generalization ability of the model.
Smart Images

Figure CN120299525A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of predicting functional sites in sequence analysis in bioinformatics, and specifically to a prediction method for identifying methylation sites in RNA sequences. Background Art
[0002] N7-methylguanosine (m7G) is a highly evolutionarily conserved RNA methylation modification that plays a key role in gene expression regulation networks and cell fate determination. This modification is widely distributed in various RNA molecules of eukaryotes, and its dynamic characteristics are involved in multiple important biological processes, including the formation of the mRNA 5'-end cap structure, the maintenance of the stability of the tRNA anticodon loop, and rRNA processing and maturation. In addition, recent studies have found that m7G also has biomarker functions in the non-cap region inside mRNA. Functional studies have shown that m7G modification can coordinate protein expression by regulating mRNA stability and nucleocytoplasmic transport efficiency, while maintaining the tRNA tertiary structure and decoding accuracy, and promoting ribosome biosynthesis by affecting rRNA subunit assembly. In addition, abnormal m7G modification has been proven to be closely related to major diseases such as malignant tumors (such as promoting the selective translation of oncogenic mRNA) and cardiovascular diseases, indicating its important significance in the study of RNA-mediated pathological mechanisms.
[0003] Currently, a variety of machine learning- and deep learning-based methods have been used for RNA methylation site identification. However, these methods usually rely on artificial feature engineering or the integration of multiple models, and their performance is greatly affected by the quality of feature design and the model integration method. In addition, existing models often have difficulty in comprehensively extracting nucleotide features in RNA sequences, resulting in limited prediction ability. To further improve the accuracy of m7G modification prediction, the accuracy of the model needs to be further improved, and it should be possible to develop better models using new deep learning frameworks. Summary of the Invention
[0004] The purpose of the present invention is to solve the problem of the accuracy of existing RNA methylation modification site prediction, and provide a multi-modal feature fusion architecture based on the attention mechanism for accurately identifying RNA methylation modification sites. This method combines feature information of different modalities and uses the attention mechanism to dynamically adjust feature weights, thereby significantly improving the prediction accuracy of RNA methylation modification sites.
[0005] The following are the technical solutions for achieving the purpose of the present invention, including the following steps: 1) Sample dataset construction: Collect positive sample sequences containing RNA methylation modification sites and negative sample sequences of non-modified sites, and divide them into a training set and an independent test set according to a preset ratio; 2) Multi-modal feature fusion: Implement feature encoding through the following four parallel channels: • Channel 1: Perform binary matrix transformation on the nucleotide sequence using one-hot encoding. • Channel 2: Based on the nucleotide chemical property (NCP) encoding, reflecting the polarity and hydrophobicity parameters of bases. • Channel 3: Apply electron-ion interaction potential (EIIP) encoding to characterize the physicochemical properties of nucleotides. • Channel 4: Obtain the distribution characteristics of the local window of the sequence through local nucleotide composition encoding (ENAC). Concatenate the features of the above four channels to form a multimodal feature fusion pathway (MRF). 3) Semantic vector embedding: Use the DNA2Vec model to perform distributed representation learning on the input sequence to generate embedding vectors containing the context semantic information of nucleotides. 4) Construction of deep neural network: The network structure includes: performing cross-modal concatenation on the output layers of the multimodal feature fusion pathway and the DNA2Vec embedding pathway, using a fully connected layer for feature dimension reduction processing, capturing long-range dependencies in the sequence through a Transformer encoder module, and finally calculating the prediction probability of methylation sites through a Sigmoid activation function. 5) Model optimization method: Use a batch normalization layer to standardize the feature distribution, configure a Dropout layer with an inactivation rate of 0.3 - 0.5 to prevent overfitting, and minimize the binary cross-entropy loss function through an adaptive moment estimation optimizer (Adam). 6) Multimodal fusion: We use the Transformer structure to fuse multimodal features and semantic vectors to achieve deep interaction of cross-modal features. 7) Establishment of verification system: Implement a five-fold cross-validation framework, retain 20% of the training data as the validation set in each round, and finally comprehensively evaluate the performance of the model with the average sensitivity (Sn), specificity (Sp), accuracy (Acc), and Matthews correlation coefficient (MCC) of five rounds of verification, and verify the generalization ability of the model on an independent test set. Beneficial effects
[0006] This model contains two feature processing pathways: the multimodal feature fusion pathway (MRF). After the outputs of the two pathways are concatenated, the feature dimension is compressed by a fully connected layer, and then long-range dependencies are captured by a Transformer encoder. Finally, the prediction probability is output by a fully connected layer. Through multi-source feature fusion and hierarchical modeling, this architecture significantly improves the parsing ability of sequence context information and provides a new method for RNA modification prediction.
[0007] Figure 1 It is a flowchart of a prediction method for identifying RNA methylation sites. Figure 2 It is the ROC curve of the model under 5-fold cross-validation; Specific implementation manners
[0008] The specific implementation manners of the present invention will be described below in conjunction with the accompanying drawings. The accompanying drawings are only for illustrative purposes and should not be construed as limiting the present invention. The accompanying drawings are only for reference and illustration, and do not constitute a limitation on the protection scope of the present invention, because many changes can be made to the present invention without departing from the spirit and scope of the present invention.
[0009] As Figure 1 shown, this figure shows the overall workflow of this study. The model includes two feature processing paths: a multimodal feature fusion path (MRF) and a DNA2Vec embedding path. The outputs of the two paths are concatenated and then the feature dimension is compressed through a fully connected layer. Subsequently, a Transformer encoder is used to capture long-range dependencies, and finally the prediction probability is output by a fully connected layer.
[0010] To build and evaluate the proposed model, we used two datasets to test its prediction ability at m7G sites and its generalization ability at other RNA sites respectively. These two-layer models use the same network structure and parameters, and the only difference lies in the input data. The specific information of the datasets is shown in Table 1, and the generalization test dataset is shown in Table 2.
[0011] Table 1 Standard dataset Train Valid Test postive negative postive negative postive negative 3511 3511 878 878 1097 1097
[0012] Table 2 Cross-tissue validation dataset Train Test postive negative postive negative hAm 1541 1541 50 50 hCm 1828 1828 50 50 hGm 1421 1421 50 50 hm1A 16296 16290 50 50 hm5C 3157 3157 50 50 hm5U 3646 3644 50 50 hm6A 65122 65048 50 50 hm6Am 2395 2320 50 50 hPsi 3087 3055 50 50 hTm 2203 2203 50 50 hAm 1541 1541 50 50
[0013] We used 5-fold cross-validation on the training set to evaluate the effectiveness of the proposed model and used an independent test set for performance verification. In addition, we further used data from other RNA methylation sites to evaluate the generalization ability of the predictor. To comprehensively measure the performance of the predictor, we used five key evaluation metrics: sensitivity (Sn), specificity (Sp), accuracy (Acc), Matthews correlation coefficient (MCC), and area under the ROC curve (AUC).
[0014] Finally, we compared the proposed method with the existing state-of-the-art methods. The comparison results on the independent test set are shown in Table 3, and the results of the generalization test are shown in Table 4.
[0015] Table 3 Performance comparison with existing methods based on the independent test set ACC SN SP MCC AUC Moss-m7G 0.8436 0.8231 0.8641 0.6879 0.9235 m7G-DLSTM 0.8154 0.8177 0.8131 0.6308 0.8907 m7GPredictor 0.7475 0.7010 0.7940 0.4971 0.8298 iRNA-m7G 0.6609 0.8687 0.4531 0.3538 0.8198 Proposed 0.9721 0.9612 0.9830 0.9445 0.9843
[0016] Table 4 Performance comparison of models on 10 RNA modification datasets AUC ACC SN SP MCC hAm 0.9721 0.9520 0.9400 0.9640 0.9046 hCm 0.9510 0.9240 0.9160 0.9320 0.8488 hGm 0.9290 0.8600 0.7640 0.9559 0.7383 hm1A 0.9819 0.9560 0.9440 0.9680 0.9125 hm5C 0.9803 0.9720 0.9640 0.9800 0.9447 hm5U 0.9804 0.9520 0.9840 0.9200 0.9071 hm6A 0.9308 0.8460 0.7800 0.9120 0.6992 hm6Am 0.9798 0.9320 0.8840 0.9800 0.8690 hPsi 0.9138 0.9160 0.9040 0.9280 0.8337 hTm 0.9712 0.9620 0.9400 0.9840 0.9256
Claims
1. A prediction method for identifying RNA methylation sites, the process of which includes the following steps: 1) Sample dataset construction: Collect positive sample sequences containing RNA methylation modification sites and negative sample sequences of non-modified sites, and divide them into a training set and an independent test set according to a preset ratio; 2) Multimodal feature fusion: Implement feature encoding through the following four parallel channels: • Channel one: Use one-hot encoding to perform binary matrix conversion on nucleotide sequences • Channel two: Based on nucleotide chemical property (NCP) encoding, reflecting the polarity and hydrophobicity parameters of bases; • Channel three: Apply electron-ion interaction potential (EIIP) encoding to characterize the physicochemical properties of nucleotides; • Channel four: Through local nucleotide composition encoding (ENAC), obtain the distribution characteristics of the local window of the sequence. Concatenate the features of the above four channels in tensors to form a multimodal feature fusion pathway (MRF); 3) Semantic vector embedding: Use the DNA2Vec model to perform distributed representation learning on the input sequence, and generate embedding vectors containing nucleotide context semantic information; 4) Deep neural network construction: Include the following hierarchical structure: Concatenate the output layers of the multimodal feature fusion pathway and the DNA2Vec embedding pathway cross-modally, perform feature dimension reduction processing through a fully connected layer, capture the long-range dependence relationship in the sequence with the help of a Transformer encoder module, and finally output the predicted methylation site probability through a fully connected layer; 5) Model optimization method: Use a batch normalization layer to standardize the feature distribution, prevent overfitting through a Dropout layer with an inactivation rate of 0.3 - 0.5, and use an adaptive moment estimation optimizer (Adam) to minimize the binary cross-entropy loss function; 6) Multimodal fusion: We use the Transformer structure to fuse multimodal features and semantic vectors to achieve deep interaction of cross-modal features; 7) Verification system establishment: Implement a five-fold cross-validation framework, retain 20% of the training data as the validation set in each round, and comprehensively evaluate the model through the average sensitivity (Sn), specificity (Sp), accuracy (Acc), and Matthews correlation coefficient (MCC) of the five rounds of validation, and verify the generalization ability of the model on the independent test set.
2. This method uses two feature processing pathways: one is the multimodal feature fusion pathway (MRF), and the other is the DNA2Vec embedding pathway. After the outputs of the two pathways are concatenated, the feature dimension is compressed through a fully connected layer, and then the long-range dependence relationship in the sequence is captured by a Transformer encoder, and finally the predicted probability is output by a fully connected layer.
3. The prediction method of an RNA methylation site according to claim 1, wherein The feature encoding scheme and model construction described in steps 2), 3), and 4), through multi-source feature fusion and hierarchical modeling, significantly improve the parsing ability of sequence context information, providing a new method for RNA modification prediction.
Citation Information
Cited By
Methylation prediction method based on structure prior and multi-scale signal decomposition
CN122417145A
Methylation prediction method based on structural prior and multi-scale signal decomposition
CN122417145B