Deep learning-based lncRNA m5C methylation site prediction method
The lncRNA m5C methylation site prediction model constructed through deep learning methods uses nucleotide and dinucleotide encoding to bind CNN and Bi-LSTM networks to solve the problem of low prediction accuracy of lncRNA m5C methylation site prediction and achieve higher prediction accuracy and model generalization capabilities.
Patent Information
- Application Number
- CN202510555268.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-08-12
AI Technical Summary
In the prior art, the accuracy of the lncRNA m5C methylation site prediction method is low, especially due to insufficient data and model overfitting, which leads to insufficient generalization ability.
Using a deep learning-based method, by obtaining the nucleotide physicochemical properties, accumulated frequency information and dinucleotide chemical properties of the lncRNA sequence to be tested, combining the CNN convolution module, Bi-LSTM network module and full-connection module, an lncRNA m5C methylation site prediction model is constructed to capture local characteristics and context-dependent characteristics, and improve prediction accuracy.
The prediction accuracy of lncRNA m5C methylation site was significantly improved, the generalization ability of the model was improved, and higher accuracy, recall and Matthews correlation coefficient were achieved.
Smart Images

Figure CN120472981A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of m5C site prediction, and in particular to a method for predicting lncRNA m5C methylation sites based on deep learning. Background Art
[0002] m5C is a key RNA modification that is widely involved in regulating RNA stability, localization, and function. As a crucial regulatory molecule, m5C modification of lncRNA not only influences its own structure and function but also participates in biological processes such as gene expression regulation, chromatin remodeling, and signaling pathway activation through interactions with proteins, DNA, or other RNAs.
[0003] In recent years, the prediction of RNA m5C sites has sparked a widespread research boom both domestically and internationally, with a series of important advances. To date, numerous computational methods based on sequence-derived information have been developed to predict m5C sites in a variety of species, including humans, house mice, and Saccharomyces cerevisiae. For example, iRNAm5C-PseDNC predicts m5C sites using a random forest algorithm. This model is built on an unbalanced and redundant dataset containing 475 m5C sites and 1425 non-m5C sites. RNAm5Cfinder is a web server based on a random forest algorithm that collects data from the GEO database and can be used to predict m5C sites in eight cell types or tissues from mice and humans. Although all m5C sites recorded in three GEO records and all other non-m5C sites in the genome were collected to train the model, the redundancy of the dataset was not well addressed. During the development of the Deepm5C model, a new benchmark dataset was constructed. A mixture of three traditional feature encoding algorithms and a feature extracted from a word embedding method was studied. Four deep learning classifier variants and four commonly used traditional classifiers were used, and four encodings were used for training. Ultimately, 32 baseline models were obtained. By integrating the prediction outputs of the optimal baseline models and training them with a one-dimensional convolutional neural network, the stacking strategy was effectively used to improve model performance.
[0004] While research on RNA m5C site prediction has achieved some success, it has focused on common RNA, particularly mRNA. There is a lack of research specifically targeting lncRNA m5C. lncRNA and mRNA m5C differ significantly in structure, function, and distribution. Applying existing prediction models for common RNA m5C sites to the task of predicting lncRNA m5C methylation sites results in low prediction accuracy. Due to the limited data available on lncRNA m5C sites, models built using traditional machine learning or deep learning methods often face overfitting, limiting the model's generalization ability and resulting in low prediction accuracy. Summary of the Invention
[0005] The purpose of the present invention is to solve the problem of low prediction accuracy of existing lncRNA m5C methylation site methods, and propose a lncRNA m5C methylation site prediction method based on deep learning.
[0006] The lncRNA m5C methylation site prediction method based on deep learning is as follows:
[0007] Step 1: Obtain the original features of the lncRNA sequence to be tested;
[0008] Step 2: Input the original features of the lncRNA sequence to be tested into the trained lncRNA m5C methylation site prediction model to obtain the m5C site prediction results;
[0009] If the m5C site prediction result is greater than the preset threshold, it means that the m5C site exists in the lncRNA sequence to be tested; otherwise, it means that the m5C site does not exist in the lncRNA sequence to be tested;
[0010] The trained lncRNA m5C methylation site prediction model is obtained by the following method:
[0011] S1. Obtain the lncRNA sequence fragments corresponding to the methylation data file in the human gene library, and use the lncRNA sequence fragments to construct a training set;
[0012] S2. Sequence encode the data in the training set to obtain the original feature vector of lncRNA and obtain the encoded training set;
[0013] S3. Use the encoded training set to train the lncRNA m5C methylation site prediction model to obtain a trained lncRNA m5C methylation site prediction model.
[0014] Furthermore, the step 1 of obtaining the original features of the lncRNA sequence to be tested is specifically as follows:
[0015] Step 11: Encode the physicochemical properties and cumulative frequency information of each nucleotide in the lncRNA sequence to be tested to obtain a first encoding feature matrix of the lncRNA sequence to be tested;
[0016] Step 1 and 2: Encode the dinucleotide chemical properties of the lncRNA sequence to be tested to obtain a dinucleotide chemical property feature vector of the lncRNA sequence to be tested;
[0017] Step 13: Concatenate the first coding feature vector and the dinucleotide chemical property feature vector of the lncRNA sequence to be tested to obtain the original feature vector of the lncRNA sequence to be tested.
[0018] Furthermore, the physical and chemical properties and cumulative frequency information of each nucleotide in the lncRNA sequence to be tested are encoded in step 1 to obtain the first encoding feature matrix of the lncRNA sequence to be tested, specifically:
[0019] First, the physical and chemical properties of each nucleotide in the lncRNA sequence to be tested are encoded as follows:
[0020] If the nucleotide is adenine A, the physicochemical property code is (1,1,1); if the nucleotide is cytosine C, the physicochemical property code is (0,1,0); if the nucleotide is guanine G, the physicochemical property code is (1,0,0); if the nucleotide is uracil U, the physicochemical property code is (0,0,1);
[0021] Then, the cumulative frequency of bases of the lncRNA sequence to be tested is obtained
[0022] Among them, f i' is the frequency of the i'th nucleotide occurring before the i'th nucleotide position, d i' is the number of times the i'th nucleotide appears before the i'th nucleotide position, i' is the nucleotide index in the lncRNA to be tested;
[0023] Finally, the nucleotide physical and chemical property codes and the corresponding base cumulative frequencies in the lncRNA sequence to be tested are combined into the nucleotide first coding feature vector, thereby obtaining the first coding feature matrix of the lncRNA sequence to be tested.
[0024] Furthermore, the dinucleotide chemical properties of the lncRNA sequence to be tested are encoded in steps one and two to obtain a dinucleotide chemical property feature vector of the lncRNA sequence to be tested, specifically:
[0025] First, obtain the PCP matrix of the lncRNA sequence to be tested:
[0026]
[0027] Where p∈[1,10], p is the PCP label, PCP p (P j P j+1 ) is the dinucleotide P j P j+1 The pth PCP value of , L is the length of the lncRNA sequence to be tested, j∈[1,L-1], j is the nucleotide index;
[0028] Then, the PCP matrix is normalized to obtain a normalized PCP matrix, and the dinucleotide autocovariance matrix A' and the cross-covariance matrix C' are obtained using the normalized PCP matrix;
[0029] Finally, the matrix A' and the matrix C' are concatenated to obtain the dinucleotide chemical property feature vector of the lncRNA sequence to be tested.
[0030] Furthermore, the normalized PCP matrix is used to obtain the dinucleotide autocovariance matrix A' and the cross-covariance matrix C', specifically:
[0031]
[0032] Where, s = 1, 2, 3, 4, λ1≠λ2, λ1 and λ2 are integers, λ1 = 1, 2, ..., 10, λ2 = 1, 2, ..., 10, is the average value of the λ1th row data in the normalized PCP matrix, is the average value of the λ2th row data in the normalized PCP matrix, PCP' p (P j P j+1 ) is the normalized dinucleotide P j P j+1 The p-th PCP value of is the average value of the p-th row data in the normalized PCP matrix.
[0033] Furthermore, the lncRNA sequence fragments corresponding to the methylation data file in the human gene library are obtained in S1, and the lncRNA sequence fragments are used to construct a training set, specifically:
[0034] S11. Obtaining the site information of lncRNA m5C in the RMBase v3.0 database or the m5C-Atlas database, and merging the site information of lncRNA m5C in the RMBase v3.0 database with the site information of lncRNA m5C in the corresponding m5C-Atlas database to obtain a total set of lncRNA m5C site information;
[0035] The site information of the lncRNA m5C includes: chromosome number, start position, end position, positive and negative strand identification;
[0036] S12. Extract multiple lncRNA sequence fragments from the human genome based on the site information of lncRNA m5C, then cut off sequence fragments of a preset length L, use the CD-HIT tool to remove sequence fragments with a similarity greater than a preset threshold, and use the remaining sequence fragments as a positive sample set;
[0037] S13. Sequence fragments of length L centered around cytosine are intercepted upstream and downstream of the positive sample site. Sequence fragments with a similarity greater than a preset threshold and fragments overlapping with the positive sample site are then removed using the CD-HIT tool. The remaining sequence fragment set is used as the negative sample set.
[0038] S14, setting labels for positive samples and negative samples, and combining the labeled positive sample set and negative sample set to obtain a training set;
[0039] Among them, the positive sample label is 1, and the negative sample label is 0.
[0040] Furthermore, the data in the training set are sequence-encoded in S2 to obtain the original feature vector of lncRNA and the encoded training set, specifically:
[0041] S21, encoding the nucleotide physical and chemical properties and cumulative frequency information of the sequence segments in the training set to obtain a first encoding feature matrix;
[0042] S22. Encode the dinucleotide chemical properties of the sequence fragment to obtain a dinucleotide chemical property feature vector of the sequence fragment;
[0043] S23. Concatenate the first encoding feature vector obtained in S21 and the dinucleotide chemical property feature vector obtained in S22 to obtain the original feature vector of the lncRNA sequence, and use the original feature vector of the lncRNA sequence and the label to form an encoded training set.
[0044] Furthermore, the lncRNA m5C methylation site prediction model includes: a CNN network module, a Bi-LSTM network module, a feature fusion module, and a fully connected module;
[0045] The CNN network module is used to perform preliminary feature extraction on the original feature vector of the lncRNA sequence to obtain local feature information F1 of the sequence fragment;
[0046] The Bi-LSTM network module uses feature F1 to obtain sequence feature F2;
[0047] The feature fusion module is used to concatenate feature F1 and feature F2 to obtain the optimal feature vector F of the lncRNA sequence;
[0048] The fully connected layer uses the optimal feature vector F of the lncRNA sequence to output the m5C site prediction result.
[0049] Furthermore, the CNN network module includes: three convolution units;
[0050] Each convolution unit includes: convolution layer, activation function layer, pooling layer in sequence;
[0051] The convolutional layer in the first convolutional unit is specifically:
[0052]
[0053] Among them, X u+m,n is the original feature vector of lncRNA, u is the output position index, k is the kernel index, conv(X) uk is the output of the convolutional layer, It is an M*N weight matrix, where M is the window size and N is the number of input channels;
[0054] The convolutional layers in other convolutional units are similar to the above formula, and the input is the output of the previous convolutional unit;
[0055] The activation function layer is as follows:
[0056]
[0057] Among them, x is the input of the activation function; the input of each activation function is x is the output of the previous convolutional layer.
[0058] Furthermore, the pooling layer is specifically:
[0059] pooling(x') u' =max(x'1,x'2,x'3,...,x' M' ) u'
[0060] Among them, x' represents the input feature, the input x' of each pooling layer is the output of the previous activation function, u' represents the output position index, M' represents the pool window size, x'1, x'2, x'3, ..., x' M' is the sequence of outputs of the previous activation function.
[0061] The beneficial effects of the present invention are:
[0062] The present invention uses a nucleotide physical and chemical property and frequency encoding method and a dinucleotide chemical property encoding method to encode RNA sequences, which are then input into a lncRNA m5C methylation site prediction model. The lncRNA m5C methylation site prediction model first uses a CNN convolution module to capture local features, then uses a Bi-LSTM to capture context-dependent features. Finally, the local features and context-dependent features are combined to obtain combined features, and the lncRNA sequence is classified through a fully connected network. The present invention improves the accuracy of lncRNA m5C methylation site prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 Flowchart of the present invention;
[0064] Figure 2 Performance comparison chart for single feature and combined features;
[0065] Figure 3 ROC curves of the model of the present invention and other learning methods;
[0066] Figure 4 PR curves of the model of the present invention and other learning methods;
[0067] Figure 5 The ROC curve diagram of the model of the present invention and other comparative methods was predicted by a five-fold cross-validation experiment;
[0068] Figure 6 This is a PR curve diagram predicted by a five-fold cross-validation experiment of the model of the present invention and other comparison methods. DETAILED DESCRIPTION
[0069] Specific implementation method 1: Figure 1 As shown, the specific process of the lncRNA m5C methylation site prediction method based on deep learning in this embodiment is as follows:
[0070] Step 1: Obtain the original features of the lncRNA sequence to be tested;
[0071] Step 11: Encode the physicochemical properties and cumulative frequency information of each nucleotide in the lncRNA sequence to be tested to obtain the first encoding feature matrix of the lncRNA sequence to be tested, specifically:
[0072] First, the physicochemical properties of each nucleotide in the lncRNA sequence to be tested are encoded:
[0073] Adenine A, cytosine C, guanine G, and uracil U are represented by the coordinates (1,1,1), (0,1,0), (1,0,0), and (0,0,1), respectively, to represent the nucleotide physical and chemical properties.
[0074] Then, the cumulative frequency of bases of the lncRNA sequence to be tested is obtained
[0075] Among them, f i' is the frequency of the i'th nucleotide occurring before the i'th nucleotide position, d i' is the number of times the i'th nucleotide appears before the i'th nucleotide position, i' is the nucleotide index in the lncRNA to be tested;
[0076] Finally, the physicochemical property code of each nucleotide in the lncRNA sequence to be tested and the corresponding base cumulative frequency are combined into the nucleotide first coding feature vector, thereby obtaining the first coding feature matrix of the lncRNA sequence to be tested;
[0077] Step 1 and 2: Encode the dinucleotide chemical properties of the lncRNA sequence to be tested to obtain the dinucleotide chemical property feature vector of the lncRNA sequence to be tested:
[0078] First, obtain the PCP matrix of the lncRNA sequence to be tested;
[0079] Then, normalize the PCP matrix to obtain the normalized PCP matrix, and use the normalized PCP matrix to obtain the autocovariance matrix A' and the cross-covariance matrix C':
[0080] Finally, the matrix A' and matrix C' are concatenated to obtain the dinucleotide chemical property feature vector of the lncRNA sequence to be tested;
[0081] Step 13: Concatenate the first coding feature vector and the dinucleotide chemical property feature vector of the lncRNA sequence to be tested to obtain the original feature vector of the lncRNA sequence to be tested.
[0082] Step 2: Input the original features of the lncRNA sequence to be tested into the trained lncRNA m5C methylation site prediction model to obtain the m5C site prediction results;
[0083] If the m5C site prediction result is greater than 0.5, it means that the m5C site exists in the lncRNA sequence to be tested; otherwise, it means that there is no m5C site;
[0084] The trained lncRNA m5C methylation site prediction model is obtained by the following method:
[0085] S1. Obtain the lncRNA sequence fragments corresponding to the methylation data file in the human gene library, and use the lncRNA sequence fragments to construct the training set and test set, specifically:
[0086] S11. Obtaining the site information of lncRNA m5C in the RMBase v3.0 database or the m5C-Atlas database, and merging the site information of lncRNA m5C in the RMBase v3.0 database with the site information of lncRNA m5C in the corresponding m5C-Atlas database to obtain a total set of lncRNA m5C site information;
[0087] The site information of the lncRNA m5C includes: chromosome number, start position, end position, positive and negative strand identification;
[0088] S12. Based on the site information of lncRNA m5C, multiple lncRNA sequence fragments were extracted from the human genome, and then sequence fragments with a length of 101 nt were intercepted. Sequence fragments with high similarity to other sequences were eliminated using the CD-HIT tool (a threshold of 0.9 was set in the CD-HIT tool in the present invention, and sequence fragments exceeding the set threshold were considered to have high similarity). The resulting set of sequence fragments was used as the positive sample set;
[0089] S13. Sequence fragments with a length of 101 nt centered on cytosine (C) are intercepted upstream and downstream of the positive sample site. Sequence fragments with high similarity to other sequences (a threshold of 0.9 is set in the CD-HIT tool in the present invention, and sequence fragments exceeding the set threshold are considered to have high similarity) and fragments overlapping with the positive sample site are eliminated using the CD-HIT tool. The resulting sequence fragment set is used as the negative sample set.
[0090] S14, setting 0 as the label of the negative sample and 1 as the label of the positive sample, then combining the positive sample set and the negative sample set after setting the label to obtain a data set, and dividing the data set into a training set and a test set;
[0091] Among them, label 0 indicates that there is no m5C site in the lncRNA sequence; label 1 indicates that there is an m5C site in the lncRNA sequence;
[0092] In this step, considering the imbalance of positive and negative samples in actual RNA sequences, the positive-negative sample ratio in the experimental dataset is approximately 1:3, and a five-fold cross-validation method is used for training on the training dataset.
[0093] S2. Sequence encode the data in the training set and test set to obtain the original feature vector of lncRNA, and obtain the encoded training set and test set, specifically:
[0094] S21. Encode the nucleotide physical and chemical properties and cumulative frequency information of the sequence fragments in the training set and the test set to obtain the first encoding feature matrix:
[0095] First, obtain the physical and chemical property encoding of each nucleotide in the sequence fragments in the training set and test set;
[0096] Adenine A, cytosine C, guanine G, and uracil U are represented by the coordinates (1,1,1), (0,1,0), (1,0,0), and (0,0,1), respectively, to represent the nucleotide physical and chemical properties.
[0097] In this step, RNA is composed of four nucleotides: A, C, G and U. According to the different chemical properties, they can be divided into three categories. In terms of the number of rings, adenine and guanine have two rings, and cytosine and uracil have one ring. In terms of chemical function, adenylate and cytosine contain amino groups, while guanine and uracil contain ketone groups. For the formation of secondary structure, guanine and cytosine have strong hydrogen bonds, while adenine and uracil have weak hydrogen bonds. The present invention uses three coordinates (x, y, z) to represent the chemical properties of the four nucleotides, the x coordinate represents the ring structure, the y coordinate represents the functional group, and the z coordinate represents the hydrogen bond, and assigns 0 and 1 to these three coordinates; the nucleotide S at position i in the RNA sequence i Use S i =(x i ,y i ,z i )coding;
[0098] Then, the cumulative base frequencies of the sequence fragments in the training set and the test set are obtained
[0099] Among them, f i is the frequency of the i-th nucleotide appearing before the i-th nucleotide position, d i is the number of times the i-th nucleotide occurs before the i-th nucleotide position;
[0100] This step encodes the nucleotide accumulation frequency, which is defined as the frequency with which the nucleotide at position i appears before position i; the accumulation frequency f of the nucleotide at position i is i =d i / i, where d i is the sum of the number of times the i-th nucleotide appears in the previous i nucleotides;
[0101] Finally, the physical and chemical property code of each nucleotide and the corresponding base cumulative frequency are combined into the first coding feature vector of the current nucleotide, thereby obtaining the first coding feature matrix corresponding to the sequence fragment;
[0102] Among them, if the sequence segment length is L, the first encoding feature matrix is obtained as a 4*L-dimensional matrix;
[0103] S22. Encode the dinucleotide chemical properties of the sequence fragment to obtain a dinucleotide chemical property feature vector of the sequence fragment:
[0104] First, get the PCP matrix of the sequence segment:
[0105]
[0106] Where p∈[1,10], p is the PCP label, PCP p (Pj P j+1 ) is the dinucleotide P j P j+1 The p-th PCP value of , L is the length of the sequence segment, j∈[1,L-1], j is the nucleotide index;
[0107] Then, the PCP matrix is normalized to obtain the normalized PCP matrix, and the normalized PCP matrix is used to obtain the dinucleotide autocovariance matrix A with a dimension of 10*4 and the cross-covariance matrix C with a dimension of 10*9*4:
[0108]
[0109] Where s=1,2,3,4, λ1 and λ2 are integers. is the average value of the λ1th row data in the normalized PCP matrix, is the average value of the λ2th row data in the normalized PCP matrix, PCP' p (P j P j+1 ) is the normalized dinucleotide P j P j+1 The p-th PCP value of is the average value of the p-th row data in the normalized PCP matrix;
[0110] Finally, matrix A and matrix C are concatenated to obtain the dinucleotide chemical property feature vector.
[0111] S23. Concatenate the first encoding feature vector obtained in S21 and the dinucleotide chemical property feature vector obtained in S22 to obtain the original feature vector of the lncRNA sequence, and use the original feature vector and the label of the lncRNA sequence to form an encoded training set and test set.
[0112] S3, using the encoded training set to train the lncRNA m5C methylation site prediction model to obtain a trained lncRNA m5C methylation site prediction model;
[0113] The lncRNA m5C methylation site prediction model includes: a CNN network module, a Bi-LSTM network module, a feature fusion module, and a fully connected module;
[0114] The CNN network module is used to perform preliminary feature extraction on the original feature vector of the lncRNA sequence, ignoring unimportant features, and obtaining local feature information F1 of the sequence fragment;
[0115] The CNN network module includes three convolution units in sequence; each convolution unit includes: a convolution layer, an activation function, and a pooling layer;
[0116] The convolutional layer of the first convolutional unit is specifically:
[0117]
[0118] Among them, X u+m,n is the original feature vector of lncRNA, u is the output position index, k is the kernel index, conv(X) uk is the output of the convolutional layer, It is an M*N weight matrix, where M is the window size and N is the number of input channels;
[0119] The convolutional layers in other convolutional units are similar to the above formula, and the input is the output of the previous convolutional unit;
[0120] The activation function ReLU is expressed as:
[0121]
[0122] Among them, x is the input of the activation function; the input of each activation function is x is the output of the previous convolutional layer;
[0123] The pooling layer retains the extracted content while reducing its feature dimension, specifically:
[0124] pooling(x') u' =max(x'1,x'2,x'3,...,x' M' ) u'
[0125] Among them, x' represents the input feature, the input x' of each pooling layer is the output of the previous activation function, u' represents the output position index, M' represents the pool window size, x'1, x'2, x'3, ..., x' M' is the sequence of outputs of the previous activation function.
[0126] The Bi-LSTM network module uses feature F1 to obtain sequence features F2 with contextual interdependence between sequence nucleotides, which contains sequence order dependency information;
[0127] Bi-LSTM extracts features from two directions, making dependency information extraction more complete. F1 undergoes further feature extraction through the Bi-LSTM network layer to obtain the sequence feature F2.
[0128] The feature fusion module is used to concatenate feature F1 and feature F2 to obtain the optimal feature vector F of the lncRNA sequence;
[0129] The fully connected layer performs a comprehensive analysis and nonlinear transformation on F through a series of weight matrices, maps it to the classification space, and outputs the predicted probability of whether the m5C site exists, thereby determining whether the input sequence has an m5C site.
[0130] If the predicted probability of the m5C site is greater than 0.5, it means that the m5C site exists in the current sequence fragment; otherwise, it means that it does not exist.
[0131] S4. Use the test set to evaluate the trained lncRNA m5C methylation site prediction model, using precision, F1-score, recall, accuracy, and Matthews correlation coefficient (MCC) as evaluation indicators. The specific calculation formula is as follows:
[0132]
[0133] TP stands for true positive, meaning the sample's actual label is positive and the predicted label is positive; TN stands for true negative, meaning the sample's actual label is negative and the predicted label is negative; FP stands for false positive, meaning the sample's actual label is negative and the predicted label is positive; and FN stands for false negative, meaning the sample's actual label is positive and the predicted label is negative. In the F1-Score, P stands for Precision and R stands for Recall. The Matthews Correlation Coefficient (MCC) indicates the correlation between the prediction and the label (MCC = 0 represents random guesswork, 1 represents a perfect model).
[0134] Example: In order to verify the beneficial effects of the present invention, the present invention conducted the following experiments:
[0135] A total of 2435 human lncRNA m5C methylation sites were collected from the RMBase v3.0 database and the m5C-Atlas database as positive samples. Based on the site information, non-m5C sites on the human genome were intercepted as negative samples. Considering that non-m5C sites are far more numerous than m5C sites in actual sequences, more non-m5C site sequences were intercepted for the experiment, resulting in a total of 5703 data samples being screened.
[0136] Taking a lncRNA sequence as an example, a 50-nt sequence was taken from both the start and end of the RNA sequence. The sequence was encoded using the physicochemical properties and cumulative frequency of nucleotides and the chemical properties of dinucleotides. The encoded results were then fed into a convolutional neural network to capture local information. A bidirectional long short-term memory (Bi-LSTM) network was then used to capture contextual features. The two were combined to predict lncRNA m5C sites.
[0137] In order to deeply explore the impact of the amount of information contained in the features on the experimental results, this example conducts a detailed comparative analysis of single features and combined features. The experimental results are clearly presented in Figure 2 In. From Figure 2 As can be intuitively observed, while both feature encoding methods have their own strengths and weaknesses in these performance metrics, both achieve good performance. However, the combined feature significantly improves all performance metrics, including ACC, Recall, Pre, F1, and MCC, compared to single features. This series of concrete data fully demonstrates that combining different types of features can aggregate information from various aspects, resulting in more comprehensive and accurate results.
[0138] In order to fully and deeply reveal the significant advantages of the present invention, a comprehensive comparison is conducted between the present invention and other learning algorithms, including logistic regression, naive Bayes and support vector machine machine learning methods, as well as CNN and Bi-LSTM deep learning methods. In order to more intuitively demonstrate the balance between accuracy and recall of different learning methods, ROC curves and PR curves are plotted. Figure 3 and Figure 4 As shown in the figure, it can be seen at a glance that the performance of other learning algorithms is relatively poor, and the experimental results obtained are at a low level. Among all the learning methods involved in the comparison, the present invention has achieved a high level in AUROC and AUPR, showing the most excellent prediction results, and can classify various types of samples more accurately.
[0139] In order to demonstrate the outstanding superiority of the model constructed by the present invention, it was compared with other existing cutting-edge prediction models published, including iRNAm5C-PseDNC, iRNA-PseTNC, RNAm5CFinder, and m5CPred-SVM. The significant advantages of the present invention are illustrated by comparison with known methods, and the comparison results are shown in Table 1. The accuracy (ACC) of the present invention is 0.8457, the precision (Pre) is 0.8260, the recall rate (Recall) reaches 0.6427, the F1 score is 0.7189, and the Matthews correlation coefficient (MCC) is 0.6186. Compared with the existing models, the present invention is 2.82%, 6.76%, 11.09%, 8.12%, and 8.11% higher than the second highest value in these performance indicators. Obviously, the present invention has made great progress. In addition to the above-mentioned indicators for evaluating the robustness of the model, the ROC curves and PR curves of the present invention and other prediction methods are also plotted to further verify the prediction performance, as shown in Table 1. Figure 5 and Figure 6By comparing the AUROC and AUPR values of the corresponding areas under the ROC curve and the PR curve, it can be seen that the prediction results of the present invention on the training data set are higher than those of the other four prediction methods, indicating that the present invention can improve the robustness of the model.
[0140] Table 1 Performance comparison with other related models
[0141]
[0142] The detailed description of the embodiments of the present invention provided above is not intended to limit the scope of the invention as claimed, but rather merely represents selected embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative work are within the scope of protection of the present invention.
Claims
1. A deep learning-based lncRNA m5C methylation site prediction method, characterized by The specific process of the method is: Step 1: Obtain the original features of the lncRNA sequence to be tested; Step 2: Input the original features of the lncRNA sequence to be tested into the trained lncRNA m5C methylation site prediction model to obtain the m5C site prediction results; If the m5C site prediction result is greater than the preset threshold, it means that the m5C site exists in the lncRNA sequence to be tested; otherwise, it means that the m5C site does not exist in the lncRNA sequence to be tested; The trained lncRNA m5C methylation site prediction model is obtained by the following method: S1. Obtain the lncRNA sequence fragments corresponding to the methylation data file in the human gene library, and use the lncRNA sequence fragments to construct a training set; S2. Sequence encode the data in the training set to obtain the original feature vector of lncRNA, thereby obtaining the encoded training set; S3. Use the encoded training set to train the lncRNA m5C methylation site prediction model to obtain a trained lncRNA m5C methylation site prediction model.
2. The method for predicting lncRNA m5C methylation sites based on deep learning according to claim 1, characterized in that: The step 1 of obtaining the original features of the lncRNA sequence to be tested is specifically as follows: Step 11: Encode the physicochemical properties and cumulative frequency information of each nucleotide in the lncRNA sequence to be tested to obtain a first encoding feature matrix of the lncRNA sequence to be tested; Step 1 and 2: Encode the dinucleotide chemical properties of the lncRNA sequence to be tested to obtain a dinucleotide chemical property feature vector of the lncRNA sequence to be tested; Step 13: Concatenate the first coding feature vector and the dinucleotide chemical property feature vector of the lncRNA sequence to be tested to obtain the original feature vector of the lncRNA sequence to be tested.
3. The method for predicting lncRNA m5C methylation sites based on deep learning according to claim 2, characterized in that: The physical and chemical properties and cumulative frequency information of each nucleotide in the lncRNA sequence to be tested are encoded in step 1 to obtain the first coding feature matrix of the lncRNA sequence to be tested, specifically: First, the physical and chemical properties of each nucleotide in the lncRNA sequence to be tested are encoded as follows: If the nucleotide is adenine A, the physicochemical property code is (1,1,1); if the nucleotide is cytosine C, the physicochemical property code is (0,1,0); if the nucleotide is guanine G, the physicochemical property code is (1,0,0); if the nucleotide is uracil U, the physicochemical property code is (0,0,1); Then, the cumulative frequency of bases of the lncRNA sequence to be tested is obtained Among them, f i' is the frequency of the i'th nucleotide occurring before the i'th nucleotide position, d i' is the number of times the i'th nucleotide appears before the i'th nucleotide position, i' is the nucleotide index in the lncRNA to be tested; Finally, the nucleotide physical and chemical property codes and the corresponding base cumulative frequencies in the lncRNA sequence to be tested are combined into the nucleotide first coding feature vector, thereby obtaining the first coding feature matrix of the lncRNA sequence to be tested.
4. The method for predicting lncRNA m5C methylation sites based on deep learning according to claim 3, characterized in that: The dinucleotide chemical properties of the lncRNA sequence to be tested are encoded in steps 1 and 2 to obtain a dinucleotide chemical property feature vector of the lncRNA sequence to be tested, specifically: First, obtain the PCP matrix of the lncRNA sequence to be tested: Where p∈[1,10], p is the PCP label, PCP p (P j P j+1 ) is the dinucleotide P j P j+1 The pth PCP value of , L is the length of the lncRNA sequence to be tested, j∈[1,L-1], j is the nucleotide index; Then, the PCP matrix is normalized to obtain a normalized PCP matrix, and the dinucleotide autocovariance matrix A' and the cross-covariance matrix C' are obtained using the normalized PCP matrix; Finally, the matrix A' and the matrix C' are concatenated to obtain the dinucleotide chemical property feature vector of the lncRNA sequence to be tested.
5. The method for predicting lncRNA m5C methylation sites based on deep learning according to claim 4, characterized in that: The normalized PCP matrix is used to obtain the dinucleotide autocovariance matrix A' and the cross-covariance matrix C', specifically: Where, s = 1, 2, 3, 4, λ1≠λ2, λ1 and λ2 are integers, λ1 = 1, 2, ..., 10, λ2 = 1, 2, ..., 10, is the average value of the λ1th row data in the normalized PCP matrix, is the average value of the λ2th row data in the normalized PCP matrix, PCP' p (P j P j+1 ) is the normalized dinucleotide P j P j+1 The p-th PCP value of is the average value of the p-th row data in the normalized PCP matrix.
6. The method for predicting lncRNA m5C methylation sites based on deep learning according to claim 5, characterized in that: The S1 step is to obtain the lncRNA sequence fragments corresponding to the methylation data file in the human gene library and construct a training set using the lncRNA sequence fragments, specifically: S11. Obtaining the site information of lncRNA m5C in the RMBase v3.0 database or the m5C-Atlas database, and merging the site information of lncRNA m5C in the RMBase v3.0 database with the site information of lncRNA m5C in the corresponding m5C-Atlas database to obtain a total set of lncRNA m5C site information; The site information of the lncRNA m5C includes: chromosome number, start position, end position, positive and negative strand identification; S12. Extract multiple lncRNA sequence fragments from the human genome based on the site information of lncRNA m5C, then cut off sequence fragments of a preset length L, use the CD-HIT tool to remove sequence fragments with a similarity greater than a preset threshold, and use the remaining sequence fragments as a positive sample set; S13. Sequence fragments of length L centered around cytosine C are intercepted upstream and downstream of the positive sample site. Sequence fragments with a similarity greater than a preset threshold and fragments overlapping with the positive sample site are then removed using the CD-HIT tool. The remaining sequence fragment set is used as the negative sample set. S14, setting labels for positive samples and negative samples, and combining the labeled positive sample set and negative sample set to obtain a training set; Among them, the positive sample label is 1, and the negative sample label is 0.
7. The method for predicting lncRNA m5C methylation sites based on deep learning according to claim 6, characterized in that: The step S2 is to perform sequence encoding on the data in the training set to obtain the original feature vector of the lncRNA and obtain the encoded training set, specifically: S21, encoding the nucleotide physical and chemical properties and cumulative frequency information of the sequence segments in the training set to obtain a first encoding feature matrix; S22. Encode the dinucleotide chemical properties of the sequence fragment to obtain a dinucleotide chemical property feature vector of the sequence fragment; S23. Concatenate the first encoding feature vector obtained in S21 and the dinucleotide chemical property feature vector obtained in S22 to obtain the original feature vector of the lncRNA sequence, and use the original feature vector of the lncRNA sequence and the label to form an encoded training set.
8. The method for predicting lncRNA m5C methylation sites based on deep learning according to claim 7, characterized in that: The lncRNA m5C methylation site prediction model includes: a CNN network module, a Bi-LSTM network module, a feature fusion module, and a fully connected module; The CNN network module is used to perform preliminary feature extraction on the original feature vector of the lncRNA sequence to obtain local feature information F1 of the sequence fragment; The Bi-LSTM network module uses feature F1 to obtain sequence feature F2; The feature fusion module is used to concatenate feature F1 and feature F2 to obtain the optimal feature vector F of the lncRNA sequence; The fully connected layer uses the optimal feature vector F of the lncRNA sequence to output the m5C site prediction result.
9. The method for predicting lncRNA m5C methylation sites based on deep learning according to claim 8, characterized in that: The CNN network module includes: three convolution units; Each convolution unit includes: convolution layer, activation function layer, pooling layer in sequence; The convolutional layer in the first convolutional unit is specifically: Among them, X u+m,n is the original feature vector of lncRNA, u is the output position index, k is the kernel index, conv(X) uk is the output of the convolutional layer, It is an M*N weight matrix, where M is the window size and N is the number of input channels; The convolutional layers in other convolutional units are similar to the above formula, and the input is the output of the previous convolutional unit; The activation function layer is as follows: Among them, x is the input of the activation function; the input of each activation function is x is the output of the previous convolutional layer.
10. The method for predicting lncRNA m5C methylation sites based on deep learning according to claim 9, characterized in that: The pooling layer is specifically: pooling(x') u' =max(x'1,x'2,x'3,...,x' M' ) u' Among them, x' represents the input feature, the input x' of each pooling layer is the output of the previous activation function, u' represents the output position index, M' represents the pool window size, x'1, x'2, x'3, ..., x' M' is the sequence of outputs of the previous activation function.