miRNA-Disease Prediction Method Based on Sparse Learning and Random Walk
Through sparse learning and multi-level random walk methods, the problem that existing algorithms fail to make full use of initial probability node information and topological structure information is solved, significantly improving the performance of miRNA-disease prediction model, and achieving more accurate potential correlation prediction.
Patent Information
- Application Number
- CN202210987756.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-17
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-08-17
AI Technical Summary
When predicting the association between miRNA and disease, existing random walk algorithms fail to fully consider the initial probability node information and the topological information of the similarity network, resulting in insufficient performance of the prediction model.
The method based on sparse learning and multi-level random walk is adopted to decompose and reconstruct the disease-miRNA association matrix through sparse learning to obtain richer association information, and a restart random walk algorithm is used to perform multi-level random walk in the similarity network to capture topological features to predict potential associations.
Through sparse learning and multi-level random walk method, the performance of miRNA-disease prediction model was significantly improved, with the AUC value reaching 0.9368, demonstrating the effectiveness of this method in inferring potential associations of miRNA diseases.
Smart Images

Figure CN115359903B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine learning, and particularly to a miRNA-disease prediction method based on sparse learning and random walk. Background Technique
[0002] MicroRNA (miRNA) is a non-coding RNA consisting of 19-22 nucleotides. It has been found that the dysregulation of miRNA may lead to the dysregulation of many cellular behaviors, and many complex human diseases are related to it. In particular, miRNA plays the role of an oncogene or tumor suppressor in the occurrence and metastasis of some cancers, including breast tumors, lung tumors, prostate tumors, etc.
[0003] However, using biological experiments to identify disease-related miRNAs is both expensive and time-consuming; therefore, there is an urgent need for a simple and effective computational prediction model to predict disease-related miRNAs.
[0004] Methods based on random walk have been proposed to predict the association between miRNAs and diseases. For example, Xu et al. proposed a random walk-based ranking algorithm MIRANK; Chen et al. constructed the IMCMDA algorithm to infer disease-miRNA correlations, integrating disease similarity and miRNA similarity into an inductive completion matrix to obtain prediction scores; however, the prediction models of these random walk algorithms do not fully consider the initial probability node information and the topological structure information of the similarity network. Summary of the Invention
[0005] Aiming at the deficiencies of existing algorithms, the present invention fully considers the impact of the initial probability node information in the restart random walk algorithm on the model performance, and performs potential association prediction based on capturing the topological structure features of the similarity network by using a multi-level random walk algorithm.
[0006] The technical solution adopted by the present invention is: the miRNA-disease prediction method based on sparse learning and random walk includes the following steps:
[0007] Step 1: Perform sparse learning on the disease-miRNA association matrix, and divide the association matrix into two parts: the first part is a linear combination of the original association matrix and the low-rank matrix, regarded as projecting noisy data into the low-rank space; the second part is a sparse matrix with most elements separated from the original matrix being 0, considered to be removing noise or outliers; after removing the sparse part, reconstruct a new disease-miRNA association matrix, and obtain an initial probability matrix through the reconstructed association matrix;
[0008] Since there are a large number of unvalidated association nodes in the known associations between the original diseases and miRNAs, and the random walk algorithm in previous studies cannot take into account all node information when determining the initial probability matrix;
[0009] Furthermore, specifically including:
[0010] S11. Decompose the disease-miRNA association matrix A into a linear combination of the original association matrix A and the sparse matrix E; respectively restrict A and E with the nuclear norm and the sparse norm; construct the Lagrangian function L, and solve to obtain the low-rank matrix X * ;
[0011] S12. Calculate to obtain a new miRNA-disease association matrix.
[0012] Furthermore, the new miRNA-disease association matrix is: A * = A × X * , where A is the original association matrix and X * is the low-rank matrix.
[0013] Furthermore, the formula for the initial probability matrix is:
[0014]
[0015] where nd is the number of diseases, nm is the number of miRNAs, m i is miRNA, and the disease is d j .
[0016] Step 2. Calculate the Gaussian kernel similarity between diseases and miRNAs using the reconstructed association matrix, construct a disease integrated similarity network based on disease similarity and disease Gaussian kernel similarity; construct a miRNA integrated similarity network based on miRNA similarity and miRNA Gaussian kernel similarity; integrate the disease integrated similarity network and the miRNA integrated similarity network into a heterogeneous network through the reconstructed association network;
[0017] Step 3. In order to capture the global network features and utilize the association information between network nodes, perform multi-level random walks for different purposes; first perform random walks on the miRNA similarity network and the disease similarity network respectively to obtain the topological structure similarity between diseases and miRNAs; combined with the topological structure features of the network, taking into account the association information between the nodes of the similarity network, then perform random walks on the disease and miRNA topological structure networks respectively to effectively predict the potential similarity between diseases and miRNAs.
[0018] Furthermore, specifically including:
[0019] First, introduce the restart random walk algorithm on the disease similarity network and the miRNA similarity network. Calculate the disease topological structure features RM = {Rm1, Rm2, …, Rm i} and the miRNA topological structure features RD = {Rd1, Rd2, …, Rd j} as follows:
[0020]
[0021]
[0022] Among them, t represents the iteration step, β is the decay factor of the random walk, and the initial probabilities and are defined as follows:
[0023]
[0024]
[0025] Normalize RM and RD:
[0026]
[0027]
[0028] Secondly, perform random walks simultaneously on the disease and miRNA topological structure feature networks and integrate them to obtain stable probabilities, thereby predicting the potential associations between miRNAs and diseases;
[0029] Based on the assumption that similar miRNAs tend to be associated with similar diseases, integrate the integrated similarity network, the disease similarity network, and the new miRNA-disease association network obtained by sparse learning decomposition and reconstruction into a heterogeneous network, and perform restart random walks on the miRNA topological structure network and the disease topological structure network respectively:
[0030] On the disease network:
[0031] DR t = (1 - γ) * R t-1 * SD + γ * Y (25)
[0032] Among them, DR is the score matrix calculated by performing random walks on the disease network, MR is the score matrix calculated by performing random walks on the miRNA network, SD is the disease topological structure feature after the above normalization, and R is the average of the scores of two random walks;
[0033] On the miRNA network:
[0034] MR t= (1 - γ) * SM * R t-1 + γ * Y (26)
[0035] Calculate the average value according to the disease network and miRNA network:
[0036]
[0037] Where t represents the iteration step, γ is the decay factor of random walk, Y is the initial probability matrix, and R0 = Y.
[0038] Advantages of the present invention:
[0039] 1. Decompose and reconstruct the miRNA-disease association matrix through the sparse learning method (SLM) to obtain richer association information, and at the same time obtain the initial probability matrix of the restart random walk model;
[0040] 2. Construct a heterogeneous network using the disease similarity network, miRNA similarity network, and miRNA-disease association network. According to the topological structure characteristics of diseases and miRNAs, obtain stable probability to predict potential miRNA-disease associations through the multi-layer random walk algorithm;
[0041] 3. Use global leave-one-out cross-validation (LOOCV) and 5-fold cross-validation to evaluate the model. The results show that the AUC value of global LOOCV is 0.9368, and the average AUC value of 5-fold is 0.9335.
[0042] 4. In the case study, the results show that the SLMRWMDA model is effective in inferring potential miRNA-disease associations;
[0043] 5. Apply SLMRWMDA to three high-risk human cancers for three case studies to prove the adaptability of the present model. The results show that the SLMRWMDA model helps to infer potential miRNA-disease correlations. Brief Description of the Drawings
[0044] Figure 1 is the flowchart of the miRNA-disease prediction method based on sparse learning and random walk of the present invention;
[0045] Figure 2 is the performance comparison between SLMRWMDA of the present invention and other algorithms in 5-fold CV;
[0046] Figure 3 is the performance comparison between SLMRWMDA of the present invention and other algorithms in global LOOCV. Detailed Embodiments
[0047] The present invention will be further described below in conjunction with the accompanying drawings and embodiments. This figure is a simplified schematic diagram, which only illustrates the basic structure of the present invention in a schematic manner. Therefore, it only shows the components related to the present invention.
[0048] As Figure 1 shown, the miRNA-disease prediction method based on sparse learning and random walk includes the following steps:
[0049] Step 1: Perform sparse learning on the disease-miRNA association matrix, and divide the association matrix into two parts: The first part is a linear combination of the original association matrix and the low-rank matrix, regarded as projecting the noisy data into the low-rank space; the second part is a sparse matrix in which most elements are 0 separated from the original matrix, considered to be removing noise or outliers; After removing the sparse part, a new disease-miRNA association matrix is reconstructed, and an initial probability matrix is obtained through the reconstructed association matrix;
[0050] Disease-miRNA association:
[0051] Using the latest dataset version HMDD v3.2, after removing duplicate relationships, it contains 18,733 associations between 1,208 miRNAs and 894 diseases; miRNAs that are in HMDD v3.2 but not in the databases miRBase or mirTarBase are deleted, and diseases that are in HMDD v3.2 but not in MeSH (Medical Subject Heading, Category C medical subject terms in the National Library of Medicine of the United States (http: / / www.nlm.nih.gov)) are deleted, obtaining 8,968 known associations between 374 diseases and 788 miRNAs; An adjacency matrix is used to store the known miRNA-disease association information; if there is a known association between a disease and a miRNA, it is set to 1, otherwise it is set to 0;
[0052]
[0053] Furthermore, it specifically includes:
[0054] First, decompose the disease-miRNA association matrix A into a linear combination of the original association matrix A and the low-rank matrix, and the matrix containing noise is projected into a lower-dimensional space with more information;
[0055] Then, add a sparse matrix, and the sparse matrix is regarded as the noise data filtered out from the original association matrix, where most elements are 0;
[0056] Finally, the sum of the above two parts is used as the new miRNA-disease association matrix;
[0057] The correlation matrix A is decomposed according to Equation (2) as follows:
[0058] A = A×X + E (2)
[0059] In Equation (2), A is low-rank and sparse, X is a low-rank matrix, and E is a sparse matrix; minimizing the trace norm of the matrix helps reduce the rank of the matrix, while the sparse norm can identify noise and outliers. By matrix decomposition, we get A×X + E, and then consider E as noise to be removed; we use the nuclear norm and the sparse norm to restrict A and E respectively, and define the formula as:
[0060]
[0061] where
[0062] ||X|| * = ∑ i σ i (σ i is the singular value of the matrix) (4)
[0063]
[0064] where α = 0.1 is used to balance the weights of the low-rank matrix and the sparse matrix; because the form of Equation (3) is similar to that of Robust Principal Component Analysis (RPCA); if the matrix A on the right side of A = A×X is set to be an identity matrix, then the equation can be transformed into:
[0065]
[0066] At this time, Equation (6) is a convex optimization problem with constraints; among the Iterative Thresholding (IT) algorithm, Accelerated Proximal Gradient (APG) algorithm, Exact Augmented Lagrangian Multiplier (EALM) algorithm, and Inexact Augmented Lagrangian Multiplier (IALM) algorithm, the IALM algorithm is used to transform the equation into an unconstrained optimization problem for solution to obtain more accurate results, and the Lagrangian function L is constructed;
[0067]
[0068] where μ (μ > 0) is the penalty parameter. In the equation, by updating the Lagrangian multipliers Y1 T , Y2 T and fixing other variables to minimize X, J, and E, the solutions of the equation are X * and E * ; if A ij represents the association between miRNA m i and disease d j , then is the similarity matrix of miRNA; after decomposition, we get X * and E* , we removed the noise matrix E * , and constructed a new disease-miRNA association matrix A through the linear combination of the original association matrix A and the low-rank matrix X * ; (A * * = A × X * );
[0069] After sparse learning, the obtained association information is richer. The nodes representing the association between miRNA and disease in the new association matrix are used as seed nodes, making the originally sparse initial probability matrix become rich. Based on the reconstructed association matrix A, the initial probability matrix Y is defined as follows:
[0070]
[0071] Step 2: Calculate the Gaussian kernel similarity between diseases and miRNAs using the reconstructed association matrix. Construct a disease integrated similarity network based on disease similarity and disease Gaussian kernel similarity; construct a miRNA integrated similarity network based on miRNA similarity and miRNA Gaussian kernel similarity; integrate the disease integrated similarity network and the miRNA integrated similarity network into a heterogeneous network through the reconstructed association network.
[0072] Disease similarity:
[0073] A directed acyclic graph (DAG) is used to represent the relationship between different diseases. The DAG of diseases is constructed according to MeSH (Medical Subject Headings) to obtain the semantic information of diseases;
[0074] The given disease D is represented as DAG = (D, T(D), E(D)), where T(D) represents all the ancestor nodes of disease D and itself, and E(D) represents the set of edges connecting the parent nodes and child nodes. Therefore, the semantic contribution degree D D (d) of disease d to disease D is defined as:
[0075]
[0076] where Δ = 0.5 is the semantic contribution attenuation factor;
[0077] Furthermore, the semantic value of disease D is defined as:
[0078] DV(D) = ∑ d∈T(D) D D (d) (10)
[0079] Based on the assumption that the more parts of the DAG shared by two diseases, the greater their semantic similarity, the disease d i and the disease d j The semantic similarity DS(d i , d j ) is defined as:
[0080]
[0081] where t belongs to the intersection of the DAGs shared by two diseases d i and d j ;
[0082] miRNA similarity:
[0083] Three different criteria are adopted to evaluate the similarity of miRNAs;
[0084] 1. MiRNA sequence similarity
[0085] The miRNA sequences are from the miRBase database. For each miRNA, the entire mature sequence is approximately 22 nucleotides "AUCG". Based on the sequence information, the pairwise sequence alignment function "pairwiseAlignment" in the Biostrings package is used to calculate the similarity score. In this function, the gap opening penalty is set to 5, the gap extension penalty is set to 2, the match score is set to 1, and the mismatch score is set to -1. After obtaining the sequence similarity score, it is normalized to the range [0,1] using the max-min normalization method, and the resulting miRNA sequence similarity matrix is
[0086] 2. MiRNA functional similarity
[0087] Based on the hypothesis that miRNAs with similar functions are more likely to be related to the same disease, the miRNA functional similarity is measured by the similarity of their associated disease DAGs; the MDA comes from HMDD, and the disease DAGs are constructed according to MeSH descriptors; the resulting miRNA functional similarity matrix is
[0088] 3. MiRNA semantic similarity
[0089] The miRNA semantic similarity is described by the miRNA target genes and gene-related Gene Ontology (GO) annotations; the miRNA target gene information comes from mirTarBase. For each pair of miRNAs, the target gene lists are maintained, and the semantic similarity between the two corresponding gene sets is calculated; the resulting miRNA semantic similarity matrix is
[0090] Gaussian interaction profile kernel similarity between diseases and miRNAs:
[0091] Based on the hypothesis that similar miRNAs are usually associated with similar diseases, the Gaussian interaction spectrum kernel similarity between diseases and miRNA pairs is calculated using a Gaussian kernel function based on the adjacency matrix; define the vector IP(d i ) to represent the interaction spectrum of diseases, and the vector IP(m j ) to represent the interaction spectrum of miRNAs; the Gaussian interaction spectrum kernel similarity matrices GD and GM for disease pairs and miRNA pairs are calculated as follows:
[0092] GD(d i , d j ) = exp(-θ d ||IP(d i ) - IP(d j )|| 2 ) (12)
[0093] GM(m i , m j ) = exp(-θ m ||IP(m i ) - IP(m j )|| 2 ) (13)
[0094] where the adjustment coefficients θ d and θ m are calculated as follows:
[0095]
[0096]
[0097] where θ′ d and θ′ m are the original bandwidths, set to a value of 1, nd is the number of diseases, and nm is the number of miRNAs.
[0098] Integrated similarity of diseases and miRNAs:
[0099] For diseases, construct the integrated similarity matrix Dsim to integrate the semantic similarity and Gaussian interaction spectrum kernel similarity of diseases:
[0100]
[0101] For miRNAs, construct the integrated similarity matrix Msim to integrate the similarity of miRNAs (see Equation 18) and the Gaussian interaction spectrum kernel similarity:
[0102]
[0103]
[0104] Step 3: To capture global network features and utilize the correlation information between network nodes, perform multi-level random walks for different purposes. First, perform random walks on the miRNA similarity network and the disease similarity network respectively to obtain the topological structure similarities of diseases and miRNAs. Combining the topological structure features of the network and fully considering the correlation information between the nodes of the similarity network, then perform random walks on the disease and miRNA topological structure networks respectively, which can more effectively predict the potential similarity between diseases and miRNAs.
[0105] Further, it specifically includes:
[0106] Different from the traditional double random walk, two different-purpose random walks are performed in the heterogeneous network composed of the disease similarity network and the miRNA similarity network through the disease-miRNA association matrix;
[0107] First, introduce the restart random walk algorithm on the miRNA similarity network and the disease similarity network to obtain the network topological feature matrices of diseases and miRNAs, and then standardize them;
[0108] Further, it specifically includes:
[0109] First, introduce the restart random walk algorithm on the disease similarity network and the miRNA similarity network. Through the restart random walk algorithm, global network features can be fully captured and the correlation information between network nodes can be fully utilized; the disease topological structure features RM = {Rm1, Rm2, …, Rm 788} and the miRNA topological structure features RD = {Rd1, Rd2, …, Rd 374} are calculated as follows:
[0110]
[0111]
[0112] where t represents the iteration step, β (0 < β < 1) is the decay factor of the random walk, and the initial probabilities and are defined as follows:
[0113]
[0114]
[0115] Perform standardization processing on RM and RD:
[0116]
[0117]
[0118] Next, perform random walks simultaneously on the disease and miRNA topological structure feature networks and integrate them to obtain stable probabilities, thereby predicting potential miRNA-disease associations;
[0119] Based on the assumption that similar miRNAs tend to be associated with similar diseases, integrate the integrated similarity network, disease similarity network, and the new miRNA-disease association network obtained after decomposition and reconstruction through sparse learning into a heterogeneous network, and perform restart random walks on the miRNA topological structure network and the disease topological structure network respectively;
[0120] On the disease network:
[0121] DR t =(1 - γ)*R t-1 *SD + γ*Y (25)
[0122] where DR is the score matrix calculated by performing random walks on the disease network, MR is the score matrix calculated by performing random walks on the miRNA network, SD is the disease topological structure feature after the above normalization, and R is the average of the scores of two random walks;
[0123] On the miRNA network:
[0124] MR t =(1 - γ)*SM*R t-1 + γ*Y (26)
[0125] Take the average:
[0126]
[0127] where t represents the iteration step, γ(0 < γ < 1) is the decay factor of random walks, Y is the initial probability matrix, and R0 = Y.
[0128] Experimental results:
[0129] Two cross-validations were used to evaluate the performance of the SLMRWMDA model of the present invention, namely, 5-fold cross-validation and global leave-one-out cross-validation; the SLMRWMDA model was compared with five other existing technical methods such as IMCMDA, MDHGI, BRWH, MCLPMD, and LPLNS; in the 5-fold cross-validation, all known miRNA-disease associations were first randomly divided into 5 subsets of the same size. Each subset was used as a test sample in turn, and the remaining four subsets were used as training samples. Then, the scores of each test sample in the test set were compared with all candidate score samples. Finally, the score rankings of the test samples were obtained. In the global leave-one-out validation, each pair of known disease-miRNA associations was used as a test sample in turn, and the remaining miRNA-diseases were used as training samples; similarly, the score rankings were finally calculated. Finally, the prediction performance was evaluated by calculating the receiver operating characteristic (ROC) curve and calculating the area under the ROC curve (AUC) (such as Figure 2 , 3 ); Table 1 shows the AUC results of the 5-fold cross-validation and the global leave-one-out cross-validation. It can be seen that the performance of this method is the best; among them, the AUC of the global leave-one-out is 0.9368, and the AUC of the 5-fold cross-validation is 0.9350;
[0130] In addition, in order to verify the effectiveness of the method of reconstructing the association matrix of the present invention to determine the initial probability matrix, ablation experiments were carried out. The AUC of the 5-fold cross-validation obtained without considering the initial probability matrix was 0.9228. From this, it can be seen that the consideration of the initial probability nodes by the SLMRWMDA model of the present invention is effective.
[0131] Table 1 Comparison of SLMRWMDA with Five Other Algorithms
[0132] Model Name 5-fold Cross-validation Leave-One-Out Cross-validation IMCMDA 0.7687 0.7767 MDHGI 0.8783 0.8955 BRWH 0.8589 0.8629 MCLPMDA 0.9291 0.9305 LPLNS 0.9183 0.9178 This Model 0.9350 0.9368
[0133] More and more evidence shows that many miRNAs play important roles in the occurrence and development of cancer, and the expression levels of certain miRNAs are closely related to the occurrence, recurrence, metastasis, etc. of tumors; therefore, in order to verify the ability of MDBRWRMDA to predict unknown diseases, three case studies were carried out for verification; the first type of case study used the known disease-miRNA associations in HMDDV3.2 as the training data set; the potential related miRNAs of important human diseases (esophageal cancer) were scored by the SLMRWMDA model, and each candidate miRNA was ranked according to the scores. Finally, the top 50 miRNAs in the prediction list were verified using the dbDEMC and miR2Disease disease-miRNA association databases.
[0134] For the first type of cases of esophageal cancer, the results show that 9 out of the top 10 miRNAs, 19 out of the top 20 miRNAs, and 41 out of the top 50 miRNAs are verified by at least one of the databases dbDEMC (abbreviated as ①) and miR2Disease (abbreviated as ②). The results are shown in Table 2.
[0135] Table 2 Verification of the top 50 predicted esophageal cancer-related miRNAs in the databases
[0136]
[0137]
[0138] To prove the prediction ability of the MDBRWRMDA model for new diseases (i.e., diseases with no known related miRNAs), lung cancer is selected for the second case study; by resetting the miRNAs known to be related to lung cancer in the training data to unknown and treating lung cancer as a new disease. Similarly, the MDBRWRMDA model scores the potentially related miRNAs of lung cancer, sorts each candidate miRNA according to the scores, and finally uses the dbDEMC, miR2Disease, and HMDDv3.2 databases to verify the top 50 miRNAs in the prediction list. The results show (as shown in Table 3, where the first and third columns of the table record the top 1 - 25 and 26 - 50 predicted miRNAs respectively) that all of the top 10 and top 20 and 49 out of the top 50 miRNAs are verified by at least one of the databases dbDEMC (abbreviated as ①), miR2Disease (abbreviated as ②), and HMDD3.2 (abbreviated as ③).
[0139] Table 3 Verification of the top 50 predicted lung cancer-related miRNAs in the databases
[0140]
[0141]
[0142] The third type of case study is based on the data in HMDD V2.0. The MDBRWRMDA model scores the potentially related miRNAs of human breast cancer, sorts each candidate miRNA according to the scores, and finally uses the dbDEMC, miR2Disease, and HMDDV2.0 disease-miRNA association database to verify the top 50 miRNAs in the prediction list.
[0143] The results show that all of the top 10 and top 20 and 47 out of the top 50 predicted miRNAs were verified by at least one of the databases dbDEMC (abbreviated as ①), miR2Disease (abbreviated as ②), and HMDD V2.0 (abbreviated as ③), and the results are shown in Table 4 below:
[0144] Table 4 Verification of the top 50 breast cancer-related miRNAs predicted in the database
[0145]
[0146]
[0147] The present invention uses global leave-one-out cross-validation (LOOCV) and 5-fold cross-validation to evaluate the model. The results show that the AUC value of global LOOCV is 0.9368, and the average AUC value of 5-fold is 0.9335.
[0148] In the case study, the results show that SLMRWMDA is effective in inferring the potential associations between miRNAs and diseases.
[0149] Three case studies were conducted by applying SLMRWMDA to three high-risk human cancers to demonstrate the adaptability of the present model. The results show that the SLMRWMDA model helps to infer the potential correlations between diseases and miRNAs.
[0150] Taking the ideal embodiments of the present invention described above as an inspiration, through the above description, relevant staff can completely make various changes and modifications without departing from the technical idea of the present invention. The technical scope of the present invention is not limited to the content in the specification, and its technical scope must be determined according to the scope of the claims.
Claims
1. A miRNA-disease prediction method based on sparse learning and random walk, characterized in that It includes the following steps: Step 1: Perform sparse learning on the disease-miRNA association matrix, and divide the association matrix into two parts: the first part is the linear combination of the original association matrix and the low-rank matrix; the second part is to remove noise or outliers; After removing the sparse part, reconstruct to obtain a new disease-miRNA association matrix, and obtain the initial probability matrix through the reconstructed association matrix; Step 2: Calculate the Gaussian kernel similarity between diseases and miRNAs using the reconstructed association matrix, construct a disease integrated similarity network based on disease similarity and disease Gaussian kernel similarity; construct a miRNA integrated similarity network based on miRNA similarity and miRNA Gaussian kernel similarity; integrate the disease integrated similarity network and the miRNA integrated similarity network into a heterogeneous network through the reconstructed association network; Step 3: In order to capture global network features and utilize the association information between network nodes, perform multi-level random walks for different purposes; first perform random walks on the miRNA similarity network and the disease similarity network respectively to obtain the topological structure similarity between diseases and miRNAs; combine the topological structure features of the network, considering the association information between the nodes of the similarity network, and then perform random walks on the disease and miRNA topological structure networks respectively to predict the potential similarity between diseases and miRNAs; First, introduce the restart random walk algorithm on the disease similarity network and the miRNA similarity network, and calculate the disease topological structure features and the miRNA topological structure features as follows: ; Among them, t represents the iteration step, is the decay factor of the random walk, and the initial probability and are defined as follows: ; Normalize RM and RD as follows: ; Secondly, perform random walks on the disease and miRNA topological structure feature networks simultaneously and integrate to obtain a stable probability, thereby predicting the potential association between miRNAs and diseases; Perform restart random walks on the miRNA topological structure network and the disease topological structure network respectively: On the disease network: ; Among them, DR is the score matrix calculated by performing random walks on the disease network, MR is the score matrix calculated by performing random walks on the miRNA network, and R is the average of the scores of the two random walks; On the miRNA network: ; Calculate the average value according to the disease network and the miRNA network: ; Among them, t represents the iteration step, is the decay factor of the random walk, Y is the initial probability matrix, and .
2. The miRNA-disease prediction method based on sparse learning and random walk according to claim 1, wherein Step 1 specifically includes: S11. Decompose the disease - miRNA association matrix A into the original association matrix A and the sparse matrix E as a linear combination; respectively use the nuclear norm and the sparse norm to A and E perform restrictions; construct the Lagrangian function L and solve to obtain the low - rank matrix ; S12. Calculate to obtain a new miRNA-disease association matrix.
3. The miRNA-disease prediction method based on sparse learning and random walk according to claim 2, wherein The new miRNA-disease association matrix is as follows: , where A is the original association matrix, is the low-rank matrix.
4. The miRNA-disease prediction method based on sparse learning and random walk according to claim 2, wherein The formula for the initial probability matrix is: ; wherein, nd is the number of diseases, nm is the number of miRNAs, is , the disease is , is the new miRNA-disease association matrix.
Citation Information
Cited By
MiRNA-disease association prediction method and system based on multi-head attention mechanism
CN116453684A