Prediction Method for Drug-Disease Association Relationships Based on Graph Regularized Matrix Factorization

Through the method of graph regularization matrix decomposition, directed acyclic graph and drug molecule similarity network are constructed, combined with graph convolution and nuclear methods, the complexity and isomerization of drug disease association prediction in the prior art is solved, and the accurate prediction of the association relationship between drugs and diseases is achieved.

CN115985520BActive Publication Date: 2025-06-13QINGDAO UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211615901.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-15
Publication Date
2025-06-13
Estimated Expiration
2042-12-15

AI Technical Summary

Technical Problem

The existing drug-disease association prediction methods based on network reasoning have complex network structures and multiple information isomerism problems, resulting in the inability to accurately extract the association relationship between drugs and diseases.

Method used

Using a graph-regular matrix decomposition method, by constructing directed acyclic graphs and drug molecular similarity networks, combining graph convolution and nuclear methods, feature decomposition of the correlation matrix between drugs and diseases is performed, and the nearest neighbor relationship of nodes is optimized to achieve accurate prediction of the correlation relationship.

Benefits of technology

Accurate prediction of the association relationship between drugs and diseases is achieved, and the geometric structure information of each disease's characteristics and drug similarity network are fully taken into account, which improves the accuracy and efficiency of prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115985520B_ABST
    Figure CN115985520B_ABST
Patent Text Reader

Abstract

The present invention discloses a prediction method for drug-disease association relationships based on graph-regularized matrix factorization. The method of the present invention extracts the semantic similarity between diseases according to the directed acyclic graph of each disease, and then combines the association relationships between each disease and drugs in the existing database. Using the graph convolution method, it extracts disease features to construct a disease feature matrix, determines the cosine similarity between diseases and fuses it with the semantic similarity. After obtaining the disease association relationship based on the directed acyclic graph, it constructs a drug feature matrix according to the drug features in the existing database, combines it with the disease feature matrix, establishes an association matrix between drug molecules and diseases, and performs eigen-decomposition on the association matrix based on the matrix factorization algorithm of graph regularization and kernel method. By constructing an objective function, it optimizes the neighbor relationships of nodes in the drug similarity network and the disease similarity network, fully utilizes disease features and drug features, and realizes the accurate prediction of the association relationship between drugs and diseases.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of predicting drug-disease associations, and particularly to a method for predicting drug-disease association relationships based on graph-regularized matrix factorization. Background Art

[0002] With the development of technologies such as computer-aided drug design, network pharmacology, bioinformatics, and artificial intelligence, applying computer technology to the research of predicting drug-disease association relationships can pre-screen drug molecules with certain activities against certain diseases among known drugs, effectively improving the success rate of drug research and development, reducing the cost of drug research and development, and accelerating the speed of drug research and development.

[0003] Batch obtaining the relationships between known drugs and diseases based on the drug-disease association relationship network and making full use of the drug-disease association relationship network to integrate other information can effectively improve the accuracy of predicting drug-disease association relationships. In addition, the emergence of various drug and disease knowledge databases has further promoted the rapid development of new algorithms.

[0004] The network-based inference method is the most widely used method at present. HeTDR adopts a method for predicting drug-disease association relationships based on heterogeneous networks and text mining, extracts drug features using drug-related networks and extracts disease features using biomedical corpora, and combines them with the known drug-disease association network to predict the correlation between drugs and diseases. MNBDR designs a drug screening method based on module networks, uses the gene expression datasets of existing drug samples and disease samples, and uses the random walk algorithm to capture the basic modules in disease development to screen potential drugs for a given disease. DRHN constructs a computational method for heterogeneous networks, uses similarity calculations and experimentally verified drug-disease association relationships to establish a drug-disease bipartite network, iteratively updates the weights of unconnected drug-disease nodes in the network until stable, and determines the final affinity of each pair of drug-disease.

[0005] With the development of deep learning, applying deep learning to the drug-disease association network graph has become the focus of current research. Xuan et al. proposed a deep learning framework based on convolutional neural network and bidirectional long short-term memory network to obtain the original features and path features of drug-disease pairs, achieving drug repositioning. Metapath2vec learns the embedded node representations in heterogeneous networks based on metapaths and random walks, and uses the heterogeneous jump graph strategy to predict drug-disease associations. Since the drug-disease association network graph is a graph structure, relevant scholars have conducted in-depth research on graph convolution. GFPred integrates the relationships between drugs and diseases, disease similarities, and drug similarities, and proposes a fully connected prediction method based on graph convolutional autoencoders to fuse the attention mechanism to predict diseases related to drugs. Yu et al. proposed a layer attention graph convolutional network. After combining the feature encodings from multiple graph convolutional layers using the attention mechanism for different networks, they observed drug-disease associations and scored them. BiFusion uses a bidirectional graph convolutional network model to fuse heterogeneous information and improves the drug-disease prediction results through the protein interaction network, providing an accurate drug repositioning algorithm.

[0006] However, the network structures adopted by existing drug-disease association relationship prediction methods based on network inference are too complex, and they do not consider the fusion of multiple elements for prediction analysis. They cannot accurately extract the association relationship between drugs and diseases, and there are problems of heterogeneity and computational complexity caused by multiple types of information. Therefore, there is an urgent need to propose a prediction method for drug-disease association relationships based on graph-regularized matrix factorization to achieve accurate prediction of drug-disease association relationships. Summary of the Invention

[0007] In view of the problem that the current drug-disease association relationship prediction method based on network inference is difficult to accurately predict drug-disease association relationships, the present invention proposes a prediction method for drug-disease association relationships based on graph-regularized matrix factorization. It obtains disease association relationships based on a directed acyclic graph, constructs an association matrix between drug molecules and diseases by combining a drug molecular similarity network, and predicts the association between drugs and diseases based on a matrix factorization algorithm of graph regularization and kernel method, achieving accurate prediction of the association relationship between drugs and diseases.

[0008] The present invention adopts the following technical solutions:

[0009] A prediction method for drug-disease association relationships based on graph-regularized matrix factorization, characterized by comprising the following steps:

[0010] Step 1: According to the classification information of diseases, construct a directed acyclic graph (DAG) for each disease respectively. Extract the semantic similarity between diseases based on the DAGs of each disease, and then combine with the existing database to obtain the association relationship between each disease and drugs. Use the graph convolution method to extract disease features in the DAG of each disease, construct a disease feature matrix, calculate the cosine similarity between diseases, and fuse the semantic similarity and cosine similarity between diseases to obtain the disease association relationship based on the DAG.

[0011] Step 2: Extract the drug features of each drug molecule in the existing database to obtain a drug feature matrix, and calculate the cosine similarity between each drug feature to obtain a drug molecule similarity network.

[0012] Step 3: Establish an association matrix between drug molecules and diseases according to the disease feature matrix and the drug feature matrix.

[0013] Step 4: Based on the matrix factorization algorithm of graph regularization and kernel method, perform eigen-decomposition on the association matrix between drug molecules and diseases. Combine with the drug-disease relationship graph network in the existing database to construct an objective function, optimize the neighbor relationship of nodes in the drug similarity network and the disease similarity network, and predict the association between drug molecules and diseases.

[0014] Preferably, in the said Step 1, it specifically includes the following steps:

[0015] Step 1.1: According to the classification information of diseases, construct a directed acyclic graph for each disease respectively, and obtain the semantic values of all nodes in the DAG of each disease.

[0016] There are multiple nodes in the said directed acyclic graph. Take the disease d itself as the child node and the diseases related to the disease d as the parent nodes. Construct a directed acyclic graph for each disease respectively. The directed acyclic graph is expressed as:

[0017] DAG(d) = (N(d), E(d)) (1)

[0018] In the formula, d is the disease name, DAG(·) is the directed acyclic graph of the disease, N(·) is the parent nodes related to the disease in the directed acyclic graph, and E(·) is the connection relationship between the parent nodes and the child nodes in the directed acyclic graph.

[0019] Calculate the semantic values of all nodes in the DAG of each disease respectively, as shown in formula (2):

[0020]

[0021] In the formula, n is the node number, n' is the child node of the node n, C d(·) is the semantic value of the node with respect to disease d, and Δ is the semantic contribution factor;

[0022] Determine the semantic value of each disease according to the semantic values of all nodes in each disease directed acyclic graph, as shown in formula (3):

[0023] DV(d) = ∑ n∈N(d) C d (n) (3)

[0024] In the formula, DV(·) is the semantic value of the disease;

[0025] Step 1.2, calculate the semantic similarity between diseases according to the semantic values of each disease, as shown in formula (4):

[0026]

[0027] In the formula, is the semantic similarity between disease i and disease j, x is the node that both the directed acyclic graphs of disease i and disease j contain, is the semantic value of node x with respect to disease i, is the semantic value of node x with respect to disease j, is the semantic value of node n with respect to disease d i DV(d i ) is the semantic value of disease i, DV(d j ) is the semantic value of disease j;

[0028] Step 1.3, obtain the drugs used to treat each disease based on the existing database, determine the association relationship between the disease and the drug for each disease respectively, extract the association relationship between the disease and the drug and use it as a descriptor (the descriptor is a binary vector, and the length of the descriptor is the number of drugs in the database). If the disease is associated with the drug, set the value of the descriptor to 1. If there is no association between the disease and the drug, set the value of the descriptor to 0;

[0029] For each disease respectively, use the graph convolution method to extract disease features in each directed acyclic graph, construct a disease feature matrix, and determine the eigenvalue of each disease feature, as shown in formula (5):

[0030]

[0031] In the formula, X d is the descriptor value, is the eigenvalue after convolution calculation of all nodes in the directed acyclic graph, GCN(·) is the graph convolution function, V d is the eigenvalue of the disease feature, Pool(·) is the aggregation function;

[0032] Step 1.4, calculate the cosine similarity between diseases according to the characteristics of each disease, as shown in formula (6):

[0033]

[0034] In the formula, is the cosine similarity between disease i and disease j, m is the number of disease characteristics, M is the total number of disease characteristics, is the eigenvalue of the m-th disease characteristic with respect to disease i, is the eigenvalue of the m-th disease characteristic with respect to disease j;

[0035] Step 1.5, fuse the semantic similarity and cosine similarity between diseases to obtain the disease association relationship based on the directed acyclic graph, as shown in formula (7):

[0036]

[0037] In the formula, is the association degree between disease i and disease j, and α is the fusion coefficient.

[0038] Preferably, in step 1.3, the descriptor is a binary vector, and the length of the descriptor is the number of drugs in the database.

[0039] Preferably, in step 2, according to the molecular structures of all drugs in the existing database, determine the drug molecules contained in the database, extract the drug characteristics of all drug molecules in the existing database using Morgan fingerprints, establish a drug characteristic matrix, and calculate the cosine similarity between the drug characteristics in the drug characteristic matrix to obtain a drug molecule similarity network.

[0040] Preferably, in step 3, the association matrix between drug molecules and diseases is:

[0041]

[0042] In the formula, Y is the association matrix between drug molecules and diseases, D is the disease characteristic matrix, and G is the drug molecule characteristic matrix.

[0043] Preferably, in step 4, use the kernel method to perform eigenvalue decomposition on the association matrix between drug molecules and diseases to obtain a drug similarity matrix A and a disease similarity matrix B. The association function between the drug similarity matrix A and the disease similarity matrix B is:

[0044] f(I,J) = ∑ m,n λ m,n κ G (<a I ,a m >)κ D(<b J , b n >)(9)

[0045] Where f is the correlation function between the drug similarity matrix A and the disease similarity matrix B, I is the row number of the vector in the correlation matrix Y, used to represent the drug molecule name in the correlation matrix, J is the column number of the vector in the correlation matrix Y, used to represent the disease name in the correlation matrix; a m is the value of the m-th row in the drug similarity matrix A, b n is the value of the n-th row in the disease similarity matrix B, κ g is the drug vector kernel function, κ D is the disease vector kernel function, a I is the value of the I-th row in the drug similarity matrix A, b J is the value of the J-th row in the disease similarity matrix B;

[0046] Based on the Kronecker least squares method, using the Kronecker product of the drug vector kernel function and the disease vector kernel function to accelerate the calculation process, perform eigenvalue decomposition on the drug vector kernel function and the disease vector kernel function respectively, and obtain the correlation function between the drug molecule and the disease as:

[0047]

[0048] Among them,

[0049]

[0050] In the formula, Q G is the drug similarity matrix, Q D is the disease similarity matrix, T is the transpose matrix;

[0051] According to the correlation between the drug molecule and the disease in the existing database drug-disease relationship graph network, merge the drug similarity matrix and the disease similarity matrix obtained by eigenvalue decomposition to construct the objective function, as shown in formula (12):

[0052]

[0053] In the formula, y m,n is the relationship between the drug molecule m and the disease n in the existing database, S G* is the neighbor similarity matrix of the drug, S D* is the neighbor similarity matrix of the disease, is the norm of the correlation function f in the Hilbert space related to the kernel K, and λ and β are both regularization parameters;

[0054] Construct a drug similarity network and a disease similarity network based on the drug similarity matrix and the disease similarity matrix. Based on graph regularization, process the drug similarity network and the disease similarity network, respectively retain the geometric structure information of the adjacent nodes of each node in the drug similarity network and the disease similarity network, and calculate the drug similarity weight and the disease similarity weight, as shown in formula (13):

[0055]

[0056] In the formula, W G is the drug similarity weight, and W D is the disease similarity weight. N p (·) is the set of nodes around node P;

[0057] Calculate the drug neighbor similarity matrix and the disease neighbor similarity matrix based on the drug similarity weight and the disease similarity weight, as shown in formula (14):

[0058]

[0059] In the formula, S G* is the drug neighbor similarity matrix, and S G is the drug similarity matrix; S D* is the disease neighbor similarity matrix, and S D is the disease similarity matrix;

[0060] Since the geometric structures of the drug neighbor similarity matrix and the drug similarity matrix, and the disease neighbor similarity matrix and the disease similarity matrix are consistent, based on the drug neighbor similarity matrix and the disease neighbor similarity matrix, use the constructed objective function to predict the association between drug molecules and diseases.

[0061] The present invention has the following beneficial effects:

[0062] The method of the present invention proposes a prediction method for the association relationship between drugs and diseases based on graph-regularized matrix factorization. Obtain the semantic similarity between diseases according to the directed acyclic graph of diseases, fuse the semantic similarity of diseases with the cosine similarity, and obtain the association relationship between diseases based on the disease characteristics and the similarity calculation method, fully considering the characteristics of each disease.

[0063] Meanwhile, the method of the present invention also decomposes the association relationship between drugs and diseases based on the graph regularization matrix, combines the matrix decomposition method based on the kernel method, fully excavates the influence of diseases and drugs on the matrix decomposition of the drug-disease association matrix, takes the neighbor relationship of nodes in the drug similarity network and the disease similarity network as the optimization objective, preserves the original geometric structure information of nodes in the network during the decomposition process, and cooperates with the Kronecker least squares method to accelerate the calculation rate, realizing the accurate prediction of the association relationship between drugs and diseases. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 It is a schematic diagram of a method for predicting the association relationship between drugs and diseases based on graph regularization matrix decomposition.

[0065] Figure 2 It is a disease association relationship graph based on a directed acyclic graph. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0066] The following further describes the specific embodiments of the present invention with reference to the drawings:

[0067] The present invention proposes a method for predicting the association relationship between drugs and diseases based on graph regularization matrix decomposition, as Figure 1 shown, which includes the following steps:

[0068] Step 1, according to the classification information of diseases, construct a directed acyclic graph (DAG) for each disease respectively, extract the semantic similarity between diseases based on the directed acyclic graph of each disease, and then combine the existing database to obtain the association relationship between each disease and drugs. Use the graph convolution method to extract disease features in the directed acyclic graph of each disease, construct a disease feature matrix, calculate the cosine similarity between diseases, and fuse the semantic similarity and cosine similarity between diseases to obtain the disease association relationship based on the directed acyclic graph, as Figure 2 shown, which specifically includes the following steps:

[0069] Step 1.1, according to the classification information of diseases, construct a directed acyclic graph for each disease respectively, and obtain the semantic values of all nodes in the directed acyclic graph of each disease.

[0070] There are multiple nodes in the directed acyclic graph. Take the disease d itself as a child node and the diseases related to the disease d as parent nodes. The unidirectional edges connecting the nodes correspond to the association relationships between diseases. Construct a directed acyclic graph for each disease respectively. The directed acyclic graph is expressed as:

[0071] DAG(d) = (N(d), E(d)) (1)

[0072] In the formula, d is the disease name, DAG(·) is the directed acyclic graph of the disease, N(·) is the parent node related to the disease in the directed acyclic graph, and E(·) is the connection relationship between the parent node and the child node in the directed acyclic graph.

[0073] Calculate the semantic values of all nodes in the directed acyclic graph of each disease respectively, as shown in formula (2):

[0074]

[0075] In the formula, n is the node number, n' is the child node of node n, C d (·) is the semantic value of the node relative to the disease d, and Δ is the semantic contribution factor.

[0076] Determine the semantic values of each disease respectively according to the semantic values of all nodes in the directed acyclic graph of each disease, as shown in formula (3):

[0077] DV(d) = ∑ n∈N(d) C d (n) (3)

[0078] In the formula, DV(·) is the semantic value of the disease.

[0079] Step 1.2, when two diseases have a large number of ancestor nodes, it proves that there is a high semantic similarity between these two diseases. Therefore, calculate the semantic similarity between diseases according to the semantic values of each disease, as shown in formula (4):

[0080]

[0081] In the formula, is the semantic similarity between disease i and disease j, x is the node that the directed acyclic graphs of disease i and disease j both contain, C di (x) is the semantic value of node x relative to disease i, is the semantic value of node x relative to disease j, is the semantic value of node n relative to disease d i of, DV(d i ) is the semantic value of disease i, DV(d j ) is the semantic value of disease j.

[0082] Step 1.3. Since the semantic similarity between diseases alone cannot deeply mine the association relationships between diseases, the technical solution of this application introduces the relationship between diseases and drugs. In this embodiment, drugs used to treat each disease are obtained based on the ComparativeToxicogenics database, which contains the association relationships between 708 known drugs and 5,603 diseases, and each drug can be associated with at least one disease. The association relationships between diseases and drugs are determined for each disease respectively, and the association relationships between diseases and drugs are extracted and used as descriptors. The descriptor is a binary vector, and the length of the descriptor is the number of drugs in the database; if a disease is associated with a drug, the value of the descriptor is set to 1, and if there is no association between the disease and the drug, the value of the descriptor is set to 0.

[0083] For each disease respectively, the disease features are extracted from each directed acyclic graph by using the graph convolution method, a disease feature matrix is constructed, and the eigenvalues of each disease feature are determined, as shown in formula (5):

[0084]

[0085] In the formula, X d is the descriptor value, is the eigenvalue after convolution calculation of all nodes in the directed acyclic graph, GCN(·) is the graph convolution function, V d is the eigenvalue of the disease feature, and Pool(·) is the aggregation function.

[0086] Step 1.4. According to each disease feature, the cosine similarity between diseases is calculated, as shown in formula (6):

[0087]

[0088] In the formula, is the cosine similarity between disease i and disease j, m is the number of disease features, M is the total number of disease features, is the eigenvalue of the m-th disease feature relative to disease i, is the eigenvalue of the m-th disease feature relative to disease j.

[0089] Step 1.5. The semantic similarity and cosine similarity between diseases are fused to obtain the disease association relationship based on the directed acyclic graph, as shown in formula (7):

[0090]

[0091] In the formula, is the association degree between disease i and disease j, and α is the fusion coefficient.

[0092] Step 2: According to the molecular structures of the drugs in the Comparative Toxicogenics database, determine the drug molecules contained in the Comparative Toxicogenics database. Use the Morgan fingerprint to extract the drug features of all drug molecules in the existing database, establish a drug feature matrix, and calculate the cosine similarity between the drug features in the drug feature matrix to obtain a drug molecule similarity network. In this embodiment, constructing a drug molecule similarity network by using the Morgan fingerprint to extract the features of drug molecules based on the existing database is a prior art in the field.

[0093] Step 3: According to the disease feature matrix and the drug feature matrix, establish an association matrix between drug molecules and diseases, as shown in formula (8):

[0094]

[0095] In the formula, Y is the association matrix between drug molecules and diseases, D is the disease feature matrix, and G is the drug molecule feature matrix.

[0096] Step 4: Perform eigenvalue decomposition on the association matrix between drug molecules and diseases based on the matrix decomposition algorithm of graph regularization and kernel method to predict the association relationship between drugs and diseases. For the standard non - negative matrix factorization, the purpose is to find two low - rank factorization matrices, and the product between them should be as close as possible to the original matrix. In the prediction of the association relationship between drugs and diseases, these two matrices are the drug similarity matrix and the disease similarity matrix respectively.

[0097] To avoid overfitting and improve the accuracy of the prediction results, the kernel method is introduced into the decomposition process of the association matrix between drug molecules and diseases. Use the kernel method to perform eigenvalue decomposition on the association matrix between drug molecules and diseases to obtain the drug similarity matrix A and the disease similarity matrix B. The association function between the drug similarity matrix A and the disease similarity matrix B is:

[0098] f(I,J)=∑ m,n λ m,n κ G (<a I ,a m >)κ D (<b J ,b n >) (9)

[0099] In the formula, f is the association function between the drug similarity matrix A and the disease similarity matrix B, I is the row number of the vector in the association matrix Y, used to represent the drug molecule name in the association matrix, J is the column number of the vector in the association matrix Y, used to represent the disease name in the association matrix; a mis the value of the m-th row in the drug similarity matrix A, b n is the value of the n-th row in the disease similarity matrix B, κ G is the drug vector kernel function, κ D is the disease vector kernel function, a I is the value of the I-th row in the drug similarity matrix A, b J is the value of the J-th row in the disease similarity matrix B.

[0100] Based on the Kronecker least squares method, the Kronecker product of the drug vector kernel function and the disease vector kernel function is used to accelerate the calculation process. Let the drug vector kernel function κ G be set as and the disease vector kernel function κ D be set as Perform eigenvalue decomposition on the drug vector kernel function and the disease vector kernel function respectively, and the correlation function between the drug molecule and the disease is obtained as:

[0101]

[0102] where

[0103]

[0104] In the formula, Q G is the drug similarity matrix, Q D is the disease similarity matrix, and T is the transpose matrix.

[0105] According to the association between the drug molecule and the disease in the drug-disease relationship graph network of the existing database, the drug similarity matrix and the disease similarity matrix obtained by eigenvalue decomposition are combined to construct the objective function, as shown in formula (12):

[0106]

[0107] In the formula, y m,n is the relationship between the drug molecule m and the disease n in the existing database, S G* is the nearest neighbor similarity matrix of the drug, S D* is the nearest neighbor similarity matrix of the disease, is the norm of the correlation function f in the Hilbert space related to the kernel K, and both λ and β are regularization parameters.

[0108] In the objective function ensures the consistency of the geometric structures of the nearest neighbor similarity matrix of the drug and the drug similarity matrix, as well as the nearest neighbor similarity matrix of the disease and the disease similarity matrix.

[0109] Construct a drug similarity network and a disease similarity network based on the drug similarity matrix and the disease similarity matrix. Based on graph regularization, process the drug similarity network and the disease similarity network, respectively retain the geometric structure information of the adjacent nodes of each node in the drug similarity network and the disease similarity network, and calculate the drug similarity weight and the disease similarity weight, as shown in formula (13):

[0110]

[0111] In the formula, W G is the drug similarity weight, and W D is the disease similarity weight, and N p (·) is the set of nodes around node P.

[0112] Calculate the drug neighbor similarity matrix and the disease neighbor similarity matrix based on the drug similarity weight and the disease similarity weight, as shown in formula (14):

[0113]

[0114] In the formula, S G* is the drug neighbor similarity matrix, and S G is the drug similarity matrix; S D* is the disease neighbor similarity matrix, and S D is the disease similarity matrix.

[0115] Since the geometric structures of the drug neighbor similarity matrix and the drug similarity matrix, and the disease neighbor similarity matrix and the disease similarity matrix are consistent, based on the drug neighbor similarity matrix and the disease neighbor similarity matrix, use the constructed objective function to predict the association between drug molecules and diseases.

[0116] Of course, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Changes, modifications, additions or substitutions made by those skilled in the art within the essence of the present invention should also fall within the protection scope of the present invention.

Claims

1. A prediction method for drug-disease association relationships based on graph regularization matrix factorization, characterized in that, it includes the following steps: Step 1, according to the classification information of diseases, construct a directed acyclic graph for each disease respectively. Extract the semantic similarity between diseases based on the directed acyclic graphs of each disease, and then combine with the existing database to obtain the association relationship between each disease and drugs. Use the graph convolution method to extract disease features in the directed acyclic graph of each disease, construct a disease feature matrix, calculate the cosine similarity between diseases, and by fusing the semantic similarity and cosine similarity between diseases, obtain the disease association relationship based on the directed acyclic graph; Step 2, extract the drug features of each drug molecule in the existing database to obtain a drug feature matrix, and obtain the drug molecule similarity network by calculating the cosine similarity between each drug feature; Step 3, according to the disease feature matrix and the drug feature matrix, establish an association matrix between drug molecules and diseases; Step 4, based on the matrix factorization algorithm of graph regularization and kernel method, perform eigen-decomposition on the association matrix between drug molecules and diseases, combine with the drug-disease relationship graph network of the existing database, construct an objective function, optimize the neighbor relationship of nodes in the drug similarity network and the disease similarity network, and predict the association between drug molecules and diseases.

2. The prediction method for drug-disease association relationships based on graph regularization matrix factorization according to claim 1, characterized in that, in the said Step 1, it specifically includes the following steps: Step 1.1, according to the classification information of diseases, construct a directed acyclic graph for each disease respectively, and obtain the semantic values of all nodes in the directed acyclic graph of each disease; There are multiple nodes set in the said directed acyclic graph. Take the disease d itself as a child node, and the diseases related to the disease d as parent nodes. Construct a directed acyclic graph for each disease respectively. The directed acyclic graph is expressed as: DAG(d)=(N(d),E(d)) (1) In the formula, d is the disease name, DAG(·) is the directed acyclic graph of the disease, N(·) is the parent node related to the disease in the directed acyclic graph, and E(·) is the connection relationship between the parent node and the child node in the directed acyclic graph; Calculate the semantic values of all nodes in the directed acyclic graph of each disease respectively, as shown in formula (2): where n is the node number, n' is the child node of node n, and C d (·) is the semantic value of the node relative to disease d, and Δ is the semantic contribution factor; Determine the semantic value of each disease respectively according to the semantic values of all nodes in the directed acyclic graph of each disease, as shown in formula (3): DV(d) = ∑ n∈N(d) C d (n) (3) In the formula, DV(·) is the semantic value of the disease; Step 1.2, according to the semantic values of each disease, calculate the semantic similarity between diseases, as shown in formula (4): In the formula, is the semantic similarity between disease i and disease j, x is the node contained in both the directed acyclic graphs of disease i and disease j, is the semantic value of node x with respect to disease i, is the semantic value of node x with respect to disease j, is the semantic value of node n with respect to disease d i , DV(d i ) is the semantic value of disease i, DV(d j ) is the semantic value of disease j; Step 1.3, based on the existing database, obtain the drugs used to treat each disease, determine the association relationship between the disease and the drug for each disease respectively, extract the association relationship between the disease and the drug and use it as a descriptor. The said descriptor is a binary vector, and the length of the descriptor is the number of drugs in the database. If the disease is associated with the drug, set the value of the descriptor to 1. If there is no association between the disease and the drug, set the value of the descriptor to 0; For each disease, the graph convolution method is used to extract disease features in each directed acyclic graph, construct a disease feature matrix, and determine the eigenvalues of each disease feature, as shown in formula (5): where X d is the descriptor value, is the eigenvalue after the convolution calculation of all nodes in the directed acyclic graph, GCN(·) is the graph convolution function, V d is the eigenvalue of the disease feature, and Pool(·) is the aggregation function; Step 1.4, calculate the cosine similarity between diseases according to each disease feature, as shown in formula (6): In the formula, is the cosine similarity between disease i and disease j, m is the number of disease features, M is the total number of disease features, is the eigenvalue of the m-th disease feature with respect to disease i, is the eigenvalue of the m-th disease feature with respect to disease j; Step 1.5, fuse the semantic similarity and cosine similarity between diseases to obtain the disease association relationship based on the directed acyclic graph, as shown in formula (7): In the formula, is the correlation degree between disease i and disease j, and α is the fusion coefficient.

3. According to the method for predicting the drug-disease association relationship based on graph-regularized matrix factorization described in claim 2, it is characterized in that in the said step 1.3, the descriptor is a binary vector, and the length of the descriptor is the number of drugs in the database.

4. According to the method for predicting the drug-disease association relationship based on graph-regularized matrix factorization described in claim 2, it is characterized in that in the said step 2, according to the molecular structures of all drugs in the existing database, the drug molecules included in the database are determined, the drug features of all drug molecules in the existing database are extracted by using Morgan fingerprints, a drug feature matrix is established, and the cosine similarity between each drug feature in the drug feature matrix is calculated to obtain a drug molecule similarity network.

5. According to the method for predicting the drug-disease association relationship based on graph-regularized matrix factorization described in claim 4, it is characterized in that in the said step 3, the association matrix between drug molecules and diseases is: In the formula, Y is the association matrix between drug molecules and diseases, D is the disease feature matrix, and G is the drug molecule feature matrix.

6. According to the method for predicting the drug-disease association relationship based on graph-regularized matrix factorization described in claim 1, it is characterized in that in the said step 4, the kernel method is used to perform eigen-decomposition on the association matrix between drug molecules and diseases to obtain a drug similarity matrix A and a disease similarity matrix B, and the association function between the drug similarity matrix A and the disease similarity matrix B is: f(I,J) = ∑ m,n λ m,n κ G (<a I , a m >)κ D (<b J , b n >) (9) Wherein, f is the correlation function between the drug similarity matrix A and the disease similarity matrix B, I is the row number of the vector in the correlation matrix Y, used to represent the drug molecule name in the correlation matrix, J is the column number of the vector in the correlation matrix Y, used to represent the disease name in the correlation matrix; a m is the value of the m-th row in the drug similarity matrix A, b n is the value of the n-th row in the disease similarity matrix B, κ G is the drug vector kernel function, κ D is the disease vector kernel function, a I is the value of the I-th row in the drug similarity matrix A, b J is the value of the J-th row in the disease similarity matrix B; Based on the Kronecker least squares method, the Kronecker product of the drug vector kernel function and the disease vector kernel function is used to accelerate the calculation process, and the eigen-decomposition is respectively performed on the drug vector kernel function and the disease vector kernel function to obtain the association function between drug molecules and diseases as: Wherein, Where Q G is the drug similarity matrix, and Q D is the disease similarity matrix, and T is the transpose matrix; According to the association between drug molecules and diseases in the drug-disease relationship graph network of the existing database, the drug similarity matrix and the disease similarity matrix obtained by eigen-decomposition are combined to construct an objective function, as shown in formula (12): where y m,n is the relationship between the drug molecule m and the disease n in the existing database, S G* is the nearest neighbor similarity matrix of drugs, S D* is the nearest neighbor similarity matrix of diseases, is the norm of the correlation function f on the Hilbert space related to the kernel K, and both λ and β are regularization parameters; According to the drug similarity matrix and the disease similarity matrix, a drug similarity network and a disease similarity network are constructed, the drug similarity network and the disease similarity network are processed based on graph regularization, the geometric structure information of the adjacent nodes of each node in the drug similarity network and the disease similarity network is respectively retained, and the drug similarity weight and the disease similarity weight are calculated, as shown in formula (13): Wherein, W G is the drug similarity weight, and W D is the disease similarity weight, and N p (·) is the set of nodes around node P; According to the drug similarity weight and the disease similarity weight, a drug neighbor similarity matrix and a disease neighbor similarity matrix are calculated, as shown in formula (14): where S G* is the drug neighbor similarity matrix, and S G is the drug similarity matrix; S D* is the disease neighbor similarity matrix, and S D is the disease similarity matrix; Since the geometric structures of the drug neighborhood similarity matrix and the drug similarity matrix, and the disease neighborhood similarity matrix and the disease similarity matrix are consistent, based on the drug neighborhood similarity matrix and the disease neighborhood similarity matrix, the association between drug molecules and diseases is predicted by constructing an objective function.

Citation Information

Patent Citations

  • MiRNA-disease association predicting method based on double random walk models

    CN109935332A

  • Method for predicting potential association between miRNA and disease based on constrained probability matrix decomposition

    CN110782948A