Drug-disease association prediction method based on layered negative sample sampling
By adopting a hierarchical negative sampling comparison learning method in drug-disease association prediction, combined with PageRank and GAT, the problem of singularity of negative sample selection strategies is solved, and the prediction performance is significantly improved.
Patent Information
- Application Number
- CN202510120200.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-25
- Publication Date
- 2025-06-20
Smart Images

Figure CN120183741A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of graph neural network models, and particularly relates to a drug-disease association prediction method based on hierarchical negative sample sampling. Background Art
[0002] With the continuous development of pharmaceutical technology, more and more drugs have been developed and shown good effects in clinical trials. However, with the continuous increase in the types of current diseases and their mutations, it is urgent to develop effective drugs. However, in traditional drug development, the development cost and cycle of drugs are relatively long, and the drug development and approval processes face many challenges. Therefore, since the drug-disease association prediction method can improve the success rate and reduce costs, it has been applied to the field of drug development.
[0003] Drug repositioning does not rely on tedious wet experimental studies, but rather identifies potential new indications for approved drugs, followed by experimental validation to identify candidate drugs. It is estimated that the cost of repositioning a drug to the market is only 1 / 10 of that of developing a new compound drug, and the time required for drug repositioning is only 1 / 3 of the time required for new drug development. In recent years, with the advancement of deep learning technology, drug repositioning has made significant progress. It has been successfully applied in pharmaceutical research and development fields such as the discovery of anticancer drugs, overcoming drug resistance, and promoting the progress of personalized medicine. The main methods for drug-disease association prediction are divided into two categories: traditional machine learning methods and deep learning methods. In the field of drug-disease association prediction, machine learning is applied to develop more accurate and effective prediction models. Prediction algorithms based on traditional machine learning are often divided into the following steps: 1. Preprocessing drug and disease information into corresponding features, 2. Building a model to predict the obtained model. For example, a method for inferring new drug indications, suitable for personalized medicine approach (PREDICT) integrates multiple similarity measures between drugs and diseases and uses a logistic regression classifier to predict drug-disease associations. Although machine learning methods have their own advantages in drug-disease association prediction, their performance can still be further improved. Therefore, in order to better handle complex network structures and enhance feature extraction, researchers have proposed prediction methods based on deep learning. For example, a method for predicting drug-disease associations using graph representation learning on heterogeneous information networks (HINGRL) captures the characteristics of drugs and diseases by considering the biological knowledge of drugs and diseases and combining drug-protein-disease topological structures. Compared with traditional experimental verification methods, these two types of methods greatly reduce the consumption of time, energy and funds, but there are still some problems to be solved. Better negative sample quality can make the model prediction results more accurate. However, the currently known negative sample selection strategies only select local or global features. Therefore, the two key factors to improve the prediction performance of deep learning methods are: the construction of feature aggregation modules and negative sample selection strategies. Summary of the invention
[0004] The object of the present invention is to address the problems of overly single feature selection methods and unbalanced negative sample selection methods; and a method for predicting drug-disease associations using hierarchical negative sampling contrast learning is proposed. To more accurately construct a similarity network, this method calculates and integrates the drug Jaccard similarity kernel and the drug GIP (Gaussian Interaction Profile) kernel similarity; as well as the semantic similarity kernel of diseases and the GIP kernel similarity of diseases. The neural network algorithm is used to process relevant data to improve the prediction accuracy of the algorithm. To further solve the problem of node feature aggregation, this method uses a graph convolutional neural network (GCN) and a graph attention network (GAT) to aggregate the known drug and disease similarity networks. To be able to select better high-quality negative samples, the present invention designs a new hierarchical negative sampling method, which comprehensively considers the global and local information of negative samples using the PageRank algorithm and the GAT attention score to improve the prediction effect of the algorithm.
[0005] The specific steps of the present invention are as follows:
[0006] (1) Prepare a drug-disease association dataset and obtain drug-disease association information;
[0007] (2) Calculate the semantic similarity between two diseases d i and d j , as well as the GIP kernel similarity between diseases d i and d j , and perform linear fusion;
[0008] (3) Calculate the Jaccard similarity between drugs r i and r j , and the GIP kernel similarity between drugs r i and r j . And perform linear fusion;
[0009] (4) Use GCN and GAT to aggregate and extract features from the similarity network;
[0010] (5) Use a Bilinear aggregator to extract features from the drug-disease similarity network;
[0011] (6) Use a hierarchical negative sampling module to select negative samples;
[0012] (7) Use a multi-layer perceptron to process the data and obtain the final prediction score.
[0013] Specifically, in step (1), download the known drug-disease association dataset from a public database. Dr = {r1, r2, r3,..., r n}\ represents the drug dataset, Ds = {d1, d2, d3,..., d m}\ disease dataset, n represents the number of drugs in the dataset, and m represents the number of diseases in the dataset; construct the association matrix Y ∈ R n×m \ represents the known drug-disease association information, where Y ij = 1 indicates that drug r i \ is associated with disease d j , and Y ij = 0 indicates that drug r i \ has no association with disease d j .
[0014] In particular, in step (2), the present invention adopts a data preprocessing method of multi-similarity fusion, which linearly fuses the semantic similarity of diseases and the GIP Gaussian kernel similarity of diseases. If disease d i \ is associated with disease d j , then the average value of the two similarities of the diseases is taken as the similarity score between the two. If disease d i \ has no association with disease d j , the GIP kernel similarity is taken as the similarity of the diseases. For the joint similarity Ds(d i , d j ) of disease d i \ and disease d j , the calculation method is as follows:
[0015]
[0016] where MD(d i , d j ) is the disease semantic similarity; GipD(d i , d j ) is the disease GIP kernel similarity. In the related calculation of MD, MeSH descriptors are used to represent diseases as directed acyclic graphs (DAGs). Then, the semantic similarity scores of these diseases in the DAG are calculated. Similar to drugs, there may be some entries in the drug-disease association (DDA) matrix where the semantic similarity of diseases is missing. Considering this situation, GIP is used as a supplement, and its form is as follows:
[0017]
[0018] where \ and \ are the interaction information of vector diseases d i \ and d j . σ d \ is the bandwidth of the GIP kernel.
[0019] Specifically, in step (3), the Jaccard similarity of drugs and the GIP kernel similarity of drugs are linearly fused. When drug r i is associated with drug r j , the mean value of the two similarities of the drug is taken. If there is no association, the GIP kernel similarity between drug r i and drug r j is taken as the similarity score between the two. The combined similarity Rs(r i , r j ) of drug r i and drug r j is defined as follows:
[0020]
[0021] where DR(r i , r j ) is the drug structure similarity; GipR(r i , r j ) is the drug Gaussian interaction profile kernel similarity. In the relevant calculation of DR, the rdkit framework is used to score the drug structure similarity of the collected drug simplified molecular linear input specification (SMILES) sequences. In the drug-disease association matrix, considering that there are some drugs with missing structural similarity entries. For this situation, GIP is used as a supplement, and its form is as follows:
[0022]
[0023] where and are the interaction information of vector drugs r i and r j . σ r is the bandwidth of the GIP kernel.
[0024] Specifically, in step (4), in order to better aggregate and extract features from the drug-disease similarity network. Among them, GCN is realized by performing a convolution operation on the surrounding nodes of the target node. This shows that in the information obtained by GCN, the features of each node contain information about the node itself and its neighbors. GAT can better capture the interaction information between nodes by dynamically assigning attention weights according to the importance of neighbor nodes. By combining GCN and GAT, the information interaction between nodes and their neighbor nodes can be more comprehensively captured.
[0025] Taking drug similarity as an example, in the present invention, the neighborhood feature information of nodes is extracted through two-layer graph convolution. Taking drug-related information as an example, for drug node r i , its feature representation is The extraction form is defined as follows:
[0026]
[0027] wherein represents the drug r i of the extended neighborhood. represents the embedding of node j at the l-th layer. represents the weight matrix of the l-th layer graph convolution, then represents the bias vector of the l-th layer. c ij is represented as the normalization factor.
[0028] Node r in the drug similarity network i is encoded by GAT. The present invention hopes to assign different attention weights to the relationship between the target node and its neighbor nodes through the multi-head attention mechanism. Finally, the final node representation is generated by combining the results of multiple attention heads. The specific steps are as follows:
[0029] For each node i and its neighbor node j, calculate the correlation score e between the node pairs through a linear transformation and the LeakyReLU activation function l-ij .
[0030]
[0031] where a represents a linear transformation vector that can be adaptively updated during training. W represents the weight matrix in the linear transformation. x i represents the neighborhood feature of the drug r i . Then, through the softmax normalization function, these scores are transformed into attention weights a ij , that is:
[0032]
[0033] where represents the extended neighborhood of r i . a ij is normalized by the attention score e l-ij , that is, the attention coefficient from node j to node i. Then, use the obtained attention weights to perform a weighted sum of the features of the neighbor nodes, and generate a new representation h' of node i through the ELU activation function i . The calculation process for each attention head k is as follows:
[0034]
[0035] where σ(·) represents the ELU activation function, represents the neighbor nodes of i. x j represents the corresponding feature representation vector, Indicates the attention weight of attention head k for node i and its neighbor node j, and W is the weight matrix of attention head k.
[0036] Single-head attention is obtained through the above calculation process. To enable the model to have better expressive power and robustness, the GAT module in this paper adopts the multi-head attention mechanism. To avoid the vector dimension becoming too large due to the connection of multiple attention heads, introducing noise and affecting the prediction ability of the model. Therefore, the outputs of multiple attention heads are averaged to obtain the node representation.
[0037]
[0038] Similarly, the representation of the disease node is obtained through the above steps. To be able to represent the characteristics of drugs and diseases more comprehensively and enhance the representation ability of graph data. The present invention adopts a graph information fusion aggregator. This aggregator combines the outputs of GCN and GAT, and aggregates the neighborhood information and neighborhood interaction information into a comprehensive neighborhood feature representation:
[0039]
[0040] Among them and are vectors representing the comprehensive neighborhood information of drugs and diseases respectively. and are vectors representing the neighborhood information of the target node. and are vectors representing the neighborhood interaction information of the target node. σ is a non-linear activation function, and λ is a hyperparameter used to balance the weight ratio assigned to neighborhood information and neighborhood interaction information.
[0041] Specifically, in step (5), to better process the node information in the heterogeneous network, the present invention uses a multi-neighborhood extraction module to aggregate and extract feature information from the drug-disease association network. Among them, GCN extends the operations of the convolutional neural network to graph data. GCN assumes that adjacent nodes are independent of each other and uses a weighted sum to represent the status representation of nodes. The GA aggregator of the target node (drug r or disease d) is represented as:
[0042]
[0043] Among them, GA represents a non-linear aggregator, represents the extended neighbors including node v itself. e i represents the feature vector of node i. a vi is the weight of neighbor i, defined as W g is the weight matrix for feature transformation. σ represents the activation function.
[0044] There may be interactions between adjacent nodes, but common graph convolutional networks often ignore this information. Although graph attention networks can adaptively aggregate adjacent nodes according to their importance, they are still unable to extract the possible interaction features between adjacent nodes. At the same time, by multiplying two vectors, the present invention can effectively model the interaction between nodes, which is achieved by strengthening the consistent information and weakening the inconsistent information. Based on this, the present invention designs a BA-type aggregator for the target node v:
[0045]
[0046] where BA represents a non-linear aggregator, represents the number of interactions of node v. This formula reduces the bias effect of node degree to a certain extent through normalization processing. ⊙ represents element-wise multiplication, and W b is the weight matrix for feature transformation.
[0047] Then, the encoder constructed in step (5) is used to transmit the information between drugs and diseases. At the same time, indirect interactions are extracted from the local mechanism. Specifically, the encoder for node v is defined as:
[0048]
[0049] where β is a hyperparameter used to adjust the respective strengths of the GA aggregator and the BA aggregator.
[0050] In particular, in step (6), in order to be able to select better-quality negative samples, the present invention develops a new hierarchical negative sample selection method based on the hierarchical negative sampling strategy of the similarity network. In the present invention, hierarchical negative sampling is carried out through the following three steps. First, the PageRank algorithm is used to score the similarity network of drugs and diseases (denoted as R1, D1). Sort in descending order according to the scores, and extract the top L biomolecule information with higher tendencies. Then, through the association information between diseases and drugs, the drug information related to the diseases in D1 is screened out. At the same time, the obtained drug information related to D1 is merged with the initial drug information set R1 to form a new drug information set (denoted as R2). At this time, R2 ∈ {1, 2,..., I - 1, I}, I > L, where R1 ∈ {1, 2,..., L}.
[0051] In order to be able to determine the reliability of the information nodes in R2, the present invention uses the GAT aggregator to score the information nodes again, so as to obtain a new representation q v (v ∈ {r, d}) of the nodes. Its calculation is defined as:
[0052]
[0053] where \(W\) represents the weight matrix, and \(a\) ij represents the normalized attention score, and \(x\) j represents the corresponding feature representation vector. To enhance the reliability of negative sampling, in this experiment, valuable self-supervised signals are mined through a joint training system. For the drug \(r\) i and \(d\) j given in step (4), the final positive and negative samples are selected based on the representations finally learned in the single neighborhood:
[0054] \(S_c\) r \(=\text{softmax}(M\) d \(q\) r )
[0055] where \(S_c\) r represents the predicted probability that each disease can be cured by the drug \(r\) in the step; \(M\) d represents the disease global feature matrix, where \(M\) d is composed of the comprehensive neighborhood feature representations in the similarity network . After that, the present invention will select the top \(L\) drugs as \(R_3\) according to the score of \(S\) c . Combine the disease set \(D_1\) and the drug set \(R_3\) to form potential drug-disease association pairs, denoted as \(RD_{data}\).
[0056] Specifically, in step (7), the known positive sample association pairs in \(RD_{data}\) calculated in step (6) are removed, and the remaining part is used as the negative sample candidate set. To ensure the robustness of the model's negative sampling strategy, in the final negative sample candidate set, \(K\) samples with the same number as the positive samples are randomly selected as reliable negative samples. Use a multi-layer perceptron to score the final drug-disease association prediction as follows:
[0057]
[0058] where is the embedding vector of the drug \(r\) i , is the embedding vector of the disease \(d\) j ; is the local feature of the drug \(r\) i , is the local feature of the disease \(d\) j . The number of layers \(l\) of the multi-layer perceptron is set to 2, the dimension of the first layer is set to 128, and the dimension of the second layer is set to 64. To better learn the information interaction between neighborhoods, the present invention uses InfoNCE to learn its lower bound. The triple loss function of the contrastive learning in the \(l\)th layer is defined as:
[0059]
[0060] where δ(·) represents the cosine similarity function, and N represents the total number of drug or disease nodes. When j ≠ i, ε = 1; otherwise, ε = 0. represents the representation of node i after aggregation in step (5), represents the representation of node i obtained after aggregation in step (4). To better predict drug-disease interactions, the present invention combines the prediction loss and the triplet loss to achieve the overall optimization of the model. The prediction loss is represented by a weighted cross-entropy loss function. As follows:
[0061]
[0062] where represents the final predicted association score, and (i, j) represents the drug and disease pair. Among them, A + represents the currently known DDA, and A - represents the unknown DDA. To prevent the model from overfitting and accelerate the model training process, L2 regularization Φ is used to smooth the weights, and the final loss calculation function is obtained.
[0063] Loss = L P + Θ1L r + Θ2Φ,
[0064] where L P and L T represent the prediction loss and the triplet loss in contrastive learning respectively, Θ1 and Θ2 represent hyperparameters used to balance the loss calculation. Φ uses trainable model parameters. To minimize the loss function, this experiment uses the Adam optimizer to select the optimal parameters, including the number of hidden layers L, the number of epochs e, the learning rate lr, and the weight decay wd. The present invention sets the optimal hyperparameters lr to 10-5, e to 100, wd to 10-5, and L to 3.
[0065] Compared with the prior art, the present invention has the following advantages: First, the present invention utilizes two types of drug similarities (drug Jaccard similarity and drug structural similarity) and two types of disease similarities (disease semantic similarity and disease structural similarity), linearly fuses the above two similarities respectively, and strengthens the similarity data in different cases; applies a multi-neighborhood aggregator and a multi-neighborhood aggregator to the heterogeneous association network and the homogeneous association network respectively, comprehensively considering the information of the node itself and the interaction information between the node neighborhoods; on this basis, introduces a hierarchical negative sampling strategy, and comprehensively considers the local information features and global information features of the nodes through the PageRank algorithm and the GAT attention mechanism; selects better K negative samples through screening, and the prediction performance is significantly improved compared with the existing methods. Brief Description of the Drawings
[0066] The present invention will be further described below with reference to the accompanying drawings.
[0067] Figure 1 is a flowchart of the method of the present invention;
[0068] Figure 2 is a flowchart of the hierarchical negative sampling method;
[0069] Figure 3 is a radar chart of the comparison of the results with other methods after ten-fold cross-validation. Detailed Embodiment
[0070] The present invention will be further elaborated below in combination with specific data sets. The data sets applied are only used to illustrate the present invention, and not to limit the scope of use of the present invention.
[0071] Embodiment 1
[0072] The specific steps of the present invention are as follows:
[0073] (1) Prepare a drug-disease association data set and obtain drug-disease association information;
[0074] In step (1), download the known drug-disease association data set from a public database. Dr = {r1, r2, r3,..., r n} represents the drug data set, Ds = {d1, d2, d3,..., d m} represents the disease data set, n represents the number of drugs in the data set, and m represents the number of diseases in the data set. Construct an association matrix Y ∈ R n×m to represent the known drug-disease association information, where Y ij = 1 indicates that there is an association between drug r i and disease d j , and Y ij = 0 indicates that there is no association between drug r i and disease dj There is no association between them.
[0075] (2) Calculate the semantic similarity between two diseases d i and d j , as well as the GIP kernel similarity between diseases d i and d j , and perform linear fusion;
[0076] In step (2), the present invention adopts a data preprocessing method of multi-similarity fusion, which linearly fuses the semantic similarity of diseases and the GIP Gaussian kernel similarity of diseases. If there is an association between disease d i and disease d j , the average value of the two similarities of the diseases is taken as the similarity score between the two. If there is no association between disease d i and disease d j , the GIP kernel similarity is taken as the similarity of the diseases. For the combined similarity Ds(d i and disease d j ) of diseases d i ,d j ), the calculation method is as follows:
[0077]
[0078] where MD(d i ,d j ) is the semantic similarity of diseases; GipD(d i ,d j ) is the GIP kernel similarity of diseases. In the relevant calculation of MD, MeSH descriptors are used to represent diseases as a directed acyclic graph (DAG). Then, the semantic similarity scores of these diseases in the DAG are calculated. Similar to drugs, there may be some entries in the drug-disease association (DDA) matrix where the semantic similarity of diseases is missing. Considering this situation, the present invention uses GIP as a supplement, and its form is as follows:
[0079]
[0080] where and are the interaction information of vector diseases d i and d j . σ d is the bandwidth of the GIP kernel.
[0081] (3) Calculate the Jaccard similarity between drugs r i and r j , and the GIP kernel similarity between drugs r i and r j , and perform linear fusion;
[0082] In step (3), the Jaccard similarity of drugs and the GIP kernel similarity of drugs are linearly fused. Among them, when drug r i is associated with drug r j , the mean value of the two similarities of the drug is taken. If there is no association, the GIP kernel similarity between drug r i and drug r j is taken as the similarity score between the two. The combined similarity Rs(r i , r j ) of drug r i and drug r j is defined as follows:
[0083]
[0084] where DR(r i , r j ) is the drug structure similarity; GipR(r i , r j ) is the drug Gaussian interaction profile kernel similarity. In the related calculation of DR, the rdkit framework is used to score the drug structure similarity of the collected drug Simplified Molecular-Input Line-Entry System (SMILES) sequences. In the drug-disease association matrix, considering that there are some drugs with missing structural similarity entries. For this situation, GIP is used as a supplement, and its form is as follows:
[0085]
[0086] where and are the interaction information of vector drugs r i and r j . σ r is the bandwidth of the GIP kernel.
[0087] (4) Use GCN and GAT to aggregate and extract features from the similarity network;
[0088] In step (4), in order to better aggregate and extract features from the drug and disease similarity networks. Among them, GCN is realized by performing a convolution operation on the surrounding nodes of the target node. This shows that in the information obtained by GCN, the features of each node contain information about the node itself and its neighbors. GAT can better capture the interaction information between nodes by dynamically assigning attention weights according to the importance of neighbor nodes. By combining GCN and GAT, the information interaction between nodes and their neighbor nodes can be more comprehensively captured.
[0089] Taking drug similarity as an example, in the present invention, the neighborhood feature information of nodes is extracted through two-layer graph convolution. Taking drug-related information as an example, for drug node r i , its feature representation is The extraction form is defined as follows:
[0090]
[0091] Where represents the extended neighborhood of drug r i . represents the embedding of node j at the l-th layer. represents the weight matrix of the l-th layer graph convolution, then represents the bias vector of the l-th layer. c ij represents the normalization factor.
[0092] Node r in the drug similarity network i is encoded through GAT. The present invention hopes to assign different attention weights to the relationship between the target node and its neighbor nodes through the multi-head attention mechanism. Finally, the final node representation is generated by combining the results of multiple attention heads. The specific steps are as follows:
[0093] For each node i and its neighbor node j, the correlation score e between the node pairs is calculated through a linear transformation and the LeakyReLU activation function l-ij .
[0094]
[0095] where a represents a linear transformation vector that can be adaptively updated during training. W represents the weight matrix in the linear transformation. x i represents the neighborhood feature of drug r i . Then, through the softmax normalization function, these scores are converted into attention weights a ij , that is:
[0096]
[0097] Where represents the extended neighborhood of r i . a ij is normalized by the attention score e l-ij , that is, the attention coefficient from node j to node i. Then, the obtained attention weights are used to weight-sum the features of the neighbor nodes, and a new representation h' of node i is generated through the ELU activation function i . The calculation process for each attention head k is as follows:
[0098]
[0099] where σ(·) represents the ELU activation function, represents the neighbor nodes of i. x j represents the corresponding feature representation vector. represents the attention weight of attention head k to node i and neighbor node j, and W is the weight matrix of attention head k.
[0100] Single-head attention is obtained through the above calculation process. To enable the model to have better expressiveness and robustness, the GAT module in this paper adopts the multi-head attention mechanism. To avoid the vector dimension becoming too large when multiple attention heads are connected, noise is introduced and the prediction ability of the model is affected. Therefore, the present invention chooses to average the outputs of multiple attention heads to obtain the node representation.
[0101]
[0102] Similarly, the representation of the disease node is obtained through the above steps. To be able to represent the characteristics of drugs and diseases more comprehensively and enhance the representation ability of graph data. The present invention adopts a graph information fusion aggregator. This aggregator combines the outputs of GCN and GAT, and aggregates the neighborhood information and neighborhood interaction information into a comprehensive neighborhood feature representation:
[0103]
[0104] where and are vectors representing the comprehensive neighborhood information of drugs and diseases respectively. and are vectors representing the neighborhood information of the target node. and are vectors representing the neighborhood interaction information of the target node. σ is a non-linear activation function, and λ is a hyperparameter used to balance the weight ratio assigned to neighborhood information and neighborhood interaction information.
[0105] (5) Use the Bilinear aggregator to extract features from the drug-disease similarity network;
[0106] In step (5), to better process the node information in the heterogeneous network, the present invention uses a multi-neighborhood extraction module to aggregate and extract feature information from the drug-disease association network. Among them, GCN extends the operations of the convolutional neural network to graph data. GCN assumes that adjacent nodes are independent of each other and uses a weighted sum to represent the status representation of nodes. The GA aggregator of the target node (drug r or disease d) is represented as:
[0107]
[0108] where GA represents a non-linear aggregator, represents the extended neighbors including the node v itself. e i represents the feature vector of node i. a vi is the weight of neighbor i, defined as W g is the weight matrix for feature transformation. σ represents the activation function.
[0109] There may be interactions between adjacent nodes, but common graph convolutional networks often ignore this information. Although graph attention networks can adaptively aggregate adjacent nodes according to their importance. However, it is still unable to extract the possible interaction features between adjacent nodes. At the same time, by multiplying two vectors, the present invention can effectively model the interaction between nodes, and this method is achieved by strengthening the consistent information and weakening the inconsistent information. Based on this, the present invention designs a BA-type aggregator for the target node v:
[0110]
[0111] where BA represents a non-linear aggregator, represents the number of interactions of node v. This formula reduces the bias effect of node degrees to a certain extent through normalization processing. ⊙ represents element-wise multiplication, and W b is the weight matrix for feature transformation.
[0112] Then, according to the encoder constructed in step (5), it is used to transmit the information between drugs and diseases, and at the same time extract the indirect interactions from the local mechanism. Specifically, the encoder for node v is defined as:
[0113]
[0114] where β is a hyperparameter used to adjust the respective strengths of the GA aggregator and the BA aggregator.
[0115] (6) Use a hierarchical negative sampling module to select negative samples;
[0116] In step (6), in order to be able to select better-quality negative samples, as shown in the attached instructions Figure 2As shown, the present invention develops a new hierarchical negative sample selection method based on the hierarchical negative sampling strategy of the similarity network. In the present invention, the hierarchical negative sampling is carried out through the following three steps. First, the similarity networks of drugs and diseases are scored using the PageRank algorithm (denoted as R1, D1). The scores are sorted in descending order, and the top L biomolecule information with higher tendencies is extracted. Then, through the association information between diseases and drugs, the drug information related to the diseases in D1 is screened out. At the same time, the obtained drug information related to D1 is merged with the initial drug information set R1 to form a new drug information set (denoted as R2). At this time, R2 ∈ {1, 2,..., I - 1, I}, I > L, where R1 ∈ {1, 2,..., L}.
[0117] In order to determine the reliability of the information nodes in R2, the present invention uses the GAT aggregator to perform a new scoring on the information nodes, thereby obtaining a new representation q of the nodes v (v ∈ {r, d}). Its calculation is defined as:
[0118]
[0119] where W represents the weight matrix, a ij represents the normalized attention score, x j represents the corresponding feature representation vector. In this embodiment, q v represents the set of both drug nodes and disease nodes; while q i represents a specific drug node i or disease node i, and it can be understood that qi is a subset of qv.
[0120] To enhance the reliability of negative sampling. In this experiment, valuable self-supervised signals are mined through a joint training system. For the drugs r i and d j given in step (4), the final positive and negative samples are selected through the representations finally learned in the single neighborhood:
[0121] Sc r = softmax(M d q r )
[0122] where Sc r represents the predicted probability that each disease can be cured by the drug r in the step; M d represents the disease global feature matrix, where M d is composed of the comprehensive neighborhood feature representations in the similarity network. After that, the present invention will pass through S cSelect the top L drugs as R3 according to the scoring situation. Combine the disease set D1 and the drug set R3 to form potential drug-disease association pairs, denoted as RDdata.
[0123] (7) Use a multi-layer perceptron to process the data and obtain the final prediction score.
[0124] In step (7), remove the known positive sample association pairs in RDdata calculated in step (6), and use the remaining part as the negative sample candidate set. To ensure the robustness of the model's negative sampling strategy. In the final negative sample candidate set, randomly select K samples with the same number as the positive samples as reliable negative samples. To reconstruct the simple drug-disease association, the present invention uses a multi-layer perceptron to predict the drug r i and disease d j The association probability between them is as follows:
[0125]
[0126] where is the embedding vector of drug r i , is the embedding vector of disease d j ; is the local feature of drug r i , is the local feature of disease d j . The number of layers l of the multi-layer perceptron is set to 2, the dimension of the first layer is set to 128, and the dimension of the second layer is set to 64. Here, using a decoder, calculate the association score of drug-disease, and the generated score represents the association probability between drug i and disease j. To be able to better learn the information interaction between neighborhoods, the present invention uses InfoNCE to learn its lower bound, and the triplet loss function of contrastive learning in the lth layer is defined as:
[0127]
[0128] where δ(·) represents the cosine similarity function, and N represents the total number of drug nodes. When j≠i, ε = 1; otherwise ε = 0. represents the representation of node i after aggregation in step (5), represents the representation of node i obtained after aggregation in step (4). To better predict drug-disease interactions, the present invention combines the prediction loss and the triplet loss to achieve the overall optimization of the model. Among them, the prediction loss is represented by a weighted cross-entropy loss function, as follows:
[0129]
[0130] where represents the final predicted association score, and (i, j) represents the drug-disease pair. Among them, A + represents the currently known DDA, A - represents the unknown DDA. To prevent the model from overfitting and accelerate the model training process, L2 regularization Φ is used to smooth the weights, and the final loss calculation function is obtained.
[0131] Loss = L P + Θ1L T + Θ2Φ,
[0132] where L P and L T represent the prediction loss and the triplet loss in contrastive learning respectively. Θ1 and Θ2 represent hyperparameters used to balance the loss calculation. Φ uses trainable model parameters. To minimize the loss function, this experiment uses the Adam optimizer to select the optimal parameters, including the number of hidden layers L, the number of epochs e, the learning rate lr, and the weight decay wd. The present invention sets the optimal hyperparameters lr to 10-5, e to 100, wd to 10-5, and L to 3.
[0133] Example 2
[0134] (1) The experiment uses python simulation processing. In ten-fold cross-validation, the maximum number of iterations is 100.
[0135] (2) A comparative analysis of the prediction performance of the present invention is carried out. And it is compared with five existing methods (the end-to-end deep learning model LAGCN based on heterogeneous graphs, the regularized least squares model MKDGRLS based on multiple kernels, the graph representation learning technology modelHINGRL based on heterogeneous graphs, the method DRWBNCF using weighted bilinear graph convolution, and the heterogeneous graph neural network model REDDA based on integrating multiple biological relationships). Figure 3 is the radar chart for the performance comparison of the method of the present invention with other methods, and Table 1 is the comparison of the evaluation indexes of the method of the present invention with other methods after ten-fold cross-validation experiments.
[0136] Table 1 Comparison of five-fold cross-validation of the method of the present invention with four other methods
[0137]
[0138] As can be seen from the simulation results in Table 1 and Figure 3 the evaluation indexes of the present invention are relatively superior to the other five methods.
[0139] (3) Case analysis was conducted on the prediction results of the present invention for two diseases (Parkinson's disease, Alzheimer's disease). The prediction scores were sorted in descending order, and the top 15 drugs were selected for analysis. Validation was carried out through DrugBank, CTD, PubChem, DrugCentral, ClinicalTrials, and published literature. The specific results are shown in Tables 2 and 3. It is not difficult to see from the tables that the present invention can predict drugs that have been confirmed to be related to diseases and can predict potential drug-disease associations.
[0140] Table 2 The top 15 drugs predicted to be related to Parkinson's disease
[0141]
[0142]
[0143] Table 3 The top 15 drugs predicted to be related to Alzheimer's disease
[0144]
[0145] The above simulation results show that the present invention can effectively predict drug-disease associations and can significantly improve the prediction accuracy compared with existing methods.
Claims
1. A drug-disease association prediction method based on stratified negative sample sampling, characterized in that: The following steps are involved: (1) Prepare a drug-disease association dataset and obtain drug-disease association information; (2) Calculate the two diseases d i and d j The semantic similarity between i and d j The GIP kernel similarity between them is calculated and linear fusion is performed; (3) Calculate drug r i and r j Jaccard similarity and drug r i and r j The GIP kernel similarity is calculated and linear fusion is performed; (4) Use GCN and GAT to aggregate similarity networks and extract features; (5) Use Bilinear aggregator to extract features from drug-disease similarity network; (6) Use the stratified negative sampling module to select negative samples; (7) Use a multi-layer perceptron to process the data and obtain the final prediction score.
2. The method according to claim 1, characterized in that In step (1), a known drug-disease association dataset is downloaded from the database, Dr = {r1, r2, r3, ..., r n } represents the drug data set, Ds = {d1, d2, d3, …, d m }Disease data set, n represents the number of drugs in the data set, m represents the number of diseases in the data set, and constructs the association matrix Y∈R n×m represents the known drug-disease association information, where Y ij =1 means drug r i With disease j There is a correlation between ij =0 means drug r i With disease j There is no correlation between them.
3. The method according to claim 1, characterized in that In step (2), a data preprocessing method of multi-similarity fusion is used, which linearly fuses the semantic similarity of the disease and the GIP Gaussian kernel similarity of the disease. i With disease j When there is an association, the mean of the two similarities of the disease is taken as the similarity score between the two. i With disease j When there is no association, GIP nuclear similarity is taken as disease similarity; For disease i and diseases j The joint similarity Ds(d i ,d j ) is calculated as follows: Among them, MD(d i ,d j ) is the disease semantic similarity; GipD(d i ,d j ) is the disease GIP kernel similarity. In the relevant calculation of MD, MeSH descriptors are used to represent diseases as directed acyclic graphs (DAGs). Then, the semantic similarity scores of these diseases in DAGs are calculated. Similar to drugs, there may be entries in the drug-disease association DDA matrix that lack semantic similarity for some diseases. Considering this situation, GIP is used as a supplement, and its form is as follows: in and is the vector disease d i and d j The interaction information, σ d is the bandwidth of the GIP core.
4. The method according to claim 1, characterized in that: In step (3), the Jaccard similarity of the drugs and the GIP kernel similarity of the drugs are linearly fused, where when the drug r i With drugs j If there is a correlation, the mean of the two similarities of the drugs is taken. If there is no correlation, the drug r is taken. i With drugs j The GIP nuclear similarity of the two is used as the similarity score; drug r i With drugs j The joint similarity Rs(r i ,r j ) is defined as follows: Among them, DR (r i ,r j ) is the drug structure similarity; GipR(r i ,r j ) is the drug Gaussian interaction profile kernel similarity; in the drug-disease association matrix, considering that some drugs lack structural similarity entries, GIP is used as a supplement for this situation, and its form is as follows: in and is the vector drug r i and r j The interaction information, σ r is the bandwidth of the GIP core.
5. The method according to claim 1, characterized in that In step (4), in order to better aggregate and extract features of drug and disease similarity networks, GCN is implemented by performing convolution operations on the surrounding nodes of the target node. By combining GCN and GAT, the information interaction between nodes and their neighboring nodes can be more comprehensively captured; For drug similarity, the neighborhood feature information of the node is extracted through two layers of graph convolution. i , whose characteristics are expressed as The extraction form is defined as follows: in Indicates drug r i The extended neighborhood of represents the embedding of node j at layer l, represents the weight matrix of the l-th layer graph convolution, It represents the bias vector of the lth layer, c ij is expressed as a normalization factor. Node r in drug similarity network i Encoding through GAT; Through the multi-head attention mechanism, different attention weights are assigned to the relationship between the target node and its neighbor nodes, and finally the final node representation is generated by combining the results of multiple attention heads. The specific steps are as follows: For each node i and its neighbor node j, the correlation score e between the node pairs is calculated by linear transformation and LeakyReLU activation function l-ij ; Where a represents a linear transformation vector that can be adaptively updated during training, W represents the weight matrix in the linear transformation, and x i Indicates drug r i Neighborhood characteristics, x j Indicates drug r j Neighborhood characteristics, x T represents the transposition of the initial neighborhood features; then the above score is converted into attention weight a through the softmax normalization function ij ,Right now: in Represents r i The extended neighborhood of a ij By the attention score e l-ij Normalization, that is, the attention coefficient from node j to node i; then, the obtained attention weight is used to weight the features of the neighboring nodes, and the new representation h' of node i is generated through the ELU activation function i , the calculation process for each attention head k is as follows: in represents the node representation of the kth attention head of drug node i, σ(·) represents the ELU activation function, j∈ represents the neighbor node of i, x j represents the corresponding feature representation vector, represents the attention weight of attention head k to node i and neighbor node j, W is the weight matrix of attention head k; The single-head attention is obtained through the above calculation process; in order to make the model have better expressiveness and robustness, the GAT module adopts a multi-head attention mechanism; in order to avoid the connection of multiple attention heads causing the vector dimension to be too large, introducing noise, and affecting the prediction ability of the model, the outputs of multiple attention heads are averaged to obtain the drug node representation; in, is the node representation of drug node i, is the node representation of the kth attention head of drug node i; x j Represents the corresponding feature representation vector. Similarly, the representation of disease nodes Through the above steps, the disease node is expressed as: is the node representation of disease node i, is the node representation of the kth attention head of disease node i; x j represents the corresponding feature representation vector; In order to more comprehensively represent the characteristics of drugs and diseases and enhance the representation ability of graph data, a graph information fusion aggregator is used; this aggregator combines the outputs of GCN and GAT to aggregate neighborhood information and neighborhood interaction information into a comprehensive neighborhood feature representation: in and are vectors representing the comprehensive neighborhood information of drugs and diseases, and A vector representing the neighborhood information of the target node, and is a vector representing the neighborhood interaction information of the target node, σ is a nonlinear activation function, and λ is a hyperparameter used to balance the weight ratio assigned to neighborhood information and neighborhood interaction information.
6. The method according to claim 1, characterized in that In step (5), in order to better process the node information in the heterogeneous network, the multi-neighborhood extraction module is used to aggregate the drug-disease association network and extract feature information. Among them, GCN extends the operation of convolutional neural network to graph data. GCN assumes that adjacent nodes are independent of each other and uses weighted sum to represent the status of the node. The GA aggregator of the target node (drug r or disease d) is expressed as: where GA represents a nonlinear aggregator, represents the extended neighbors including node v itself, e i represents the feature vector of node i, a vi is the weight of neighbor i, defined as W g is the weight matrix used for feature transformation, and σ represents the activation function; BA type aggregator of target node v: in represents the output features of the target node v in the nonlinear aggregator (BA), represents the number of interactions of node v, e i and e j represents the feature representation of the neighborhood nodes i and j of the target node v; this formula reduces the bias of node degree to a certain extent through standardization, ⊙ represents element-by-element multiplication, and W b is the weight matrix for feature transformation; Then, the encoder constructed in step (5) is used to transfer the information between drugs and diseases, while extracting indirect interactions from local mechanisms. Through GA and BA aggregators, the final feature representation of the target node v is for: Where β is a hyperparameter used to adjust the strength of the GA aggregator and the BA aggregator.
7. The method according to claim 1, characterized in that In step (6), stratified negative sampling is performed through the following three steps: first, the PageRank algorithm is used to score the similarity network of drugs and diseases, denoted as R1, D1, and the scores are sorted in descending order to extract the top L biological molecule information with higher tendency; then, the drug information related to the disease in D1 is screened out through the association information between diseases and drugs, and the obtained drug information related to D1 is merged with the initial drug information set R1 to form a new drug information set, denoted as R2, at this time R2∈{1,2,...,I-1,I},I>L, where R1∈{1,2,...,L}. In order to determine the reliability of the information nodes in R2, the GAT aggregator is used to re-score the information nodes to obtain the new representation q of the nodes. v (v∈{r,d}). For the target node v, the representation of node i is defined as: Where W represents the weight matrix, a ij represents the normalized attention score, x j Represents the corresponding feature representation vector. In order to enhance the reliability of negative sampling, the valuable self-supervisory signal is mined through the joint training system. For the drug r given in step (4), i and d j , the final positive and negative samples are selected by the final learned representation in a single neighborhood: Sc r =softmax(M d q r ) Among them Sc r Indicates the predicted probability that each disease can be cured by drug r in the step; M d Represents the disease global feature matrix, where M d Represented by the neighborhood features synthesized in similar networks Composition, then, through S c Based on the scores, the top L drugs are selected as R3, and the disease set D1 and the drug set R3 are combined into potential drug-disease association pairs, which are recorded as RDdata.
8. The method according to claim 1, characterized in that In step (7), the known positive sample association pairs in RDdata calculated in step (6) are removed, and the remaining part is used as the negative sample candidate set. In order to ensure the robustness of the model's negative sampling strategy, K samples with the same number as the positive samples are randomly selected as reliable negative samples in the final negative sample candidate set. The final drug-disease association prediction is scored using a multilayer perceptron as shown below: in Indicates drug r i and diseases j The probability score of association between For drugs i The embedding vector of For disease j The embedding vector of For drugs i The local characteristics of For disease j The local features of the multilayer perceptron are set to 2, the first layer dimension is set to 128, and the second layer dimension is set to 64. In order to better learn the information interaction between neighbors, InfoNCE is used to learn its lower bound. The triple loss function of the lth layer contrast learning is defined as: Where δ(·) represents the cosine similarity function, N represents the total number of drug nodes, τ represents the temperature hyperparameter that adjusts the smoothness of the similarity distribution, and when j≠i, ε=1; otherwise ε=0; represents the representation of node i after aggregation in step (5), represents the representation of node i obtained after aggregation in step (4), In order to better predict drug-disease interactions, the prediction loss and triple loss are combined to achieve the overall optimization of the model, where the prediction loss is expressed by the weighted cross entropy loss function as shown below: Where N and M represent the number of drug and disease nodes, respectively. represents the final predicted association score, A ij represents the label data in the disease-drug association dataset, (i, j) represents the drug and disease pair, where A + Indicates the currently known DDA, A - Represents the unknown DDA. In order to prevent the model from overfitting and accelerate the model training process, L2 regularization Φ is used to smooth the weights to obtain the final loss calculation function. Loss=L P +Θ1L T +Θ2Φ, Where L P and L T Respectively represent the prediction loss and triple loss in contrastive learning, Θ1 and Θ2 represent the hyperparameters used for balanced loss calculation, Φ uses trainable model parameters, and in order to minimize the loss function, the Adam optimizer is used to select the optimal parameters, including the number of hidden layers L, the number of epochs e, the learning rate lr, and the weight decay wd. The present invention sets the optimal hyperparameter lr to 10-5, e to 100, wd to 10-5, and L to 3.
Citation Information
Cited By
Similarity fusion and mixed gating-based circular RNA and disease association prediction method
CN120748486A