Hypergraph-based collaborative drug prediction method
By constructing hypergraphs and extracting the characteristics of drugs and cell lines, the characteristics selection dependence and generalization of drug combination prediction in the prior art are solved, and higher prediction accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510103197.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-22
AI Technical Summary
The prior art has problems with strong feature selection dependence, limited generalization and neglecting the intrinsic association of cell lines and drug pairs when identifying drug combinations with synergistic effects.
By constructing a hypergraph, the features of drugs and cell lines are extracted using hypergraph node encoder and hypergraph hyper-edge encoder, and inputting them into the hypergraph decoder to reconstruct the hypergraph, calculate the loss function to learn the optimal characteristics, and finally perform link prediction.
Improves the accuracy and robustness of synergistic drug combination predictions, enabling higher prediction accuracy without using any drug and cell line characteristics.
Smart Images

Figure CN119943439A_ABST
Abstract
Description
Technical Field
[0001] The invention discloses a hypergraph-based synergistic drug prediction method, belonging to the technical field of drug synergistic combination prediction. Background Art
[0002] Monotherapy does play an important role in the treatment of many diseases. However, monotherapy is often subject to the dual constraints of drug resistance and limited efficacy when facing complex diseases, and it is difficult to meet clinical needs. Therefore, combination drug therapy has gradually become an important strategy. Through the synergistic effect of multiple drugs, it can not only improve the efficacy, but also reduce the side effects and drug resistance risks caused by single drugs to a certain extent. However, how to identify drug combinations with synergistic effects remains a huge challenge. Although comprehensive clinical trials can verify the effectiveness of drug combinations, their high cost, time-consuming process and large experimental scale make it difficult to achieve in actual operations. Therefore, there is an urgent need to develop efficient and accurate drug combination prediction methods to assist scientists and clinicians in making faster and more scientific decisions in drug development and treatment optimization.
[0003] With the development of machine learning and deep learning, algorithms for predicting synergistic drug combinations continue to emerge and show certain potential. Current methods usually rely on specific feature selection processes. For example, drug features may be represented by SMILES (Simplified molecular input line entry specification) to generate molecular graphs or fingerprint features, or descriptors may be extracted based on the physicochemical properties of drugs; cell line features are mostly based on gene expression data. Subsequently, these features are input into the model to predict the synergistic effects of drug combinations. For example, DeepSynergy uses drug fingerprints generated by jCompoundMapper, physicochemical properties calculated by ChemoPy, and binary toxicant features as drug inputs, while using gene expression profiles to characterize cell line features and simulating synergistic effects between drugs through pyramidal layers. MatchMaker uses drug chemical descriptors and gene expression features calculated by ChemoPy to learn the representation of drugs on specific cell lines through two parallel sub-networks.
[0004] However, such methods still have shortcomings in application. First, they are too dependent on feature selection, and the performance of the model is highly dependent on the quality and completeness of the selected features. Once a drug or cell line feature is missing, it is difficult for the model to predict synergistic drug combinations, resulting in limited generalization. Secondly, these methods mainly focus on the relationship between features and ignore the intrinsic correlation between cell lines and drug pairs. However, this correlation is essentially a kind of graph structure information, which can more comprehensively reflect the complex interactions between cell lines and drug pairs. If this association relationship can be constructed as a graph structure and combined with existing methods, it is expected to significantly enhance the model's ability to capture synergy, thereby further improving the accuracy and robustness of the prediction. Summary of the invention
[0005] The purpose of the present invention is to provide a hypergraph-based synergistic drug prediction method, aiming to solve the technical problem that it is difficult to effectively identify drug combinations with synergistic effects when multiple drugs act synergistically.
[0006] To achieve the above object, the technical solution of the present invention is: a synergistic drug prediction method based on a hypergraph, the specific steps are:
[0007] Step 1: Based on the relationship between drug combinations and cell lines, a hypergraph is constructed with drug pairs as nodes and cell lines as hyperedges.
[0008] Step 2: Use the hypergraph node encoder to extract drug features;
[0009] Step 3: Use the hypergraph hyperedge encoder to extract cell line features;
[0010] Step 4: Input drug features and cell line features into the hypergraph decoder, reconstruct the hypergraph and calculate the loss function to learn the optimal drug features and cell line features;
[0011] Step 5: Perform link prediction based on the learned optimal drug features and cell line features.
[0012] The Step 1 is specifically as follows:
[0013] Step 1.1: Read the synergistic drug dataset to obtain the triplet of (drug 1, drug 2, cell line) and the corresponding synergistic and non-synergistic labels. The synergistic and non-synergistic labels are the results of the drug combination analysis model based on the preset threshold. The values greater than the preset threshold are considered as synergistic, and the values less than or equal to the preset threshold are considered as non-synergistic.
[0014] Step 1.2: Treat the triplet (drug 1, drug 2, cell line) as a binary pair (drug pair, cell line) and construct the association matrix H∈R of the hypergraph G N×C, where N is the number of drug pairs, C is the number of cell lines, and the columns are regarded as hyperedges and the rows as nodes;
[0015] Step 1.3: Randomly select a preset proportion of positive samples and the same number of negative samples in the hypergraph, use the combination of positive and negative samples as the training set and clear them to zero in the hypergraph.
[0016] The Step 2 is specifically as follows:
[0017] Step 2.1: Map the hypergraph’s incidence matrix H to the latent space Calculated by the following formula:
[0018]
[0019] Among them, W v and b v are the learned weights and biases;
[0020] Step 2.2: Get the mean μ through the linear module v and variance σ v , calculated by the following formula:
[0021]
[0022] in, and are two learnable weights, and It is a deviation;
[0023] Step 2.3: Reparameterize to get the final drug pair representation Z v , calculated by the following formula:
[0024] Z v =μ v +σ v ⊙∈
[0025] in, represents the generation of random values from a standard normal distribution and ⊙ represents the element-wise product.
[0026] The Step 3 is specifically as follows:
[0027] Step 3.1: Map the hypergraph’s incidence matrix H to the latent space Calculated by the following formula:
[0028]
[0029] Where T represents the matrix transpose operation, W E and b E are the learned weights and biases;
[0030] Step 3.2: Get the mean μ through the linear module E and variance σ E , calculated by the following formula:
[0031]
[0032] in, and are two learnable weights, and It is a deviation;
[0033] Step 3.3: Reparameterize to get the final cell line representation Z E , calculated by the following formula:
[0034] Z E =μ E +σ E ⊙∈
[0035] in, ⊙ represents element-wise product.
[0036] The Step 4 is specifically as follows:
[0037] Step 4.1: Use the extracted drug features and cell line features to reconstruct the hypergraph G re , calculated by the following formula:
[0038] G re =Sigmoid(Z v (Z E ) T )
[0039] Among them, Sigmoid is the activation function, and T represents the matrix transposition operation;
[0040] Step 4.2: Calculate the binary cross entropy loss between the reconstructed graph and the original hypergraph Calculated by the following formula:
[0041]
[0042] Among them, BCE represents the binary cross entropy function, which is in the form of:
[0043]
[0044] Among them, x i represents the predicted value of the i-th sample after reconstructing the hypergraph; y i Represents the true value of the i-th sample of the hypergraph;
[0045] Step 4.3: Calculate the KL divergence loss of drugs and cell lines using the following formula:
[0046]
[0047] in, represents the KL divergence of the drug, represents the KL divergence of the cell line, and represents the variance and mean of the ith drug, and represents the variance and mean of the ith cell line, Indicates expected value;
[0048] Step 4.3: Binary cross entropy loss Add the KL divergence loss to get the total loss, which is calculated by the following formula:
[0049]
[0050] in, is the total loss.
[0051] The Step 5 is specifically as follows:
[0052] Step 5.1: Final representation of drugs and cell lines based on training and Where d is the embedding dimension, for each pair of links (a, b), the link feature is calculated using the Hadamard product Calculated by the following formula:
[0053] f a,b =e a ⊙e b
[0054] Among them, e a represents the embedding vector of drug a, e b represents the embedding vector of cell line b;
[0055] Step 5.2: Use the link features to train the classifier based on the gradient boosting tree model, with the goal of minimizing the loss function of the binary classification task Calculated by the following formula:
[0056]
[0057] Where M is the number of training samples, y i is the true value of the i-th sample, is the predicted probability of the model for the i-th sample.
[0058] The beneficial effects of the present invention are as follows: the present invention first constructs a hypergraph based on the relationship between drug combinations and cell lines; secondly, uses a hypergraph node encoder to learn drug features; then uses a hypergraph hyperedge encoder to learn cell line features; then the two learned features are input into a hypergraph decoder to reconstruct the hypergraph and calculate the loss to learn the optimal features; finally, the learned features are used to calculate the Hadamard product for link prediction. Compared with the traditional method, the experimental results of the present invention show that the method proposed by the present invention can improve the accuracy of synergistic drug combination prediction without using any drug and cell line features, which helps to open up new horizons in the field of drug synergistic combination prediction technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 It is the overall framework diagram of the present invention;
[0060] Figure 2 It is a flow chart of link prediction of the present invention. DETAILED DESCRIPTION
[0061] The present invention will be further described below in conjunction with the accompanying drawings and specific implementation methods.
[0062] Example 1: Figure 1 As shown, a synergistic drug prediction method based on a hypergraph, the method steps are as follows:
[0063] Step 1: Based on the relationship between drug combinations and cell lines, a hypergraph is constructed with drug pairs as nodes and cell lines as hyperedges.
[0064] like Figure 1 (a), specifically:
[0065] Step 1.1: Read the synergistic drug dataset in CSV format, and obtain the triples of (drug 1, drug 2, cell line) and the corresponding synergistic and non-synergistic labels. The synergistic and non-synergistic labels are the result of the Loewe additivity model with a threshold of 30. Values greater than 30 are considered synergistic, and values less than or equal to 30 are considered non-synergistic.
[0066] Step 1.2: Treat the triplet (drug 1, drug 2, cell line) as a binary pair (drug pair, cell line) and construct the association matrix H∈R of the hypergraph G N×C Where N is the number of drug pairs, C is the number of cell lines, and the columns are regarded as hyperedges and the rows as nodes.
[0067] Step 1.3: Randomly select 20% of the positive samples and the same number of negative samples in the hypergraph, combine the positive and negative samples as the training set and clear them to zero in the hypergraph.
[0068] Step 2: Use the hypergraph node encoder to extract drug features.
[0069] like Figure 1 (b), specifically:
[0070] Step 2.1: Map the hypergraph’s incidence matrix H to the latent space Calculated by the following formula:
[0071]
[0072] Among them, W v and b v are the learned weights and biases;
[0073] Step 2.2: Get the mean μ through the linear module v and variance σ v , calculated by the following formula:
[0074]
[0075] in, and are two learnable weights, and It is a deviation;
[0076] Step 2.3: Reparameterize to get the final drug pair representation Z v , calculated by the following formula:
[0077] Z v =μ v +σ v ⊙∈
[0078] in, represents the generation of random values from a standard normal distribution and ⊙ represents the element-wise product.
[0079] Step 3: Use the hypergraph hyperedge encoder to extract cell line features.
[0080] like Figure 1 (c), specifically:
[0081] Step 3.1: Map the hypergraph’s incidence matrix H to the latent space Calculated by the following formula:
[0082]
[0083] Where T represents the matrix transpose operation, W E and b E are the learned weights and biases;
[0084] Step 3.2: Get the mean μ through the linear module Eand variance σ E , calculated by the following formula:
[0085]
[0086] in, and are two learnable weights, and It is a deviation;
[0087] Step 3.3: Reparameterize to get the final cell line representation Z E , calculated by the following formula:
[0088] Z E =μ E +σ E ⊙∈
[0089] in, ⊙ represents element-wise product.
[0090] Step 4: Input the drug features and cell line features into the hypergraph decoder, reconstruct the hypergraph and calculate the loss function to learn the optimal drug features and cell line features.
[0091] like Figure 1 (d), specifically:
[0092] Step 4.1: Use the extracted drug features and cell line features to reconstruct the hypergraph G re , calculated by the following formula:
[0093] G re =Sigmoid(Z v (Z E ) T )
[0094] Among them, Sigmoid is the activation function, and T represents the matrix transposition operation;
[0095] Step 4.2: Calculate the binary cross entropy loss between the reconstructed graph and the original hypergraph Calculated by the following formula:
[0096]
[0097] Among them, BCE represents the binary cross entropy function, which is in the form of:
[0098]
[0099] Among them, x i represents the predicted value of the i-th sample after reconstructing the hypergraph; y i Represents the true value of the i-th sample of the hypergraph;
[0100] Step 4.3: Calculate the KL divergence loss of drugs and cell lines using the following formula:
[0101]
[0102] in, represents the KL divergence of the drug, represents the KL divergence of the cell line, and represents the variance and mean of the ith drug, and represents the variance and mean of the ith cell line, Indicates expected value;
[0103] Step 4.3: Binary cross entropy loss Add the KL divergence loss to get the total loss, which is calculated by the following formula:
[0104]
[0105] in, is the total loss.
[0106] Step 5: Perform link prediction based on the learned optimal drug features and cell line features.
[0107] like Figure 2 As shown, specifically:
[0108] Step 5.1: Final representation of drugs and cell lines based on training and Where d is the embedding dimension, which is 100. For each pair of links (a, b), the link feature is calculated using the Hadamard product. Calculated by the following formula:
[0109] f a,b =e a ⊙e b
[0110] Among them, e a represents the embedding vector of drug a, e b represents the embedding vector of cell line b;
[0111] Step 5.2: Use the link features to train the classifier based on the gradient boosting tree model, with the goal of minimizing the loss function of the binary classification task Calculated by the following formula:
[0112]
[0113] Where M is the number of training samples, y i is the true value of the i-th sample, is the predicted probability of the model for the i-th sample.
[0114] In order to test the effectiveness of the present invention, it is applied to the O'Neill dataset. In this embodiment, 37 drugs and 31 cell lines are collected from the O'Neill dataset, each cell line has 546 samples, a total of 16,926 samples, including 1,622 positive samples and 15,304 negative samples.
[0115] To evaluate the performance of the model, it was compared with 2 classic machine learning methods: decision tree and random forest.
[0116] In the O'Neill dataset, all data are randomly divided into training and test sets, using 80% of the data for training and the remaining 20% for testing. Then, five-fold cross validation is used on the training set to optimize the model hyperparameters. The final result is the average result of the five-fold cross validation.
[0117] The gradient boosting tree model was selected as the base learner used in the experiment, and the early stopping strategy was used. In order to comprehensively evaluate the performance of the model proposed in this invention and all the comparative methods, a variety of classification performance indicators were selected, including the area under the receiver operating characteristic curve (AUC), AUPR (area under the precision-recall curve) and precision. These indicators can reflect the predictive ability and generalization performance of the model from different angles. The specific calculation method is as follows:
[0118]
[0119] in, represents the true positive rate, represents the false positive rate. P represents precision, and R represents recall. TP and FP represent the number of true positive cases and the number of false positive cases, respectively.
[0120] The results in Table 1 show that the present invention achieves the best classification results in all cases of the three evaluation indicators. The bold parts of the text in Table 1 are the best results.
[0121] Table 1 Comparison of experimental results
[0122]
[0123]
[0124] In summary, after comparison with other prediction methods, the effectiveness of the hypergraph-based synergistic drug prediction method was demonstrated.
[0125] The specific implementation modes of the present invention are described in detail above in conjunction with the accompanying drawings, but the present invention is not limited to the above implementation modes, and various changes can be made within the knowledge scope of ordinary technicians in this field without departing from the purpose of the present invention.
Claims
1. A hypergraph-based synergistic drug prediction method, characterized in that: The steps include: Step 1: Based on the relationship between drug combinations and cell lines, a hypergraph is constructed with drug pairs as nodes and cell lines as hyperedges. Step 2: Use the hypergraph node encoder to extract drug features; Step 3: Use the hypergraph hyperedge encoder to extract cell line features; Step 4: Input drug features and cell line features into the hypergraph decoder, reconstruct the hypergraph and calculate the loss function to learn the optimal drug features and cell line features; Step 5: Perform link prediction based on the learned optimal drug features and cell line features.
2. The hypergraph-based collaborative drug prediction method according to claim 1, characterized in that: The Step 1 is specifically as follows: Step 1.1: Read the synergistic drug dataset to obtain the triplet of (drug 1, drug 2, cell line) and the corresponding synergistic and non-synergistic labels. The synergistic and non-synergistic labels are the results of the drug combination analysis model based on the preset threshold. The values greater than the preset threshold are considered as synergistic, and the values less than or equal to the preset threshold are considered as non-synergistic. Step 1.2: Treat the triplet (drug 1, drug 2, cell line) as a binary pair (drug pair, cell line) and construct the association matrix H∈R of the hypergraph G N×C , where N is the number of drug pairs, C is the number of cell lines, and the columns are used as hyperedges and the rows as nodes; Step 1.3: Randomly select a preset proportion of positive samples and the same number of negative samples in the hypergraph, use the combination of positive and negative samples as the training set and clear them to zero in the hypergraph.
3. The hypergraph-based collaborative drug prediction method according to claim 1, characterized in that: The Step 2 is specifically as follows: Step 2.1: Map the hypergraph’s incidence matrix H to the latent space Calculated by the following formula: Among them, W v and b v are the learned weights and biases; Step 2.2: Get the mean μ through the linear module v and variance σ v , calculated by the following formula: in, and are two learnable weights, and It is a deviation; Step 2.3: Reparameterize to get the final drug pair representation Z v , calculated by the following formula: Z v =μ v +s v ⊙∈ in, represents the generation of random values from a standard normal distribution and ⊙ represents the element-wise product.
4. The hypergraph-based collaborative drug prediction method according to claim 1, characterized in that: The Step 3 is specifically as follows: Step 3.1: Map the hypergraph’s incidence matrix H to the latent space Calculated by the following formula: Where T represents the matrix transpose operation, W E and b E are the learned weights and biases; Step 3.2: Get the mean μ through the linear module E and variance σ E , calculated by the following formula: in, and are two learnable weights, and It is a deviation; Step 3.3: Reparameterize to get the final cell line representation Z E , calculated by the following formula: Z E =μ E +s E ⊙∈ in, ⊙ represents element-wise product.
5. The hypergraph-based collaborative drug prediction method according to claim 1, characterized in that: The Step 4 is specifically as follows: Step 4.1: Use the extracted drug features and cell line features to reconstruct the hypergraph G re , calculated by the following formula: G re =Sigmoid(Z v (Z E ) T ) Among them, Sigmoid is the activation function, and T represents the matrix transposition operation; Step 4.2: Calculate the binary cross entropy loss between the reconstructed graph and the original hypergraph Calculated by the following formula: Among them, BCE represents the binary cross entropy function, which is in the form of: Among them, x i represents the predicted value of the i-th sample after reconstructing the hypergraph; y i Represents the true value of the i-th sample of the hypergraph; Step 4.3: Calculate the KL divergence loss of drugs and cell lines using the following formula: in, represents the KL divergence of the drug, represents the KL divergence of the cell line, and represents the variance and mean of the ith drug, and represents the variance and mean of the ith cell line, Indicates expected value; Step 4.3: Binary cross entropy loss Add the KL divergence loss to get the total loss, which is calculated by the following formula: in, is the total loss.
6. The hypergraph-based collaborative drug prediction method according to claim 1, characterized in that: The Step 5 is specifically as follows: Step 5.1: Final representation of drugs and cell lines based on training and Where d is the embedding dimension, for each pair of links (a, b), the link feature is calculated using the Hadamard product Calculated by the following formula: f a,b =e a ⊙e b Among them, e a represents the embedding vector of drug a, e b represents the embedding vector of cell line b; Step 5.2: Use the link features to train the classifier based on the gradient boosting tree model, with the goal of minimizing the loss function of the binary classification task Calculated by the following formula: Where M is the number of training samples, y i is the true value of the i-th sample, is the predicted probability of the model for the i-th sample.
Citation Information
Patent Citations
Hypergraph-based drug-target-disease interaction prediction method
CN113066526A
Anticancer drug collaborative prediction method based on multivariate data enhanced hypergraph neural network
CN115295163A
Training method and device of drug intermolecular interaction prediction model
CN117133382A
Prediction model construction method, prediction method and device for drug combination synergistic effect
CN118412146A