GAT double attention-based anticancer drug collaborative prediction method

By adopting the GAT dual attention mechanism method in drug synergistic prediction, the problems of high time-consuming and limited prediction performance in traditional methods are solved, and more efficient and explainable drug synergistic prediction is achieved.

CN120183549APending Publication Date: 2025-06-20YUNNAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510266316.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

Traditional drug combination screening and verification processes are time-consuming and costly, and existing computing models are difficult to effectively capture the complex interactive information between drug combinations, resulting in limited predictive performance.

Method used

The coordinated prediction method of anti-cancer drugs based on GAT dual attention is adopted. The graph attention network model combined with the dual attention mechanism is used to capture the local structural characteristics and global interaction information of the drug molecular graph, and the interaction information of the drug combination is weighted and aggregated through the dual attention mechanism, and finally synergistic prediction is performed through the full connection layer.

Benefits of technology

It significantly improved the accuracy and consistency of drug synergy prediction, especially in the KAPPA indicators, and provided a key molecular structure explanation of drug synergy through attention mechanisms, enhancing the interpretability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120183549A_ABST
    Figure CN120183549A_ABST
Patent Text Reader

Abstract

The invention discloses an anti-cancer drug collaborative prediction method based on GAT double attention, which is used for predicting the synergistic effect of anti-cancer drugs and comprises the following steps: acquiring the molecular structure and cell line characteristic data of the anti-cancer drugs; constructing and modifying a graph attention network (GAT) model, introducing a double-attention mechanism into the model, and respectively capturing local structure features and global interaction information of the drug molecular graph; performing weighted aggregation on the interaction information of the medicine combination through a double-attention mechanism; and fusing the extracted features with cell line data, and finally performing synergistic effect prediction through a full connection layer. According to the method, the capacity of the model for capturing the complex relation between drug molecules is remarkably improved, particularly, the KAPPA index is remarkably improved, and the consistency and reliability of the model are enhanced. The key molecular structure explanation of the drug synergistic effect is provided through the double attention mechanism, and a new technical means is provided for the combination design of anti-cancer drugs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of drug discovery, and particularly relates to a method for predicting the synergy of anti-cancer drugs based on GAT double attention. Background Art

[0002] In the field of cancer treatment, single-drug therapies often struggle to address the complex biological mechanisms of tumors. Therefore, combination drug therapies have gradually become the mainstream treatment strategy. Combination drug therapies can not only significantly improve the treatment effect but also effectively reduce toxic side effects and overcome drug resistance. However, the traditional process of screening and validating drug combinations is time-consuming and costly. Especially when performing high-throughput screening through experimental methods, it faces huge time and resource pressures. In addition, drug combinations may produce antagonistic effects or severe adverse reactions, further increasing the complexity and risk of drug development.

[0003] In recent years, with the rapid development of machine learning and deep learning technologies, researchers have proposed various computational models to predict the synergy of drug combinations. Traditional machine learning methods, such as support vector machines (SVM), random forests (RF), etc., rely on manually designed feature engineering and are difficult to capture the complex interaction information between drug combinations, resulting in limited prediction performance. Although these methods have reduced the screening space of drug combinations to a certain extent, their prediction accuracy and generalization ability still need to be improved, especially when dealing with large-scale and high-dimensional drug data.

[0004] With the progress of deep learning technology, especially the successful application of graph neural networks (GNNs) and attention mechanisms in the field of drug discovery, researchers have begun to explore deep learning-based methods for predicting drug synergy. In previous studies, a variety of models have been proposed and achieved certain results. For example, "DTSyn: a dual-transformer-based neural network to predict synergistic drug combinations" proposed a deep neural network model DTSyn based on the multi-head attention mechanism, with an accuracy of 0.81, but the TPR performance was unstable on other independent datasets; "SynergyX: a multi-modality mutual attention network for interpretable drug synergy prediction" proposed a multi-modal mutual attention network, significantly improving the prediction performance of antitumor drug synergy, but due to memory limitations, it was unable to attract attention at single-gene resolution in cell lines, reducing the interpretability accuracy of the model; "SynerGNet: A Graph Neural Network Model to Predict Anticancer Drug Synergy" developed a graph neural network model with an accuracy of 0.748 on the original data and augmented datasets, but its performance metrics were not outstanding compared to attention-based models; "AttenSyn: An Attention-Based Deep Graph Neural Network for Anticancer Synergistic Drug Combination Prediction" proposed an attention-based deep graph neural network, combining GCN and LSTM and introducing the attention mechanism, with an accuracy of 0.84, but it only used molecular structure information and cell line characteristics for prediction and failed to fully utilize the global interaction information of drug combinations. Summary of the Invention

[0005] To overcome the problems in the background art, the present invention provides an anticancer drug synergy prediction method based on GAT double attention.

[0006] To achieve the above object, the present invention is implemented through the following technical solutions:

[0007] An anticancer drug synergy prediction method based on GAT double attention, comprising the following steps:

[0008] S1: Obtain the molecular structure of anticancer drugs and cell line characteristic data and construct a dataset;

[0009] S2: Construct a graph attention network model and introduce a dual attention mechanism into the model to capture the local structural features and global interaction information of the drug molecule graph respectively;

[0010] S3: Weightedly aggregate the interaction information of the drug combination through the dual attention mechanism;

[0011] S4: Fuse the extracted features with the cell line data and finally perform synergy prediction through a fully connected layer.

[0012] Further, in the S1, the anti-cancer drug molecular structure data includes the SMILES representation of the drug, and the cell line feature data includes the gene expression data of the cancer cell line. The data is sourced from DrugBank and the Cancer Cell Line Encyclopedia.

[0013] Further, the construction of the dataset in the S1 includes the following steps:

[0014] S1a: Use O'Neil's drug combination dataset as the benchmark dataset. The dataset contains 23,052 triples, and each triple consists of two drugs and one cancer cell line;

[0015] S1b: Calculate the synergy score for each pair of drugs through the Combenefit tool. Select 10 as the threshold to classify the drug-cell line triples. Triples with a score higher than 10 represent that the drug combination shows synergy in this cell line, and triples with a score lower than 0 indicate that the drug combination shows antagonism in this cell line;

[0016] S1c: Finally, obtain 13,243 unique triples, covering 38 drugs and 31 cell lines.

[0017] Further, in the S2, the graph attention network model adopts a multi-head attention mechanism with 4 heads. The number of layers of the graph convolution module is 2. The input dimension of the first layer is 78, and the output dimension is 128. The input dimension of the second layer is 128, and the output dimension is 128.

[0018] Further, in the S2, the dual attention mechanism includes two independent attentions, which are respectively used to process local information and global information. The input dimension of each dual attention layer is 128, the output dimension is 128, the number of attention heads is 4, and the dimension of each head is 32.

[0019] Further, in the S3, the dual attention mechanism generates a pooled feature representation by calculating the interaction attention weights between the two drug molecular features, specifically including:

[0020] S3a: Calculate the query, key, and value through linear transformation;

[0021] S3b: Calculate the attention coefficients between two graphs;

[0022] S3c: Perform softmax normalization on the attention coefficients;

[0023] S3d: Use the attention weights to perform weighted summation on the values to generate the pooled feature representation.

[0024] Furthermore, in S4, the fully connected layer maps the high-dimensional features to a low-dimensional space and finally outputs the classification result. The output dimension is 2, respectively representing the probabilities of drug synergy.

[0025] Advantages of the present invention:

[0026] 1. The present invention uses a deep learning model to automatically capture the features of drug molecules and cell lines, avoiding the cumbersome manual feature engineering in traditional machine learning methods, simplifying the model structure, and at the same time improving the efficiency and accuracy of feature extraction.

[0027] 2. The anti-cancer drug synergy prediction method based on GAT double attention proposed by the present invention combines the advantages of GAT and the double attention mechanism, aiming to simultaneously capture the local structural features and global interaction information of the drug molecular graph. Through the double attention mechanism, the model can dynamically allocate weights between drug combinations, significantly improving the accuracy and consistency of prediction, especially achieving a significant improvement in the KAPPA index. In addition, the present invention provides a key molecular structure explanation for drug synergy through the attention mechanism, enhancing the interpretability of the model and providing a new technical means for anti-cancer drug combination design.

[0028] 3. The present invention combines GAT with the double attention mechanism, overcomes the limitations of existing models in dealing with complex drug interaction information, and provides an efficient and reliable solution for the field of drug synergy prediction.

[0029] 4. The present invention can predict the synergy effect of drug combinations according to drug combinations and cancer types, providing an efficient and reliable tool for the screening and optimization of drug combinations, helping to accelerate the research and development process of anti-cancer drugs, reducing experimental costs, and providing a scientific basis for personalized cancer treatment. Description of the Drawings

[0030] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0031] Figure 1 is the flowchart of the model usage of the present invention;

[0032] Figure 2 is the schematic diagram of the spatial attention module in the dual attention of the present invention;

[0033] Figure 3 is the schematic diagram of the channel attention module in the dual attention of the present invention;

[0034] Figure 4 is the flowchart of the GAT dual attention processing of the present invention;

[0035] Figure 5 is the structural diagram of the GAT dual attention model of the present invention;

[0036] Figure 6 is the comparison chart of each index in the GAT dual attention model and the AttenSyn model of the present invention; Specific Embodiments

[0037] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0038] Embodiment 1

[0039] A method for predicting the synergy of anti-cancer drugs based on GAT dual attention disclosed by the present invention improves the accuracy of predicting drug synergy while ensuring that the prediction speed is not significantly lost. It mainly uses a deep learning model for drug synergy based on GAT and DualAttention. The deep learning model includes a molecular feature extraction module, a graph convolution module, a dual attention pooling module, and an output module; the drug interaction feature extraction module is used to extract features from the input triple, the graph convolution module is used to perform convolution operations on the triple, the dual attention pooling module is used to perform pooling operations on the convolved features, and the output module is used to generate the final prediction result.

[0040] In this embodiment, the centers of the drug synergy feature extraction module, the graph convolution module, the dual attention pooling module, and the output module are in the same computational process, and the input triple data is processed sequentially.

[0041] In this embodiment, the drug synergy feature extraction module consists of two independent feature extraction sub-modules, which are respectively used to process the two drugs in the input triple. Each sub-module includes a linear layer, a ReLU activation function, and a Dropout layer. The input dimension of the linear layer is 78, and the output dimension is 2048. After passing through the ReLU activation function and the Dropout layer, it is further reduced to 512, and finally a 256-dimensional feature vector is output.

[0042] In this embodiment, the graph convolution module consists of multiple graph attention convolution layers (GATConv). The input dimension of each graph attention convolution layer is the same as the output dimension of the previous layer, and each convolution layer adopts a multi-head attention mechanism with 4 heads. The number of layers of the graph convolution module is 2. The input dimension of the first layer is 78, and the output dimension is 128. The input dimension of the second layer is 128, and the output dimension is 128. The output of the graph convolution module is subjected to sequence modeling through an LSTM layer, and the hidden state dimension of the LSTM layer is 128.

[0043] In this embodiment, the dual attention pooling module consists of two dual attention layers (DualAttention). The input dimension of each dual attention layer is 128, and the output dimension is 128. The dual attention layer generates a pooled feature representation by calculating the interactive attention weights between the two drug molecule features. The number of attention heads of the dual attention layer is 4, and the dimension of each head is 32.

[0044] In this embodiment, the output module consists of multiple linear layers and ReLU activation functions. The input dimension is 4×128 + 256, and after being reduced in dimension by the linear layer to 64, the final output dimension is 2. The ReLU activation function and the Dropout layer are used to connect between the linear layers of the output module, and the dropout rate of the Dropout layer is 0.2.

[0045] In this embodiment, the input of the drug synergy feature extraction module is the feature representations of the two drug molecules in the triple, and the dimension of the feature representation is 78. The feature extraction module maps the input features to a high-dimensional space through a linear layer, and then performs feature selection and dimensionality reduction through the ReLU activation function and the Dropout layer, and finally outputs a 256-dimensional feature vector.

[0046] In this embodiment, the input of the graph convolution module is the feature vector output by the drug synergy feature extraction module and the edge index information of the drug molecular graph. The graph convolution module aggregates the node features in the drug molecular graph through the multi-head attention mechanism to extract the local and global information of the drug molecular graph. The convolved features are subjected to sequence modeling through the LSTM layer to generate the final graph convolution feature representation.

[0047] In this embodiment, the input of the dual attention pooling module is the feature representation output by the graph convolution module and the mask information in the triple. The dual attention layer generates the pooled feature representation by calculating the interactive attention weights between two drug molecular features. The pooled feature representation is used to capture the interaction information between drug molecules.

[0048] In this embodiment, the input of the output module is the feature representation output by the dual attention pooling module and the feature vector output by the drug synergy feature extraction module. The output module splices multiple feature representations and performs dimensionality reduction through a multi-layer neural network, and finally outputs the prediction result. The final output dimension of the output module is 2, respectively representing the probability of drug synergy.

[0049] In this embodiment, the model is trained using the cross-entropy loss function, the optimizer uses the Adam optimizer, the learning rate is 0.0005, the training batch size is 128, and the number of training epochs is 500. During the training process, the learning rate is reduced to 0.7 times every 100 epochs to optimize the model convergence effect. The training logs and result files are dynamically generated, recording the key metrics during the training process, including AUC, ACC, BACC, PREC, TPR, KAPPA, RECALL, Precision, and F1 score.

[0050] Compared with the prior art, the present invention has the following significant advantages:

[0051] Automated feature extraction: The present invention uses a deep learning model to automatically capture the features of drug molecules and cell lines, avoiding the cumbersome manual feature engineering in traditional machine learning methods, simplifying the model structure, and at the same time improving the efficiency and accuracy of feature extraction.

[0052] Model performance improvement: The present invention improves on the basis of the AttenSyn model, using the graph attention network (GAT) combined with the dual attention mechanism (DualAttention), which improves the model performance. The specific manifestations are as follows:

[0053] ACC metric: The ACC of AttenSyn is 0.84, while GAT-Dual is improved to 0.86, indicating that the present invention has a partial improvement in classification accuracy.

[0054] PR_AUC metric: The PR_AUC of AttenSyn is 0.91, and GAT-Dual maintains the same 0.91, indicating that the present invention maintains a high level in terms of the area under the precision-recall curve.

[0055] BACC metric: The BACC of AttenSyn is 0.84, while GAT-Dual is improved to 0.86, indicating that the present invention is more robust on class-imbalanced data.

[0056] KAPPA coefficient: The KAPPA coefficient of AttenSyn is 0.67, while GAT-Dual is improved to 0.72, indicating that the consistency between the classification results of the model of the present invention and the true labels is significantly improved.

[0057] F1 Score metric: The F1 Score of AttenSyn is 0.83, and GAT-Dual is improved to 0.85, indicating that the present invention achieves a better balance between precision and recall.

[0058] Example 2

[0059] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples.

[0060] The present invention does not mention the hardware facilities for training because there will be a certain error value in the results generated by each complete training during the training process of the model. Therefore, any hardware facilities can be used to implement this method, but the error of the results needs to be understood in advance.

[0061] The present application provides an anti-cancer drug combination prediction method based on a GAT dual-attention model, and the method includes:

[0062] S1: Establish a data set: Collect the data on the synergistic effects of different types of drugs on cancer cell lines.

[0063] Among them, S1 includes two sub-steps:

[0064] S1a: Use the drug combination data set constructed by O'Neil et al. as the benchmark data set. This is an anti-cancer synergistic drug combination data set, which contains the molecular structure information of drugs and their synergistic effect data in different cell lines. The data set contains 23,052 triples, where each triple consists of two drugs and one cancer cell line. There are 39 cancer cell lines and 38 unique drugs in the data set. These drugs consist of 24 FDA-approved drugs and 14 experimental drugs.

[0065] The SMILES (Simplified Molecular Input Line Entry System) of the drugs were obtained from DrugBank, and the gene expression data of the cancer cell lines were from the Cancer Cell Line Encyclopedia (CCLE).

[0066] S1b: The synergy scores for each pair of drugs were calculated using the Combenefit tool, which is specifically designed to analyze and quantify the synergistic effects of drug combinations. It supports multiple classical synergy models, including Loewe, Bliss, and HSA (Highest Single Agent), and can handle batch data from single experiments or high-throughput screens. Based on previous studies, a threshold of 10 was chosen to classify drug-cell line triples. Triples with scores higher than 10 represent drug combinations that exhibit synergy in that cell line; triples with scores lower than 0 indicate drug combinations that exhibit antagonism in that cell line.

[0067] After scoring, a total of 13,243 unique triples were obtained, covering 38 drugs and 31 cell lines. These drugs include 24 FDA-approved drugs and 14 experimental drugs, ensuring the diversity and comprehensiveness of the dataset.

[0068] S2: Implement a dataset class and a data loading tool for processing graph-structured data.

[0069] Step S2 consists of three sub-steps:

[0070] S2a: Implement a dataset class. MyDataset is a custom dataset class that inherits from torch.utils.data.Dataset and is used to load and manage graph-structured data.

[0071] If no data is provided, the data is loaded from the file. / data / data.pt. If data is provided, the provided data is used directly. The length of the dataset is returned by the __len__ method; a single data sample is returned according to the index by the __getitem__ method; a new MyDataset object is returned according to the slice by the get_data method.

[0072] The data file data.pt is a list containing multiple samples, and each sample is a tuple in the format (graph1, graph2, cell, label):

[0073] graph1 and graph2 are torch_geometric.data.Data objects representing the SMILES of two drugs. cell represents the gene expression data of the cancer cell line. label is the label of the sample, indicating whether it is a synergistic or antagonistic effect.

[0074] S2b: Implement collate. Collate is a custom data loading function used to pack multiple samples into a batch. Input data, where data is a list containing multiple samples, and each sample is a tuple (graph1, graph2, cell, label).

[0075] The processing logic of this function is as follows: Traverse each sample, store graph1, graph2, and label into d1_list, d2_list, and label_list respectively; assign cell to graph1.cell for accessing global features in the model; use Batch.from_data_list to pack the graph data in d1_list and d2_list into a batch; convert label_list to torch.tensor.

[0076] After processing, the function will return three objects: Batch.from_data_list(d1_list): The packed batch of graph1. Batch.from_data_list(d2_list): The packed batch of graph2. torch.tensor(label_list): The packed label tensor.

[0077] S3: Build a drug synergy prediction method based on the improved model of GAT and DualAttention. The overall structure of the model is as Figure 5 shown, including Graph Convolution Layers, Recurrent Neural Network Layer, Fully Connected Layers, Attention Mechanism Layers, Normalization Layers, and Dropout layer.

[0078] Among them, step S3 specifically includes:

[0079] S3a: Construct Graph Convolution Layers. One of the core parts of the model is the Graph Convolution Layers, and the convolution operation of the Graph Attention Network (GAT) is used. GAT is a graph convolution method based on the attention mechanism, which can assign different weights to each node in the graph, thus better capturing the relationships between nodes. Each convolution layer adopts the multi-head attention mechanism, and the number of heads is 4. The number of layers of the graph convolution module is 2. The input dimension of the first layer is 78, and the output dimension is 128. The input dimension of the second layer is 128, and the output dimension is 128.

[0080] The input of the graph convolution layer is the feature vector xi of each node and the edge information (adjacency matrix or edge index) of the graph.

[0081] The attention mechanism of GAT calculates the attention coefficient between node i and node j through the following formula:

[0082] e ij = LeakyReLU(a T [Wx i ||Wx j )

[0083] Among them, W is a learnable weight matrix, a is the parameter vector of the attention mechanism, and ∥ represents the concatenation operation of vectors. After the attention coefficient eij is normalized by softmax, the weight αij is obtained:

[0084]

[0085] Finally, the output feature of node i is the weighted sum of its neighbor nodes:

[0086]

[0087] Among them, σ is the activation function.

[0088] The parameter molecule_channels is the dimension of the input features; hidden_channels: the dimension of the hidden layer; heads: the number of heads of the multi-head attention, which is used to enhance the expression ability of the model; layer_count: the stacking times of the GAT layers; through the multi-layer GAT convolution, the model can gradually extract the high-order features of the nodes in the graph and capture the complex relationships between nodes.

[0089] S3b: Construct a Recurrent Neural Network Layer. After each layer of graph convolution, the model uses LSTM (Long Short-Term Memory) to further process the hidden states of the nodes. LSTM is a type of recurrent neural network that can capture long-term dependencies in sequential data.

[0090] The input to this layer is: the output feature hi′ of the graph convolution layer.

[0091] The core of LSTM includes an input gate, a forget gate, an output gate, and a candidate state:

[0092] i t = σ(W i [h t-1 , x t + b i )

[0093] f t = σ(W f [h t-1 , x t + b f )

[0094] o t = σ(W o [h t-1 , x t + b o )

[0095]

[0096]

[0097] h t = o t ⊙ tanh(C t )

[0098] where σ is the sigmoid function, and ⊙ represents element-wise multiplication.

[0099] Parameters hidden_channels: the dimension of the hidden state of LSTM; num_layers: the number of layers of LSTM (here it is 1).

[0100] LSTM can capture the dynamic changes of node features in the sequence, further enhancing the expressive power of the model.

[0101] S3c: Construct a Dual Attention mechanism

[0102] The model introduces a custom dual attention mechanism for interacting and fusing the features of two input graphs. For example Figure 2, Figure 3 As shown, the node features of two graphs are input and summed. The input dimension of each dual attention layer is 128, and the output dimension is 128. The dual attention layer generates a pooled feature representation by calculating the interactive attention weights between the features of two drug molecules. The number of attention heads in the dual attention layer is 4, and the dimension of each head is 32.

[0103] The specific process is as Figure 4 shown, and the process is as follows: First, calculate the query (Query), key (Key), and value (Value) through linear transformation:

[0104] q1 = ReLU(W q x1), k1 = ReLU(W k x1), v1 = ReLU(W v x1)

[0105] Similarly, calculate q2, k2, v2.

[0106] Calculate the attention coefficients between the two graphs:

[0107]

[0108] Perform softmax normalization on the attention coefficients:

[0109] α1 = softmax(a1), α2 = softmax(a2)

[0110] Use the attention weights to perform weighted summation on the values:

[0111] output1 = α1υ1, output2 = α2υ2

[0112] Parameter dim: The dimension of the input features. num_heads: The number of heads in the multi-head attention. dropout_rate: The probability of Dropout, used to prevent overfitting.

[0113] The dual attention mechanism can capture the interactive information between the two graphs, thus better fusing their features.

[0114] S3d: Construct fully connected layers and dimensionality reduction (Fully Connected Layers and Dimensionality Reduction). At the end of the model, a series of fully connected layers are used to reduce the dimension and transform the features.

[0115] Input the features processed by graph convolution, LSTM, and the attention mechanism.

[0116] The calculation method of the fully connected layer is:

[0117] y = σ(Wx + b)

[0118] Where W is the weight matrix, b is the bias vector, and σ is the activation function.

[0119] Parameters hidden_channels: the dimension of the hidden layer. middle_channels: the dimension of the middle layer. out_channels: the dimension of the output layer. The fully connected layer maps high-dimensional features to a low-dimensional space and finally outputs the classification result.

[0120] S3e: Perform normalization and Dropout:

[0121] To improve the stability and generalization ability of the model, layer normalization (LayerNorm) and Dropout are used in the model.

[0122] Calculation method of layer normalization:

[0123]

[0124] Where μ and σ are the mean and standard deviation of the input respectively, and γ and β are learnable parameters.

[0125] Through layer normalization, the training is accelerated and the convergence of the model is improved.

[0126] Dropout is to randomly set the outputs of some neurons to 0 during the training process. Its calculation method is:

[0127] y = dropout(x, p)

[0128] Where p is the probability of Dropout. Through Dropout, overfitting is prevented. The dropout rate of the Dropout layer is 0.2.

[0129] S3f: The final output of the model is obtained by concatenating all features and then passing through the fully connected layer:

[0130] output = Linear(Concat(gcn_hidden_left, gcn_hidden_right, rnn_pooled_left, rnn_pooled_right, cell))

[0131] Where Linear represents the fully connected layer and Concat represents the feature concatenation operation.

[0132] S4: Write code to train the model.

[0133] Where S4 includes the following sub-steps.

[0134] S4a: Set the random seed. To ensure the reproducibility of the experiment, the code sets the random seed. These settings ensure that the random number generation, data loading, and model initialization are consistent each time the code is run.

[0135] S4b: Set up logging. The code uses the logging module to record information during training: the log information will be output to the console and a specified log file simultaneously. The log content includes training progress, loss values, performance metrics, etc.

[0136] S4c: Write the model training function train. The train function is used to train the model.

[0137] The inputs to this function are: model: the model to be trained; device: the training device (such as CPU or GPU); train_loader: the training data loader; optimizer: the optimizer; epoch: the current training epoch; scheduler: the learning rate scheduler.

[0138] S4d: Write the model prediction function predicting. The predicting function is used to evaluate the model performance on the validation data. The inputs to this function are: model: the model to be evaluated; device: the evaluation device (such as CPU or GPU); loader_test: the validation data loader.

[0139] S4e: Conduct model training and evaluation. The main logic part of the code implements model training and evaluation. Set hyperparameters. The batch size is 128, the initial learning rate is 0.0005, and train for 500 epochs. Load the dataset and train five times for five-fold cross-validation. Initialize the model, optimizer, and learning rate scheduler. Set the learning rate to decay, reducing the learning rate to 0.7 times the original every 100 epochs. In each training epoch, call the train function to train the model and call the predicting function to evaluate the model. Calculate performance metrics (such as AUC, accuracy, F1 score, etc.), and save the improved results in the result file. The graphics card used is 4090. Use the PyTorch deep learning framework. Figure 6 Shows the comparison of the effects of our model method with other model methods.

[0140] S5: The model usage process is as Figure 1 shown.

[0141] Among them, S5 contains the following sub-steps.

[0142] S5a: Data input: The input data includes molecular structure and cell line feature data for subsequent model training and prediction.

[0143] S5b: Obtain molecular structure and cell line characteristic data: Extract the molecular structure information of the drug from the database or experiment, including chemical structure, functional groups, etc. Obtain the characteristic data of the cancer cell line, including gene expression, protein expression, metabolic characteristics, etc.

[0144] S5c: Construct a dual-attention GAT model: Build a model framework based on the Graph Attention Network (GAT), and utilize the characteristics of graph-structured data for feature extraction. At the same time, introduce a dual-attention mechanism to enhance the model's ability to capture the interaction information between the drug and the cell line.

[0145] S5d: Capture the local features and global interaction information of the drug: Through the graph convolutional layer of the GAT model, extract the local features of the drug molecule, such as chemical bonds, functional groups, etc. Utilize the dual-attention mechanism to capture the global interaction information between the drug and the cell line, including the impact of the drug on the cell line and the response of the cell line to the drug.

[0146] S5e: The dual-attention mechanism weights and aggregates the interaction information: The dual-attention mechanism weights the interaction information between the drug and the cell line to highlight the key features. Then, through the calculation of the attention weights, aggregate the interaction information between the drug and the cell line to generate a comprehensive feature representation.

[0147] S5f: Fuse the GAT and dual-attention features: Fuse the local features extracted by GAT and the global interaction features generated by the dual-attention mechanism. Integrate the two types of features through feature concatenation or weighted summation to form the final feature representation.

[0148] S5g: The fully connected layer reduces the dimension and concatenates the features: Input the fused features into the fully connected layer for dimensionality reduction to reduce the feature dimension. Through the neuron connections of the fully connected layer, perform a non-linear transformation on the features to extract a higher-level feature representation.

[0149] S5h: Predict the synergistic effect result: Input the dimension-reduced features into the output layer to predict the synergistic effect. Output the prediction result, including the probability or score of the synergistic effect, for evaluating the synergistic effect of the drug combination.

[0150] Analysis of the GAT dual-attention model prediction method: As Figure 6As shown, the experimental results indicate that the model has been improved in multiple key metrics: the KAPPA coefficient has increased from 0.67 to 0.72, the ACC and BACC metrics have both increased from 0.84 to 0.86, the F1 Score has increased from 0.83 to 0.85, and the PR_AUC metric remains at a high level of 0.91. In addition, the model enhances interpretability through a dual attention mechanism, capable of providing key molecular structure explanations for drug synergy, offering a new technical means for the design of anticancer drug combinations. The model dynamically generates training logs and result files during the training process, recording key metrics such as AUC, ACC, BACC, PREC, TPR, KAPPA, RECALL, Precision, and F1 score. Finally, the training and prediction processes of the model can be implemented on any hardware facility, but the error of the results needs to be understood in advance.

[0151] The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of this specification. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the present invention, so that those skilled in the art can well understand and utilize the present invention. The present invention is only limited by the claims and their full scope and equivalents.

Claims

1. A GAT dual attention based anticancer drug synergy prediction method, characterized in that: The following steps are involved: S1: Obtain anticancer drug molecular structure and cell line characteristic data and construct a data set; S2: Construct a GAT model and introduce a dual attention mechanism into the model to capture the local structural features and global interaction information of the drug molecule graph respectively; S3: Weighted aggregation of interactive information of drug combinations through dual attention mechanism; S4: The extracted features are fused with the cell line data and finally the synergy prediction is performed through a fully connected layer.

2. The anticancer drug synergy prediction method based on GAT dual attention according to claim 1, characterized in that: In S1, the anticancer drug molecular structure data includes the SMILES representation of the drug, and the cell line characteristic data includes the gene expression data of the cancer cell line, and the data are derived from DrugBank and the Cancer Cell Line Encyclopedia.

3. The anticancer drug synergy prediction method based on GAT dual attention according to claim 1, characterized in that: The construction of the dataset in S1 includes the following steps: S1a: O'Neil's drug combination dataset is used as a benchmark dataset, which contains 23,052 triplets, each of which consists of two drugs and a cancer cell line. S1b: The synergy score of each drug pair was calculated by Combefit tool, and 10 was selected as the threshold to classify drug-cell line triplets. Triplets with scores higher than 10 indicated that the drug combination showed synergistic effects in the cell line, and triplets with scores lower than 0 indicated that the drug combination showed antagonistic effects in the cell line. S1c: Finally, 13,243 unique triplets were obtained, covering 38 drugs and 31 cell lines.

4. The anticancer drug synergy prediction method based on GAT dual attention according to claim 1, characterized in that: In S2, the graph attention network model adopts a multi-head attention mechanism with 4 heads, 2 layers of the graph convolution module, an input dimension of 78 and an output dimension of 128 for the first layer, and an input dimension of 128 and an output dimension of 128 for the second layer.

5. The anticancer drug synergy prediction method based on GAT dual attention according to claim 1, characterized in that: In S2, the dual attention mechanism includes two independent attentions, which are used to process local information and global information respectively. The input dimension of each dual attention layer is 128, the output dimension is 128, the number of attention heads is 4, and the dimension of each head is 32.

6. The anticancer drug synergy prediction method based on GAT dual attention according to claim 1, characterized in that: In S3, the dual attention mechanism generates a pooled feature representation by calculating the interactive attention weights between the two drug molecule features, specifically including: S3a: Calculates query, key and value through linear transformation; S3b: Calculate the attention coefficient between two graphs; S3c: perform softmax normalization on the attention coefficient; S3d: Use attention weights to perform weighted summation of values ​​to generate pooled feature representations.

7. The anticancer drug synergy prediction method based on GAT dual attention according to claim 1, characterized in that: In S4, the fully connected layer maps high-dimensional features to low-dimensional space and finally outputs a classification result with an output dimension of 2, which respectively represent the probability of drug synergy.