A Drug-Target Interaction Prediction Method Based on TransGAT
By combining Transformer and GAT models, using the dual attention mechanism to perform drug-target feature fusion, the problems of inaccurate feature extraction and unexplainable model in the existing methods are solved, and the efficiency and transparency of drug-target interaction prediction are achieved.
Patent Information
- Application Number
- CN202310302892.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-24
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-03-24
AI Technical Summary
The existing drug-target interaction prediction methods have problems such as inaccurate manual feature extraction, difficulty in interpreting deep learning models, insufficient topological modeling, and insufficient feature fusion, which leads to time-consuming and labor-intensive drug discovery and opaque prediction results.
Combining Transformer and graph attention network (GAT) models, using dual attention mechanisms for feature fusion, and modeling drugs through GAT graph data. Transformer encodes protein sequences, and uses dual attention feature fusion module to improve prediction accuracy and interpretability.
The accuracy and transparency of drug-target interaction predictions were significantly improved on the public data set, reducing the time and cost of drug discovery, and enhancing the interpretability of the model.
Smart Images

Figure CN116312808B_ABST
Abstract
Description
[0001] Technical Field: The present invention relates to the field related to artificial intelligence and drug discovery. Specifically, a method for predicting drug-target interactions is invented, which combines Transformer, Graph Attention Network (GAT), and a dual-attention feature fusion module. Background Art
[0002] Analysis of drug-target interactions, as an important part of drug discovery, plays an irreplaceable role. To find safe and effective drugs, traditional Drug-Target Interaction (DTI) analysis often requires testing thousands of compounds, which consumes a large amount of human, material resources and time costs, and has a relatively high risk of failure in the test. In recent years, computer-aided drug design has attracted more and more attention from drug researchers. By using technologies such as artificial intelligence, potential Drug-Target Pairs (DTPs) that may interact are screened out from a large amount of drug and protein-related data for further experimental verification by drug scientists. This method can not only reduce the waste of human and material resources in the drug discovery process, but also greatly shorten the time of drug discovery. In recent years, due to the development of artificial intelligence, especially deep learning technology, and in particular, the increasingly wide application of Transformer technology in various fields recently, good results have been achieved in the research related to drug-target interaction prediction. The present invention designs a novel method for predicting drug-target interactions and has achieved fruitful results through verification on public datasets.
[0003] 1. Professional Terms
[0004] (1) Deep Learning. In recent years, deep learning technology has achieved extremely brilliant achievements. This technology has evolved from multi-layer neural networks. Its essence is to build a machine learning model with a large number of neural network hidden layers, and through a large amount of training data, train to learn more representative features to increase the accuracy of classification. Different from traditional machine learning methods, deep learning usually has more hidden layers. Through feature interaction between layers, the original feature representation is transformed into a new feature space, and then the loss function and optimization function are used to optimize the training effect through the feature interaction information, thereby optimizing the model.
[0005] (2) Transformer. A Transformer is a deep learning model based on the self-attention mechanism. Different from traditional deep convolutional networks, the Transformer encoder consists of multi-head self-attention modules. Each multi-head self-attention module is composed of multiple self-attention modules, and residual connections are used between layers. The self-attention mechanism module is used to extract the tensor parameters of the input data, mainly including query (Q), key (K), and value (V). Among them, Q is used to interact with other key vectors to obtain the weights of the vectors; K is used to interact with the query vector to assist other vectors in outputting results; V is the result of summing the weights generated by Q and other Ks with its own weight. The self-attention mechanism is expressed by the following formula:
[0006]
[0007] where d k is the dimension of Q and K.
[0008] (3) Graph Attention Network (GAT). The Graph Attention Network model (GAT) is developed on the basis of the Graph Neural Network (GNN). Due to the structural characteristics of drugs, the graph neural network has natural advantages in processing drug chemical structures. It can effectively model drug chemical data and is convenient for processing topological structure information between chemical formulas. And GAT adds an attention mechanism on the basis of the original GNN, making it convenient to calculate the weights of different nodes in drug chemical structure data, helping to discover the key structural parts in the input features and improving the performance of model prediction.
[0009] 2. Analysis of the Research Status at Home and Abroad
[0010] At present, the methods for predicting drug-target interactions based on machine learning at home and abroad mainly fall into three branches: using traditional machine learning methods, drug-target interaction prediction methods based on deep learning, and drug-target interaction prediction methods based on graph neural networks. Among them, the drug-target interaction prediction method based on traditional machine learning methods needs to manually extract the features of drugs and proteins (proteins are the main molecular targets of drugs), and then input them into a classifier for drug-target interaction prediction. This way of extracting features has two main disadvantages. One is that it is affected by human subjective factors, which may lead to inaccurate extracted features. The second is that the number of features extracted manually is limited, and many features that are difficult to observe with the naked eye may be ignored, thus affecting the accuracy of drug-target interaction prediction. In recent years, the drug-target interaction prediction method based on deep learning has also received extensive attention. At present, it mainly uses deep neural networks (DNNs) as the main feature extraction approach. This method can obtain a large number of drug and protein features, and there has been a significant improvement in prediction accuracy compared with traditional machine learning methods. However, this method often operates in a "black box" state, which makes it difficult for drug scientists to deeply analyze the operation mechanism of the model, and there is a certain deviation from the high safety requirements of drug design, and the prediction results are difficult to gain the trust of pharmacists. Secondly, since the chemical formula structure of drugs is essentially a graph structure, it is difficult for deep neural networks to model and analyze the topological structure of the graph, which also limits the development of deep neural networks in drug-target interaction prediction. As mentioned above, the graph structure can effectively model the molecular chemical formula, pay more attention to the topological structure information of drugs, and has a certain promoting significance for the drug-target interaction prediction method based on graph neural networks.
[0011] Problems existing in the currently disclosed drug-target interaction prediction methods
[0012] The above various methods have certain reference significance for understanding the research progress of drug-target interaction prediction. However, these methods also have some problems, mainly including:
[0013] (1) Traditional machine learning methods are limited by the inefficiency and subjectivity of manual feature extraction and are difficult to be widely promoted; due to the disadvantages of its poor interpretability and inconvenience in modeling the topological structure, the operation mechanism of the deep learning-based method is not transparent enough for drug scientists; the graph neural network method can effectively model the molecular structure, but ignores the distance relationship between drugs and proteins, and the distance information in long-sequence protein sequences is not accurately obtained.
[0014] (2) Existing methods are difficult to simultaneously consider the spatial characteristics of drug molecules and the sequence characteristics of proteins, and a new model needs to be designed to effectively encode the characteristics of both drugs and proteins.
[0015] (3) Existing methods often pay more attention to the characteristics of DTP, and simply concatenate the encoded characteristics. This method cannot obtain the key characteristics affecting drug-protein interactions and lacks an efficient feature fusion method.
[0016] Problems to be solved by the present invention
[0017] In view of the problems existing in the research status at home and abroad, the present invention designs a drug-target interaction prediction method based on the combination of Transformer and GAT. This method not only achieves ideal results on the publicly available drug-target dataset, but also has strong model interpretability. Specifically, the present invention mainly solves the following problems:
[0018] (1) The present invention designs a drug-target interaction prediction method based on TransGAT. This method uses the graph attention network (GAT) to model drugs and uses the Transformer Encoder to encode protein sequences. It effectively combines the advantages of GAT in processing graph data and the characteristics of Transformer in handling long text sequences.
[0019] (2) A feature fusion method based on a dual attention mechanism is used to extract the key features after feature fusion, further enhancing the accuracy of drug-target interaction prediction.
[0020] (3) The method disclosed in the present invention has good interpretability, which helps drug scientists deeply understand the model operation mechanism and ensure the safety of drug design. Summary of the invention
[0021] The purpose of the present invention is to solve the problems existing in the above-mentioned drug-target interaction prediction method, and invent a drug-target interaction prediction method based on TransGAT, named TransGAT. The main content of the invention is as follows:
[0022] (1) A drug-target interaction prediction method based on TransGAT is invented. This method combines the advantages of two model architectures, Transformer and GAT, and uses a dual attention mechanism feature fusion method to input the fused features into a classifier for drug-target interaction prediction.
[0023] (2) Select drug information, protein information, and DTP information from the drug-target database. Among them, the drug information is in the SMILES (Simplified molecular input line entry system) format, and the protein information is in the FASTA format.
[0024] (3) Input the drug data into the GAT for encoding. Consider the drug chemical formula structure as graph data, where each atom in the graph is represented by a 74-dimensional integer vector. This vector describes 8 properties, including atom type, degree, number of implicit Hs, formal charge, number of radical electrons, atom hybridization, total Hs, and whether the atom is aromatic. The drug encoder is written in the form of a three-layer GAT-block. It updates the atomic feature vector by aggregating the corresponding neighboring atom sets connected by chemical bonds. This propagation mechanism captures the substructure information of the molecule and aggregates the neighbor nodes through the self-attention mechanism, achieving an adaptive weight matching for different neighbors and retaining the node-level drug representation for a subsequent clear understanding of the local interaction with the protein fragment. The drug chemical formula structure is encoded to obtain the encoded drug feature set.
[0025] (4) Input the protein sequence into the Transformer Encoder for encoding. The protein feature encoder consists of 6 layers of Transformer Encoder, which converts the input protein sequence into a matrix representation in the latent feature space. Each row of the matrix represents the subsequence representation in the protein. To achieve this process, the present invention adopts the concept of word embedding and initializes all amino acids into a learnable embedding matrix. This embedding matrix converts amino acids into vector representations, and these vectors are used as the input to the Encoder. By using Transformer, it is possible to capture the local features in the protein sequence at different scales, thus obtaining a better representation. Finally, the protein feature encoder converts the protein sequence into a matrix representation, where each row represents a subsequence representation in the protein. The encoded protein feature set is obtained.
[0026] (5) Use the dual attention mechanism network to perform feature fusion on the obtained drug feature set and protein feature set above. This module consists of two layers: a bilinear interaction mapping for capturing pairwise attention weights, and a bilinear pooling layer for extracting the joint drug-target representation on the interaction mapping. The bilinear interaction mapping can obtain single-head pairwise interactions, and these elements represent the interaction intensity of the corresponding drug-target substructure pairs and are mapped to potential binding sites and molecular substructures. By introducing a bilinear pooling layer on the interaction mapping, a joint representation vector is obtained. Multi-head interactions have better performance than single-head interactions. Finally, the model can explicitly learn the pairwise local interactions between drugs and proteins. Among the drug-target pairs after feature fusion, the drug-target pairs formed by drug-targets with known relationships are positive, and the remaining drug-target pairs are negative.
[0027] (6) Randomly divide each experimental dataset into a training set, a validation set, and a test set according to the ratio of 7:1:2. Randomly select the same number of positive and negative drug-target pairs. Part of them are used as the training set to train the model, and part of them are used as test cases to evaluate the model effect.
[0028] (7) Use the training set to train the TransGAT model. Use GAT to extract drug features, and the Transformer Encoder to extract protein sequence features. Then input the extracted protein features and drug features into the dual attention feature fusion module for feature fusion. Finally, extract the features of the fused drug-target pairs and input them into a Multilayer perceptron (MLP) for binary classification to output the prediction results of the drug-target interaction relationship. Set reasonable hyperparameters until the model no longer converges, then stop training and save the trained model.
[0029] (8) Use the test cases to evaluate the trained model. If the accuracy during training cannot be achieved, the parameters need to be reset for training until the test results meet the ideal requirements.
[0030] The beneficial effects of the present invention are as follows: The present invention designs a drug-target interaction prediction method based on TransGAT. This method combines the advantages of Transformer in processing long sequence text data, GAT in processing graph structure data, and the dual attention feature fusion module. Multiple evaluation metrics on the public dataset are significantly better than existing drug-target interaction prediction methods. Brief Description of the Drawings
[0031] Figure 1 : Flowchart of the drug-target interaction prediction model based on TransGAT of the present inventionFigure 1 a is the overall network process architecture, Figure 1 b is the drug coding module based on GAT, Figure 1 c is the protein coding module based on Transformer.
[0032] Figure 2 : The dual-attention feature fusion module. Among them, H d represents the drug feature representation, H p represents the protein feature representation, and U and V respectively represent two attention weight matrices. Specific implementation manner
[0033] To better understand the purpose, technical solution and advantages of the present invention more clearly, the present invention will be further described below in conjunction with the accompanying drawings and specific example implementation manners. Those skilled in the art can easily understand the advantages and effects of the present invention from the content disclosed in this specification, but the present invention is not limited in any form. It should be noted that for those of ordinary skill in the art, without departing from the idea of the present invention, several changes and improvements can still be made, and these all belong to the protection scope of the present invention. The following will describe in detail some implementation manners of the specific examples of the present invention in conjunction with the accompanying drawings. Without conflict, the following implementation manners can be extended to all drug-target interaction predictions.
[0034] According to a drug-target interaction prediction method based on TransGAT provided by the present invention, its main process is shown in Figure 1 , and the main implementation steps include:
[0035] Step 1: Data acquisition. Select drug information, protein information and DTP information from the drug-target database. The data in this example comes from three databases, namely BindingDB, BioSNAP and Human. Among them, the drug information is in SMILES format, and the protein information is in FASTA format.
[0036] Step 2: Data preprocessing. The SMILES format adopted for drugs in this example is a two-dimensional structure diagram, expressed as: D = (v, ξ), where D represents the drug, and v and ξ respectively represent the nodes (vertices) and edges of the graph. In this example, the nodes and edges respectively represent the atoms and chemical bonds of the drug molecule. During the feature extraction process, for the convenience of calculation, each drug molecular structure is represented by a feature matrix and an adjacency matrix representation. Among them, N iis the i-th atom of the drug molecule, and K represents the characteristic dimension of the atom. In this example, the protein is represented in the FASTA format, which is a long sequence form and is expressed as: ρ = (ρ1, ρ2,..., ρ n ), where ρ i represents the i-th amino acid in the protein.
[0037] Step 3: Design the TransGAT model. The present invention comprehensively utilizes the advantages of three modules: Transformer, Graph Attention Network (GAT), and dual attention feature fusion. Among them, Transformer has strong advantages in protein sequence encoding, the graph attention network can more deeply extract drug features beneficial to the prediction result, and the dual attention feature fusion module highlights DTP feature information, further increasing the prediction accuracy of the model. This model is mainly divided into the following three functions:
[0038] (1) Feature encoding module. Among them, the protein encoding model creates a learnable embedding matrix containing 23 amino acid types, and D P represents the dimension of the matrix. By looking up the embedding matrix, each protein sequence is initialized as a corresponding feature matrix where θ is the maximum allowable length of the protein sequence. This sequence is set to align different protein lengths and enable batch training. To process protein sequences of different lengths, the present invention cuts each protein sequence into segments within the maximum allowable length and pads the shorter segments with zeros to ensure that all protein sequences have the same length and are processed in batches during training. The Transformer Encoder extracts local feature information from the protein feature matrix. A protein sequence is a string of amino acids that make up a protein. Instead of considering each amino acid separately, they are divided into overlapping groups of three. This grouping can analyze the protein sequence more comprehensively because it takes into account the interactions between adjacent amino acids. The protein encoding layer inputs each group of three data of size 3 into 6 layers of the Transformer Encoder, and outputs the protein feature encoding after encoding. The drug encoder uses a simple linear transformation, multiplies the node feature matrix by the weight matrix and transposes it to obtain the drug structure matrix. 6 layers of GAT-Block are used to learn the graph representation of the drug compound. GAT updates the atomic feature vector by aggregating its corresponding set of neighboring atoms, and these neighboring atoms are connected by chemical bonds. The structure of the drug encoder is as follows: where e ij represents the attention value between node i and node j, LeakyReLU is the activation function, and W are learnable parameters, and the learnable parameters of each layer are the same, represents the feature vector of node i.
[0039] (2) Feature Fusion Module. Please refer to Figure 2 , the feature fusion module uses a dual attention mechanism to deeply fuse the previously input feature encodings, increasing the feature weights beneficial to the prediction results to improve the prediction accuracy. The bilinear interaction graph captures the pairwise attention weights between drugs and proteins, while the bilinear pooling layer extracts the joint drug-target representation. The hidden protein and drug representations at layer 6 are obtained using separate Transformer and GAT encoders. The number of substructures encoded in the protein and the number of atoms in the drug are denoted by M and N respectively. The bilinear interaction mapping can obtain single-headed pairwise interactions This is a matrix of size N×M. This matrix represents the pairwise interactions between drugs and proteins. The bilinear attention network module captures these pairwise local interactions between drugs and proteins, which is very important for better predicting and interpreting drug-target interactions.
[0040] (3) Drug-Target Interaction Prediction Module. This module inputs the fused features into an MLP for binary classification to predict whether there is a relationship between the input target drug and protein. To calculate the interaction probability, the features are input into a decoder, which is a fully connected classification layer followed by a sigmoid function that maps the output to a probability value between 0 and 1, representing the likelihood of a drug-target interaction.
[0041] Step 4: Set the TransGAT model parameters. In this example, to enable the designed TransGAT model to achieve an ideal prediction effect, the following hyperparameters need to be set:
[0042]
[0043] (1) Activation function. The activation function used in the training process of this example is the ReLU function. This function is a piecewise function that sets all negative values to 0 while leaving positive values unchanged. This function only activates neurons with positive values, which can increase the computational efficiency and there is no problem of gradient disappearance. The ReLU function can be expressed as:
[0044]
[0045] (2) Loss function. The loss function used in the training process of this example is the Binary Cross-Entropy loss function, which can be expressed as:
[0046]
[0047] The purpose is as follows: when the sample is positive, y = 1, and at this time, Loss = -log(P(y)). When P(y) is larger, Loss is smaller. The most ideal situation is when P(y) = 1, Loss = 0. When the sample is a negative example, y = 0, and at this time, Loss = -log(1 - P(y)). When P(y) is smaller, Loss is smaller. The most ideal situation is when P(y) = 0, Loss = 0.
[0048] (3) Optimization function. The optimization function used in the training process of this example is adam. The adam optimization algorithm uses momentum and adaptive learning rate to accelerate the convergence speed. It adds momentum on the basis of the RMSprop optimization algorithm, which further speeds up the learning efficiency of the model.
[0049] Step 5: Comparison of experimental results. The evaluation metrics adopted in this example include AUROC (Area under the receiver operating characteristic curve), AUPRC (Area under the precision-recall curve), Recall, Precision, Accuracy, etc. The experimental results of all or some of the above evaluation metrics on three public datasets are compared as follows:
[0050] (1) Comparison results on the BindingDB dataset, where the bold data are the optimal results
[0051]
[0052] (2) Comparison results on the Human dataset, where the bold data are the optimal results
[0053]
[0054] (3) Comparison results on the BioSNAP dataset, where the bold data are the optimal results
[0055]
[0056] It is found through comparison that the performance of the drug-target interaction relationship prediction method adopted in the present invention is significantly better than existing methods such as GCN-DTI, GraphDTA, DeepConv-DTI, TransformerCPI, MolTrans, etc. in multiple evaluation metrics.
[0057] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, including but not limited to implementing the same program in the form of systems, devices and their respective modules designed based on the method provided by the present invention, such as logic gates, switches, application specific integrated circuits, programmable logic controllers, and embedded microcontrollers, etc., which does not affect the essence of the present invention.
Claims
1. A method for predicting drug-target interactions based on TransGAT, characterized in that, Including the following steps: S1: Data acquisition. Select drug information, protein information, and drug-target pair information from a drug-target database, where the drug information is in SMILES format and the protein information is in FASTA format. S2: Data preprocessing. The SMILES format used for drugs in this example is a two-dimensional structure diagram, expressed as: D = (ν, ξ), where D represents the drug, ν and ξ represent the vertices and edges of the graph, respectively. In this example, nodes and represent the atoms and chemical bonds of drug molecules, respectively. In the feature extraction process, for the convenience of calculation, each drug molecule structure is represented by a feature matrix and an adjacency matrix Indicates that N i is the i-th atom of the drug molecule, K represents the characteristic dimension of the atom. In this example, the protein is represented in FASTA format, which is a long sequence format, expressed as: ρ = (ρ1, ρ2, ..., ρ n ), where shake represents the i-th amino acid in the protein; S3: Design the TransGAT model, which comprehensively utilizes the advantages of three modules: Transformer, graph attention network, and dual-attention feature fusion. Among them, Transformer has strong advantages in protein sequence encoding. The graph attention network can more deeply extract drug features beneficial to the prediction results. The dual-attention feature fusion module highlights the feature information of drug-target pairs, further increasing the prediction accuracy of the model. The model is mainly divided into the following three functions: (1) Feature encoding module. The protein encoding method is as follows: The module creates a learnable embedding matrix which contains 23 amino acid types, D P represents the dimension of the matrix. By looking up the embedding matrix, each protein sequence is initialized as a corresponding feature matrix where θ is the maximum allowed length of the protein sequence. Then each protein sequence is cut into segments within the maximum allowed length, and the shorter segments are padded with zeros and processed in batches during training. The Transformer encoder extracts local feature information from the protein feature matrix. The protein encoding layer inputs every three groups of data of size 3 into the 6-layer Transformer encoder, and outputs the protein feature encoding after encoding. The drug encoding method is as follows: The drug chemical formula structure is regarded as graph data, where each atom is represented by a 74-dimensional integer vector, which describes 8 properties, namely: atom type, degree, number of implicit Hs, formal charge, number of radical electrons, atom hybridization, total Hs, and whether the atom is aromatic. A total of 6 layers of graph attention network GAT are used to learn the graph representation of the drug compound. The graph attention network GAT updates the atomic feature vector by aggregating its corresponding neighborhood atom set, and these neighborhood atoms are connected by chemical bonds. The structure of the drug encoder is represented as where e ij represents the attention value between node i and node j, and LeakyReLU is the activation function and W are learnable parameters, and the learnable parameters of each layer are the same represents the feature vector of node i (2) Feature fusion module. The feature fusion module uses a dual attention mechanism to deeply fuse the previously input feature encodings. This module consists of two layers: a bilinear interaction mapping for capturing pairwise attention weights, and a bilinear pooling layer for extracting a joint drug-target representation on the interaction mapping. The bilinear interaction mapping can obtain single-head pairwise interactions, and these elements represent the interaction intensities of the corresponding drug-target substructure pairs and are mapped to potential binding sites and molecular substructures. By introducing a bilinear pooling layer on the interaction mapping, a joint representation vector is obtained. Multi-head interactions have better performance than single-head interactions. Finally, the model can explicitly learn the pairwise local interactions between drugs and proteins. Among the drug-target pairs after feature fusion, the drug-target pairs formed by drug-targets with known relationships are positive, and the remaining drug-target pairs are negative. (3) Drug-target interaction relationship prediction module. This module inputs the fused features into a multi-layer perceptron for classification to predict whether there is a relationship between the input target drug and protein. The multi-layer perceptron uses the sigmoid function to map the output to a probability value between 0 and 1, representing the likelihood of drug-target interaction. S4: Set the parameters of the TransGAT model. The activation function uses the ReLU function, the loss function uses the Binary Cross-Entropy loss function, the optimization function uses adam, the number of iterations is 200, the number of Transformer layers is 6, the number of GAT layers is 6, the number of heads in the feature fusion attention mechanism is 2, and the embedding size of the feature fusion module is 768. S5: Train the TransGAT model using the training set. Train the model with the hyperparameters set in step S4 until the model no longer converges, then stop training and save the trained model. S6: Model encapsulation to form a drug-target interaction prediction model based on TransGAT.
2. The method for predicting drug-target interaction based on TransGAT according to claim 1, wherein This method uses the Graph Attention Network (GAT) to model drugs and the Transformer encoder to encode protein sequences.
3. The method for predicting drug-target interaction based on TransGAT according to claim 1, characterized in that The TransGAT model described in this method is mainly divided into three parts, namely: feature encoding part, feature fusion part, and drug-target interaction relationship prediction part.
4. The method for predicting drug-target interaction based on TransGAT according to claim 1, wherein This method uses a dual attention mechanism for feature fusion.
5. The method for predicting drug-target interaction based on TransGAT according to claim 1, wherein, This method can be used for predicting drug-target interaction relationships.
Citation Information
Patent Citations
Drug ATCCode prediction method based on graph transformation network
CN114420310A
Adversarial framework for molecular conformation space modeling in internal coordinates
WO2022259185A1