Drug-target interaction prediction method and device based on multi-view feature fusion and medium
Through the multi-view feature fusion method, the characteristics of drugs and targets are extracted using GraphTransformer and cross-attention mechanism, solving the problem of insufficient interpretability of drug-target interaction prediction in the prior art, and achieving higher prediction accuracy and interpretability.
Patent Information
- Application Number
- CN202510230115.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-28
AI Technical Summary
Existing drug-target interaction prediction methods fail to fully explore the characteristics of protein and drug molecular structure, resulting in insufficient interpretability of the predicted results.
Using a multi-view feature fusion method, the characteristics of drugs and targets were extracted through GraphTransformer, combined with one-dimensional amino acid sequence and three-dimensional protein structure, and deep fusion was performed using cross-attention mechanisms and bilinear attention networks to calculate the probability of interaction between drugs and targets.
It significantly improves the prediction accuracy of drug-target interactions, enhances the interpretability of prediction results, and can locate specific binding sites in drug-target interactions.
Smart Images

Figure CN120072036A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of bioinformatics, and relates to a method, device and medium for predicting drug-target interactions based on multi-perspective feature fusion. Background Art
[0002] Drugs play an irreplaceable and important role in maintaining human health and treating diseases. In disease treatment, whether it is antibiotics used to control bacterial infections or antiviral drugs used to inhibit virus transmission, drugs serve as key tools that affect the confrontation between the human immune system and pathogens. For example, imatinib is a drug effective in treating chronic myeloid leukemia, which inhibits the activity of specific enzymes in tumor cells by targeting, thereby preventing the progression of cancer. In addition, another type of anti-cancer drug, PD-1 inhibitor, shows wide application value in the treatment of various cancers such as melanoma and lung cancer by activating the patient's immune system to enable it to recognize and attack tumor cells. These examples fully demonstrate the importance of drugs in modern medicine and their profound impact on improving human health.
[0003] Drug research and development is a complex and long process, including multiple key links from determining drug targets to new drug development. First, the target needs to be identified. Secondly, virtual screening technology is used to screen out potential lead compounds from a large number of compounds. Finally, drug-target interaction prediction technology is used to evaluate the binding ability and potential biological activity of these compounds with the target, so as to achieve the discovery and development of new drugs. Traditional virtual screening methods mainly rely on biochemical experiments, by screening a large number of compounds and verifying one by one whether they interact with the target. The advantage of this method is that it can discover new molecules with biological activity, but its disadvantages are also obvious: low screening efficiency, high cost, and often unable to clarify the specific mechanism of action between the molecule and the target. To overcome the limitations of traditional methods, emerging technologies such as computer-aided drug design have gradually been introduced into the field of drug research and development, especially widely used in the task of predicting drug-target interactions. These technologies can not only efficiently assist in virtual screening of candidate drugs, but also significantly improve the efficiency of drug discovery, and at the same time provide deeper insights into understanding the mechanism of action between molecules and targets. According to the computational models for predicting the interactions of drug-target pairs, they can be roughly divided into 3 categories:
[0004] (1) Computational models based on molecular docking. The computational model for predicting drug-target interactions based on molecular docking uses computational simulation to predict the interactions between drug molecules (ligands) and specific targets (usually proteins). The core of this method is to calculate the optimal spatial position and interaction mode of the drug binding to the target to determine their affinity and binding mode.
[0005] (2) Ligand-based computational models. Ligand-based drug-target interaction prediction models mainly rely on the structural information of known active ligands (i.e., small molecule compounds known to interact with targets), and predict the interaction between new ligands and targets by analyzing and comparing the common features of these ligands. The core of this type of model is to construct a "pharmacophore" or characteristic map that can effectively represent the activity of the drug, thereby predicting potential drug-target binding in the absence of an unknown target or structural information.
[0006] (3) Computational models based on deep learning. Lee et al. used amino acid sequences and drug fingerprint extended connectivity fingerprint (ECFP4) as input, and used convolutional neural networks on protein sequences of different lengths to capture local residue patterns, and proposed the DeepConvDTI model. Based on the assumption that the binding sites between drugs and targets appear in their specific substructures, Huang et al. proposed a DTI model MolTrans that uses the frequent subsequence algorithm (FCS) to decompose protein and drug molecules into common subsequences. et al. proposed the DeepDTA model, which applied a multi-layer CNN module to calculate the binding affinity between drug SMILES and protein sequences. Considering that the molecular structure provides detailed information about how drugs and proteins bind at the atomic level, Torng et al. proposed the CPI-GNN model, which converts each two-dimensional molecular graph into a graph with nodes composed of its atoms and edges composed of chemical bonds, and then used CNN and GNN to represent protein sequences and 2D drug molecular graphs, respectively, and then used the attention mechanism to generate the final prediction.
[0007] In summary, when predicting potential drug-target interactions, it is crucial to effectively explore the biochemical properties of drug molecular structures and protein structures. However, existing computational models have not fully explored the characteristics of protein and drug molecular structures, so the interpretability of their prediction results still needs to be improved. Summary of the invention
[0008] In view of the fact that current drug-target interaction prediction methods lack sufficient characterization of protein structure information and ignore the shortcomings of special substructure patterns in molecules (such as subfragments and functional groups), the present invention proposes a drug-target interaction prediction method, device and medium based on multi-perspective feature fusion, which can effectively mine the biochemical property characteristics in drug molecules and protein structures, significantly improve the prediction accuracy of drug-target interactions, and have stable prediction performance.
[0009] In order to achieve the above technical objectives, the present invention adopts the following technical solutions:
[0010] A drug-target interaction prediction method based on multi-view feature fusion, comprising:
[0011] Step 1, extract drug features and target features;
[0012] (1) Extract drug features: First, generate a corresponding 2D molecular graph according to the SMILES sequence of the drug, and construct the graph structure of the drug based on the characteristics of each atom in the drug molecule and the chemical bond relationship between them; then, obtain the sub-graph structure corresponding to each molecular fragment by extracting the molecular fragments of the drug; finally, use a GraphTransformer as a molecular graph encoder to encode the sub-graph structures corresponding to each molecular fragment of the drug to obtain drug features;
[0013] (2) Extract target features: First, extract the protein pocket of the target, and construct the graph structure of the target based on each residue in the protein pocket and the chemical bond between them; then, divide the target into multiple sub-pockets to obtain the sub-graph structure corresponding to each sub-pocket; then use a GraphTransformer as a pocket encoder to encode the sub-graph structures corresponding to each sub-pocket of the target to obtain the pocket features of the target; in addition, use a self-attention module to extract the sequence features of the target from the one-dimensional amino acid sequence of the target; finally, use cross-attention to fuse the pocket features and sequence features of the target to obtain target features;
[0014] Step 2, use a bilinear attention network as a classifier to deeply fuse the drug features and target features, and calculate the interaction probability between the drug and the target through a multi-layer perceptron.
[0015] Furthermore, in the graph structure of the drug, each node represents an atom, and its feature vector includes the following attributes: the type of the atom, the degree of the atom, the formal charge of the atom, the number of radical electrons of the atom, the hybridization state of the atom, the aromaticity of the atom, the total number of hydrogen atoms connected to the atom, and the chirality of the atom; the edges between the nodes represent the chemical bonds between the atoms, and its features include the following attributes: the type of the bond, whether the bond is a conjugated bond, whether the bond belongs to a ring structure, and whether the bond has stereochemical characteristics;
[0016] In the graph structure of the target, each node represents a different residue in the three-dimensional structure of the target protein, and the edges between the nodes represent the chemical bonds between the corresponding residues; the feature vectors of each node residue include the residue type, the residue self-distance, and the residue dihedral angle, and the features of the edges between the node residues include: residue connectivity, residue CA distance, residue center distance, and residue maximum distance.
[0017] Further, the RDKit method library is used to generate a 2D molecular graph of a drug based on the SMILES sequence of the drug; the Macfrag method is used to process the 2D molecular graph of the drug to extract molecular fragments of the drug that conform to the preset chemical meaning.
[0018] Further, by combining computational geometry methods and protein clustering techniques, protein pockets are extracted from the three-dimensional structure of target proteins. The specific steps include:
[0019] First, the protein space is voxelized to construct a grid structure of the protein;
[0020] Then, the convex hull of the protein is calculated, and its surface is divided into multiple triangles, each triangle representing a potential binding site;
[0021] Next, a protein pocket is defined by the volume generated by the triangle vertices, and the empty voxels closest to the protein atoms are identified to determine the possible ligand atom positions, serving as candidate protein pockets;
[0022] Finally, based on the overlap between the candidate protein pockets and the biochemical and physical properties of each candidate protein pocket, several protein pockets with high confidence are selected from all candidate protein pockets as the finally extracted protein pockets. Further, GraphTransformer is used as a molecular graph encoder to encode the subgraph structures corresponding to each molecular fragment of the drug, specifically:
[0023] (1) The feature representation of the i-th node in the subgraph structure corresponding to each molecular fragment of the drug is denoted as x i , and the feature representation of the edge between the i-th and j-th nodes i is denoted as y ij ;
[0024] (2) Linear transformations are performed on the node feature x i and the edge feature y ij in the input subgraph structure:
[0025]
[0026] Among them, and are the weight parameters of the linear layer in GraphTransformer, and are the bias parameters of the linear layer in GraphTransformer; represents the node feature obtained by linearly transforming x i , and represents the edge feature obtained by linearly transforming y ij ;
[0027] (3) Use the eigenvectors of the Laplacian operator as the node position encoding:
[0028]
[0029] Among them, represents the Laplacian matrix of the subgraph structure; I is the N-dimensional identity matrix with a diagonal of 1, A is the N×N adjacency matrix, which is constructed from the connection relationships between the nodes in the subgraph structure; D is the degree matrix, representing the degree of each node; Λ and U represent the diagonal matrix of eigenvalues and the orthogonal matrix of eigenvectors respectively; according to the distribution of eigenvalues after the Laplacian matrix decomposition, select the eigenvectors corresponding to several of the smallest non-zero eigenvalues as the multi-dimensional position encoding λ of the nodes i ; is a weight matrix; is the bias vector; represents the initial position encoding of node i, which is obtained by linearly transforming the Laplacian eigenvector λ i ; represents the node feature with added position information;
[0030] (4) Fuse the node features and edge features through the multi-head attention mechanism of GraphTransformer;
[0031] Among them, the detailed update steps for the l-th layer of the GraphTransformer network are as follows:
[0032]
[0033] Among them, represent the features of node i and node j in the l-th layer respectively, represents the edge between node i and node j in the l-th layer; Q k,l , K k,l , V k,l and E k,l are the linear layer parameters of the k-th head in the l-th layer, represent the query, key, and value obtained by normalizing the node features and edge features of the k-th head in the l-th layer through the normalization function Norm() respectively, represents the edge feature between node i and node j in the l-th layer; is the attention weight calculated jointly by the node features and edge features;
[0034] (5) Integrate the features learned by multiple attention heads through the Concat splicing operation to update the node features and edge features:
[0035]
[0036] Among them, is the feature of node i in the (l + 1)-th layer, N i is all neighbor nodes of node i, is the edge feature between node i and node j in the (l + 1)-th layer, represents the linear layer parameters in the multi-head attention mechanism, h gt represents the number of attention heads;
[0037] (6) Input the node and edge features after multiple fusions and updates of node features and edge features through the multi-head attention mechanism into a feed-forward network with residual connections and normalization operations:
[0038]
[0039] where, and and and are weight matrices for transforming the node features and edge features in the (l + 1)-th layer to the hidden layer; are respectively and the node features and edge features updated through residual and normalization operations; ReLU represents the activation function in the neural network;
[0040] The node features including edge features obtained through the above process The features including edge features of n nodes in each subgraph structure constitute the features of the subgraph structure are represented as Finally, the features of all subgraph structures of the drug are composed into the feature R of the drug d ;
[0041] Using GraphTransformer as a pocket encoder to encode the subgraph structures corresponding to each sub-pocket of the target to obtain the pocket features of the target is the same as using GraphTransformer as a molecular graph encoder to encode the subgraph structures corresponding to each molecular fragment of the drug.
[0042] Furthermore, using a self-attention module to extract the sequence features of the target from the one-dimensional amino acid sequence of the target specifically includes:
[0043] (1) Perform one-hot encoding on the one-dimensional amino acid sequence of the target to obtain P s ∈R n×d ; where, n represents the sequence length, and d represents the encoding dimension;
[0044] (2) Send the encoded P s into the self-attention module to learn the sequence encoding Ps The correlation between the insides is as follows:
[0045]
[0046] In the formula, is the target sequence feature obtained by learning the correlation between the insides of the sequence through the self-attention mechanism, Q s , K s , V s are the parameters of the self-attention mechanism respectively; are the query weight matrix, key weight matrix, and value weight matrix in the self-attention module respectively; Attention s represents the self-attention mechanism, softmax is the activation function, d k is the scaling factor.
[0047] Furthermore, the pocket feature and sequence feature of the target are fused using cross-attention to obtain the target feature, which is expressed as:
[0048]
[0049] In the formula, W Q , W K , W P represent the query weight matrix, key weight matrix, and value weight matrix respectively, and are used to convert the sequence data into the query vector Q P , the key vector K P and the value vector V P ; R p is the target feature obtained by fusion; is the target sequence feature extracted by the self-attention module, is the target pocket feature obtained by encoding the subgraph structure corresponding to each sub-pocket of the target by the pocket encoder; Attention c represents the cross-attention mechanism.
[0050] Furthermore, in step 2, a bilinear attention network is used as a classifier to deeply fuse the drug feature and the target feature, and the interaction probability between the drug and the target is calculated through a multi-layer perceptron, including:
[0051] First, calculate the bilinear attention mapping to generate the attention map A;
[0052]
[0053] In the formula, q T is the learnable weight vector used to calculate the interaction intensity; M d , are learnable parameter matrices, which are respectively used to map the drug feature R d and the target feature R p into the shared space;
[0054] Then, based on the attention map A, the joint feature f of the target feature and the drug feature is calculated:
[0055]
[0056] In the formula, σ represents the activation function, M d ′, M p ′ are the same as M d and in function, and represent the parameter matrices used to map the drug feature R d and the target feature R p into the shared space;
[0057] Finally, the final attention score is calculated through the multi-layer perceptron MLP
[0058]
[0059] In the formula, W represents the learnable weight matrix, b represents the learnable bias vector, and Sigmoid represents the activation function; the finally calculated attention score is the interaction probability between the drug and the target.
[0060] An electronic device includes a memory and a processor. When the computer program stored in the memory is executed by the processor, the processor implements the method for predicting drug-target interaction based on multi-view feature fusion described in any one of the above.
[0061] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for predicting drug-target interaction based on multi-view feature fusion described in any one of the above is implemented.
[0062] Beneficial effects
[0063] The present invention proposes a method (hereinafter referred to as MotifGT-DTI), a device, and a medium for predicting drug-target interactions based on multi-view feature fusion, which are used to accurately predict drug-target interactions. For target proteins, the present invention uses a cross-attention mechanism to fuse two closely related view features, namely the one-dimensional amino acid sequence and the three-dimensional protein structure, so as to achieve a more comprehensive feature representation of target proteins. For drugs, considering that the complex molecular structure of drugs is composed of nodes and edges, the present invention uses GraphTransformer, which can simultaneously focus on node features and edge features, as the encoder of the molecular graph. To accurately identify the key substructures that bind to proteins, the method of the present invention also adopts a motif-based molecular representation method. The final visualization results show that in addition to predicting interactions, the present method can also locate the specific binding sites in drug-target interaction pairs. Experimental results show that the present invention can effectively reveal the molecular fragments responsible for the interaction, providing higher interpretability for the prediction results. Description of the Drawings
[0064] Figure 1 is a framework diagram of the method for predicting drug-target interactions based on multi-view feature fusion according to an embodiment of the present invention;
[0065] Figure 2 is a flowchart of the method for predicting drug-target interactions based on multi-view feature fusion according to an embodiment of the present invention;
[0066] Figure 3 is an analysis of the linkage parameters of the number of attention heads and the feature dimension in GraphTransformer in an embodiment of the present invention;
[0067] Figure 4 is the experimental result of cold-start generalization in an embodiment of the present invention;
[0068] Figure 5 is the visualization encoder effect of tsne dimensionality reduction in an embodiment of the present invention;
[0069] Figure 6 is the model interpretability analysis in an embodiment of the present invention. Detailed Embodiments
[0070] The following details the embodiments of the present invention. Based on the technical solution of the present invention, detailed implementation manners and specific operation processes are given, further explaining the technical solution of the present invention.
[0071] This embodiment provides a drug-target interaction prediction method based on multi-view feature fusion, including: extracting drug features and target features, and then deeply fusing the drug features and target features and calculating the interaction probability between the drug and the target. The specific implementation process refers to Figure 1 , Figure 2 as shown, and is divided into multiple parts and introduced as follows:
[0072] I. Extracting drug features
[0073] 1. Constructing a drug molecular graph and graph structure
[0074] To construct the motif subgraph of each drug molecule, first obtain the SMILES data of the drug from three public databases, namely Human, Biosnap, and Drugbank; then call the RDKit library function to construct the corresponding 2D molecular graph including atom and bond information by parsing the SMILES sequence; and then process the 2D molecular graph of the drug through the Macfrag method to extract the molecular fragment subgraph with important chemical meaning in the drug molecule, which is called the molecular motif.
[0075] Construct the graph structure of the drug according to the characteristics of each atom in the drug molecule and the chemical bond relationship between them; then obtain the subgraph structure corresponding to each molecular fragment according to the extracted drug molecular fragments. The subgraph structure of each molecular fragment can be expressed as V d represents the atomic node, and E d represents the edge (chemical bond) between nodes.
[0076] For each node in the subgraph structure, represent its feature as a 41-dimensional vector x i , which is composed of atomic type, atomic degree, atomic formal charge, atomic radical electron number, atomic hybridization, atomic aromatic number, total number of H atoms connected to the atom, and atomic chirality. For the edge between any node i and node j in the subgraph structure, represent its feature as a 10-dimensional vector y ij , including 4 features: bond type, whether the bond is conjugated, whether the bond is cyclic, and whether the bond is stereochemical.
[0077] 2. Using GraphTransformer to extract drug features
[0078] When the drug chemical structure is abstracted as a 2D undirected graph, both atoms and edges are very important and indispensable for describing chemical properties. For example, in the prediction task of the present invention, the nodes in the drug motif graph usually represent different atoms in the drug molecule, and the edges represent the chemical bonds between atoms.
[0079] One of the greatest advantages of the Graphtransformer architecture lies in its extension of the representation of edge features to node features. By considering both node features and edge features, it can more accurately describe the molecular structure. On the other hand, compared with traditional graph structure encoders such as GCN, it aggregates global information based on the attention mechanism, breaking the limitation of local neighborhood information and alleviating the problem of local smoothness to a certain extent.
[0080] After constructing the subgraph structure of each molecular fragment of the drug, the present invention uses Graphtransformer as a molecular graph encoder to encode the subgraph structures corresponding to the molecular fragments of the drug to obtain drug features. The method includes the following steps:
[0081] (1) Represent the feature of the i-th node in the subgraph structure corresponding to each molecular fragment of the drug as x i , and represent the feature of the edge between the i-th and the j-th nodes as y ij .
[0082] (2) Perform a linear transformation on the node feature x i and the edge feature y ij in the input subgraph structure:
[0083]
[0084] Among them, and are the weight parameters of the linear layer in Graphtransformer, and are the bias parameters of the linear layer in Graphtransformer; represents the node feature obtained by linearly transforming x i , represents the edge feature obtained by linearly transforming y ij .
[0085] (3) In order to better learn the topological structure between nodes, use the eigenvectors of the Laplacian operator as node position encodings:
[0086]
[0087] Among them, represents the Laplacian matrix of the graph, I is the N-dimensional identity matrix, A is the N×N adjacency matrix, D is the degree matrix, and Λ, U represent the diagonal matrix of eigenvalues and the orthogonal matrix of eigenvectors respectively. According to the distribution of eigenvalues after the Laplacian matrix decomposition, select the first k smallest non-trivial eigenvectors (corresponding to the first k smallest non-zero eigenvalues) as the multi-dimensional position encodings of the nodes, denoted as λ i .
[0088] (4) The fusion of node and edge features is achieved through the multi-head attention mechanism of the GraphTransformer.
[0089] For node i and node j in the l-th layer, the node features and edge features are normalized into queries keys values Inject the node features and edge features for update to obtain accurate global attention. Specifically, the attention score between node i and node j in the k-th head of the l-th layer is calculated based on the node features and edge features. The detailed update steps of the l-th layer are as follows:
[0090]
[0091] Among them, equations 6-8 are used to calculate the node features, where Q k,l , k k,l and V k,l are linear layer parameters, and equation (9) calculates the edge features, and then equation (10) is used in the attention mechanism to fuse the node features and edge features.
[0092] (5) To obtain a richer feature representation, the features learned by multiple attention heads are integrated to update the node features and edge features. The specific formula is as follows:
[0093]
[0094] Among them, represents the linear layer parameter, and h gt represents the number of attention heads.
[0095] (6) Input the node and edge features after multiple iterations into a feed-forward network with residual connections and normalization operations:
[0096]
[0097] Among them, and and and are linear layer parameters.
[0098] Through the above process, the node features fused with the edge information of the drug molecule can be obtained, denoted as R d .
[0099] II. Extracting target features
[0100] 1. Constructing the target protein pocket graph
[0101] To construct the pocket subgraph of each target protein, first obtain the pids of the target proteins from three public databases, namely Human, Biosnap, and Drugbank, respectively. Then, according to the protein pids, obtain the three-dimensional protein structure files from the Protein Data Bank (PDB). After that, the present invention uses a method based on computational geometry and protein clustering to extract the corresponding protein pockets, and constructs the corresponding protein pocket graphs for each protein pocket. Finally, each protein is represented as multiple pocket subgraphs, where each subgraph is represented as G p =(V p, E p ), where V p represents residues and E p represents the edges (chemical bonds) between residues.
[0102] To initialize the initial features of the protein pocket subgraph, each residue node is represented as a 41-dimensional vector characterizing the residue type, residue self-distance, and residue dihedral angle. The edge between two residues is represented as a 5-dimensional vector of residue connectivity, residue CA distance, residue center distance, and residue maximum distance.
[0103] 2. Use GraphTransformer to extract the pocket features of the target
[0104] After constructing the subgraph structures corresponding to each sub-pocket of the target protein, the present invention uses GraphTransformer as a molecular graph encoder to encode the subgraph structures corresponding to each sub-pocket of the drug, and obtains the pocket features of the target, which are the three-dimensional structural features of the target.
[0105] The method of the present invention for using GraphTransformer to extract the pocket features of the target is the same as the method of using GraphTransformer as a molecular graph encoder to encode the subgraph structures corresponding to each molecular fragment of the drug.
[0106] 3. Use the self-attention mechanism to extract the one-dimensional sequence features of the target
[0107] There is a close relationship between the one-dimensional sequence and three-dimensional structure of a protein. For example, the one-dimensional amino acid sequence of a protein determines its three-dimensional spatial structure, and the three-dimensional structure determines the function and properties of the protein, such as the binding site. Based on this, the embodiments of the present invention fuse the one-dimensional structure (sequences) and three-dimensional structure (pocktes) to characterize the complex protein structure.
[0108] Before performing the target protein feature fusion, first process the amino acid sequence features of the protein. Specifically, after one-hot encoding the one-dimensional amino acid sequence of each target protein, P s ∈Rn×d , each amino acid is represented as a 20 - dimensional vector. Then, the initialized sequence features are fed into the self - attention module to learn the correlations within the sequence, and the process is as follows:
[0109]
[0110] In the formula, is the target sequence feature obtained by learning the correlations within the sequence through the self - attention mechanism, that is, the one - dimensional sequence feature of the target protein; Q s , K s , V s are the parameters of the self - attention mechanism respectively.
[0111] 4. Use cross - attention to fuse protein 1D and 3D features
[0112] Based on the excellent performance of the attention mechanism in various fields, this embodiment uses cross - attention as the feature fusion method. Cross - attention makes full use of the advantages of two information sources to capture more complex relationships between the sequence and the spatial structure, and the process is as follows:
[0113]
[0114] Finally, the target protein feature R that contains both sequence information and three - dimensional structure information can be obtained p .
[0115] III. Use a bilinear attention network as a classifier to predict the probability of drug - target interaction
[0116] The bilinear attention network (BAN) architecture is an attention mechanism commonly used in visual question - answering tasks. BAN can effectively capture the complex bilinear interactions between two modalities. It introduces bilinear pooling to express the relationship between the question (text) and image (visual) features in a finer - grained way. Here, this embodiment uses it to capture the fine - grained interaction features between protein and drug molecule pairs.
[0117] Specifically, the above - encoded protein feature R p and drug feature R d are input into the BAN module. First, the bilinear attention map is calculated to generate the attention map A, which reflects the interaction between the two different modality inputs of the drug molecule and the protein molecule; then, based on the attention map A generated in step one, the joint feature f of the protein feature and the drug molecule feature is calculated, and finally, the final correlation score is calculated through the MLP layer. The specific process is as follows:
[0118]
[0119]
[0120] IV. Experimental Verification
[0121] 1. Experimental Data
[0122] The following data were mainly used in the experiment: (1) Protein sequence data. Since the sequence structure of proteins is the basis of complex spatial structures and the biological functions of proteins are implemented by individual or continuous amino acid fragments, the sequence structure cannot be ignored when exploring the biological functions of proteins. To obtain protein sequences, three commonly used and publicly available datasets, namely Human, Biosnap, and Drugbank, were downloaded in this experiment, which contain 1,787, 1,971, and 4,255 protein sequence data respectively; (2) Protein structure data. Considering that the interaction of drug-target pairs is based on a specific region in the protein (i.e., the pocket structure), the pocket structure features are crucial for the model to explore potential drug-target pairs. To obtain pocket features, the structure files of proteins (.pdb files) need to be obtained from the Protein Data Bank (PDB) database first. The structure files contain the atomic structure coordinates of each amino acid in the protein and the connection relationships between atoms; (3) Drug molecule data. The three commonly used and publicly available datasets of Human, Biosnap, and Drugbank downloaded contain 2,726, 4,510, and 6,647 drug molecule data respectively. The drug molecules in the datasets mainly exist in the form of SMILES sequences, and then relevant library functions can be called to convert them into 2D molecular graphs; (4) Drug-target interaction data. In the three commonly used and publicly available datasets of Human, Biosnap, and Drugbank downloaded, each row of data contains: 1. Drug data (in the form of SMILES); 2. Protein amino acid sequence data; 3. Labels (indicating whether the drug-protein pair actually interacts). Among them, the SMILES sequences of drugs and the protein amino acid sequences are used as model inputs, and the label information is used to calculate the accuracy of the model prediction ability.
[0123] 2. Evaluation Metrics
[0124] To verify the effectiveness of this method, five evaluation metrics were used in this method, and comparative experiments, ablation experiments, and cold start experiments were designed to test the prediction performance of the MotifGT-DTI method.
[0125] The Receiver Operating Characteristic (ROC) curve is used to evaluate the classification ability of a binary classifier when its discrimination threshold changes. This metric describes the relationship between the True Positive Rate (TPR, sensitivity) and the False Positive Rate (FPR, 1 - specificity) according to different discrimination thresholds. TPR refers to the proportion of correctly predicted positive samples among all positive samples. FPR refers to the proportion of incorrectly predicted positive samples among all negative samples. A prediction value higher than the discrimination threshold is judged as a positive sample, and conversely, a prediction value lower than the discrimination threshold is judged as a negative sample. TP refers to the number of samples judged as positive and actually being positive, FP refers to the number of samples judged as positive but actually being negative, TN refers to the number of samples judged as negative and actually being negative, and FN refers to the number of samples judged as negative but actually being positive.
[0126]
[0127] Calculate the Area Under the Curve (AUC) of the ROC curve as one of the metrics to measure the model performance.
[0128] In an imbalanced dataset, AUPRC can effectively measure the model performance. It shows the changing relationship between the Precision and Recall of the model at different thresholds. The larger the value, the better the overall performance of the model. Among them, Precision refers to the proportion of actually positive samples among the positive class samples predicted by the model, reflecting the performance of the model in terms of false positives. Recall refers to the proportion of actually positive samples that are correctly predicted as positive, which is used to evaluate the performance of the model in terms of false negatives.
[0129]
[0130] To more comprehensively evaluate the model performance, this paper also uses the F1 - score as an evaluation metric. The F1 - score is the harmonic mean of the recall rate (Recall) and the precision rate. Especially in the case of class imbalance, the comprehensive performance of the model can be evaluated by calculating the F1Score. The formula is as follows:
[0131]
[0132] 3. Comparison with Other Methods
[0133] To evaluate the performance of MotifGT-DTI, it was compared with the current four state-of-the-art computational methods, including TransformerCPI, MolTrans, DrugBAN, and FragXsiteDTI. For a fair and reasonable comparison, all models were evaluated based on the same benchmark dataset, and the parameters were the recommended optimal parameters. On the Human dataset, MotifGT-DTI achieved the highest AUROC of 0.995 and AUPRC of 0.994, and Precision and Recall also reached 0.954 and 0.962 respectively, showing the most excellent prediction performance on this dataset. On the Biosnap dataset, MotifGT-DTI also achieved the optimal AUROC of 0.918 and AUPRC of 0.926, and the performance of Precision, F1Score, and Recall also significantly outperformed other models, reaching 0.844, 0.852, and 0.861 respectively, indicating its high generalization ability and stability. On the Drugbank dataset, MotifGT-DTI led again with AUROC 0.896 and AUPRC 0.902, and the Precision and Recall metrics also showed relatively high performance, further confirming its excellent performance on multiple datasets and different evaluation metrics. Overall, the MotifGT-DTI model showed significant advantages in multiple evaluation metrics, especially leading significantly in the AUROC and AUPRC metrics, reflecting its strong generalization ability and robustness in the drug-target prediction task and making it the most competitive model on the three datasets.
[0134] 4. Validate the model generalization ability
[0135] When actually applying the DTI model, most drugs and proteins in the training set usually do not appear in the test set. Whether the model can still provide accurate predictions when facing unseen drugs or unseen proteins is a key issue in evaluating the model's generalization ability and robustness. Therefore, this experiment conducted a cold start experiment on the Human dataset to evaluate the performance of MotifGT-DTI in these special cases. Specifically, three different data partitioning strategies were adopted for the experiment:
[0136] Drug cold start: Remove any drug that appears in the training set from the test set.
[0137] Target cold start: Remove any protein that appears in the training set from the test set.
[0138] Pair cold start: Remove any drug or target that appears in the training set from the test set.
[0139] The drug cold start and target cold start scenarios respectively evaluate the sensitivity of the model when facing new drugs and new targets. The paired cold start measures the generalization ability of the model when facing completely unknown drug-target pairs. As can be seen in Figure 4 , MotifGT-DTI achieved the highest AUROC value in all scenarios, indicating its excellent ability to capture the features of new entities. The paired cold start scenario brings greater challenges by introducing new drugs and new targets. Notably, MotifGT-DTI always showed the best performance, demonstrating its strong robustness under various conditions.
[0140] 5. Case Analysis
[0141] To evaluate the interpretability of the model in the method of the present invention for drug-target interaction prediction, a series of visualization experiments were conducted. Human lactate dehydrogenase A (LDH-A) is a key enzyme that promotes cancer cell growth, and inhibiting its activity is of great significance for cancer treatment. The complex formed by LDH-A and ligand 9YA was downloaded from PDB (PDBID: 5W8L), which did not appear in the training data and belongs to a typical paired cold start case. Notably, the model in the method of the present invention successfully identified this unseen protein-ligand interaction pair, demonstrating the performance of the method of the present invention when predicting data outside the dataset. The bilinear attention map visualizes each substructure, providing key insights and explanations at the molecular level for drug design work. As shown in Figure 6 (b), the four drug structures predicted in the attention map had the highest attention scores and were marked with dashed circles. Compared with the known true protein and ligand binding sites, MotifGT-DTI correctly predicted three of them. As shown in Figure 6 (c), the remaining one was found to be highly correlated with the binding site. Specifically, the key binding structure - the carboxylic acid group was correctly identified (interacting with residues Arg168, His192, and Thr247 as a hydrogen bond acceptor). The thiazole structure interacting with residue His192 was also correctly identified. The biphenyl ring structure interacting with residues Arg105 and Asn137 was also correctly predicted as a drug structure that plays an important role in the binding process. Although the predicted oxazole structure is not currently a known binding site, studies have shown that it is a key chemical feature in LDH-A related inhibitors and is closely related to reducing LDH-A enzyme activity and inhibiting cancer cell growth. All results indicate that MotifGT-DTI can identify functional molecular motifs that contribute to the interaction and provide interpretability for the prediction results.
[0142] The above embodiments are the preferred embodiments of the present application. Those of ordinary skill in the art can also make various transformations or improvements based on this. Without departing from the general concept of the present application, these transformations or improvements should all fall within the scope of protection required by the present application.
Claims
1. A drug-target interaction prediction method based on multi-view feature fusion, characterized in that: include: Step 1, extracting drug features and target features; (1) Extracting drug features: First, generate the corresponding 2D molecular graph based on the SMILES sequence of the drug, and construct the drug graph structure based on the characteristics of each atom in the drug molecule and the chemical bond relationship between them; then, extract the molecular fragments of the drug to obtain the subgraph structure corresponding to each molecular fragment; finally, use a Graph Transformer as a molecular graph encoder to encode the subgraph structure corresponding to each molecular fragment of the drug to obtain the drug features; (2) Extracting target features: First, extract the protein pockets of the target and construct the graph structure of the target based on the residues in the protein pockets and the chemical bonds between them. Then, divide the target into multiple sub-pockets to obtain the sub-graph structure corresponding to each sub-pocket. Then, use a Graph Transformer as a pocket encoder to encode the sub-graph structure corresponding to each sub-pocket of the target to obtain the pocket features of the target. In addition, a self-attention module is used to extract the sequence features of the target from the one-dimensional amino acid sequence of the target. Finally, cross-attention is used to fuse the target pocket features and sequence features to obtain the target features; In step 2, a bilinear attention network is used as a classifier to deeply fuse drug features with target features, and the interaction probability between drug and target is calculated through a multi-layer perceptron.
2. The drug-target interaction prediction method based on multi-view feature fusion according to claim 1, characterized in that: In the drug graph structure, each node represents an atom. Its eigenvector These properties include: the type of atom, the degree of the atom, the formal charge of the atom, the number of radical electrons of the atom, the hybridization state of the atom, the aromaticity of the atom, the total number of hydrogen atoms attached to the atom, and the chirality of the atom; The edges between nodes represent chemical bonds between atoms, and their characteristics include the following properties: the type of bond, whether the bond is conjugated, whether the bond belongs to a ring structure, and whether the bond has stereochemical characteristics; In the graph structure of the target, each node represents a different residue in the three-dimensional structure of the target protein, and the edges between the nodes represent the chemical bonds between the corresponding residues; the characteristic vector of each node residue includes the residual type, residual self-distance and residual dihedral angle, and the characteristics of the edges between node residues include: residual connectivity, residual CA distance, residual center distance and residual maximum distance.
3. The drug-target interaction prediction method based on multi-view feature fusion according to claim 1, characterized in that: The RDKit method library is used to generate a 2D molecular map of the drug based on the SMILES sequence of the drug; the Macfrag method is used to process the 2D molecular map of the drug and extract molecular fragments of the drug that meet the preset chemical meaning.
4. The drug-target interaction prediction method based on multi-view feature fusion according to claim 1, characterized in that: By combining computational geometry methods with protein clustering technology, protein pockets are extracted from the three-dimensional structure of the target protein. The specific steps include: First, the protein space is voxelized to construct the protein grid structure; The convex hull of the protein is then calculated, dividing its surface into triangles, each representing a potential binding site; Then, the protein pocket is defined by the volume generated by the triangle vertices, and the empty voxels closest to the protein atoms are identified to determine the possible ligand atom positions as candidate protein pockets; Finally, based on the overlap between the candidate protein pockets and the biochemical and physical properties of each candidate protein pocket, several protein pockets with high confidence are screened out from all the candidate protein pockets as the final extracted protein pockets.
5. The drug-target interaction prediction method based on multi-view feature fusion according to claim 1, characterized in that: Use Graph Transformer as the molecular graph encoder to encode the subgraph structure corresponding to each molecular fragment of the drug, specifically: (1) The feature of the i-th node in the subgraph structure corresponding to each molecular fragment of the drug is represented as x i , the feature of the edge between the i-th and j-th nodes i is represented by y ij ; (2) Input the node feature x in the subgraph structure i With edge feature y ij Perform a linear transformation: in, and is the linear layer weight parameter in Graph Transformer, and is the linear layer bias parameter in GraphTransformer; Represents x i The node features obtained by linear transformation, Represents y ij Edge features obtained through linear transformation; (3) Use the eigenvector of the Laplacian operator as the node position encoding: in, The Laplacian matrix representing the subgraph structure; I is an N-dimensional unit matrix with a diagonal of 1, A is an N×N adjacency matrix constructed by the connection relationship between the nodes in the subgraph structure; D is a degree matrix, representing the degree of each node; Λ, U represent the diagonal matrix of eigenvalues and the orthogonal matrix of eigenvectors respectively; according to the distribution of eigenvalues after Laplacian matrix decomposition, select the eigenvectors corresponding to the smallest non-zero eigenvalues as the multidimensional position code λ of the node i ; is a weight matrix; is the bias vector; Represents the initial position encoding of node i, by transforming the Laplace eigenvector λ i Linear transformation is obtained; Indicates the node feature with added location information; (4) Graph Transformer’s multi-head attention mechanism is used to fuse node features and edge features; The detailed update steps for the lth layer of the Graph Transformer network are as follows: in, Represent the features of node i and node j in layer l respectively, represents the edge between node i and node j in layer l; Q k,l , K k,l 、V k,l and E k,l is the parameter of the kth linear layer in the lth layer, They represent the query, key, and value of the k-th head node feature and edge feature of the l-th layer after the normalization function Norm(), respectively. Represents the edge features between nodes i and j in the lth layer; The attention weight is calculated by jointly calculating the node features and edge features; (5) The features learned by multiple attention heads are integrated through the Concat operation to update the node features and edge features: in, is the feature of node i in the l+1th layer, N i are all neighbor nodes of node i, is the edge feature between node i and node j in the l+1th layer, represents the linear layer parameter in the multi-head attention mechanism, h gt Represents the number of attention heads; (6) The node and edge features after multiple fusion and updating of the node features and edge features through the multi-head attention mechanism are input into the feedforward network with residual connection and normalization operation: in, and as well as and is the weight matrix used to convert the l+1th layer node features and edge features to the hidden layer; They are and The node features and edge features after the residual and normalization operations are updated; ReLU represents the activation function in the neural network; The node features including edge features are obtained through the above process The features of n nodes in each subgraph structure, including edge features Features that make up the subgraph structure Expressed as Finally, the features of all the subgraph structures of the drug are combined into the feature R of the drug d ; Using Graph Transformer as a pocket encoder to encode the subgraph structure corresponding to each sub-pocket of the target to obtain the pocket features of the target is the same as using Graph Transformer as a molecular graph encoder to encode the subgraph structure corresponding to each molecular fragment of the drug.
6. The drug-target interaction prediction method based on multi-view feature fusion according to claim 1, characterized in that: The method of using a self-attention module to extract the sequence features of the target from the one-dimensional amino acid sequence of the target specifically includes: (1) One-hot encoding of the target's one-dimensional amino acid sequence is performed to obtain P s ∈R n×d ; Where n represents the sequence length and d represents the encoding dimension; (2) Code P s Send to the self-attention module to learn the sequence encoding P s Internal correlation, the process is as follows: In the formula, Q is the target sequence feature obtained by learning the correlation between sequences through the self-attention mechanism. s , K s , V s are the parameters of the self-attention mechanism respectively; They are the query weight matrix, key weight matrix, and value weight matrix in the self-attention module respectively; Attention s represents the self-attention mechanism, softmax is the activation function, d k is the scaling factor.
7. The drug-target interaction prediction method based on multi-view feature fusion according to claim 1, characterized in that: Use cross-attention to fuse the target pocket features and sequence features to obtain the target features, expressed as: Where W Q , W K , W P Represent the query weight matrix, key weight matrix, and value weight matrix, respectively, which are used to convert sequence data into query vector Q P , key vector K P and the value vector V P ; R p is the target feature obtained by fusion; is the target sequence feature extracted by the self-attention module, The target pocket feature is obtained by encoding the subgraph structure corresponding to each sub-pocket of the target by the pocket encoder; Attention c stands for Cross-Attention Mechanism.
8. The drug-target interaction prediction method based on multi-view feature fusion according to claim 1, characterized in that: Step 2 uses a bilinear attention network as a classifier to deeply fuse drug features with target features, and calculates the interaction probability between drugs and targets through a multi-layer perceptron, including: First, the bilinear attention map is calculated to generate the attention map A; In the formula, q T is a learnable weight vector used to calculate the interaction strength; M d , is a learnable parameter matrix, which is used to transform the drug features R d and target feature R p Mapped into a shared space; Then, the joint feature f of the target feature and the drug feature is calculated based on the attention map A: In the formula, σ represents the activation function, M d ′、M p ′ is the same as M d as well as The function is the same, indicating that it is used to convert the drug feature R d And the target feature R p The parameter matrix mapped to the shared space; Finally, the final attention score is calculated by the multi-layer perceptron MLP In the formula, W represents the learnable weight matrix, b represents the learnable bias vector, and Sigmoid represents the activation function; the final calculation results in the attention score is the probability of interaction between drug and target.
9. An electronic device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: When the computer program is executed by the processor, the processor is caused to implement the method according to any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
TransGAT-based drug-target interaction prediction method
CN116312808A
Drug-target interaction prediction method based on bidirectional Intention network
CN117746990A
DTI prediction method based on hierarchical multi-modal self-attention graph neural network
CN118173161A
Drug target and drug key pharmacophore detection method based on gene and drug atom interaction relation extraction method
CN118298910A
Multi-modal characterization fusion drug target binding affinity prediction method
CN118471324A
Cited By
Drug target affinity prediction method and system based on three-dimensional molecule fragmentation
CN120877844A
Method and system for predicting drug target affinity based on three-dimensional molecular fragmentation
CN120877844B
Medicine appearance defect detection method and related equipment
CN121599973A
A method for detecting defects in the appearance of a pharmaceutical product and related apparatus
CN121599973B
Drug target interaction prediction method, system and equipment and storage medium
CN121641166A