Method, model, equipment and medium for predicting drug-target interaction
Through the combination of graph attention sampling, NDLS, graph SAGE and bilayer GTN, the problems of limited receptive fields and excessive smoothness in the prior art are solved, and the prediction ability of drug-target interaction prediction models are improved.
Patent Information
- Application Number
- CN202510072516.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-17
AI Technical Summary
When predicting drug-target interactions, shallow GNNs with limited receptive fields can only aggregate incomplete information in neighborhoods and cannot consider the receptive fields at different levels of network nodes, resulting in insufficient prediction capabilities.
Through graph attention sampling and multi-neighborhood interactive attention fusion, NDLS and graph SAGE modules are used for adaptive iterative optimization and deep sampling aggregation, and information aggregation is combined with double-layer GTN to aggregate node multi-source information to improve prediction capabilities.
Effectively aggregate more relevant node information from neighborhoods, improving the prediction ability of drug-target interaction prediction model, and solving the problems of limited receptive field and oversmoothing.
Smart Images

Figure CN119993258A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of predicting drug-target interaction, and in particular to a method, model, device and medium for predicting drug-target interaction. Background Art
[0002] In the characterization of drugs and proteins, a large number of studies have inferred potential drug-target interactions (DTI) based on the sequence information of drugs and proteins. Potential drug-target interactions (DTI) are a key component of drug discovery and new uses. This type of method has the advantages of simplicity, high efficiency and a wide range of applications.
[0003] The sequence structures currently used are often learned only from the hidden representation of a single drug, ignoring the valuable information provided by the interaction between the two drugs for the substructure.
[0004] For example: The invention application with application number 202110382488.9 discloses a drug-target interaction prediction model method based on deep embedding learning of molecular graphs and sequences. This method establishes a graph neural network based on an attention mechanism and an attention-guided bidirectional LSTM to predict interactions. On the one hand, in terms of drug molecules, better spatial features can be learned based on molecular graphs; on the other hand, the amount of protein sequence data is large, which can cover a larger protein space and improve generalization capabilities.
[0005] However, the scheme it published also has the following problems: when extracting the sequence features of drugs and proteins respectively, the characteristic information between drugs and proteins is relatively independent, and the hidden characterization information that interactions with other drugs and proteins can provide for substructures will be ignored.
[0006] In addition, although Graph Neural Network (GNN) is widely used in the field of DTI, it still has other limitations. In existing methods, shallow GNNs with limited receptive fields can only aggregate incomplete information in the neighborhood, fail to consider the receptive fields of different levels of network nodes, and cannot effectively aggregate more relevant node information from the neighborhood.
[0007] For example: The invention application with application number 202211080659.3 discloses a drug and target prediction method based on graph attribute neural network. The scheme published in this application utilizes graph attribute neural network to reduce the dependence of the deep learning model for drug-target interaction prediction on training samples and improve the prediction performance.
[0008] However, the published solution also has the following problems: the receptive field is limited, GNN can only aggregate incomplete information within the neighborhood, cannot take into account the receptive fields of different levels of network nodes, and cannot effectively aggregate more relevant node information from the neighborhood.
[0009] Therefore, in reality, a method for predicting drug-target interactions is needed, which can dynamically extract potential interaction information from other drugs (targets); and take into account the receptive fields of different levels of network nodes, by integrating their respective complementary advantages, effectively aggregate more relevant node information from the neighborhood, and improve the predictive ability of the DTI prediction model. Summary of the invention
[0010] In view of the above-mentioned problems, the purpose of the present invention is to provide a method, model, device and medium for predicting drug-target interaction, focusing on the receptive fields of different levels of graph nodes, and performing adaptive iterative optimization of graph node features.
[0011] Embodiments of the present invention provide a method, model, device and medium for predicting drug-target interaction.
[0012] A first aspect: A method for predicting drug-target interactions, comprising:
[0013] S1, encode the sequence information of drugs and proteins, and generate drug and protein feature vector quantum graphs respectively;
[0014] S2, apply graph attention sampling to the feature vector subgraph to obtain the interaction information between drugs or proteins, and update the feature vectors of the central nodes of the drug and protein subgraphs;
[0015] S3, using NDLS to perform adaptive iterative optimization of drug and protein feature vectors; at the same time, using graph SAGE to update the feature vector embedding information of the central node and obtain multi-source information of the node;
[0016] S4. Use a double-layer GTN to aggregate multi-source information of nodes, use the fully connected layer as the classification module, output the prediction value, and determine whether there is an interaction relationship between drugs and proteins based on the prediction value.
[0017] Furthermore, when updating the feature vectors of the central nodes of the drug and protein subgraphs in S2:
[0018] The feature vectors of drug and protein center nodes are updated by using randomly sampled sets of similar attribute nodes and weight combinations.
[0019] Furthermore, the graph attention sampling in S2 is expressed as follows:
[0020] Attention=SoftMax(C V1 *C V2 T )*C V2 (1)
[0021] C V1 '=αC V1 +βAttention (2)
[0022] Among them, V1 is the central node and V2 is a set of similar attribute nodes.
[0023] Furthermore, the S3 uses the graph SAGE to update the feature vector embedding information of the central node, including:
[0024] The first-order neighbor nodes of the drug and protein central nodes are randomly sampled, and then their first-order neighbors are randomly selected with these neighbor nodes as the starting point. Finally, starting from the most edge node, the feature embedding information of the central node is updated layer by layer from the neighbor nodes.
[0025] Furthermore, the graph SAGE is used in S3 to update the feature vector embedding information of the central node, and the formula is expressed as follows:
[0026]
[0027] H agg (v) k =UPDATE(H v ,H agg (v)) (4)
[0028] Where v is a node, N(v) is a set of neighboring nodes, and H agg (v) k Updated embeddings for the graph SAGE model layers.
[0029] Furthermore, NDLS is used in S3 to perform adaptive iterative optimization of drug and protein feature vectors, and the formula is expressed as:
[0030] NDLS(v i ,ε)= min{k:||Q i (∞) -Q i (k) ||2<ε} (11)
[0031]
[0032] Among them, X i is node v i The eigenvector of , |‖‖|2 is the two-norm, Q i (k) Indicates Q (k) The i-th row indicates the k-th layer other nodes v i The influence distribution of is, ε is the distance parameter.
[0033] Furthermore, when the information of multiple sources of nodes is aggregated through the double-layer GTN in S4, it includes:
[0034] Multi-head attention calculation is performed on each edge from the distant node j to the source node i, and the formula is expressed as:
[0035]
[0036] e ij =W e e ij +b e (15)
[0037]
[0038] After obtaining the multi-head attention of the graph, information aggregation is performed from the distant node j to the source node i. The formula is expressed as:
[0039]
[0040] in, For a given feature node, is a trainable parameter, is the exponential scale dot product function, d is the hidden layer size of each head, is the source node feature, It is the feature of long distance nodes.
[0041] The second aspect: a model for predicting drug-target interactions, including:
[0042] The sequence information encoding module generates drug and protein feature vector quantum graphs respectively according to the sequence information encoding of drugs and proteins;
[0043] The graph attention sampling module applies graph attention sampling to the drug and protein subgraphs respectively to obtain the interaction information between drugs or proteins and updates the feature vectors of the central nodes of the drug and protein subgraphs;
[0044] NDLS and Graph SAGE modules use NDLS to perform adaptive iterative optimization of drug and protein feature vectors; at the same time, Graph SAGE is used to update the feature vector embedding information of the central node and obtain multi-source information of the node;
[0045] The GTN fusion prediction module aggregates multi-source information of nodes through a double-layer GTN, uses a fully connected layer as a classification module, and outputs the prediction value.
[0046] A third aspect: An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method provided in the first aspect when executing the program.
[0047] A fourth aspect: A non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method provided in the first aspect.
[0048] Beneficial effects of the present invention:
[0049] The present invention effectively improves the performance of the model through sampling ratio and weight distribution, uses the NDLS module to determine the specific propagation depth of the full-graph nodes, and performs adaptive iterative optimization of node features; at the same time, uses graph SAGE to deeply sample and aggregate neighborhood nodes, focusing on the receptive fields of network nodes at different levels; through GTN, the complementary advantages of each are integrated, so as to effectively aggregate more relevant node information from the neighborhood. The method of the present invention takes into account the receptive fields of different levels of network nodes, and by integrating the complementary advantages of each, more relevant node information is effectively aggregated from the neighborhood, further improving the prediction ability of the DTI prediction model. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 A schematic diagram of the process of the method for predicting drug-target interaction of the present invention;
[0051] Figure 2 This is a schematic diagram of the structure of the drug-target interaction prediction model of the present invention;
[0052] Figure 3 This is a performance comparison diagram of the ablation experiment of the model of the present invention;
[0053] Figure 4 It is a schematic diagram of the structure of the electronic device of the present invention. DETAILED DESCRIPTION
[0054] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar symbols throughout represent the same or similar elements or elements with the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be understood as limiting the present invention.
[0055] When applying GNN models in the field of existing drug-target interaction (DTI), it is easy to encounter smoothing problems when learning the potential representations of drugs and targets, and the shallow GNN receptive field is limited and can only aggregate incomplete information within the neighborhood.
[0056] In view of the above problems, the present invention provides a method for predicting drug-target interaction. Figure 1 A schematic diagram of a method for predicting drug-target interaction provided in an embodiment of the present invention, the method comprising:
[0057] S1. Encode the sequence information of drugs and proteins and generate drug and protein feature vector quantum graphs respectively.
[0058] The present invention uses two public data sets, DrugBank and Davis, to train and evaluate the proposed DTI prediction model. The DrugBank data set contains 6708 drugs and 4410 proteins, with 18655 known DTI information; the Davis data set contains 68 drugs and 379 proteins, with 7320 DTI information.
[0059] like Figure 2 As shown, using the above dataset, the sequence information of drugs and proteins is encoded to generate drug and protein feature vector quantum graphs respectively.
[0060] S2. Graph Attention Sampling is applied to drug and protein subgraphs respectively to obtain the interaction information between drugs or proteins. The feature vectors of drug and protein center nodes are updated by using randomly sampled sets of similar attribute nodes and weight combinations. The experimental research has greater flexibility and scalability by customizing the sampling ratio and weight distribution.
[0061] Graph Attention Sampling is a method for sampling nodes with different attributes in heterogeneous graphs (feature vector subgraphs).
[0062] The attention sampling method is applied to the central node. By considering the mutual information between it and the nodes of the same attribute, the potential correlation information between the various attribute nodes is explored to explore and optimize the prediction performance of the DTI prediction model.
[0063] In a heterogeneous graph, the central node is V1, and the set of nodes with the same attributes that contain the specified number of samples is V2. The graph attention sampling calculation formula is expressed as:
[0064] Attention=SoftMax(C V1 *C V2 T )*C V2 (1)
[0065] C V1 '=αC V1 +βAttention (2)
[0066] Using the above formula, we first calculate the attention score between the feature vector of the central node and the sampled node vector of the same attribute using matrix multiplication, and then normalize the attention score using the softmax function to obtain the sampled node attention score.
[0067] Then, a new vector representation of the central node is obtained based on the weight combination method. The attention sampling of the graph nodes updates the feature information of the central node by using random sampling and weight combination. Through customized sampling ratio and weight distribution, it has greater flexibility and scalability.
[0068] S3. NDLS is used to perform adaptive iterative optimization of drug and protein feature vectors. At the same time, graph SAGE is used to update the feature vector embedding information of the central node to obtain multi-source information of the node.
[0069] GraphSAGE is used to randomly sample the first-order neighbor nodes of the drug and protein central nodes, and then these neighbors are used as the starting point to randomly select their first-order neighbors. Finally, starting from the most edge nodes, the feature embedding information of the central node is updated layer by layer from the neighbor nodes. At the same time, NDLS is used to solve the over-smoothing or under-smoothing problem of the GNN propagation process, determine the specific propagation depth for the full graph nodes according to the graph structure, perform adaptive iterative optimization of drug and protein feature vectors, and perform multi-neighborhood feature embedding learning.
[0070] Among them, Graph Sample and Aggregate (GraphSAGE) is an inductive learning framework for learning graph nodes based on the feature information of neighboring nodes; Graph SAGE does not consider the complete k-hop neighborhood of the central node, but first samples the computational graph of the k-hop neighborhood to generate a feature representation of the central node.
[0071] For the input network G = (V, E), where V is the set of nodes and E is the set of links or edges in the graph. The node feature matrix Y∈R V×D , D is the dimension of node features. For example: the adjacency matrix of graph A is represented by R V×V , if two nodes u and v are connected, the value of A is equal to 1, otherwise it is 0.
[0072] The main purpose of using graph SAGE is to learn the graph nodes H v Feature embedding for each node v∈V. For each node v, create a set of neighbor nodes N(v) from its neighborhood, ensuring that each node has the same number of neighbors.
[0073] For each node v, the embeddings of its neighbors are aggregated by using the AGGREGATE function, as follows:
[0074]
[0075] H agg (v) k =UPDATE(H v ,H agg(v)) (4)
[0076] The aggregate embedding H of the SAGE model in the computational graph agg (v) Then, by using the UPDATE function to update the embedding of node v, H agg (v) k is the updated embedding of the SAGE model layer, and then the updated embedding is passed as the input to the next stage of the model.
[0077] Node-dependent Local Moothing (NDLS) is mainly used to propose a solution strategy to the over-smoothing and under-smoothing problems that may arise in the GNN information propagation process.
[0078] In the process of information aggregation in existing graph neural networks, the neighbor order used for information aggregation is the same for all nodes. Due to the differences in the local structures of nodes, aggregating neighbor information of different nodes may lead to under-smoothing or over-smoothing problems.
[0079] According to the classic GNN model, the feature representation of drugs and targets at the kth layer is denoted by X (k) , can be obtained by feedforward propagation recursion. This propagation process is described as:
[0080]
[0081] Where D = [d ij ] is the diagonal node degree matrix of G, r∈[0,1] is the convolution coefficient, W is the trainable weight matrix of the kth layer, and σ(·) is the activation function.
[0082] The oversmoothing problem is mainly caused by the multiplication of A and X. k By derivation, we can make σ(·) and W the unit function and unit matrix respectively, then formula X can be rewritten as:
[0083]
[0084] Among them, X (0) is equivalent to the initial representation matrix of C. Through the infinite depth (k→∞) propagation process, X (k) After smoothing, we get the final representation of the drug and target:
[0085]
[0086] in, is the final adjacency matrix of G, so Indicates v i and v j The weight between i,v j ∈V).
[0087] Assume X i Yes i The eigenvector of is also the i-th row of X. When calculating Q, X j (0) The change of X i (k) From this, we can determine the degree of influence of q ij (k) Values:
[0088]
[0089] Among them, Q i (k) Indicates Q (k) The i-th row indicates the k-th layer other nodes v i The NDLS strategy is then used to determine the minimum value of node-specific k, and the distance parameter ε is an arbitrarily small constant to control the smoothing effect.
[0090] The formula of NDLS is:
[0091] NDLS(v i ,ε)= min{k:||Q i (∞) -Q i (k) ||2<ε} (11)
[0092] Where |‖‖|2 is the second norm, NDLS(v i dr ,ε)>0, once the learning X is determined i k minimum values, an average operation is applied to aggregate the values from v i Sufficient neighborhood information within k hops of i The update rule is expressed as follows:
[0093]
[0094] Through the above formula, the feature vector of each node in G can be obtained, thereby determining X.
[0095] For a matrix Q, each element of it, such as q ij , which represents the mutual influence between nodes i and j. Even when both i and j are drug nodes, q can be obtained by ij The value of . It can extract distinguishable representation features by determining the optimal propagation depth for each node and further avoid over-smoothing. In addition, the features within k hops are aggregated and then averaged to better capture neighborhood information effectively.
[0096] S4. Use a double-layer GTN to aggregate multi-source information of nodes, use the fully connected layer as the classification module, output the prediction value, and determine whether there is an interaction relationship between drugs and proteins based on the prediction value.
[0097] The Graph Transformer Network (GTN) introduces the attention mechanism of the Transformer model into graph structure learning, which significantly improves the model training speed and can more effectively capture the relevant information between feature nodes and the entire graph.
[0098] Specifically, given a feature node The multi-head attention for each edge from distant node j to source node i is calculated as follows:
[0099]
[0100] e ij =W e e ij +b e (15)
[0101]
[0102] in, is an exponentially scaled dot product function, and d is the hidden layer size of each head. First, using different trainable parameters Convert source node features and distant node features into and Then the edge feature e ij Encoded and added as supplementary information to each layer Calculate the attention of each edge from the distant node j to the source node i.
[0103] After obtaining the attention of the graph, information is aggregated from distant node j to source node i:
[0104]
[0105] The present invention also provides a model for predicting drug-target interaction, such as Figure 2 As shown, the model includes:
[0106] The sequence information encoding module generates drug and protein feature vector quantum graphs respectively according to the sequence information encoding of drugs and proteins;
[0107] The graph attention sampling module applies graph attention sampling to the drug and protein subgraphs respectively to obtain the interaction information between drugs or proteins and updates the feature vectors of the central nodes of the drug and protein subgraphs;
[0108] NDLS and Graph SAGE modules use NDLS to perform adaptive iterative optimization of drug and protein feature vectors; at the same time, Graph SAGE is used to update the feature vector embedding information of the central node and obtain multi-source information of the node;
[0109] The GTN fusion prediction module aggregates multi-source information of nodes through a double-layer GTN, uses a fully connected layer as a classification module, and outputs the prediction value.
[0110] The DTI prediction model constructed by the present invention is encoded according to the sequence information of drugs and proteins through the sequence information encoding module, and the corresponding feature vectors are generated respectively. The graph attention sampling module is used to apply attention sampling to the drug and protein subgraph nodes respectively to consider the interactive information between drugs and proteins. The feature vectors of the drug and protein central nodes are updated by using a randomly sampled node set and a weight combination. The experimental research has greater flexibility and scalability by customizing the sampling ratio and weight distribution. Then, the NDLS and graph SAGE modules are used to solve the over-smoothing or under-smoothing problem of the GNN propagation process by adopting the NDLS strategy, and the specific propagation depth is determined for the full graph node according to the graph structure, and the adaptive iterative optimization of the drug and protein feature vectors is performed; at the same time, the first-order neighbor nodes of the drug and protein central nodes are randomly sampled by using the graph SAGE, and then their first-order neighbors are randomly selected with these neighbors as the starting point, and finally starting from the most edge nodes, the feature embedding information of the central node is updated layer by layer from the neighbor nodes. Finally, the GTN fusion prediction module is used to fuse multi-source information through a double-layer GTN, and the fully connected layer is used as the classification module. The binary cross entropy is selected as the model loss function, and the output is a prediction value between 0 and 1. The prediction value is used to determine whether there is an interaction relationship between drug proteins.
[0111] The evaluation of the model of the present invention can use the five-fold cross-validation method, and use the AUC and AUPR indicators to evaluate the model performance. The AUC curve is a curve with FPR as the horizontal axis and TPR as the vertical axis. AUC represents the area under the ROC curve. The AUPR curve is a curve with Recall as the horizontal axis and Precision as the vertical axis. AUPR is the area under the PR curve. Among them, TP represents the number of true positive samples, TN represents the number of true negative samples, FP represents the number of false positive samples, and FN represents the number of false negative samples. The formula is expressed as follows:
[0112] TCP(Recall)=TP / (TP+FN)
[0113] FPR=FP / (FP+TN)
[0114] Precision=TP / (TP+FP)
[0115] Setting parameter experiments,In order to explore the impact of the number of attention samples on the experimental results, on the DrugBank dataset, a parameter range with a data interval of 50 and a sampling number ranging from 100 to 500 was selected for performance evaluation. The experimental results are shown in Table 1:
[0116] Table 1 Performance comparison of the number of attention samples
[0117]
[0118]
[0119] From the results in Table 1, it can be seen that before the sampling number is 350, the performance of the model also shows an overall upward trend with the increase of the sampling number. After that, the prediction performance gradually decreases with the increase of the sampling number. It can be seen that the model achieves the best performance when the attention sampling number is 350.
[0120] In order to verify the effectiveness of GraphSAGE and NDLS modules for DTI prediction, ablation experiments were conducted on the DrugBank dataset. Specifically, the performance of each new model was evaluated by eliminating modules, and the importance of different modules on model performance was explored in combination with different numbers of attention samples. The experimental results are shown in Table 2:
[0121] Table 2 Performance comparison of ablation experiments
[0122]
[0123] It can be seen from Table 2 that in all value ranges of the number of attention samples, the AUC and AUPR values decrease after the GraphSAGE or NDLS module is removed.
[0124] In addition, in order to more intuitively show the impact of module ablation, the data are also counted in Figure 3 As shown in the figure, NDLS is the most important part affecting the model performance, and the performance of the model after its removal decreases most significantly. After the GraphSAGE module is removed, the model performance also decreases to varying degrees with the selection of different attention sampling numbers. From the perspective of model design, the ablation experiment shows that combining NDLS and GraphSAGE modules is effective. At the same time, each model achieves the best performance when the sampling number is 350, which is consistent with the conclusion of the parameter experiment. This may be because the feature complementarity extracted by the NDLS and GraphSAGE modules is best under this sampling number.
[0125] In order to further verify the effectiveness of GTN for feature fusion, feature fusion modules such as GCN, TAG, GAT and SuperGAT are selected for comparison. The experimental results are shown in Table 3:
[0126] Table 3 Performance comparison of different feature fusion modules in AUC and AUPR
[0127] GCN TAG GAT SuperGAT GTN AUC 0.926 0.930 0.959 0.933 0.972 AUPR 0.916 0.925 0.949 0.930 0.964
[0128] It can be seen from the results in Table 3 that compared with other advanced graph neural network methods, the use of GTN can effectively integrate the complementary information from the respective advantages of GraphSAGE and NDLS modules, enabling the model to achieve the best performance.
[0129] In order to further evaluate the performance of the model, six advanced DTI prediction methods, including MMDG-DTI, DeepConv-DTI, MCANet, TransformerCPI, Moltrans and HyperAttentionDTI, were selected, and the five-fold cross-validation method was used for comparative experiments on two public datasets, DrugBank and Davis. Tables 4 and 5 show the AUC and AUPR index values obtained after five-fold cross-validation of the model of the present invention and other models:
[0130] Table 4 Comparison with existing methods on DrugBank dataset
[0131] MMDG-DTI DeepConv-DTI MCANet TransformerCPI Moltrans HyperAttentionDTI Our AUC 0.901 0.836 0.921 0.865 0.862 0.889 0.973 AUPR 0.897 0.831 0.928 0.868 0.862 0.884 0.964
[0132] Table 5 Comparison with existing methods on the Davis dataset
[0133]
[0134]
[0135] It can be seen from the results in Table 4 and Table 5 that the model of the present invention has the best performance in the DTI prediction task. Compared with the other six models, the model of the present invention is more effective than other advanced methods in the DTI prediction task.
[0136] The present invention proposes a DTI prediction model based on graph attention sampling and multi-neighborhood interactive attention fusion to predict drug-target interactions. For the GNN-based DTI prediction model, the shallow GNN with limited receptive field can only aggregate incomplete information within the neighborhood, with strong structural bias and noise; the deep GNN has an over-smoothing problem and may aggregate a large amount of irrelevant information. The NDLS module is used to determine the specific propagation depth of the full-graph nodes and perform adaptive iterative optimization of node features. At the same time, GraphSAGE is used to deeply sample and aggregate neighborhood nodes, focusing on the receptive fields of network nodes at different levels. Finally, GTN fuses the complementary advantages of each other to effectively aggregate more relevant node information from the neighborhood.
[0137] In addition, the present invention further explores the role of feature information encoded by different drug sequences on feature information encoded by a single drug sequence; similarly, protein sequence features also incorporate feature information of other proteins. The results of parameter analysis experiments and ablation experiments confirm the positive effects of NDLS and GraphSAGE modules on the model. At the same time, the performance of the model can be effectively improved through sampling ratio and weight distribution, which also reflects to a certain extent that drugs (proteins) can obtain feature information that is beneficial to prediction from other drugs (proteins).
[0138] The present invention also provides an electronic device, Figure 4 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention, such as Figure 4 As shown, the electronic device may include: a processor, a communications interface, a memory, and a communications bus, wherein the processor, the communications interface, and the memory communicate with each other via the communications bus. The processor may call the logic instructions in the memory, for example, to execute the following method:
[0139] S1, encode the sequence information of drugs and proteins, and generate drug and protein feature vector quantum graphs respectively;
[0140] S2, apply graph attention sampling to the feature vector subgraph to obtain the interaction information between drugs or proteins, and update the feature vectors of the central nodes of the drug and protein subgraphs;
[0141] S3, using NDLS to perform adaptive iterative optimization of drug and protein feature vectors; at the same time, using graph SAGE to update the feature vector embedding information of the central node and obtain multi-source information of the node;
[0142] S4. Use a double-layer GTN to aggregate multi-source information of nodes, use the fully connected layer as the classification module, output the prediction value, and determine whether there is an interaction relationship between drugs and proteins based on the prediction value.
[0143] In addition, the logic instructions in the above-mentioned memory can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on such an understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0144] An embodiment of the present invention further provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the method provided in each of the above embodiments is implemented, for example, including:
[0145] S1, encode the sequence information of drugs and proteins, and generate drug and protein feature vector quantum graphs respectively;
[0146] S2, apply graph attention sampling to the feature vector subgraph to obtain the interaction information between drugs or proteins, and update the feature vectors of the central nodes of the drug and protein subgraphs;
[0147] S3, using NDLS to perform adaptive iterative optimization of drug and protein feature vectors; at the same time, using graph SAGE to update the feature vector embedding information of the central node and obtain multi-source information of the node;
[0148] S4. Use a double-layer GTN to aggregate multi-source information of nodes, use the fully connected layer as the classification module, output the prediction value, and determine whether there is an interaction relationship between drugs and proteins based on the prediction value.
[0149] The model embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, i.e., they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Those of ordinary skill in the art may understand and implement it without creative effort.
[0150] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0151] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for predicting drug-target interactions, characterized in that: include: S1, encode the sequence information of drugs and proteins, and generate drug and protein feature vector quantum graphs respectively; S2, apply graph attention sampling to the feature vector subgraph to obtain the interaction information between drugs or proteins, and update the feature vectors of the central nodes of the drug and protein subgraphs; S3, using NDLS to perform adaptive iterative optimization of drug and protein feature vectors; at the same time, using graph SAGE to update the feature vector embedding information of the central node and obtain multi-source information of the node; S4. Use a double-layer GTN to aggregate multi-source information of nodes, use the fully connected layer as the classification module, output the prediction value, and determine whether there is an interaction relationship between drugs and proteins based on the prediction value.
2. The method according to claim 1, characterized in that When updating the feature vectors of the central nodes of the drug and protein subgraphs in S2: The feature vectors of drug and protein center nodes are updated by using randomly sampled sets of similar attribute nodes and weight combinations.
3. The method according to claim 2, characterized in that The graph attention sampling in S2 is expressed as follows: Attention=SoftMax(C V1 *C V2 T )*C V2 (1) C V1 '=αC V1 +βAttention (2) Among them, V1 is the central node and V2 is a set of similar attribute nodes.
4. The method according to claim 1, characterized in that In S3, the feature vector embedding information of the central node is updated by using graph SAGE, including: The first-order neighbor nodes of the drug and protein central nodes are randomly sampled, and then their first-order neighbors are randomly selected with these neighbor nodes as the starting point. Finally, starting from the most edge node, the feature embedding information of the central node is updated layer by layer from the neighbor nodes.
5. The method according to claim 4, characterized in that In S3, graph SAGE is used to update the feature vector embedding information of the central node, and the formula is expressed as: H agg (v) k =UPDATE(H v ,H agg (v)) (4) Where v is a node, N(v) is a set of neighboring nodes, and H agg (v) k Updated embeddings for the graph SAGE model layers.
6. The method according to claim 1, characterized in that In the S3, NDLS is used to perform adaptive iterative optimization of drug and protein feature vectors, and the formula is expressed as follows: NDLS(v i ,ε)= min{k:||Q i (∞) -Q i (k) ||2<ε} (11) Among them, X i is node v i The eigenvector of , |‖‖|2 is the two-norm, Q i (k) Indicates Q (k) The i-th row indicates the k-th layer other nodes v i The influence distribution of is, ε is the distance parameter.
7. The method according to claim 1, characterized in that When the multi-source information of nodes is aggregated through the double-layer GTN in S4, it includes: Multi-head attention calculation is performed on each edge from the distant node j to the source node i, and the formula is expressed as: have been ij =W e have been ij +b e (15) After obtaining the multi-head attention of the graph, information aggregation is performed from the distant node j to the source node i. The formula is expressed as: in, For a given feature node, is a trainable parameter, is the exponential scale dot product function, d is the hidden layer size of each head, is the source node feature, It is the feature of long distance nodes.
8. A model for predicting drug-target interactions based on the method according to any one of claims 1 to 7, characterized in that: include: The sequence information encoding module generates drug and protein feature vector quantum graphs respectively according to the sequence information encoding of drugs and proteins; The graph attention sampling module applies graph attention sampling to the drug and protein subgraphs respectively to obtain the interaction information between drugs or proteins and updates the feature vectors of the central nodes of the drug and protein subgraphs; NDLS and Graph SAGE modules use NDLS to perform adaptive iterative optimization of drug and protein feature vectors; at the same time, Graph SAGE is used to update the feature vector embedding information of the central node and obtain multi-source information of the node; The GTN fusion prediction module aggregates multi-source information of nodes through a double-layer GTN, uses a fully connected layer as a classification module, and outputs the prediction value.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the method for predicting drug-target interaction according to any one of claims 1 to 7 are implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for predicting drug-target interaction according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Drug-target interaction prediction model method based on deep embedding learning of molecular graph and sequence
CN113327644A
A drug and target prediction method based on graph attribute neural network
CN115440297B
Substance characteristic prediction method, terminal and storage medium
CN114512198A
Silent sub prediction algorithm based on bidirectional gating recurrent neural network
CN114863995A
Prediction method and device for drug target interaction and readable storage medium
CN116504328A