Drug target affinity activity prediction method based on contrast fusion graph features
Through the drug target affinity activity prediction method based on the contrast fusion graph features, the graph convolution network and bidirectional LSTM are used to extract features, and the comparison fusion graph features are extracted through the graph isomorphic network, the problem of insufficient accuracy of drug target affinity activity prediction in the prior art is solved, and more efficient data utilization and interactive information capture are achieved, improving the accuracy of prediction.
Patent Information
- Application Number
- CN202510572473.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-05-06
AI Technical Summary
The prior art is difficult to effectively utilize large-scale and diversified data to capture the deep interactive information relationship between drugs and targets, resulting in insufficient accuracy in predicting drug target affinity activity.
The drug target affinity activity prediction method based on the contrast fusion graph features is adopted, the topological features of the drug are extracted through the graph convolution network, the context features of the target are extracted from the two-way LSTM, and the contrast fusion graph module is designed. The contrast fusion graph features are extracted through the graph isomorphic network, and the three features are finally spliced and dimensionally reduced to achieve prediction.
It improves the accuracy of predicting affinity activity of drug targets, can more effectively utilize large-scale data, capture deep interactive information between drugs and targets, and improves the accuracy of prediction.
Smart Images

Figure CN120089191A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of predicting the affinity activity of drug targets, and specifically relates to a method for predicting the affinity activity of drug targets based on contrast-fused graph features. Background Art
[0002] The prediction of drug-target affinity (DTA) usually relies on the chemical structure features of drug molecules and the biological structure features of targets, and combines methods such as deep learning to establish a model to quantify the binding constant between a drug and a target. With the rapid development of deep learning and artificial intelligence technologies, the performance of DTA models has been gradually improved, and more complex data features and deep models can be used for more accurate predictions.
[0003] Early DTA prediction methods mainly used traditional machine learning methods. Considering the limitations of traditional machine learning methods in DTA prediction, especially in dealing with complex and non-linear drug-target relationships, researchers began to explore deep learning methods. Existing deep learning methods are mainly divided into two categories: sequence-based methods and graph-based methods. Existing models still regard targets as one-dimensional sequences, and these models usually process the features of drugs and targets separately. For example, drugs are regarded as molecular graphs, and targets are regarded as one-dimensional sequences or contact graphs, which fail to effectively capture the complex interaction relationships between drugs and targets. Secondly, existing models often rely on limited training data and cannot deeply explore the potential deep characteristics of drugs and targets.
[0004] Therefore, there is a need for a method for predicting the affinity activity of drug targets based on contrast-fused graph features that can effectively utilize large-scale and diverse data and capture the deep interaction information relationship between drugs and targets. Summary of the Invention
[0005] The main object of the present invention is to provide a method for predicting the affinity activity of drug targets based on contrast-fused graph features to solve the problem in the prior art that large-scale and diverse data cannot be effectively utilized and the deep interaction information relationship between drugs and targets cannot be captured.
[0006] To achieve the above object, the present invention provides a method for predicting the affinity activity of drug targets based on contrast-fused graph features, specifically including the following steps: S1, Use a graph convolutional network to extract features from the SMILES string of a drug to obtain drug features.
[0007] S2, Use a bidirectional LSTM to extract features from the target sequence to obtain target features.
[0008] S3, Design a contrast-fused graph module, and based on the contrast-fused graph module, extract contrast-fused graph features.
[0009] S4. Concatenate the drug features, target features, and contrast fusion graph features and reduce the dimension to obtain a prediction score.
[0010] Further, step S1 specifically includes the following steps: S1.1. Convert the SMILES string of the drug into a graph representation through the tool RDKit, where the nodes in the graph are drug atoms.
[0011] S1.2. Input the drug atom nodes into a graph convolutional network to hierarchically aggregate the drug atom node features, then fuse all the node features into the drug global feature through a global addition operation, and finally map it into the drug feature through a linear layer.
[0012] Further, step S2 specifically includes the following steps: S2.1. Input the amino acid sequence of the target into the forward LSTM to capture the dependence of the current amino acid on the previous sequence, and input the amino acid sequence of the target into the backward LSTM to capture its association with the subsequent sequence.
[0013] S2.2. Add the hidden states of the forward LSTM and the backward LSTM to generate the target feature.
[0014] Further, step S3 specifically includes the following steps: S3.1. Calculate the Euclidean distance between two amino acids to determine whether there is contact between the two amino acids, and generate the target graph adjacency matrix.
[0015] S3.2. Create a fusion graph of the drug and the target through the central node.
[0016] S3.3. Perform self-supervised learning on the fusion graph based on the graph isomorphism network GIN to extract the contrast fusion graph features.
[0017] Further, step S3.1 specifically includes the following steps: S3.1.1. The distance calculation formula between amino acid A and amino acid C is: ; where is the distance between amino acid A ( ) and amino acid C ( ).
[0018] S3.1.2. If is less than N angstroms, it is considered that there is contact, and the corresponding position in the target graph adjacency matrix is recorded as 1; otherwise, it is recorded as 0.
[0019] Further, step S3.2 specifically includes the following steps: S3.2.1. Unify the drug graph dimension and the target graph dimension through a linear layer: ; Wherein, and are the features after linear transformation of the drug and the target respectively, and the dimensions are both , and are the weight matrices of the drug graph and the target graph respectively, and are the initial feature matrices of the drug and the target respectively, and are the bias vectors of the drug and the target respectively.
[0020] S3.2.2, Initialize the central node as the bridge connecting the drug graph and the target graph. The central node is represented as: .
[0021] S3.2.3, Update the adjacency matrix of the fusion graph : ; Wherein, and are the adjacency matrices of the drug graph and the target graph respectively.
[0022] Furthermore, the central node randomly connects several nodes in the drug and target graphs respectively, and the number of connecting edges is the same. The total number of nodes in the fusion graph is Wherein, and represent the number of nodes in the drug graph and the target graph respectively.
[0023] Furthermore, step S3.3 specifically includes the following steps: S3.3.1, Randomly mask some edges in each fusion graph to obtain a masked graph.
[0024] S3.3.2, For each fusion graph and masked graph, use the graph isomorphism network GIN to extract the node features in the fusion graph and the masked graph.
[0025] Furthermore, step S3.3.2 specifically includes the following steps: S3.3.2.1, Normalize the node features of the input fusion graph or masked graph through batch normalization: ; Wherein, is the normalized node feature, is the batch normalization operation.
[0026] S3.3.2.2. The node features of the fusion graph or the masking graph are processed by graph convolution: ; Among them, represents the multi-layer perceptron for aggregating neighbor node features, is the feature of node at the layer, is a learnable parameter, represents the set of neighbor nodes of node .
[0027] S3.3.2.3. After passing through multiple layers of GIN, a global pooling operation is performed on each graph to aggregate the node-level features into graph-level features. The pooling operation adopts the global addition pooling method: ; Among them, is the graph-level feature of the fusion graph or the masking graph, is the set of all nodes in the graph, is the node feature.
[0028] S3.3.2.4. Define the feature of the positive sample pair as the fusion graph feature and the masking graph feature of the positive sample pair, and the feature of the negative sample pair is and the masking graph feature of the negative sample pair. The contrastive loss is used for optimization: ; Among them, is the logarithmic function, is the exponential function, is the similarity function, is the temperature hyperparameter, is the feature of the fusion graph or the masking graph.
[0029] Furthermore, step S4 specifically includes the following steps: S4.1. Concatenate the three features to generate the final feature representation : ; Among them, is the concatenation operation, is the drug graph feature, is the target graph feature, is the contrastive fusion graph feature.
[0030] S4.2. Perform feature dimensionality reduction on through the first-layer linear transformation: ; in, and are the weights and biases of the first layer linear transformation, is the activation function, It is the feature after the first layer of dimensionality reduction.
[0031] S4.3, generate the final prediction value through the second layer of linear transformation: ; in, and are the weights and biases of the second layer linear transformation, is the predicted interaction score.
[0032] S4.4, using mean square error MSE loss function Update the model: ; in, is the total number of samples, and Respectively represent The predicted and true values of samples.
[0033] The present invention has the following beneficial effects: (1) The present invention uses a graph convolutional network to extract molecular topological features from the SMILES representation of the drug, and extracts contextual features of the target sequence through a bidirectional long short-term memory network.
[0034] (2) The present invention designs a contrast fusion graph module, which constructs a unified graph of drug molecule graph and target 2D contact graph, randomly generates a masking graph, and uses contrastive learning to optimize the feature representation of the unified graph.
[0035] (3) The present invention achieves accurate prediction of drug-target activity scores by narrowing the feature representation of positive samples in the pre-training stage and combining the concatenation of three features and linear layer mapping. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the specific implementation of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the specific implementation or the prior art description. Obviously, the drawings described below are some implementations of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. In the drawings: Figure 1The flowchart of a method for predicting the affinity activity of drug targets based on contrast fusion graph features of the present invention is shown.
[0037] Figure 2 The contrast fusion graph learning process is shown. Specific implementation manners
[0038] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention belong to the protection scope of the present invention.
[0039] As Figure 1 A method for predicting the affinity activity of drug targets based on contrast fusion graph features shown, specifically includes the following steps: S1. Use a graph convolutional network to extract features from the SMILES string of the drug to obtain drug features.
[0040] S2. Use a bidirectional LSTM to extract features from the target sequence to obtain target features.
[0041] S3. Design a contrast fusion graph module, and based on the contrast fusion graph module, extract contrast fusion graph features.
[0042] S4. Concatenate and reduce the dimension of the drug features, target features, and contrast fusion graph features to obtain a prediction score.
[0043] The present invention first uses a graph convolutional network to extract features from the SMILES string of the drug to analyze the topological structure of the drug; at the same time, uses a bidirectional LSTM to extract features from the target sequence to capture its context dependence. Subsequently, two basic graphs are constructed based on the structural information of the drug and the target: the drug molecular graph and the target 2D contact graph. On this basis, a contrast fusion graph module is designed, which connects the two basic graphs into a unified graph by introducing a central node and generates a masked graph with randomly masked partial edges to enhance the robustness of the features. The unified graph and the masked graph are respectively input into the graph isomorphism network for information transmission, the global feature representation is extracted through average pooling, and contrast learning is used to optimize the positive sample features of the drug and the target. Finally, the extracted drug topological features, target sequence features, and global features of the fusion graph are concatenated and input into a linear layer for low-dimensional mapping, and finally the activity prediction score of the drug-target is obtained.
[0044] Specifically, the SMILES string of a drug is converted into a graph representation that can be processed by a deep learning model through the tool RDKit. Graph feature extraction is a key step in extracting effective information from the constructed molecular graph. Here, a graph convolutional network is used to hierarchically aggregate the topological information of drug atom nodes. Finally, all node features are fused into the drug global feature through a global addition operation, and then mapped into the final topological feature of the drug through a linear layer.
[0045] Step S1 specifically includes the following steps: S1.1, Convert the SMILES string of the drug into a graph representation through the tool RDKit, where the nodes in the graph are drug atoms.
[0046] S1.2, Input the drug atom nodes into the graph convolutional network to hierarchically aggregate the drug atom node features, then fuse all node features into the drug global feature through a global addition operation, and finally map it into the drug feature through a linear layer.
[0047] Specifically, the amino acid sequence of a target is a complex linear data, which contains rich context information with forward and backward associations. For the amino acid sequence, a bidirectional long short-term memory network (LSTM) model is used. Two independent LSTMs scan the amino acid sequence from the forward and backward directions respectively. The forward LSTM captures the dependence of the current amino acid on the previous sequence, and the backward LSTM captures its association with the subsequent sequence. Finally, the hidden states in the two directions are added together to generate the global semantic representation of the target sequence.
[0048] Step S2 specifically includes the following steps: S2.1, Input the amino acid sequence of the target into the forward LSTM to capture the dependence of the current amino acid on the previous sequence, and input the amino acid sequence of the target into the backward LSTM to capture its association with the subsequent sequence.
[0049] S2.2, Add the hidden states of the forward LSTM and the backward LSTM together to generate the target feature.
[0050] Specifically, step S3 specifically includes the following steps: S3.1, Calculate the Euclidean distance between two amino acids, determine whether the two amino acids are in contact, and generate the target graph adjacency matrix.
[0051] S3.2, Create a fusion graph of the drug and the target through the central node.
[0052] S3.3, Perform self-supervised learning on the fusion graph based on the graph isomorphism network (GIN) to extract the contrastive fusion graph features.
[0053] Specifically, the present invention obtains three-dimensional coordinate data from a target sequence structure database (such as PDBe, PDBBind) to construct a 2D distance map of the target. If it has not been found in the database, high-confidence atomic spatial position data can also be generated through prediction models such as AlphaFold. After obtaining the three-dimensional coordinate information of each amino acid, the contact relationship between amino acids is determined through spatial distance calculation. Step S3.1 specifically includes the following steps:
[0054] S3.1.1, the distance calculation formula between amino acid A and amino acid C is: ; wherein, is the distance between amino acid A ( ) and amino acid C ( ).
[0055] S3.1.2, if is less than N angstroms, it is regarded as having contact, and the corresponding position in the target graph adjacency matrix is recorded as 1; otherwise, it is recorded as 0. By calculating all amino acid pairs, a complete adjacency matrix can be generated to describe the contact graph information of the target.
[0056] The present invention creates a fusion graph of the drug and the target to achieve effective interaction of information between the drug and the target. Previous methods usually considered the characteristics of the drug and the target separately, ignoring the interaction between the two. By constructing a fusion graph, the characteristics of the drug and the target can be integrated into a unified graph, thereby better capturing the potential relationship between them.
[0057] Specifically, step S3.2 specifically includes the following steps:
[0058] S3.2.1, before generating the fusion graph, two basic graphs, the drug molecule graph and the target 2D contact graph, need to be processed. Since the molecular structure of the drug and the feature dimensions of the target sequence are different, before fusion, the dimensions of the drug graph and the target graph need to be unified through a linear layer: ; wherein, and are the features of the drug and the target after linear transformation respectively, and the dimensions are both , and are the weight matrices of the drug graph and the target graph respectively, and are the initial feature matrices of the drug and the target respectively, and are the bias vectors of the drug and the target respectively.
[0059] S3.2.2. To connect the two basic graphs, initialize the central node as a bridge connecting the drug graph and the target graph. The central node is represented as: . Since the specific action sites of drugs and targets are unknown, the introduction of the central node can integrate the drug molecular graph and the target graph in a randomly connected manner.
[0060] S3.2.3. While constructing the fusion graph, its adjacency matrix also needs to be updated. Update the adjacency matrix of the fusion graph : ; where and are the adjacency matrices of the drug graph and the target graph respectively.
[0061] Specifically, the central node randomly connects several nodes in the drug and target graphs respectively, and the number of connecting edges is the same. The total number of nodes in the fusion graph is where and represent the number of nodes in the drug graph and the target graph respectively.
[0062] The present invention designs a pre-training process to capture the potential relationship between drugs and targets through self-supervised learning of a large number of fusion graphs. As Figure 2 shown, specifically, 100 drugs and 100 targets are collected from a public database as the pre-training dataset. Using the above method, each drug is combined with each target one by one to generate 10,000 fusion graphs, providing a data basis for the pre-training of the model.
[0063] After generating the fusion graphs, in order to further simulate the uncertainty and missing information in the graph structure, a certain proportion of edges in each fusion graph are randomly masked to obtain the corresponding 10,000 masked graphs. These masked graphs not only provide a simplified graph structure but also enhance the model's attention ability to key features in the graph through contrastive learning. Specifically, the positive sample pairs are composed of the same fusion graph and its corresponding masked graph, while the negative sample pairs are composed of combinations of different fusion graphs.
[0064] Specifically, step S3.3 specifically includes the following steps:
[0065] S3.3.1. Randomly mask some edges in each fusion graph to obtain a masked graph.
[0066] S3.3.2. For each fusion graph and masked graph, use the graph isomorphism network GIN to extract the node features in the fusion graph and the masked graph.
[0067] Specifically, for each fusion graph and mask graph, a graph isomorphism network (GIN) is used to extract node features in the fusion graph and mask graph. The key to GIN is that it uses a multi-layer perceptron (MLP) to aggregate the features of each node and its neighbors, thereby efficiently capturing the structural information of the graph. Step S3.3.2 specifically includes the following steps:
[0068] S3.3.2.1. Node features of the input fusion graph or mask graph Normalization via batch normalization: ; in, is the standardized node feature, It is a batch normalization operation.
[0069] S3.3.2.2, the node features of the fusion graph or mask graph are processed by graph convolution: ; in, represents a multilayer perceptron for aggregating neighbor node features, Is a node In the The characteristics of the layer, is a learnable parameter, Representation Node Through multiple layers of iteration, the representation of the node will continuously aggregate information from the neighborhood, thereby effectively capturing the structural characteristics of the graph.
[0070] S3.3.2.3, after passing through multiple layers of GIN, a global pooling operation is performed on each graph to aggregate node-level features into graph-level features. The pooling operation adopts the global additive pooling method: ; in, is the graph-level feature of the fusion graph or mask graph, is the set of all nodes in the graph, is the node feature; finally, the graph-level feature is a 1×128 vector representing the global characteristics of the entire graph.
[0071] S3.3.2.4, in order to use these features for contrastive learning, the loss function of positive and negative samples is defined. The specific goal is to shorten the feature distance of positive sample pairs and increase the feature distance of negative sample pairs. The features of positive sample pairs are defined as fusion graph features and the mask map features of the positive sample pair , the characteristics of the negative sample pair are and the mask map features of the negative sample pairs , using contrast loss Optimize: ; Among them, is the logarithmic function, is the exponential function, is the similarity function, is the temperature hyperparameter, used to control the smoothness of the similarity measure, is the feature of the fusion graph or the masked graph.
[0072] Through the self-supervised pre-training method, the model can learn more robust graph feature representations, laying a solid foundation for subsequent drug-target interaction prediction. In the pre-training stage, GIN is optimized to efficiently capture the node and structural features in the fusion graph, ensuring that the model has strong feature extraction capabilities.
[0073] Such as Figure 2 shown, in the formal training stage, the pre-trained and optimized GIN model is embedded as the first layer into the main model. The node features of the fusion graph are first processed through this GIN layer optimized by contrastive learning to extract the initial feature representation. Next, the fusion graph will continue to go through several layers of GIN for recursive calculation, and each layer will aggregate the information of the nodes and their neighbors to gradually generate higher-order feature representations. Specifically, after several layers of GIN, the node features of the fusion graph are integrated into graph-level features through the global pooling operation Pool . This method, through the seamless connection between pre-training and formal training, can not only effectively utilize prior knowledge to enhance the performance of the model in the initial stage, but also further explore the complex relationships between drugs and targets through the feature extraction process of multiple layers of GIN.
[0074] Specifically, step S4 specifically includes the following steps: S4.1, To achieve the final prediction of drug-target interaction, three features are concatenated to generate the final feature representation : ; Among them, is the connection operation, is the drug graph feature, is the target graph feature, is the contrastive fusion graph feature. The concatenated feature combines the information of drug molecules, target sequences and their interaction relationships, providing a comprehensive and deep representation for the prediction task.
[0075] S4.2, The model further processes the concatenated features through two linear layers to complete the prediction. Through the first linear transformation, the dimensionality of is reduced: ; wherein, and are the weight and bias of the first - layer linear transformation respectively, is the activation function, is the feature after the first - layer dimensionality reduction.
[0076] S4.3, generate the final predicted value through the second - layer linear transformation: ; wherein, and are the weight and bias of the second - layer linear transformation respectively, is the predicted interaction score.
[0077] S4.4, adopt the mean - square error MSE loss function to update the model: ; wherein, is the total number of samples, and respectively represent the predicted value and the true value of the th sample.
[0078] The model provided by the present invention includes a drug feature extraction module, a target feature extraction module, a contrast fusion graph module, and a prediction module. The drug feature extraction module extracts drug features, the target feature extraction module extracts target features, the contrast fusion graph module extracts contrast fusion features, the drug features, target features, and contrast fusion graph features are concatenated and dimensionally reduced, and the prediction score is obtained through the prediction module.
[0079] To evaluate the performance of the model Tri - DTA proposed by the present invention, a performance comparison of DTA prediction was conducted with the current state - of - the - art models on the Davis and KIBA datasets. The models involved are DeepDTA, MT - DTA, GraphDTA, AttentionDTA, and GRA - DTA. On the Davis dataset, the model proposed by the present invention shows advantages over the baseline model in all evaluation metrics. As shown in Table 1, especially in terms of the mean - square error loss MSE, the model Tri - DTA provided by the present invention achieved an improvement of 0.004, indicating that this method has improved in prediction accuracy. At the same time, it is also superior to the baseline model in terms of CI and . CI is the coefficient of consistency, which is used to measure whether the predicted affinity values of two randomly selected drug - target pairs are consistent with the ranking of their actual values. is used to evaluate the fitting degree of the model.
[0080] Table 1 Performance Comparison with Other Baseline Models in Davis
[0081] The results on the KIBA dataset also show that the proposed method performs excellently, as shown in Table 2. Compared with the baseline models, the MSE of the model proposed in the present invention is reduced by 0.005, specifically verifying its advantage in accuracy. Although it is comparable to the optimal benchmark model in terms of the CI index, it is slightly inferior to GRA-DTA in terms. Generally speaking, it is particularly prominent in the improvement of MSE, showing its potential in optimizing the prediction accuracy of DTA.
[0082] Table 2 Performance Comparison with Other Baseline Models in KIBA
[0083] In the present invention, the contrastive fusion graph module can effectively enhance the information interaction between drugs and targets, capture richer features, thereby improving the understanding of the drug-target interaction of the model of the present invention; secondly, through self-supervised pre-training, the model can learn the potential relationship between drugs and targets without artificial labels, which further improves the accuracy of DTA prediction. Self-supervised learning can not only enhance the expression ability of features, but also provide effective training signals when the data is scarce, so as to achieve more efficient feature extraction and information integration.
[0084] Of course, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Changes, modifications, additions or substitutions made by those skilled in the art within the scope of the essence of the present invention should also fall within the protection scope of the present invention.
Claims
1. A method for predicting drug target affinity activity based on comparative fusion graph features, characterized in that: The specific steps include: S1, use graph convolutional network to extract features from the SMILES string of the drug to obtain drug features; S2, use bidirectional LSTM to extract features of the target sequence to obtain target features; S3, designing a contrast fusion graph module, and extracting contrast fusion graph features based on the contrast fusion graph module; S4, concatenates and reduces the dimension of drug features, target features, and comparison fusion graph features to obtain the prediction score.
2. A method for predicting drug target affinity activity based on comparative fusion graph features according to claim 1, characterized in that: Step S1 specifically includes the following steps: S1.1, convert the SMILES string of the drug into a graph representation using the tool RDKit, where the nodes in the graph are drug atoms; S1.2, input the drug atomic nodes into the graph convolutional network to hierarchically aggregate the drug atomic node features, then fuse all node features into drug global features through global addition operations, and finally map them into drug features through linear layers.
3. A method for predicting drug target affinity activity based on comparative fusion graph features according to claim 1, characterized in that: Step S2 specifically includes the following steps: S2.1, input the amino acid sequence of the target into the forward LSTM to capture the dependency between the current amino acid and the previous sequence, and input the amino acid sequence of the target into the backward LSTM to capture its association with the following sequence; S2.2, add the hidden states of the forward LSTM and the backward LSTM to generate the target features.
4. The method for predicting drug target affinity activity based on comparative fusion graph features according to claim 1, characterized in that: Step S3 specifically includes the following steps: S3.1, calculate the Euclidean distance between two amino acids, determine whether the two amino acids are in contact, and generate the target graph adjacency matrix; S3.2, creating a fusion graph of drugs and targets through central nodes; S3.3, based on the graph isomorphism network GIN, self-supervised learning is performed on the fusion graph to extract comparative fusion graph features.
5. A method for predicting drug target affinity activity based on comparative fusion graph features according to claim 4, characterized in that: Step S3.1 specifically includes the following steps: S3.1.1, the distance calculation formula between amino acid A and amino acid C is: ; in, is amino acid A ( ) and amino acid C ( ) between; S3.1.2 If If it is less than N angstroms, it is considered that there is contact, and the corresponding position in the target graph adjacency matrix is recorded as 1; otherwise it is recorded as 0.
6. A method for predicting drug target affinity activity based on comparative fusion graph features according to claim 4, characterized in that: Step S3.2 specifically includes the following steps: S3.2.1, unify the drug map dimension and target map dimension through a linear layer: ; in, and They are the features of drugs and targets after linear transformation, and their dimensions are , and are the weight matrices of the drug graph and target graph, respectively. and are the initial feature matrices of drugs and targets, respectively. and are the bias vectors for drug and target, respectively; S3.2.2, initialize the central node as a bridge connecting the drug graph and the target graph. The central node is represented as: ; S3.2.3, Update the adjacency matrix of the fused graph : ; in, and are the adjacency matrices of the drug graph and target graph, respectively.
7. A method for predicting drug target affinity activity based on comparative fusion graph features according to claim 6, characterized in that: The central node is randomly connected to several nodes in the drug and target graphs, and the number of connecting edges is the same. The total number of nodes in the fusion graph is ,in, and Represents the number of nodes in the drug graph and target graph respectively.
8. The method for predicting drug target affinity activity based on comparative fusion graph features according to claim 4, characterized in that: Step S3.3 specifically includes the following steps: S3.3.1, randomly mask some edges in each fusion graph to obtain a masked graph; S3.3.2, for each fusion graph and mask graph, use the graph isomorphism network GIN to extract node features in the fusion graph and mask graph.
9. A method for predicting drug target affinity activity based on comparative fusion graph features according to claim 8, characterized in that: Step S3.3.2 specifically includes the following steps: S3.3.2.
1. Node features of the input fusion graph or mask graph Normalization via batch normalization: ; in, is the standardized node feature, It is a batch normalization operation; S3.3.2.2, the node features of the fusion graph or mask graph are processed by graph convolution: ; in, represents a multilayer perceptron for aggregating neighbor node features, Is a node In the The characteristics of the layer, is a learnable parameter, Representation Node The set of neighbor nodes of S3.3.2.3, after passing through multiple layers of GIN, a global pooling operation is performed on each graph to aggregate node-level features into graph-level features. The pooling operation adopts the global additive pooling method: ; in, is the graph-level feature of the fusion graph or mask graph, is the set of all nodes in the graph, is the node feature; S3.3.2.4, define the features of the positive sample pair as fusion graph features and the mask map features of the positive sample pair , the characteristics of the negative sample pair are and the mask map features of the negative sample pairs , using contrast loss To optimize: ; in, is a logarithmic function, is an exponential function, is the similarity function, is the temperature hyperparameter, is the feature of the fusion map or mask map.
10. The method for predicting drug target affinity activity based on comparative fusion graph features according to claim 1, characterized in that: Step S4 specifically includes the following steps: S4.1, concatenate the three features to generate the final feature representation : ; in, For the connection operation, is the drug graph feature, is the target map feature, To compare the fusion graph features; S4.2, through the first layer of linear transformation Perform feature dimensionality reduction: ; in, and are the weights and biases of the first layer linear transformation, is the activation function, It is the feature after the first layer of dimensionality reduction; S4.3, generate the final prediction value through the second layer of linear transformation: ; in, and are the weights and biases of the second layer linear transformation, is the predicted interaction score; S4.4, using mean square error MSE loss function Update the model: ; in, is the total number of samples, and Respectively represent The predicted and true values of samples.
Citation Information
Patent Citations
Drug-target affinity prediction system based on graph convolutional neural network, computer equipment and storage medium
CN114743590A
Drug target binding affinity prediction method based on three-branch CNN
CN116189795A
TransVAE-based drug-target binding affinity prediction method
CN116825183A
Drug target binding affinity prediction method based on drug bimodal characteristics
CN118298908A
Drug-target affinity prediction method based on multi-scale mixed attention network
CN119649898A
Cited By
Roadway anchor net cable support parameter optimization design method based on neural network
CN120316885A