A method for predicting the affinity activity of drug targets based on the features of contrast fusion graphs

Through the method of combining graph convolution network and bidirectional LSTM with a comparison fusion graph module, the problem of insufficient utilization of data for drug target affinity activity prediction in the prior art is solved, and more accurate drug target interaction relationship prediction is achieved.

CN120089191BActive Publication Date: 2025-07-11QINGDAO UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510572473.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-07-11
Estimated Expiration
2045-05-06

AI Technical Summary

Technical Problem

The prior art is difficult to effectively utilize large-scale and diversified data to capture the deep interactive information relationship between drugs and targets, resulting in inaccurate prediction of drug target affinity activity.

Method used

The graph convolution network is used to extract the SMILES string features of drugs, combine with bidirectional LSTM to extract the target sequence features, and a fusion map between drugs and targets is constructed by comparing the fusion map module. The graph isomorphic network is used for self-supervised learning, optimize the feature representation, and finally predict it through a linear layer.

Benefits of technology

It improves the accuracy and accuracy of predicting affinity activity of drug targets, and enhances the understanding and capture of the interaction between drugs and targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120089191B_ABST
    Figure CN120089191B_ABST
Patent Text Reader

Abstract

The present invention provides a method for predicting the affinity activity of a drug target based on contrastive fusion graph features, which relates to the field of predicting the affinity activity of drug targets, and specifically includes the following steps: using a graph convolutional network to extract features from the SMILES string of a drug to obtain drug features; using a bidirectional LSTM to extract features from the target sequence to obtain target features; designing a contrastive fusion graph module, and based on the contrastive fusion graph module, extracting contrastive fusion graph features; splicing and dimension-reducing the drug features, target features and contrastive fusion graph features to obtain a prediction score. The technical solution of the present invention overcomes the problems in the prior art that large-scale and diverse data cannot be effectively utilized and the deep interaction information relationship between drugs and targets cannot be captured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of predicting the affinity activity of drug targets, and particularly to a method for predicting the affinity activity of drug targets based on contrast-fused graph features. Background Art

[0002] The prediction of drug-target affinity (DTA) usually relies on the chemical structure features of drug molecules and the biological structure features of targets, and combines methods such as deep learning to establish a model to quantify the binding constant between a drug and a target. With the rapid development of deep learning and artificial intelligence technologies, the performance of DTA models has been gradually improved, and more complex data features and deep models can be used for more accurate predictions.

[0003] Early DTA prediction methods mainly used traditional machine learning methods. Considering the limitations of traditional machine learning methods in DTA prediction, especially when dealing with complex and non-linear drug-target relationships, researchers began to explore deep learning methods. Existing deep learning methods are mainly divided into two categories: sequence-based methods and graph-based methods. Existing models still regard targets as one-dimensional sequences, and these models usually process the features of drugs and targets separately. For example, drugs are regarded as molecular graphs, and targets are regarded as one-dimensional sequences or contact graphs, which fail to effectively capture the complex interaction relationships between drugs and targets. Secondly, existing models often rely on limited training data and cannot deeply explore the potential deep characteristics of drugs and targets.

[0004] Therefore, there is a need for a method for predicting the affinity activity of drug targets based on contrast-fused graph features that can effectively utilize large-scale and diverse data and capture the deep interaction information relationship between drugs and targets. Summary of the Invention

[0005] The main object of the present invention is to provide a method for predicting the affinity activity of drug targets based on contrast-fused graph features to solve the problem in the prior art that large-scale and diverse data cannot be effectively utilized and the deep interaction information relationship between drugs and targets cannot be captured.

[0006] To achieve the above object, the present invention provides a method for predicting the affinity activity of drug targets based on contrast-fused graph features, which specifically includes the following steps:

[0007] S1, using a graph convolutional network to extract features from the SMILES string of a drug to obtain drug features.

[0008] S2, using a bidirectional LSTM to extract features from the target sequence to obtain target features.

[0009] S3, designing a contrast-fused graph module, and based on the contrast-fused graph module, extracting contrast-fused graph features.

[0010] S4. Concatenate the drug features, target features, and contrast fusion graph features and reduce the dimension to obtain the prediction score.

[0011] Further, step S1 specifically includes the following steps:

[0012] S1.1. Convert the SMILES string of the drug into a graph representation through the tool RDKit, where the nodes in the graph are drug atoms.

[0013] S1.2. Input the drug atom nodes into a graph convolutional network to hierarchically aggregate the drug atom node features, then fuse all the node features into the drug global feature through a global addition operation, and finally map it into the drug feature through a linear layer.

[0014] Further, step S2 specifically includes the following steps:

[0015] S2.1. Input the amino acid sequence of the target into the forward LSTM to capture the dependence of the current amino acid on the previous sequence, and input the amino acid sequence of the target into the backward LSTM to capture its association with the subsequent sequence.

[0016] S2.2. Add the hidden states of the forward LSTM and the backward LSTM to generate the target features.

[0017] Further, step S3 specifically includes the following steps:

[0018] S3.1. Calculate the Euclidean distance between two amino acids to determine whether there is contact between the two amino acids, and generate the target graph adjacency matrix.

[0019] S3.2. Create a fusion graph of the drug and the target through the central node.

[0020] S3.3. Perform self-supervised learning on the fusion graph based on the graph isomorphism network GIN to extract the contrast fusion graph features.

[0021] Further, step S3.1 specifically includes the following steps:

[0022] S3.1.1. The distance calculation formula between amino acid A and amino acid C is:

[0023] ;

[0024] where is the distance between amino acid A ( ) and amino acid C ( ).

[0025] S3.1.2. If is less than N angstroms, it is considered to have contact, and the corresponding position in the target graph adjacency matrix is recorded as 1; otherwise, it is recorded as 0.

[0026] Furthermore, step S3.2 specifically includes the following steps:

[0027] S3.2.1, unify the dimensions of the drug graph and the target graph through a linear layer:

[0028] ;

[0029] Among them, and are the features of the drug and the target after linear transformation respectively, and the dimensions are both , and are the weight matrices of the drug graph and the target graph respectively, and are the initial feature matrices of the drug and the target respectively, and are the bias vectors of the drug and the target respectively.

[0030] S3.2.2, initialize the central node as a bridge connecting the drug graph and the target graph, and the central node is expressed as: .

[0031] S3.2.3, update the adjacency matrix of the fusion graph :

[0032] ;

[0033] Among them, and are the adjacency matrices of the drug graph and the target graph respectively.

[0034] Furthermore, the central node randomly connects several nodes in the drug and target graphs respectively, and the number of connecting edges is the same. The total number of nodes in the fusion graph is , among which, and represent the number of nodes in the drug graph and the target graph respectively.

[0035] Furthermore, step S3.3 specifically includes the following steps:

[0036] S3.3.1, randomly mask some edges in each fusion graph to obtain a masked graph.

[0037] S3.3.2, for each fusion graph and masked graph, use the graph isomorphism network GIN to extract the node features in the fusion graph and the masked graph.

[0038] Furthermore, step S3.3.2 specifically includes the following steps:

[0039] S3.3.2.1, Node features of the input fusion graph or mask graph Normalize through batch normalization:

[0040] ;

[0041] Among them, is the normalized node feature, is the batch normalization operation.

[0042] S3.3.2.2, The node features of the fusion graph or mask graph are processed through graph convolution:

[0043] ;

[0044] Among them, represents the multi-layer perceptron for aggregating neighbor node features, is the feature of node at the layer, is a learnable parameter, represents the set of neighbor nodes of node .

[0045] S3.3.2.3, After passing through multiple layers of GIN, perform global pooling operation on each graph to aggregate node-level features into graph-level features. The pooling operation adopts the global sum pooling method:

[0046] ;

[0047] Among them, is the graph-level feature of the fusion graph or mask graph, is the set of all nodes in the graph, is the node feature.

[0048] S3.3.2.4, Define the feature of the positive sample pair as the fusion graph feature and the mask graph feature of the positive sample pair, and the feature of the negative sample pair is and the mask graph feature of the negative sample pair. Use the contrastive loss for optimization:

[0049] ;

[0050] Among them, is the logarithmic function, is the exponential function, is the similarity function, is the temperature hyperparameter, is the feature of the fusion graph or mask graph.

[0051] Furthermore, step S4 specifically includes the following steps:

[0052] S4.1, Concatenate the three features to generate the final feature representation :

[0053] ;

[0054] Among them, is the concatenation operation, is the drug graph feature, is the target graph feature, is the contrast fusion graph feature.

[0055] S4.2, Perform feature dimensionality reduction on through the first-layer linear transformation:

[0056] ;

[0057] Among them, and are the weight and bias of the first-layer linear transformation respectively, is the activation function, is the feature after the first-layer dimensionality reduction.

[0058] S4.3, Generate the final predicted value through the second-layer linear transformation:

[0059] ;

[0060] Among them, and are the weight and bias of the second-layer linear transformation respectively, is the predicted interaction score.

[0061] S4.4, Adopt the mean square error MSE loss function to update the model:

[0062] ;

[0063] Among them, is the total number of samples, and respectively represent the predicted value and the true value of the th sample.

[0064] The present invention has the following beneficial effects:

[0065] (1) The present invention uses a graph convolutional network to extract molecular topological features from the SMILES representation of drugs, and extracts context features of target sequences through a bidirectional long short-term memory network.

[0066] (2) The present invention designs a contrast fusion graph module. By constructing a unified graph of a drug molecule graph and a target 2D contact graph, and randomly generating a masked graph, contrast learning is used to optimize the feature representation of the unified graph.

[0067] (3) In the pre-training stage of the present invention, the feature representations of positive samples are pulled closer. By combining the splicing of three features and linear layer mapping, accurate prediction of drug-target activity scores is achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. In the drawings:

[0069] Figure 1 A flowchart of a method for predicting drug-target affinity activity based on contrast fusion graph features of the present invention is shown.

[0070] Figure 2 A contrast fusion graph learning process is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0071] The technical solutions of the present invention will be clearly and completely described below with reference to the drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.

[0072] As Figure 1 shown, a method for predicting drug-target affinity activity based on contrast fusion graph features specifically includes the following steps:

[0073] S1. Use a graph convolutional network to extract features from the SMILES string of the drug to obtain drug features.

[0074] S2. Use a bidirectional LSTM to extract features from the target sequence to obtain target features.

[0075] S3. Design a contrast fusion graph module, and based on the contrast fusion graph module, extract contrast fusion graph features.

[0076] S4. Splice and reduce the dimensions of the drug features, target features, and contrast fusion graph features to obtain a prediction score.

[0077] The present invention first uses a graph convolutional network to extract features from the SMILES string of a drug and analyze the topological structure of the drug; at the same time, a bidirectional LSTM is used to extract features from the target sequence to capture its context dependence. Subsequently, based on the structural information of the drug and the target, two basic graphs are constructed: a drug molecular graph and a target 2D contact graph. On this basis, a contrast fusion graph module is designed, which connects the two basic graphs into a unified graph by introducing a central node and generates a masked graph with randomly masked partial edges to enhance the robustness of the features. The unified graph and the masked graph are respectively input into a graph isomorphism network for information transmission, global feature representations are extracted through average pooling, and contrast learning is used to optimize the positive sample features of the drug and the target. Finally, the extracted drug topological features, target sequence features, and global features of the fusion graph are concatenated and input into a linear layer for low-dimensional mapping, and finally the activity prediction score of the drug-target is obtained.

[0078] Specifically, the SMILES string of a drug is converted into a graph representation that can be processed by a deep learning model through the tool RDKit. Graph feature extraction is a key step in extracting effective information from the constructed molecular graph. Here, a graph convolutional network is used to hierarchically aggregate the topological information of drug atom nodes, and finally all node features are fused into a drug global feature through a global addition operation, and then mapped into the final topological feature of the drug through a linear layer.

[0079] Step S1 specifically includes the following steps:

[0080] S1.1, convert the SMILES string of the drug into a graph representation through the tool RDKit, and the nodes in the graph are drug atoms.

[0081] S1.2, input the drug atom nodes into the graph convolutional network to hierarchically aggregate the drug atom node features, then fuse all node features into a drug global feature through a global addition operation, and finally map it into a drug feature through a linear layer.

[0082] Specifically, the amino acid sequence of a target is a complex linear data, which contains rich context information related before and after. For the amino acid sequence, a bidirectional long short-term memory network LSTM model is used, and two independent LSTMs are used to scan the amino acid sequence from the forward and reverse directions respectively. The forward LSTM captures the dependence of the current amino acid on the previous sequence, and the backward LSTM captures its association with the subsequent sequence. Finally, the hidden states in the two directions are added to generate the global semantic representation of the target sequence.

[0083] Step S2 specifically includes the following steps:

[0084] S2.1. Input the amino acid sequence of the target into the forward LSTM to capture the dependence of the current amino acid on the previous sequence, and input the amino acid sequence of the target into the backward LSTM to capture its association with the subsequent sequence.

[0085] S2.2. Add the hidden states of the forward LSTM and the backward LSTM to generate the target feature.

[0086] Specifically, step S3 specifically includes the following steps:

[0087] S3.1. Calculate the Euclidean distance between two amino acids, determine whether the two amino acids are in contact, and generate the target graph adjacency matrix.

[0088] S3.2. Create a fusion graph of the drug and the target through the central node.

[0089] S3.3. Perform self-supervised learning on the fusion graph based on the graph isomorphism network GIN, and extract the contrastive fusion graph features.

[0090] Specifically, the present invention obtains three-dimensional coordinate data from the target sequence structure database (such as PDBe, PDBBind) to construct a 2D distance map of the target. If it has not been found in the database, high-confidence atomic spatial position data can also be generated through prediction models such as AlphaFold. After obtaining the three-dimensional coordinate information of each amino acid, the contact relationship between amino acids is determined through spatial distance calculation. Step S3.1 specifically includes the following steps:

[0091] S3.1.1. The distance calculation formula between amino acid A and amino acid C is:

[0092] ;

[0093] Where is the distance between amino acid A ( ) and amino acid C ( ).

[0094] S3.1.2. If is less than N angstroms, it is considered to be in contact, and the corresponding position in the target graph adjacency matrix is recorded as 1; otherwise, it is recorded as 0. By calculating all amino acid pairs, a complete adjacency matrix can be generated to describe the contact graph information of the target.

[0095] The present invention creates a fusion graph of the drug and the target to achieve effective interaction of information between the drug and the target. Previous methods usually considered the features of the drug and the target separately, ignoring the interaction between the two. By constructing a fusion graph, the features of the drug and the target can be integrated into a unified graph, thereby better capturing the potential relationship between them.

[0096] Specifically, step S3.2 specifically includes the following steps:

[0097] S3.2.1, Before generating the fusion graph, two basic graphs, the drug molecular graph and the target 2D contact graph, need to be processed. Since the characteristic dimensions of the molecular structure of the drug and the target sequence are different, before fusion, the dimensions of the drug graph and the target graph need to be unified through a linear layer:

[0098] ;

[0099] Among them, and are the features of the drug and the target after linear transformation, respectively, and the dimension is , and are the weight matrices of the drug graph and the target graph, respectively, and are the initial feature matrices of the drug and the target, respectively, and are the bias vectors of the drug and the target, respectively.

[0100] S3.2.2, In order to connect the two basic graphs, a central node is initialized as a bridge to connect the drug graph and the target graph. The central node is represented as: . Since the specific action sites of the drug and the target are unknown, the introduction of the central node can integrate the drug molecular graph and the target graph in a randomly connected manner.

[0101] S3.2.3, While constructing the fusion graph, its adjacency matrix also needs to be updated. Update the adjacency matrix of the fusion graph :

[0102] ;

[0103] Among them, and are the adjacency matrices of the drug graph and the target graph, respectively.

[0104] Specifically, the central node randomly connects several nodes in the drug and target graphs respectively, and the number of connecting edges is the same. The total number of nodes in the fusion graph is , where and represent the number of nodes in the drug graph and the target graph, respectively.

[0105] The present invention designs a pre-training process to capture the potential relationship between drugs and targets through self-supervised learning of a large number of fusion graphs. As Figure 2As shown, specifically, 100 drugs and 100 targets were collected from a public database as a pre-training dataset. Using the above method, each drug was combined with each target one by one to generate 10,000 fusion graphs, providing a data basis for the pre-training of the model.

[0106] After generating the fusion graphs, to further simulate the uncertainty and missing information in the graph structure, a certain proportion of edges in each fusion graph were randomly masked to obtain 10,000 corresponding masked graphs. These masked graphs not only provide a simplified graph structure but also enhance the model's ability to focus on key features in the graph through contrastive learning. Specifically, the positive sample pairs are composed of the same fusion graph and its corresponding masked graph, while the negative sample pairs are composed of combinations of different fusion graphs.

[0107] Specifically, step S3.3 specifically includes the following steps:

[0108] S3.3.1, Randomly mask some edges in each fusion graph to obtain a masked graph.

[0109] S3.3.2, For each fusion graph and masked graph, use the graph isomorphism network GIN to extract the node features in the fusion graph and masked graph.

[0110] Specifically, for each fusion graph and masked graph, the graph isomorphism network (GIN) is used to extract the node features in the fusion graph and masked graph. The key of GIN is that it uses a multi-layer perceptron (MLP) to aggregate the features of each node and its neighbors, thus efficiently capturing the structural information of the graph. Step S3.3.2 specifically includes the following steps:

[0111] S3.3.2.1, The node features of the input fusion graph or masked graph Are standardized through batch normalization:

[0112] ;

[0113] Where Is the standardized node feature, Is the batch normalization operation.

[0114] S3.3.2.2, The node features of the fusion graph or masked graph are processed through graph convolution:

[0115] ;

[0116] Where Represents the multi-layer perceptron used to aggregate the features of neighbor nodes, Is the node At the Layer feature, Is a learnable parameter, Representation Node Through multiple layers of iteration, the representation of the node will continuously aggregate information from the neighborhood, thereby effectively capturing the structural characteristics of the graph.

[0117] S3.3.2.3, after passing through multiple layers of GIN, a global pooling operation is performed on each graph to aggregate node-level features into graph-level features. The pooling operation adopts the global additive pooling method:

[0118] ;

[0119] in, is the graph-level feature of the fusion graph or mask graph, is the set of all nodes in the graph, is the node feature; finally, the graph-level feature is a 1×128 vector representing the global characteristics of the entire graph.

[0120] S3.3.2.4, in order to use these features for contrastive learning, the loss function of positive and negative samples is defined. The specific goal is to shorten the feature distance of positive sample pairs and increase the feature distance of negative sample pairs. The features of positive sample pairs are defined as fusion graph features and the mask map features of the positive sample pair , the characteristics of the negative sample pair are and the mask map features of the negative sample pairs , using contrast loss To optimize:

[0121] ;

[0122] in, is a logarithmic function, is an exponential function, is the similarity function, is the temperature hyperparameter, which is used to control the smoothness of the similarity metric. is the feature of the fusion map or mask map.

[0123] Through the self-supervised pre-training method, the model can learn a more robust graph feature representation, laying a solid foundation for the subsequent prediction of drug-target interaction. During the pre-training stage, GIN is optimized to efficiently capture the node and structural features in the fusion graph, ensuring that the model has strong feature extraction capabilities.

[0124] like Figure 2As shown, in the formal training stage, the pre-trained and optimized GIN model is embedded as the first layer into the main model. The node features of the fusion graph are first processed through this GIN layer optimized by contrastive learning to extract the initial feature representation. Next, the fusion graph continues to undergo recursive calculations through several layers of GIN. Each layer aggregates the information of the nodes and their neighbors, gradually generating higher-order feature representations. Specifically, after several layers of GIN, the node features of the fusion graph are integrated into graph-level features through the global pooling operation Pool. This method, through the seamless connection between pre-training and formal training, can not only effectively utilize prior knowledge to enhance the performance of the model in the initial stage, but also further explore the complex relationships between drugs and targets through the feature extraction process of multiple layers of GIN.

[0125] Specifically, step S4 specifically includes the following steps:

[0126] S4.1, To achieve the final prediction of drug-target interaction, three types of features are concatenated to generate the final feature representation :

[0127] ;

[0128] Among them, is the concatenation operation, is the drug graph feature, is the target graph feature, is the contrastive fusion graph feature. The concatenated feature combines the information of drug molecules, target sequences, and their interaction relationships, providing a comprehensive and in-depth representation for the prediction task.

[0129] S4.2, The model further processes the concatenated features through two linear layers to complete the prediction. Through the first linear transformation, the feature dimensionality of is reduced:

[0130] ;

[0131] Among them, and are respectively the weights and biases of the first linear transformation, is the activation function, is the feature after the first dimensionality reduction.

[0132] S4.3, The final prediction value is generated through the second linear transformation:

[0133] ;

[0134] Among them, and They are the weights and biases of the second - layer linear transformation respectively, which is the predicted interaction score.

[0135] S4.4. The mean - square error (MSE) loss function is adopted to update the model:

[0136] ;

[0137] wherein, is the total number of samples, and respectively represent the predicted value and the true value of the th sample.

[0138] The model provided by the present invention includes a drug feature extraction module, a target feature extraction module, a contrast - fusion graph module, and a prediction module. After the drug feature extraction module extracts drug features, the target feature extraction module extracts target features, and the contrast - fusion graph module extracts contrast - fusion features, the drug features, target features, and contrast - fusion graph features are concatenated and dimension - reduced, and the prediction score is obtained through the prediction module.

[0139] To evaluate the performance of the model Tri - DTA proposed by the present invention, performance comparisons of DTA prediction were made with the current state - of - the - art models on the Davis and KIBA datasets. The models involved are DeepDTA, MT - DTA, GraphDTA, AttentionDTA, and GRA - DTA. On the Davis dataset, the model proposed by the present invention shows advantages over the baseline models in all evaluation metrics. As shown in Table 1, especially in terms of the mean - square error loss (MSE), the model Tri - DTA provided by the present invention achieved an improvement of 0.004, which indicates that this method has improved in prediction accuracy. At the same time, it is also superior to the baseline model in terms of CI and . CI is the coefficient of consistency, which is used to measure whether the predicted affinity values of two randomly selected drug - target pairs are consistent with the order of their actual values. is used to evaluate the fitting degree of the model.

[0140] Table 1 Performance comparison with other baseline models in Davis

[0141]

[0142] The results on the KIBA dataset also show that the proposed method performs excellently, as shown in Table 2. Compared with the baseline models, the MSE of the model proposed by the present invention is reduced by 0.005, specifically verifying its advantage in accuracy. Although it is comparable to the best - performing baseline model in terms of the CI metric, in It is slightly inferior to GRA-DTA in terms of aspects. Generally speaking, it is particularly prominent in the improvement of MSE, showing its potential in optimizing the prediction accuracy of DTA.

[0143] Table 2 Performance comparison with other baseline models in KIBA

[0144]

[0145] In the present invention, the contrast fusion graph module can effectively enhance the information interaction between drugs and targets, capture more abundant features, thereby improving the understanding of the drug-target interaction by the model of the present invention; secondly, through self-supervised pre-training, the model can learn the potential relationship between drugs and targets without artificial labels, which further improves the accuracy of DTA prediction. Self-supervised learning can not only enhance the expression ability of features, but also provide effective training signals when data is scarce, so as to achieve more efficient feature extraction and information integration.

[0146] Of course, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Changes, modifications, additions or substitutions made by those skilled in the art within the substantial scope of the present invention should also fall within the protection scope of the present invention.

Claims

1. A method for predicting the affinity activity of a drug target based on the features of a contrast fusion graph, characterized in that, Specifically, it includes the following steps: S1. Extract features of the drug from its SMILES string using a graph convolutional network to obtain drug features; S2. Extract features of the target sequence using a bidirectional LSTM to obtain target features; S3. Design a contrastive fusion graph module, and based on this module, extract contrastive fusion graph features; S4. Concatenate and reduce the dimensions of the drug features, target features, and contrastive fusion graph features to obtain a prediction score; Step S3 specifically includes the following steps: S3.

1. Calculate the Euclidean distance between two amino acids to determine whether they are in contact, and generate an adjacency matrix of the target graph; S3.

2. Create a fusion graph of the drug and the target through a central node; S3.

3. Perform self-supervised learning on the fusion graph based on the graph isomorphism network GIN to extract contrastive fusion graph features; Step S3.2 specifically includes the following steps: S3.2.

1. Unify the dimensions of the drug graph and the target graph through a linear layer: ; Among them, and are the features of the drug and the target after linear transformation, respectively, and the dimensions are both , and are the weight matrices of the drug graph and the target graph, respectively, and are the initial feature matrices of the drug and the target, respectively, and are the bias vectors of the drug and the target, respectively; S3.2.2, Initialize the central node, which serves as a bridge connecting the drug graph and the target graph. The central node is represented as: ; S3.2.3, Update the adjacency matrix of the fusion graph : ; Among them, and are the adjacency matrices of the drug graph and the target graph, respectively; The central node randomly connects to several nodes in the drug graph and the target graph respectively, and the number of connected edges is the same. The total number of nodes in the fusion graph is , where and represent the number of nodes in the drug graph and the target graph respectively; Step S3.3 specifically includes the following steps: S3.3.

1. Randomly mask some edges in each fusion graph to obtain a masked graph; S3.3.

2. For each fusion graph and masked graph, use the graph isomorphism network GIN to extract the node features in the fusion graph and the masked graph; Define the features of the positive sample pair as the fusion graph features and the masked graph features of the positive sample pair, and the features of the negative sample pair as the fusion graph features and the masked graph features of the negative sample pair, and optimize using a contrastive loss.

2. The drug target affinity activity prediction method based on the characteristics of the contrast fusion graph according to claim 1, wherein Step S1 specifically includes the following steps: S1.

1. Convert the SMILES string of the drug into a graph representation through the tool RDKit, where the nodes in the graph are drug atoms; S1.

2. Input the drug atom nodes into the graph convolutional network to hierarchically aggregate the drug atom node features, then fuse all node features into drug global features through a global addition operation, and finally map them into drug features through a linear layer.

3. The drug target affinity activity prediction method based on the contrast fusion graph features according to claim 1, characterized in that, Step S2 specifically includes the following steps: S2.

1. Input the amino acid sequence of the target into the forward LSTM to capture the dependence of the current amino acid on the previous sequence, and input the amino acid sequence of the target into the backward LSTM to capture its association with the subsequent sequence; S2.

2. Add the hidden states of the forward LSTM and the backward LSTM to generate target features.

4. The method for predicting the affinity activity of a drug target according to claim 1 based on the features of a contrast fusion graph, wherein, Step S3.1 specifically includes the following steps: S3.1.

1. The distance calculation formula between amino acid A and amino acid C is: ; Among them, is the distance between amino acid A ( ) and amino acid C ( ); S3.1.2, if is less than N angstroms, it is regarded as having contact, and the corresponding position in the target point map adjacency matrix is recorded as 1; otherwise, it is recorded as 0.

5. A method for predicting the affinity activity of a drug target based on the features of a contrast fusion graph according to claim 1, wherein Step S3.3.2 specifically includes the following steps: S3.3.2.1, Node features of the input fusion graph or masking graph Normalize through batch normalization: ; Among them, is the standardized node feature, is the batch normalization operation; S3.3.2.

2. Process the node features of the fusion graph or the masked graph through graph convolution: ; Among them, represents a multi-layer perceptron for aggregating neighbor node features, is the node at the layer feature, is a learnable parameter, represents the set of neighbor nodes of the node ; S3.3.2.

3. After passing through multiple layers of GIN, perform a global pooling operation on each graph to aggregate the node-level features into graph-level features, and the pooling operation adopts the global addition pooling method: ; Among them, is the graph-level feature of the fusion graph or the masking graph, is the set of all nodes in the graph, is the node feature; S3.3.2.4, define the features of the positive sample pair as the fused graph features and the masked graph features of the positive sample pair , the features of the negative sample pair are and the masked graph features of the negative sample pair , and use contrastive loss for optimization: ; Among them, is a logarithmic function, is an exponential function, is a similarity function, is a temperature hyperparameter, is the feature of the fusion graph or the masking graph.

6. The method for predicting the affinity activity of a drug target according to claim 1, characterized in that Step S4 specifically includes the following steps: S4.1, Generate the final feature representation by splicing the three features : ; Among them, is the connection operation, is the drug graph feature, is the target graph feature, is the comparison and fusion graph feature; S4.2, perform feature dimensionality reduction on through the first-layer linear transformation: ; Among them, and are the weights and biases of the first-layer linear transformation respectively, is the activation function, is the feature after the first-layer dimensionality reduction; S4.

3. Generate the final prediction value through a second-layer linear transformation: ; Among them, and are the weights and biases of the second-layer linear transformation respectively, is the predicted interaction score; S4.4, adopt the mean squared error (MSE) loss function Update the model: ; Among them, is the total number of samples, and respectively represent the predicted value and the true value of the th sample.

Citation Information

Patent Citations

  • Drug-target affinity prediction system based on graph convolutional neural network, computer equipment and storage medium

    CN114743590A

  • Drug target binding affinity prediction method based on three-branch CNN

    CN116189795A