Drug-target affinity prediction method based on molecular map and target two-dimensional information

Through a prediction method based on molecular maps and two-dimensional information of the target, GAT and GCN models are used to extract features of drugs and targets, which solves the problems of low prediction accuracy and long running time in the prior art, and realizes the prediction of high-efficiency drug-target binding affinity for drug development.

CN120279982APending Publication Date: 2025-07-08TAIYUAN UNIVERSITY OF TECHNOLOGY
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510406357.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The drug-target binding affinity prediction method based on deep learning in the prior art has insufficient feature extraction capabilities, resulting in low prediction accuracy and long running time, making it difficult to apply to actual drug development.

Method used

Using a prediction method based on molecular graphs and target two-dimensional information, the graph attention model GAT and graph convolution model GCN are used to extract the two-dimensional information of drugs and targets, and combined with pre-training and data enhancement technology, a prediction module is built to predict drug-target binding affinity.

Benefits of technology

It improves the accuracy and operation efficiency of drug-target binding affinity prediction, can speed up the drug development process and reduce the screening time for drug side effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279982A_ABST
    Figure CN120279982A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence assisted drug design, and particularly relates to a drug-target affinity prediction method based on a molecular map and target two-dimensional information, which comprises the following steps: generating two-dimensional information: respectively processing an input drug sequence and a target sequence to obtain two-dimensional information representation; model construction: constructing a graph attention model GAT and a graph convolution model GCN which are respectively used for processing the two-dimensional information representation of the drug and the target so as to obtain feature vectors; and predicting the binding affinity: predicting the drug-target binding affinity by using a prediction module according to the obtained feature vector, and evaluating a prediction result by using a plurality of indexes. According to the method, the operation efficiency can be improved, the prediction accuracy can be improved, in actual drug development, the prediction speed of the drug-target binding affinity can be increased by using the method, and the method plays a great role in screening drugs, reducing side effects of the drugs and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence-assisted drug design, and particularly relates to a method for predicting drug-target affinity based on molecular graphs and two-dimensional information of targets. Background Art

[0002] Drug-target binding affinity (DTA) refers to the strength of the binding between a drug and its target, and is a key indicator for evaluating the ability of a drug to bind to a target. The strength of the affinity between a drug and a target can usually be expressed by dissociation constant (K d ), association constant (K a ), inhibition constant (K i ), half maximal inhibitory concentration (IC 50 ), half maximal effective concentration (EC 50 ) and other indicators. These indicators describe the strength and stability of the binding between a drug and a target, and can be obtained through experimental determination or computational prediction. The level of affinity directly determines the activity and selectivity of a drug. A drug must have a sufficiently high affinity to bind to a target and exert a therapeutic effect. Understanding drug-target binding affinity also helps to optimize the design and synthesis of drug molecules. By adjusting the structure of drug molecules, the affinity between them and the target can be changed, thereby improving the efficacy of the drug and reducing side effects. Since there are a large number of molecules and targets in the real world, it is unrealistic to screen them one by one, which requires a very high cost. Therefore, using deep learning to learn the characteristics of targets and molecules and achieve DTA prediction has gradually become the mainstream.

[0003] Deep learning (DL) and machine learning (ML) have achieved great success in natural language processing and medical imaging, etc. Similarly, they also play a huge role in the field of bioinformatics. So far, a large number of methods have been applied to the prediction of DTA. They can analyze existing datasets, learn patterns and rules from them, and quickly predict new DTA information based on these learning results. They can also discover new drug-target association relationships that cannot be predicted by traditional methods, thereby greatly improving the prediction efficiency, helping to accelerate the drug R & D process, reduce costs and risks, and providing a new way for exploring the binding affinity of unknown drug targets in the future.

[0004] When using deep learning methods to predict DTA, in most current feature extraction methods, mainly two-dimensional information of drugs and one-dimensional fragment information of targets are used. This one-dimensional fragment is difficult to accurately express this regression information, which will lead to a low accuracy of the prediction model; while the three-dimensional data is limited, and its acquisition will greatly extend the running time of the model, resulting in poor model efficiency and making it difficult to be applied in practice. In addition, the feature extraction ability of existing models also needs to be improved. Therefore, it is necessary to provide a new method to predict DTA. Summary of the Invention

[0005] In view of the technical problems of low efficiency and low accuracy of the prediction model in the above-mentioned existing methods, the present invention provides a drug-target affinity prediction method based on molecular graphs and two-dimensional information of targets, so as to accurately predict the binding affinity between drugs and targets.

[0006] In order to solve the above technical problems, the technical solution adopted by the present invention is as follows:

[0007] A drug-target affinity prediction method based on molecular graphs and two-dimensional information of targets, comprising the following steps:

[0008] S1. Generate two-dimensional information: Use the input drug sequence and target sequence to process and obtain their two-dimensional information representations respectively;

[0009] S2. Model construction: Construct a graph attention model GAT and a graph convolutional model GCN, which are respectively used to process the two-dimensional information representations of drugs and targets to obtain feature vectors;

[0010] S3. Predict the binding affinity: Use the prediction module to predict the drug-target binding affinity based on the above-obtained feature vectors, and use multiple indicators to evaluate the prediction results.

[0011] The method for generating two-dimensional information in S1 is as follows: For the SMILES sequence of a drug, convert the one-dimensional sequence data into two-dimensional molecular graph data; for the amino acid sequence of a target, obtain its PSSM matrix, amino acid property feature vector and contact map information respectively, and finally obtain its two-dimensional graph representation.

[0012] The method for obtaining feature vectors in S2 is as follows: First, use 20,000 SMILES data obtained from the ZINC 2015 dataset for pre-training, and then use the trained graph attention model GAT as the model for subsequent drug processing; then input the two-dimensional representations of the drugs and targets obtained in S1 into the graph attention model GAT and the graph convolutional model GCN respectively for feature learning work, and finally obtain the feature vectors of the drugs and targets.

[0013] The method for constructing the Graph Attention Network (GAT) in S2 is as follows: Graph Attention Network (GAT): Use Contrastive Learning of molecular graphs to pre-train the GAT model for processing SMILES data. Convert the SMILES data into molecular graphs, where the nodes in the molecular graph represent atoms and the edges represent chemical bonds. Then, perform data augmentation on the molecular graph. The data augmentation method is to randomly replace some edges to generate positive sample pairs, and generate negative sample pairs between different SMILES strings, thereby pre-training the GAT model. The GAT model contains three GAT layers, and the output dimension of each layer is set to 80 dimensions. Then, connect the output vectors of the three layers to obtain the final drug vector. After that, connect a fully connected (FC) layer to convert the drug vector into 128 dimensions.

[0014] The method for constructing the Graph Convolutional Network (GCN) in S2 is as follows: The Graph Convolutional Network (GCN) consists of three GCN network layers, a global pooling layer, and two FC layers. Each GCN network layer contains a GCN Layer and a ReLU activation function layer. The first layer converts the original feature dimension into 128 dimensions, the second layer converts 128 dimensions into 256 dimensions, and the third layer converts 256 dimensions into 512 dimensions. After that, connect a global max pooling layer and two FC layers. The FC layers respectively convert 512 dimensions into 1024 dimensions and 1024 dimensions into 128 dimensions. Therefore, the final target feature vector is also 128 dimensions.

[0015] The method for predicting binding affinity in S3 is as follows: The prediction module consists of four FC layers, and each of the first three layers is followed by a ReLU activation function and a dropout layer. When the drug is connected to the target, it forms a 256-dimensional feature vector. The first FC layer converts 256 dimensions into 1024 dimensions, the second layer converts 1024 dimensions into 512 dimensions, the third layer converts 512 dimensions into 64 dimensions, and the last layer converts 64 dimensions into 1 dimension, which is the finally predicted affinity score.

[0016] The beneficial effects of the present invention compared with the prior art are:

[0017] The present invention uses the two-dimensional information of drugs and targets, and uses the designed GCN module and GAT module to extract information from the target and the drug respectively. At the same time, in order to reduce the running time of the model and enhance the feature extraction ability of the model, 20,000 drug data are used to pre-train the GAT model. Finally, it is sent to the prediction module to realize the prediction of DTA. The present invention can improve the running efficiency and prediction accuracy. In actual drug development, by using the present invention, the prediction speed of drug-target binding affinity can be accelerated, which has great effects in aspects such as drug screening and reducing drug side effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only exemplary. For those of ordinary skill in the art, without creative efforts, other implementation drawings can also be obtained based on the provided drawings.

[0019] The structures, ratios, sizes, etc. illustrated in this specification are only used to cooperate with the content disclosed in the specification for those familiar with this technology to understand and read, and are not used to limit the limiting conditions for the implementation of the present invention. Therefore, they do not have substantial technical significance. Any modification of the structure, change in the proportional relationship, or adjustment of the size, without affecting the effects that the present invention can produce and the purposes that can be achieved, should still fall within the scope covered by the technical content disclosed in the present invention.

[0020] Figure 1 It is a schematic flowchart of the prediction method of the present invention;

[0021] Figure 2 It is a schematic diagram of the model structure of the present invention;

[0022] Figure 3 It is a schematic flowchart of the target processing process in the present invention. Specific Embodiments

[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. These descriptions are only to further illustrate the features and advantages of the present invention, rather than a limitation on the claims of the present invention; based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by the present application.

[0024] The following will further describe in detail the specific embodiments of the present invention in conjunction with the drawings and embodiments. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.

[0025] This embodiment provides a drug-target affinity prediction method based on molecular graphs and target two-dimensional information, as Figure 1 shown, including the following steps:

[0026] Step 1: Generate two-dimensional information: Use the input drug sequence and target sequence to process and obtain their two-dimensional information representations respectively.

[0027] In Step 1, the process of generating two-dimensional information includes: as Figure 3As shown, for the SMILES sequence of a drug, the one-dimensional sequence data is converted into two-dimensional molecular graph data; for the amino acid sequence of a target, its PSSM matrix, amino acid property feature vector, and contact map information are obtained respectively, and finally its two-dimensional graph representation is obtained.

[0028] Step 2: Model construction: Construct a graph attention model GAT and a graph convolutional model GCN, which are used to process the two-dimensional information representations of drugs and targets respectively to obtain feature vectors.

[0029] In Step 2, to reduce the running time of the model, as Figure 2 shown, 20,000 SMILES data obtained from the ZINC 2015 dataset are used for pre-training in advance, and then the trained graph attention model GAT is used as the model for subsequent drug processing; then the two-dimensional representations of the drugs and targets obtained in Step 1 are input into the graph attention model GAT and the graph convolutional model GCN respectively for feature learning, and finally the feature vectors of the drugs and targets are obtained. Step 2 includes the following Steps 2-1 and 2-2:

[0030] Step 2-1: Graph attention model GAT: Use molecular graph contrastive learning to pre-train the GAT model to process SMILES data. Specifically, convert the SMILES data into a molecular graph (nodes represent atoms, edges represent chemical bonds), then perform data augmentation on the molecular graph (randomly replace some edges) to generate positive sample pairs, and generate negative sample pairs between different SMILES strings. The GAT model is pre-trained through the above method; the GAT model contains three GAT layers, the output dimension of each layer is set to 80 dimensions, and then the output vectors of the three layers are connected to obtain the final drug vector. Then connect an FC layer to convert the drug vector into 128 dimensions.

[0031] Step 2-2: Graph convolutional model GCN: The GCN model consists of three GCN network layers, a global pooling layer, and two FC layers. Each GCN network layer contains a GCN Layer and a ReLU activation function layer. The first layer converts the original feature dimension into 128 dimensions, the second layer converts 128 dimensions into 256 dimensions, the third layer converts 256 dimensions into 512 dimensions, and then connect a global max pooling layer and two FC layers. The FC layers convert 512 dimensions into 1024 dimensions and 1024 dimensions into 128 dimensions respectively. Therefore, the final target feature vector is also 128 dimensions.

[0032] Step 3: Predict the binding affinity: Use the prediction module to predict the drug-target binding affinity based on the above-obtained feature vectors, and use multiple metrics to evaluate the prediction results.

[0033] In step 3, the prediction module consists of four layers of FC. Each of the first three layers is followed by a ReLU activation function and a dropout layer. The connection between the drug and the target is a 256-dimensional feature vector. The first layer of FC converts 256 dimensions into 1024 dimensions, the second layer converts 1024 dimensions into 512 dimensions, the third layer converts 512 dimensions into 64 dimensions, and the last layer converts 64 dimensions into 1 dimension, which is the finally predicted affinity score.

[0034] The prediction method of this embodiment uses the mean square error MSE, the concordance index CI, and the adjusted coefficient of determination for evaluation.

[0035] The loss function used in this embodiment is MSE, and the formula is:

[0036]

[0037] The method is implemented using Pytorch and trained using the Adam optimizer. After adjusting the parameters through multiple experiments, the learning rate is set to 0.001, the batch-size is 256, and each experiment runs for 200 epochs in this embodiment.

[0038] The MSE values of this embodiment on two common public datasets reach 0.201 and 0.121 respectively, proving that the drug-target binding affinity prediction method based on the drug molecular graph and the two-dimensional information of the target proposed in this embodiment has good prediction ability.

[0039] Only the preferred embodiments of the present invention are described in detail above. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those of ordinary skill in the art, various changes can be made without departing from the spirit of the present invention, and all such changes should be included within the protection scope of the present invention.

Claims

1. A method for predicting drug-target affinity based on molecular graphs and two-dimensional target information, characterized in that It includes the following steps: S1. Generate two-dimensional information: Use the input drug sequence and target sequence to process and obtain their two-dimensional information representations respectively; S2. Model construction: Construct a graph attention model GAT and a graph convolutional model GCN, which are respectively used to process the two-dimensional information representations of drugs and targets to obtain feature vectors; S3. Predict the binding affinity: Use the prediction module to predict the drug-target binding affinity based on the obtained feature vectors above, and use multiple metrics to evaluate the prediction results.

2. The drug-target affinity prediction method based on molecular graphs and two-dimensional target information according to claim 1, wherein The method for generating two-dimensional information in S1 is as follows: For the SMILES sequence of the drug, convert the one-dimensional sequence data into two-dimensional molecular graph data; for the amino acid sequence of the target, obtain its PSSM matrix, amino acid property feature vector and contact map information respectively, and finally obtain its two-dimensional graph representation.

3. The drug-target affinity prediction method based on molecular graph and two-dimensional information of the target according to claim 1, wherein The method for obtaining feature vectors in S2 is as follows: First, use 20,000 SMILES data obtained from the ZINC 2015 dataset for pre-training, and then use the trained graph attention model GAT as the model for subsequent drug processing; then input the two-dimensional representations of the drugs and targets obtained in S1 into the graph attention model GAT and the graph convolutional model GCN respectively for feature learning, and finally obtain the feature vectors of the drugs and targets.

4. The method for predicting drug-target affinity based on molecular graphs and two-dimensional target information according to claim 3, characterized in that The method for constructing the graph attention model GAT in S2 is as follows: Graph attention model GAT: Use contrastive learning of molecular graphs to pre-train the GAT model to process SMILES data, convert the SMILES data into molecular graphs, where the nodes in the molecular graph represent atoms and the edges represent chemical bonds, and then perform data augmentation on the molecular graph. The data augmentation method is to randomly replace some edges to generate positive sample pairs, and generate negative sample pairs between different SMILES strings, so as to pre-train the GAT model; the GAT model contains three GAT layers, and the output dimension of each layer is set to 80 dimensions, and then the output vectors of the three layers are connected to obtain the final drug vector, and then connect an FC layer to convert the drug vector into 128 dimensions.

5. The method for predicting drug-target affinity based on molecular graphs and two-dimensional target information according to claim 3, wherein The method for constructing the graph convolutional model GCN in S2 is as follows: The graph convolutional model GCN consists of three GCN network layers, a global pooling layer and two FC layers. Each GCN network layer contains a GCN Layer and a ReLU activation function layer; the first layer converts the original feature dimension into 128 dimensions, the second layer converts 128 dimensions into 256 dimensions, the third layer converts 256 dimensions into 512 dimensions, and then connects a global max pooling layer and two FC layers. The FC layers respectively convert 512 dimensions into 1024 dimensions and 1024 dimensions into 128 dimensions, so the final target feature vector is also 128 dimensions.

6. The drug-target affinity prediction method based on molecular graphs and two-dimensional target information according to claim 1, wherein The method for predicting the binding affinity in S3 is as follows: The prediction module consists of four layers of FC. After each of the first three layers, there is a ReLU activation function and a dropout layer; the connection between the drug and the target is a 256-dimensional feature vector. The first layer of FC converts the 256 dimensions into 1024 dimensions, the second layer converts the 1024 dimensions into 512 dimensions, the third layer converts the 512 dimensions into 64 dimensions, and the last layer converts the 64 dimensions into 1 dimension, which is the finally predicted affinity score.

Citation Information

Patent Citations

  • Protein residue contact map prediction method

    CN113257357A

  • Drug target binding affinity prediction method and system

    CN117594116A

  • Drug-target interaction prediction method based on bidirectional Intention network

    CN117746990A

  • Method and system for predicting capacity of small open reading window coding polypeptide in non-coding RNA (Ribonucleic Acid)

    CN118038995A

  • Drug-target affinity prediction method based on multi-scale convolutional neural network

    CN119580820A