A method for predicting the binding affinity between a drug molecule and a target protein

Through the combination of deep neural network models and graph neural networks, the problem of unavailable 3D structures in the prior art is solved, and efficient and accurate affinity prediction is achieved.

CN114783514BActive Publication Date: 2025-06-10CHONGQING FEINKE BIOTECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210538547.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-18
Publication Date
2025-06-10
Estimated Expiration
2042-05-18

AI Technical Summary

Technical Problem

Prior art methods for predicting binding affinity of drug molecules to target proteins require the use of unavailable protein 3D structures, and calculation methods are time-consuming and resource-intensive.

Method used

The deep neural network model is used to combine the protein network module, the ligand network module, the convolution pooling layer and the fully connected layer. Through data integration, encoding and the use of graph neural networks, the binding affinity of drug molecules and target proteins is predicted.

Benefits of technology

It improves the accuracy of prediction of binding affinity between drug molecules and target proteins, reduces drug screening time, and does not need to rely on protein 3D structure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114783514B_ABST
    Figure CN114783514B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for predicting the binding affinity between a drug molecule and a target protein. A novel protein characterization method is used to integrate a large amount of sequence data of proteins and ligands, improve the degree of feature aggregation, narrow the search space of the neural network, and adopt a novel neural network architecture. At the same time, methods such as recurrent neural network and joint attention mechanism are used to further process the highly aggregated data, and the implicit mapping relationship in the data is learned through continuous iterative learning, so as to achieve the prediction of binding affinity. The present invention can help drug discovery personnel such as medicinal chemistry to quickly screen ligand molecules for their own targets, obtain potential active compounds, and accelerate drug discovery.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of drug research and development, and specifically relates to a method for predicting the binding affinity between a drug molecule and a target protein. Background Art

[0002] The interaction between proteins and small ligand molecules is at the core of many fundamental biological processes. Endogenous small molecules act as messengers in many signaling pathways of cells, and exogenous small molecule drugs regulate their functions by interacting with target proteins through signal cascades. Understanding the interaction between receptor proteins and ligands also helps to understand cell - cell communication. Cell - cell communication can coordinate biological development, maintain homeostasis in the body, and single - cell functions. When cells cannot interact correctly with each other or cells misdecode molecular information, diseases will occur. Understanding the protein - ligand interaction is of great significance for understanding many biological systems and assisting drug development efforts.

[0003] Therefore, predicting the binding affinity between proteins and ligands plays a crucial role in drug discovery and development. However, experimentally determining the protein - ligand binding affinity is very time - consuming and resource - intensive. The research on computational methods for protein - ligand intermolecular binding affinity can be mainly divided into four categories, including computational methods based on ligand similarity, molecular dynamics simulation methods for calculating binding free energy, traditional scoring functions in molecular docking, and computational methods based on machine learning algorithms.

[0004] As a new algorithm of artificial intelligence, deep learning has achieved success in many fields. By analogy with the application of these algorithms in image recognition, applying neural networks in the recognition of protein - ligand intermolecular interactions is a pattern recognition of biological problems. Currently, many computational methods have been proposed to predict binding affinity, but most of them usually require the use of protein 3D structures that are usually not available. Summary of the Invention

[0005] In view of the deficiencies described in the above - mentioned prior art, the present invention provides a method for predicting the binding affinity between a drug molecule and a target protein.

[0006] The technical solution adopted by the present invention is as follows:

[0007] A method for predicting the binding affinity between a drug molecule and a target protein, the steps are as follows:

[0008] Data integration;

[0009] The integrated data includes the sequence of the target protein and the SMILES of the ligand molecule;

[0010] Integrate the sequence of the target protein and the SMILES of the ligand molecule into an integrated dataset;

[0011] Data Encoding:

[0012] Encode the target protein sequences and the SMILES of ligand molecules in the integrated dataset respectively to obtain an encoded dataset;

[0013] Specifically, the encoding of the target protein is as follows:

[0014] Each target protein is represented by a character dataset of a set length; the character dataset includes a special character representing the start, a four-letter group of the letters characterizing the target protein sequence, and a special character representing the end; and when the length of the target protein sequence is less than the set length, a special letter representing padding is used for padding and placeholder;

[0015] The four-letter group of letters characterizing the protein features includes secondary structure class, whether it is exposed in the solvent, physicochemical properties, and length;

[0016] The encoding rule of the four-letter group of letters is:

[0017]

[0018] Ligand Molecule Encoding:

[0019] It is represented by a ligand dataset of a set length, and the ligand dataset includes a special character representing the start, the SMILES of the ligand molecule, and a special character representing the end; when the length of the SMILES of the ligand molecule is less than the set length, a special letter representing padding is used for padding and placeholder; when the length of the SMILES of the ligand molecule is greater than the set length, it is directly truncated.

[0020] Affinity Prediction:

[0021] Input the encoded dataset into the affinity prediction model batch by batch to obtain the affinity prediction result.

[0022] The affinity prediction model is a deep neural network model, including a protein network module, a ligand network module, a convolutional pooling layer, and a fully connected layer; the outputs of the protein network module and the ligand network module are aggregated in the convolutional pooling layer and then input into the fully connected layer, and the fully connected layer outputs the prediction result;

[0023] The protein network module includes a protein embedding layer, a protein RNN layer, and a protein attention layer.

[0024] The protein embedding layer converts the input protein data into a fixed-length vector. The protein RNN layer uses a gated recurrent unit (GRU) to capture the non-linear joint dependencies between protein residues that are far apart from each other sequentially; the protein attention layer uses a joint attention model, and the attention model is jointly trained for proteins and the RNN / CNN part. Their learning parameters include the attention weights of all letters of a given string. Compared with unsupervised learning, each attention model here outputs a vector as the input to its corresponding subsequent CNN. The joint attention model not only improves the prediction performance but also enables the model to be interpreted at the level of "letters" (SSEs in proteins) and their pairs.

[0025] The ligand network module includes a ligand embedding layer, a ligand RNN layer, and a ligand attention layer. The ligand embedding layer converts the input ligand molecule data into a fixed-length vector. The ligand RNN layer uses a gated recurrent unit (GRU) to capture the non-linear joint dependencies between ligand atoms that are far apart from each other sequentially; the ligand attention layer uses a joint attention model, and the attention model is jointly trained for ligand atoms and the RNN / CNN part. Their learning parameters include the attention weights of all letters of a given string. Compared with unsupervised learning, each attention model here outputs a vector as the input to its corresponding subsequent CNN. The joint attention model not only improves the prediction performance but also enables the model to be interpreted at the level of "letters" (atoms in ligands) and their pairs.

[0026] As a preferred embodiment of the present invention, when training the affinity prediction model, the outputs of the protein network module and the ligand network module are aggregated in the convolutional pooling layer and then combined with the output of the graph neural network and input into the fully connected layer, and the input of the graph neural network is graph structure data.

[0027] The training process of the affinity prediction model is as follows:

[0028] Obtain training data:

[0029] Obtain protein sequences, SMILES of ligand molecules, and corresponding affinity data from a public database;

[0030] Data integration:

[0031] Integrate the protein sequences, SMILES of ligand molecules, and corresponding affinity data into an integrated dataset;

[0032] Data encoding:

[0033] Encode the protein sequences and SMILES of ligand molecules in the integrated dataset to obtain an encoded dataset, and store each encoded dataset in the training database;

[0034] Dataset Division:

[0035] Divide the data in the training database into the total training set and the test set, and further divide the total training set into the training set and the validation set;

[0036] Obtain the graph structure data of the training set;

[0037] Using proteins and molecules as nodes, extract the graph structure data of the interactions between molecules;

[0038] For each protein molecule, using the protein and its ligand molecule as nodes, and using the corresponding Binding affinity (pIC50 or pEC50) as a measure of weight, to extract the graph structure data of the interactions between molecules. There are two specific extraction methods, and the better one is preferred.

[0039] One is to take a threshold based on the IC50 value of the binding between the protein and the ligand molecule, compare the actual IC50 value with the threshold to divide it into valid edges and invalid edges, and only retain the valid edges and assign a weight of 1 to the valid edges.

[0040] The other is to retain all the edges with values, and then use the magnitude of IC50 as the weight.

[0041] Loop Iterative Training:

[0042] The data in the training set is input into the deep neural network model for training in batches. In the same batch, the protein representation data is input into the protein network module, and the ligand representation data is input into the ligand network module. The outputs of the protein network module and the ligand network module are aggregated in the convolutional pooling layer;

[0043] The graph structure data corresponding to the current batch is input into the graph neural network. After the output of the graph neural network is concatenated with the output of the convolutional pooling layer, it is input into the fully connected layer. The fully connected layer outputs the affinity prediction result, and calculates the loss function between the affinity prediction result and the corresponding existing affinity data and performs gradient backpropagation to update the parameters of the deep neural network model;

[0044] After completing 1 Epoch of training, obtain the deep neural network model after the current loop training;

[0045] Use the validation set to validate the deep neural network model after the current loop training:

[0046] The data in the validation set is input into the deep neural network model after current cycle training in batches to obtain the affinity prediction results. Calculate the root mean square error and Pearson coefficient between the affinity prediction results and the original affinity data in the validation set. When the root mean square error < 0.7 and the Pearson coefficient > 0.6, it indicates that the deep neural network model trained in the current cycle meets the basic requirements; make changes according to the actual situation. The root mean square error < 0.7 and the Pearson coefficient > 0.6 are only the minimum standards, and the model can be trained to a better result according to actual requirements.

[0047] After the loop ends, a trained deep neural network model is obtained;

[0048] The test set tests the trained deep neural network model;

[0049] The data in the test set is input into the trained deep neural network model in batches to obtain the affinity prediction results. Calculate the root mean square error and Pearson coefficient between the affinity prediction results and the original affinity data in the test set. When the root mean square error < 0.7 and the Pearson coefficient > 0.6, it indicates that the trained deep neural network model meets the basic requirements.

[0050] When training the affinity prediction model of the present invention, a graph neural network and graph structure data are added. The outputs of the protein network module and the ligand network module after convolution and pooling layer aggregation are combined together to improve the prediction accuracy of the affinity prediction model. The present invention makes full use of the sequence data of proteins and ligands, integrates a large amount of sequence data to improve the degree of feature aggregation, and then inputs the features into a deep neural network model for training. At the same time, methods such as recurrent neural network and joint attention mechanism are used to further process the highly aggregated data, and the implicit mapping relationship in the data is learned through continuous iteration, so as to achieve the purpose of combining affinity prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained without creative efforts based on these drawings.

[0052] Figure 1 It is the schematic diagram when training the affinity prediction model of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0053] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0054] A method for predicting the binding affinity between a drug molecule and a target protein, the specific steps are as follows:

[0055] Data integration;

[0056] The integrated data includes the sequence of the target protein and the SMILES of the ligand molecule;

[0057] Integrate the sequence of the target protein and the SMILES of the ligand molecule into an integrated dataset;

[0058] Data encoding:

[0059] Encode the target protein sequence and the SMILES of the ligand molecule in the integrated dataset respectively to obtain an encoded dataset;

[0060] Target protein encoding:

[0061] Each target protein is characterized by a character dataset of a set length; the character dataset includes a special character representing the start, a letter quadruple representing the target protein sequence, and a special character representing the end; and when the length of the target protein sequence is less than the set length, a special letter representing padding is used for padding;

[0062] The letter quadruple representing the protein characteristics includes secondary structure category, whether it is exposed in the solvent, physicochemical properties, and length;

[0063] The encoding rule of the letter quadruple is:

[0064]

[0065] As shown in the table, 4 separate characters of 3, 2, 4, and 3 letters are defined, which are used to represent the secondary structure category, whether it is exposed in the solvent, physicochemical properties, and length respectively, and the letters from 4 letters are combined in the above order to create 72 characters (quadruples) to describe the protein. Then the quadruples are flattened, and three special characters of start, end, and padding are added, thus defining a protein character table of 75 characters.

[0066] This coding representation method is more concise, providing higher-resolution sequences and structural details for more challenging regression tasks, more discrimination between proteins in the same family, and more explanatory predictions about which protein fragments are responsible. All of these are achieved through a much smaller alphabet of size 75, resulting in a compact representation of protein sequences that is approximately 100 times higher than the baseline.

[0067] Ligand molecule coding:

[0068] Characterize with a ligand dataset of a set length, the ligand dataset includes a special character representing the start, the SMILES of the ligand molecule, and a special character representing the end; when the length of the SMILES of the ligand molecule is less than the set length, fill in the placeholder with a special letter representing padding; when the length of the SMILES of the ligand molecule is greater than the set length, directly truncate.

[0069] Affinity prediction:

[0070] Input the encoded dataset into the affinity prediction model in batches to obtain the affinity prediction results.

[0071] The affinity prediction model is a deep neural network model, including a protein network module, a ligand network module, a convolutional pooling layer, and a fully connected layer; the outputs of the protein network module and the ligand network module are aggregated in the convolutional pooling layer and then input into the fully connected layer, and the fully connected layer outputs the prediction results.

[0072] The protein network module includes a protein embedding layer, a protein RNN layer, and a protein attention layer.

[0073] The protein embedding layer converts the input protein data into a fixed-length vector. The protein RNN layer uses a gated recurrent unit (GRU) to capture the non-linear joint dependencies between protein residues that are far apart sequentially; the protein attention layer uses a joint attention model, and the attention model is jointly trained for the protein and the RNN / CNN part. Their learning parameters include the attention weights of all letters of a given string. Compared with unsupervised learning, each attention model here outputs a vector as the input to its corresponding subsequent CNN. The joint attention model not only improves the prediction performance but also enables the model to be interpretable at the level of "letters" (SSEs in proteins) and their pairs.

[0074] The ligand network module includes a ligand embedding layer, a ligand RNN layer, and a ligand attention layer. The ligand embedding layer converts the input ligand molecular data into a fixed-length vector. The ligand RNN layer uses a gated recurrent unit (GRU) to capture the non-linear joint dependencies between ligand atoms that are far apart sequentially. The ligand attention layer uses a joint attention model, and the attention model is jointly trained for ligand atoms and the RNN / CNN part. Their learning parameters include the attention weights of all letters of a given string. Compared with unsupervised learning, each attention model here outputs a vector as the input to its corresponding subsequent CNN. The joint attention model not only improves the prediction performance but also enables the model to be interpreted at the level of "letters" (atoms in the ligand) and their pairs.

[0075] Convolutional pooling layer (CNN layer): Use a 1D-CNN model and a Maxpooling layer to downsample the upper-layer data to make the data more aggregated.

[0076] When training the affinity prediction model, a graph neural network is added to optimize the affinity prediction model. Specifically, the outputs of the protein network module and the ligand network module are aggregated in the convolutional pooling layer and then combined with the output of the graph neural network and input into the fully connected layer. The input of the graph neural network is graph-structured data.

[0077] The graph neural network used during training: Using graph-structured data, which includes protein sequences and molecular fingerprint data of small molecules. Taking protein sequences and molecular fingerprint data of small molecules as inputs, first input the molecular fingerprint data into a multi-layer perceptron for dimensionality reduction, then normalize the weights according to the weight assignment rules in the previous section, and update the node information. To prevent overfitting, the dropout method is used, and then the data is output. The output data will be concatenated with the output data of the CNN layer and input into the fully connected layer. The fully connected layer: Integrates the output data of the CNN layer and the graph neural network, and then reduces the dimensionality of the integrated data to obtain the predicted binding affinity data.

[0078] As Figure 1 shown, the specific training process is as follows:

[0079] Obtain training data:

[0080] Obtain protein sequences, SMILES of ligand molecules, and corresponding affinity data from a public database; in this embodiment, molecular data from three public datasets are used: labeled compound-protein binding data from BindingDB, compound data in SMILES format from STITCH, and protein amino acid sequences from UniRef.

[0081] Data integration:

[0082] Directly splice and integrate the protein sequence, the SMILES of the ligand molecule, and the corresponding affinity data in the integrated dataset;

[0083] Data encoding:

[0084] Encode the protein sequence and the SMILES of the ligand molecule in the integrated dataset to obtain the encoded dataset, and store each encoded dataset in the training database;

[0085] The specific encoding method is carried out according to the encoding method described above:

[0086] Dataset division:

[0087] Divide the data in the training database into the total training set and the test set, with the division ratio of 8:2, and then divide the total training set into the training set and the validation set, with the division ratio of 9:1.

[0088] Obtain the graph structure data of the training set;

[0089] Adopt the protein and the ligand molecule as nodes to extract the graph structure data of the intermolecular interaction;

[0090] Run the graph neural network on this graph structure data, and use the information of the molecular set interacting with the target protein to enhance the features of this protein.

[0091] For each protein molecule, take the protein and its ligand molecule as nodes, and use the corresponding Binding affinity (pIC50 or pEC50) as the measure of the weight to extract the graph structure data of the intermolecular interaction. There are two ways to assign weights, and the better one is selected.

[0092] One is to take a threshold according to the IC50 value of the binding of the protein and the ligand molecule, compare the actual IC50 value with the threshold to divide it into valid edges and invalid edges, and only retain the valid edges and assign the weight of the valid edges as 1.

[0093] The other is to retain all the edges with values, and then use the size of the IC50 as the weight.

[0094] Loop iterative training:

[0095] The data in the training set is input into the deep neural network model for training in batches. In the same batch, the protein representation data is input into the protein network module, the ligand representation data is input into the ligand network module, and the outputs of the protein network module and the ligand network module are aggregated in the convolutional pooling layer;

[0096] The graph structure data corresponding to the current batch is input into the graph neural network. After the graph neural network outputs, it is concatenated with the output of the convolutional pooling layer and then input into the fully connected layer. The fully connected layer outputs the affinity prediction result, and calculates the loss function between the affinity prediction result and the corresponding existing affinity data and performs gradient backpropagation to update the parameters of the deep neural network model;

[0097] After completing 1 epoch of training, the deep neural network model after the current loop training is obtained;

[0098] The validation set is used to validate the deep neural network model after the current loop training:

[0099] The data in the validation set is input into the deep neural network model after the current loop training in batches to obtain the affinity prediction result. Calculate the root mean square error and Pearson coefficient between the affinity prediction result and the original affinity data in the validation set. When the root mean square error < 0.7 and the Pearson coefficient > 0.6, it indicates that the deep neural network model of the current loop training meets the basic requirements; make changes according to the actual situation. The root mean square error < 0.7 and the Pearson coefficient > 0.6 are only the minimum standards, and it can be trained to a better result according to the actual requirements.

[0100] After the loop ends, the trained deep neural network model is obtained;

[0101] The test set is used to test the trained deep neural network model;

[0102] The data in the test set is input into the trained deep neural network model in batches to obtain the affinity prediction result. Calculate the root mean square error and Pearson coefficient between the affinity prediction result and the original affinity data in the test set. When the root mean square error < 0.7 and the Pearson coefficient > 0.6, it indicates that the trained deep neural network model meets the basic requirements.

[0103] And the root mean square error of the prediction result of the test set is 0.7847, and the Pearson coefficient is 0.6693. The root mean square error of the prediction result on a single target (taking a druggable target in GPCR as an example) is 0.7056, and the Pearson coefficient is 0.7778. This indicates that the trained deep neural network model has excellent prediction performance, has a good binding affinity prediction effect, and can greatly reduce the screening time of drug molecules.

[0104] After the training is completed, only by inputting the new protein and ligand sequences into the trained deep neural network model can its binding affinity be predicted.

[0105] During the training process, the parameters corresponding to the deep neural network model are:

[0106] Protein embedding layer: Linear(75, 256) (referring to the dimensions of input and output);

[0107] Protein RNN layer: GRU(256, 256), with a depth of 2 layers;

[0108] Small molecule embedding layer: Embedding(70, 128);

[0109] Small molecule RNN layer: GRU(128, 128), with a depth of 2 layers;

[0110] Graph neural network:

[0111] Linear(881, 256), with the activation function being LeakyReLU and Dropout(p = 0.2);

[0112] Linear(256, 256) with the activation function being LeakyReLU and Dropout(p = 0.2);

[0113] Convolutional pooling layer:

[0114] Conv1d(1, 64, kernel_size=(4,), stride=(2,), padding=(1,));

[0115] MaxPool1d(kernel_size = 4, stride = 4, padding = 0, dilation = 1);

[0116] Fully connected layer:

[0117] Linear(2304, 600), with the activation function being LeakyReLU and Dropout(p = 0.2);

[0118] Linear(2304, 600), with the activation function being LeakyReLU and Dropout(p = 0.2);

[0119] Linear(300, 1).

[0120] In the description of this specification, the descriptions referring to the reference terms "one embodiment", "example", "specific example", etc. mean that the specific features, structures, materials, or characteristics described in connection with that embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0121] As described above, it is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes should be covered within the protection scope of the present invention.

Claims

1. A method for predicting the binding affinity between a drug molecule and a target protein, characterized in that, the steps are as follows: Data integration; The integrated data includes the sequence of the target protein and the SMILES of the ligand molecule; Integrate the sequence of the target protein and the SMILES of the ligand molecule into an integrated dataset; Data encoding: Encode the sequence of the target protein and the SMILES of the ligand molecule in the integrated dataset respectively to obtain an encoded dataset; The encoding method for the sequence of the target protein is: Each target protein is characterized by a character dataset with a set length; the character dataset includes a special character representing the start, a letter quadruple representing the sequence of the target protein, and a special character representing the end; and when the length of the target protein sequence is less than the set length, a special letter representing padding is used for padding; The letter quadruple representing the protein characteristics includes secondary structure category, whether it is exposed in the solvent, physicochemical properties, and length; The encoding rule for the letter quadruple is: The encoding method for the ligand molecule is: Characterized by a ligand dataset with a set length, the ligand dataset includes a special character representing the start, the SMILES of the ligand molecule, and a special character representing the end; when the length of the SMILES of the ligand molecule is less than the set length, a special letter representing padding is used for padding; when the length of the SMILES of the ligand molecule is greater than the set length, it is directly truncated; Affinity prediction: Input the encoded dataset into the affinity prediction model in batches to obtain the affinity prediction result; The affinity prediction model is a deep neural network model, including a protein network module, a ligand network module, a convolutional pooling layer, and a fully connected layer; the outputs of the protein network module and the ligand network module are aggregated in the convolutional pooling layer and then input into the fully connected layer, and the fully connected layer outputs the prediction result; The protein network module includes a protein embedding layer, a protein RNN layer, and a protein attention layer; The ligand network module includes a ligand embedding layer, a ligand RNN layer, and a ligand attention layer; When training the affinity prediction model, a graph neural network is added to optimize the affinity prediction model; When training the affinity prediction model, the outputs of the protein network module and the ligand network module are aggregated in the convolutional pooling layer and combined with the output of the graph neural network and then input into the fully connected layer, and the input of the graph neural network is graph-structured data; Obtain the graph-structured data of the training set; use proteins and molecules as nodes to extract the graph-structured data of the intermolecular interactions; For each protein molecule, using the protein and its ligand molecule as nodes and the corresponding Binding affinity (pIC50 or pEC50) as a measure of weight, the graph structure data of intermolecular interactions is extracted; there are two specific extraction methods, and the better one is preferred; one is to take a threshold for the IC50 value of the binding between the protein and the ligand molecule, compare the actual IC50 value with the threshold to divide it into valid edges and invalid edges, and only retain the valid edges and assign the weight of the valid edges as 1; the other is to retain all the edges with values, and then use the magnitude of the IC50 as the weight. The graph neural network used during training: Using the graph structure data, which includes the protein sequence and the molecular fingerprint data of small molecules, taking the protein sequence and the molecular fingerprint data of small molecules as inputs, first input the molecular fingerprint data into a multi-layer perceptron for dimensionality reduction, then normalize the weights according to the weight assignment rules in the previous section, and update the node information. To prevent overfitting, the dropout method is used, and then the data is output; the output data will be concatenated with the output data of the CNN layer and input into the fully connected layer. The fully connected layer: Integrates the output data of the CNN layer and the graph neural network, and then reduces the dimensionality of the integrated data to obtain the predicted binding affinity data.

2. The method for predicting the binding affinity between a drug molecule and a target protein according to claim 1, characterized in that, the training process of the affinity prediction model is: Obtain training data: Obtain the protein sequence, the SMILES of the ligand molecule and the corresponding affinity data from a public database; Data integration: Integrate the protein sequence, the SMILES of the ligand molecule and the corresponding affinity data into an integrated dataset; Data encoding: Encode the protein sequence and the SMILES of the ligand molecule in the integrated dataset to obtain an encoded dataset, and store each encoded dataset in the training database; Dataset division: Divide the data in the training database into a total training set and a test set, and further divide the total training set into a training set and a validation set; Obtain the graph structure data of the training set; Using the protein and the molecule as nodes, extract the graph structure data of intermolecular interactions; Loop iterative training: The data in the training set is input into the deep neural network model for training in batches. In the same batch, the protein characterization data is input into the protein network module, and the ligand characterization data is input into the ligand network module. The outputs of the protein network module and the ligand network module are aggregated in the convolutional pooling layer; The graph structure data corresponding to the current batch is input into the graph neural network. After the graph neural network outputs, it is concatenated with the output of the convolutional pooling layer and input into the fully connected layer. The fully connected layer outputs the affinity prediction result, and calculates the loss function between the affinity prediction result and the corresponding existing affinity data and performs gradient backpropagation to update the parameters of the deep neural network model; After completing 1 Epoch of training, obtain the deep neural network model after the current loop training; Use the validation set to validate the deep neural network model after the current loop training: After the loop ends, obtain the trained deep neural network model; The test set is used to test the trained deep neural network model.

3. The method for predicting the binding affinity between a drug molecule and a target protein according to claim 2, characterized in that the verification process of the validation set is as follows: The data in the validation set is input into the deep neural network model after the current loop training in batches to obtain the affinity prediction results. Calculate the root mean square error and Pearson coefficient between the affinity prediction results and the original affinity data in the validation set. When the root mean square error < 0.7 and the Pearson coefficient > 0.6, it indicates that the trained deep neural network model meets the basic requirements; The test process of the test set is as follows: The data in the test set is input into the trained deep neural network model in batches to obtain the affinity prediction results. Calculate the root mean square error and Pearson coefficient between the affinity prediction results and the original affinity data in the test set. When the root mean square error < 0.7 and the Pearson coefficient > 0.6, it indicates that the trained deep neural network model meets the basic requirements.