Protein-ligand binding affinity prediction method based on double virtual node message passing
By introducing a dual virtual node message delivery mechanism in the protein-ligand complex, the contribution of virtual nodes is dynamically adjusted, and the information submersion and noise interference problems of non-covalent interaction learning in the prior art are solved, improving the accuracy and comprehensiveness of prediction.
Patent Information
- Application Number
- CN202510080280.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art has problems of information submersion and noise interference when learning non-covalent interactions in protein-ligand complexes, resulting in insufficient accuracy and comprehensiveness of the prediction results.
Using a dual virtual node messaging method, by constructing a composite graph and initializing a virtual node, combining summation pooling technology and multi-layer perceptrons, the contribution of the virtual nodes is dynamically adjusted to enhance the learning of non-covalent interactions.
It improves the learning accuracy and comprehensiveness of non-covalent interactions in the complex, reduces information loss and noise interference, and improves the robustness and generalization capabilities of the model.
Smart Images

Figure CN120015108A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a protein-ligand binding affinity prediction method based on dual virtual node message transmission, and belongs to the technical field of biological information processing of artificial intelligence. Background Art
[0002] The binding affinity between proteins and ligands is one of the key factors in drug design. Accurately predicting this binding affinity can help researchers quickly screen potential drug molecules, thereby accelerating the drug development process. In recent years, with the development of computational biology and machine learning technology, graph neural network-based methods have been widely used to extract features from protein-ligand complexes and predict binding affinity, which has greatly accelerated the process of drug design.
[0003] Currently, the mainstream idea for using graph neural networks to predict protein-ligand binding affinity is to predict affinity by learning covalent and non-covalent interactions in protein-ligand complexes. Common methods are mainly divided into two categories: one type of method models the complex as a graph, regards covalent and non-covalent interactions as the same interaction, and directly inputs it into the graph neural network to extract features and predict affinity. However, since the number of non-covalent interaction edges is much larger than the covalent interaction edges, this type of method may cause non-covalent interaction information to overwhelm covalent interaction information. The other type constructs the complex as a protein graph, a ligand graph, and a protein-ligand interaction graph. Figure 3 The covalent and non-covalent interactions in the complex are extracted separately through two-stage message passing to prevent the covalent interaction information from being overwhelmed.
[0004] However, existing methods are insufficient in learning non-covalent interactions in complexes. Since non-covalent bonds in complexes are not explicitly present, it is necessary to first manually set a threshold to determine the non-covalent edges, and then learn the non-covalent interactions through message passing. This method will lead to two problems: first, it will introduce erroneous non-covalent edges, causing noise during message passing, which will affect the accuracy of the prediction results; the other point is that some long-range interactions between atoms will be ignored because these interactions may not meet the preset distance threshold conditions and cannot be correctly identified and processed. Summary of the invention
[0005] The present invention proposes a protein-ligand binding affinity prediction method based on dual virtual node message passing to improve the accuracy and comprehensiveness of non-covalent interaction learning in the prior art, thereby improving the prediction performance of protein-ligand binding affinity. The specific technical solution of the present invention is:
[0006] The steps of the protein-ligand binding affinity prediction method based on dual virtual node message passing are as follows:
[0007] Step 1: Obtain a dataset containing protein-ligand 3D complexes and corresponding affinity tags, preprocess the data, and construct and save it as a binary complex file;
[0008] Step 2: constructing a complex graph based on the binary complex file, the graph including node features, covalent interaction edges, and non-covalent interaction edges;
[0009] Step 3: Construct a dual virtual node message passing module to learn the covalent and non-covalent interactions in the complex graph;
[0010] Step 4: Update the virtual node by using the node features updated in the previous round of aggregation;
[0011] Step 5: The final complex graph is processed using the sum-pooling technique and fed into a multi-layer perceptron to predict affinity.
[0012] Specifically, the detailed process of obtaining the data set and constructing the binary complex file in step 1 is as follows: after obtaining the original data set through the public database, the protein file and the ligand file in the complex are loaded using the molecular visualization and analysis tool, and the distance ligand is selected after removing the water molecules in the protein and the hydrogen atoms in the ligand. The amino acid residues within are taken as the protein pocket, and the pocket is saved as a protein pocket file. Then the generated protein pocket file and ligand file are loaded, packaged into a tuple and serialized and saved as a binary complex file.
[0013] Specifically, the detailed process of constructing the composite graph in step 2 includes:
[0014] Step 2.1: Extract the atomic information, covalent bond and non-covalent bond information in the protein pocket and ligand in each complex file and convert it into a complex graph G = (V, E), where V is the set of all nodes in the graph and E contains two types of edges: covalent interaction edge E cov and non-covalent interaction edge E ncov The covalent interaction edge is defined by whether there is a covalent bond between the two nodes, and the non-covalent interaction edge is defined by whether the interatomic distance between the protein and the ligand is less than This threshold is defined;
[0015] Step 2.2: For edge feature e ij, calculated based on the relative distance between two nodes in three-dimensional space and input into the radial basis function; for the node feature h, it is calculated based on six types of information (sign, atomic degree, implicit valence, degree of hybridization, aromaticity and number of adjacent hydrogen atoms); then the degree information of each node under the non-covalent edge condition is logarithmically transformed and additionally saved in the complex graph data.
[0016] Specifically, the detailed process of constructing the dual virtual node message transmission module in step 3 includes:
[0017] Step 3.1: Initialize two virtual nodes in the complex graph, representing the entire protein v v_p and the entire ligand v v_l , and set the embedding vectors of the two virtual nodes to zero vectors;
[0018] Step 3.2: Learn covalent interactions through ordinary message passing; for the learning of non-covalent interactions, first concatenate the non-covalent degree information with the original node features, and then calculate the fusion ratio of each protein and ligand node with the virtual ligand node and virtual protein node through a fully connected layer. The specific calculation formula is as follows:
[0019]
[0020] in, represents the previous layer feature of node i, d i represents the non-covalent degree information of node i, σ represents the activation function, MLP gate It is a multi-layer perceptron, which is used to calculate the virtual node fusion ratio, and finally fuse it with the virtual node according to the calculated fusion ratio to obtain the fused node features. The specific calculation formula is as follows:
[0021]
[0022] Among them, v v_target(i) is the target virtual node (for protein node i, the target is the virtual ligand node v v_l , and vice versa), and then another message passing is performed. Finally, the results of covalent message passing and dual virtual node message passing are input into two fully connected layers respectively and added together, and the final message passing is completed by adding the node features of the previous round through residual connection. The specific calculation formula is as follows:
[0023]
[0024] Among them, MLP cov and MLP ncov Multilayer Perceptrons are used to handle covalent and non-covalent interactions, respectively.
[0025] Specifically, in step 4, the detailed process of updating the virtual node is as follows: after each round of message transmission, the virtual node is updated. First, the feature averages of the protein and ligand nodes at the current stage are calculated. Then, these values are processed by the multilayer perceptron and the update amplitude is adjusted by the scaling factor. v_p and v v_l The specific update formula is:
[0026]
[0027] Among them, s p and l are learnable scaling factors used to control the update amplitude of virtual nodes.
[0028] Specifically, in step 5, the detailed process of predicting affinity is as follows: when predicting affinity, four rounds of operations of step 3 and step 4 are required to fully learn the covalent and non-covalent interactions in the protein-ligand complex, and then the complex graph is summed and pooled, and the result is input into the multi-layer perceptron to predict the binding affinity between the protein and the ligand.
[0029] Based on the above, the present invention proposes a protein-ligand binding affinity prediction method based on dual virtual node message passing. By introducing a dual virtual node message passing mechanism, the model can integrate global information at each iteration, rather than just relying on local interactions. In this way, even if some local interactions are not correctly identified, global information can still help correct these errors, improving the model's understanding and learning ability of complex interaction patterns; and by dynamically adjusting the contribution of virtual nodes to enhance the model's understanding and learning of complex interaction patterns, it can more accurately capture non-covalent interactions in the complex, reduce information loss or noise interference caused by fixed thresholds, and improve the robustness and generalization ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 It is a flow chart of the method of the present invention
[0031] Figure 2 This is the architecture diagram of the dual virtual node message transmission model proposed by the present invention
[0032] Figure 3 This is a schematic diagram of the virtual node fusion module of the present invention. DETAILED DESCRIPTION
[0033] The present invention is further described in detail below in conjunction with the accompanying drawings.
[0034] like Figure 1As shown, the present invention provides a protein-ligand binding affinity prediction method based on dual virtual node message transmission, which specifically includes the following steps:
[0035] Step 1: Obtain a dataset containing protein-ligand 3D complexes and corresponding affinity tags, preprocess the data, and construct and save it as a binary complex file;
[0036] Step 2: constructing a complex graph based on the binary complex file, the graph including node features, covalent interaction edges, and non-covalent interaction edges;
[0037] Step 3: Construct a dual virtual node message passing module to learn the covalent and non-covalent interactions in the complex graph;
[0038] Step 4: Update the virtual node by using the node features updated in the previous round of aggregation;
[0039] Step 5: The final complex graph is processed using the sum-pooling technique and fed into a multi-layer perceptron to predict affinity.
[0040] Specifically, the detailed process of obtaining the data set and constructing the binary complex file in step 1 is to obtain the original data set through a public database (such as PDBbind, etc.), use professional molecular visualization and analysis tools (such as PyMOL, Chimera or RDKit) to load the protein file and ligand file in the complex, remove unnecessary interference factors in affinity prediction, such as water molecules in proteins and hydrogen atoms in ligands, and then select the distance ligand. The amino acid residues within are taken as the protein pocket, and the pocket is saved as a protein pocket file. Then the generated protein pocket file and ligand file are loaded, packaged into a tuple and serialized and saved as a binary complex file.
[0041] In step 2, when constructing the composite graph, we first need to determine the definition of the nodes in the graph, that is, the atoms in the composite graph; then we need to continue to determine the edges in the graph. Here we define two types of edges to learn different types of interactions. The detailed process includes the following two sub-steps:
[0042] Step 2.1: Extract the atomic information, covalent bond and non-covalent bond information in the protein pocket and ligand in each complex file and convert it into a complex graph G = (V, E), where V is the set of all nodes in the graph and E contains two types of edges: covalent interaction edge E cov and non-covalent interaction edge E ncovThe covalent interaction edge is directly defined by whether there is a covalent bond between the two nodes. Since there is no explicit non-covalent bond between the protein and the ligand, the non-covalent interaction edge is defined by whether the atomic distance between the protein and the ligand is less than This threshold is defined;
[0043] Step 2.2: For edge feature e ij , the relative distance is calculated according to the coordinates of the two nodes in three-dimensional space and input into the radial basis function for calculation; for the node feature h, it is calculated according to six types of information (sign, atomic degree, implicit valence, degree of hybridization, aromaticity and number of adjacent hydrogen atoms); then the degree information of each node under the non-covalent edge condition is logarithmically transformed and additionally saved in the complex graph data.
[0044] In step 3, the dual virtual node message transmission model is constructed as shown in the figure Figure 2 As shown, the detailed process includes:
[0045] Step 3.1: First, initialize two virtual nodes in the complex graph, representing the entire protein v v_p and the entire ligand v v_l , and before the first round of message passing, the embedding vectors of the two virtual nodes are set to zero vectors;
[0046] Step 3.2: Learn covalent interactions through ordinary message passing; for the learning of non-covalent interactions, first concatenate the non-covalent degree information with the original node features, and then calculate the fusion ratio of each protein and ligand node with the virtual ligand node and virtual protein node through a fully connected layer. The specific calculation formula is as follows:
[0047]
[0048] in, represents the previous layer feature of node i, d i represents the non-covalent degree information of node i, σ represents the activation function, MLP gate It is a multi-layer perceptron, which is used to calculate the virtual node fusion ratio, and finally fuse it with the virtual node according to the calculated fusion ratio to obtain the fused node characteristics; in this way, each protein and ligand node can fuse the virtual node information according to its own needs to avoid the introduction of deviations caused by direct fusion; the specific calculation formula is as follows:
[0049]
[0050] Among them, v v_target(i) is the target virtual node (for protein node i, the target is the virtual ligand node v v_l, and vice versa), and then a message is passed according to the non-covalent edge. Finally, the results of covalent message passing and dual virtual node message passing are input into two fully connected layers respectively and added together. The final message passing is completed by adding the node features of the previous round through residual connection. The specific calculation formula is as follows:
[0051]
[0052] Among them, MLP cov and MLP ncov Multilayer Perceptrons are used to handle covalent and non-covalent interactions, respectively.
[0053] In step 4, the details of updating the virtual nodes are as follows: the virtual nodes are updated after each round of message passing to ensure that the model can dynamically integrate the latest global information; first, the feature averages of the protein and ligand nodes after this round of message passing need to be calculated, and then these features are mapped to a higher-dimensional space using a multi-layer perceptron to extract richer representations; and the update amplitude is adjusted by a learnable scaling factor, v v_p and v v_l The specific update formula is:
[0054]
[0055] Among them, s p and l are learnable scaling factors used to control the update amplitude of virtual nodes.
[0056] In step 5, when finally predicting the affinity, it is necessary to go through four rounds of steps 3 and 4 to allow the model to fully learn the covalent and non-covalent interactions in the protein-ligand complex, then sum and pool the complex graph, and input the result into the multi-layer perceptron to predict the binding affinity between the protein and the ligand.
[0057] Through the above steps, the characteristics of virtual nodes can be fully utilized to make up for the information loss or noise interference problems in learning non-covalent interactions caused by manually defined thresholds in previous methods, and effectively improve the robustness and generalization ability of the model.
[0058] The above is only a preferred embodiment of the present invention and is not intended to limit the present invention. It should be noted that for those skilled in the art, several improvements and modifications made without departing from the principles of the present invention should also be considered as the protection scope of the present invention.
Claims
1. A protein-ligand binding affinity prediction method based on dual virtual node message passing, characterized in that: The following steps are involved: Step 1: Obtain a dataset containing protein-ligand 3D complexes and corresponding affinity tags, preprocess the data, and construct and save it as a binary complex file; Step 2: constructing a complex graph based on the binary complex file, the graph including node features, covalent interaction edges, and non-covalent interaction edges; Step 3: Construct a dual virtual node message passing module to learn the covalent and non-covalent interactions in the complex graph; Step 4: Update the virtual node by using the node features updated in the previous round of aggregation; Step 5: The final complex graph is processed using the sum-pooling technique and fed into a multi-layer perceptron to predict affinity.
2. The protein-ligand binding affinity prediction method based on dual virtual node message passing according to claim 1, characterized in that: In step 1, after obtaining the original data set through the public database, the protein file and ligand file in the complex are loaded using the molecular visualization and analysis tool, and the distance ligand is selected after removing the water molecules in the protein and the hydrogen atoms in the ligand. The amino acid residues within are taken as the protein pocket, and the pocket is saved as a protein pocket file. Then the generated protein pocket file and ligand file are loaded, packaged into a tuple and serialized and saved as a binary complex file.
3. The protein-ligand binding affinity prediction method based on dual virtual node message passing according to claim 1, characterized in that: In step 2, the specific steps of constructing the composite graph include: Step 2.1: Extract the atomic information, covalent bond and non-covalent bond information in the protein pocket and ligand in each complex file and convert it into a complex graph G = (V, E), where V is the set of all nodes in the graph and E contains two types of edges: covalent interaction edge E cov and non-covalent interaction edge E ncov The covalent interaction edge is defined by whether there is a covalent bond between the two nodes, and the non-covalent interaction edge is defined by whether the interatomic distance between the protein and the ligand is less than This threshold is defined; Step 2.2: For edge feature e ii , calculated based on the relative distance between two nodes in three-dimensional space and input into the radial basis function; for the node feature h, it is calculated based on six types of information (sign, atomic degree, implicit valence, degree of hybridization, aromaticity and number of adjacent hydrogen atoms); then the degree information of each node under the non-covalent edge condition is logarithmically transformed and additionally saved in the complex graph data.
4. The protein-ligand binding affinity prediction method based on dual virtual node message passing according to claim 1, characterized in that: In step 3, the specific steps of constructing the dual virtual node message transmission module include: Step 3.1: Initialize two virtual nodes in the complex graph, representing the entire protein v v_p and the entire ligand v v_l , and set the embedding vectors of the two virtual nodes to zero vectors; Step 3.2: Learn covalent interactions through ordinary message passing; for the learning of non-covalent interactions, first concatenate the non-covalent degree information with the original node features, and then calculate the fusion ratio of each protein and ligand node with the virtual ligand node and virtual protein node through a fully connected layer. The specific calculation formula is as follows: in, represents the previous layer feature of node i, d i represents the non-covalent degree information of node i, σ represents the activation function, MLP gate It is a multi-layer perceptron, which is used to calculate the virtual node fusion ratio, and finally fuse it with the virtual node according to the calculated fusion ratio to obtain the fused node features. The specific calculation formula is as follows: Among them, v v_target(i) is the target virtual node (for protein node i, the target is the virtual ligand node v v_l , and vice versa), and then another message passing is performed. Finally, the results of covalent message passing and dual virtual node message passing are input into two fully connected layers respectively and added together, and the final message passing is completed by adding the node features of the previous round through residual connection. The specific calculation formula is as follows: Among them, MLP cov and MLP ncov Multilayer perceptrons are used to handle covalent and non-covalent interactions, respectively.
5. The protein-ligand binding affinity prediction method based on dual virtual node message passing according to claim 1, characterized in that: In step 4, the virtual node is updated after each round of message transmission. First, the feature averages of the protein and ligand nodes at the current stage are calculated, and then these values are processed by the multi-layer perceptron and the update amplitude is adjusted by the scaling factor, v v_p and v v_l The specific update formula is: Among them, s p and l are learnable scaling factors used to control the update amplitude of virtual nodes.
6. The protein-ligand binding affinity prediction method based on dual virtual node message passing according to claim 1, characterized in that: In step 5, four rounds of steps 3 and 4 are required for affinity prediction. After fully learning the covalent and non-covalent interactions in the protein-ligand complex, the complex graph is summed and pooled, and the result is input into a multi-layer perceptron to predict the binding affinity between the protein and the ligand.
Citation Information
Patent Citations
Prediction model training method, binding affinity prediction method, device and equipment
CN115171776A
Three-dimensional protein-ligand activity prediction method based on attention mechanism
CN115512785A
Protein and ligand affinity prediction method based on graph neural network decoupling
CN116312758A
Multi-modal protein-ligand binding affinity prediction method based on cross-channel fusion
CN117912545A
Systems and methods for dynamic-backbone protein-ligand structure prediction with multiscale generative diffusion models
WO2024187031A2