Method and device for determining metalloprotein-ligand binding conformation and binding affinity thereof

By using a molecular docking model based on deep learning, the problem of accuracy in predicting the binding mode and affinity of metalloproteins to ligands has been solved, achieving high-precision prediction and evaluation of binding conformations, and improving the efficiency and accuracy of drug design.

CN121922191APending Publication Date: 2026-04-24ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG UNIV
Filing Date
2026-01-08
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately predict the binding modes and affinity between metalloproteins and ligands, especially in multi-metal systems where predictive capabilities are limited.

Method used

A deep learning-based molecular docking model is adopted, including a dual-scale metalloprotein map encoding module, a dual-scale molecular map encoding module, a metalloprotein-ligand interaction map construction module, a coordinating atom prediction module, a binding conformation prediction module, and a binding affinity assessment module. High-precision modeling is achieved through feature extraction, coordinating atom prediction, and binding conformation assessment.

Benefits of technology

This technology enables high-precision prediction of binding modes and assessment of binding affinity for metalloprotein systems, solving the problems of prediction errors and insufficient scalability in existing technologies, and improving the accuracy and efficiency of drug design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121922191A_ABST
    Figure CN121922191A_ABST
Patent Text Reader

Abstract

The invention discloses a metal protein-ligand binding conformation and a method and device for determining binding affinity of the metal protein-ligand binding conformation. According to the method, a molecular map corresponding to a ligand to be docked and a metal protein map corresponding to metal protein are obtained and input into a molecular docking model; the metal protein map and the molecular map are processed in the molecular docking model, node feature representation of the metal protein map and node feature representation of the molecular map are obtained respectively, then a metal protein-ligand interaction map is obtained, predicted ligand coordination atoms are obtained through a coordination atom prediction module, and the metal protein-ligand interaction map is obtained. And processing the metal protein-ligand interaction diagram and ligand coordination atoms to obtain a predicted molecular binding conformation between the ligand and the metal protein, and scoring to obtain a prediction result of the binding affinity of the predicted molecular binding conformation to the metal protein. According to the invention, metal protein molecular docking can be realized, and the binding mode and binding affinity of the drug micromolecule and the metal protein can be accurately and rapidly predicted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biomedical technology, specifically to a method and apparatus for determining the metalloprotein-ligand binding conformation and its binding affinity. Background Technology

[0002] Nearly 40% of proteins in the human body rely on one or more metal ions to maintain their function and structural stability; these proteins are called metalloproteins. Metalloproteins play an indispensable role in many biological processes, and their dysfunction is often closely related to the onset and progression of various diseases. Therefore, targeting metalloproteins as drug targets has become a promising therapeutic strategy, which can be used to develop highly selective inhibitors.

[0003] Molecular docking technology, as an important tool for early lead compound discovery, can describe metalloprotein-ligand interactions by predicting the binding mode and binding affinity of compounds to metalloproteins. This effectively guides the rational design of drugs targeting metalloproteins, accelerates the new drug development process, and reduces experimental costs.

[0004] However, current molecular docking methods for metalloprotein systems still have many limitations. On the one hand, only a few docking software programs, such as AutoDock4Zn, GM-DockZn, and MpsDockZn, are specifically designed for metalloprotein systems, but most of them are limited to zinc ion coordination systems. These programs typically perform limited optimization of the zinc binding site by introducing empirical parameters or modifying potential terms, such as adding metal-donor distance constraints or electrostatic correction terms to the scoring function. However, they lack universality and scalability when dealing with a wider range of metals or complex multi-metal systems, and it is difficult to accurately describe the coordination chemistry characteristics of different metal ions under different chemical environments.

[0005] On the other hand, traditional molecular docking tools (such as AutoDock, Glide, and Gold) face significant challenges in modeling metalloprotein-ligand interactions. Unlike hydrogen bonds or hydrophobic interactions in the binding pocket of conventional proteins, metal-ligand coordination involves not only spatial geometric constraints but also exhibits pronounced directional, polarimetric, and electronic coupling characteristics. The coordination geometry of metal ions is often influenced by their oxidation state, spin state, and the electron-donating properties of the ligands, resulting in a variety of spatial configurations ranging from tetrahedral and octahedral to planar tetracoordinate. This diversity leads to significant flexibility in coordination number and bond angle, which traditional scoring models based on static potential energy surfaces or empirical rules struggle to accurately capture.

[0006] Furthermore, the interactions between metal ions and surrounding amino acid residues and ligands are often accompanied by strong polarization effects and significant charge transfer behavior. Traditional molecular force fields, typically described using fixed-charge models or simplified Lennard-Jones potentials, cannot accurately reproduce the impact of electron density redistribution on the strength and directionality of coordination bonds. This leads to systematic biases in both energy assessment and geometric constraints in traditional docking algorithms. These multi-layered complexities make it difficult for traditional molecular docking tools to explore binding modes with correct coordination geometry in chemical space and accurately calculate their binding free energies in metal ion coordination systems.

[0007] In recent years, with the rapid development of deep learning technology in structural biology and medicinal chemistry, deep learning-based molecular docking models have shown unprecedented potential in modeling protein-ligand interactions. These methods, by utilizing massive amounts of three-dimensional structural data and high-dimensional feature learning capabilities, can automatically capture the complex spatial conformations, electrostatic interactions, and hydrophobic interactions between proteins and ligands, thereby significantly improving docking accuracy and conformation prediction efficiency. Representative works include models such as DeepDock, DiffDock, FlexPose, KarmaDock, and SurfDock.

[0008] However, despite significant progress in predicting the binding of small molecule ligands to conventional protein targets, these models still have considerable limitations when dealing with metalloprotein systems. On the one hand, most existing methods are not specifically optimized for metal ions and their unique coordination characteristics. In model design, metal atoms are often treated as ordinary heavy atoms, ignoring their valence electron structure, coordination geometry diversity, and the influence of variable valence states on the local potential energy surface. On the other hand, metal coordination bonds differ from covalent bonds or van der Waals interactions, exhibiting significant directionality and charge transfer characteristics. This strong polarization and electrochemical coupling effect cannot be effectively captured by traditional molecular representations or simple distance features. This also makes it difficult for these deep learning models to accurately reconstruct the complex metal-ligand interaction network in the metalloprotein binding pocket, especially when multi-metal cooperative coordination is involved, further limiting their predictive ability.

[0009] Therefore, there is an urgent need to develop a novel molecular docking method that can accurately model the unique coordination characteristics of metalloprotein systems, thereby overcoming the limitations of existing technologies in predicting metalloprotein-ligand binding modes and assessing their binding affinity, and providing new computational solutions for the rational drug design of metal-related targets. Summary of the Invention

[0010] The technical problem to be solved by the embodiments of the present invention is to provide a method and apparatus for predicting the metalloprotein-ligand binding conformation and binding affinity based on deep learning, so as to achieve high-precision modeling of metal coordination systems and efficient prediction of binding conformation.

[0011] In a first aspect, the present invention provides a method for determining the metalloprotein-ligand binding conformation and its binding affinity, the method comprising:

[0012] Obtain the metalloprotein map representation of the target metalloprotein and the molecular map representation of the ligand to be docked;

[0013] The metalloprotein map and the molecular map are input into the molecular docking model, which includes: a dual-scale metalloprotein map encoding module, a dual-scale molecular map encoding module, a metalloprotein-ligand interaction map construction module, a coordination atom prediction module, a binding conformation prediction module, and a binding affinity assessment module.

[0014] The metalloprotein map is feature extracted by the dual-scale metalloprotein map encoding module to obtain the node feature representation of the metalloprotein map;

[0015] The molecular graph is subjected to feature extraction by the dual-scale molecular graph encoding module to obtain the node feature representation of the molecular graph;

[0016] Based on the node feature representations of the metalloprotein graph and the molecular graph, the predicted ligand coordinating atoms are obtained through the coordinating atom prediction module.

[0017] The node feature representations of the metalloprotein graph and the molecular graph are processed by the metalloprotein-ligand interaction graph construction module to obtain the metalloprotein-ligand interaction graph.

[0018] The predicted molecular binding conformation between the ligand and the metalloprotein is obtained by calculating the metalloprotein-ligand interaction diagram and the predicted ligand coordinating atoms using the binding conformation prediction module.

[0019] The binding affinity assessment module is used to evaluate the predicted molecule binding conformation to obtain the predicted binding affinity of the ligand to the metalloprotein.

[0020] In a second aspect, the present invention provides an apparatus for determining the metalloprotein-ligand binding conformation and its binding affinity, comprising:

[0021] The metalloprotein-molecular map acquisition module is used to acquire the molecular map corresponding to the ligand to be docked and the metalloprotein map corresponding to the metalloprotein, and input the molecular map and the metalloprotein map into the molecular docking model. The molecular docking model includes: a dual-scale metalloprotein map encoding module, a dual-scale molecular map encoding module, a metalloprotein-ligand interaction map construction module, a coordination atom prediction module, a binding conformation prediction module, and a binding affinity assessment module.

[0022] The node feature acquisition module is used to call the dual-scale molecular graph encoding module to extract features from the molecular graph to obtain the node feature representation of the molecular graph, and to call the dual-scale metalloprotein graph encoding module to extract features from the metalloprotein graph to obtain the protein node feature representation of the metalloprotein graph.

[0023] The coordination atom prediction module is used to call the coordination atom prediction module to predict the ligand coordination atoms that form metal bonds with the metal based on the node feature representation of the metalloprotein map and the node feature representation of the molecular map.

[0024] The metalloprotein-ligand interaction graph acquisition module is used to call the metalloprotein-ligand interaction graph construction module to process the metalloprotein node feature representation and the molecular node feature representation to obtain the metalloprotein-ligand interaction graph;

[0025] The conformation acquisition module is used to call the binding conformation prediction module to calculate the metalloprotein-ligand interaction diagram and the predicted ligand coordinating atoms to obtain the predicted molecular binding conformation between the ligand and the metalloprotein.

[0026] The affinity acquisition module is used to call the affinity assessment module to evaluate the predicted molecule binding conformation and obtain the predicted binding affinity of the ligand to the metalloprotein.

[0027] Thirdly, the present invention provides a computer device, including a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein when the computer program is executed by the processor, it enables a method for determining the metalloprotein-ligand binding conformation and its binding affinity.

[0028] Fourthly, the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, enables a method for determining the metalloprotein-ligand binding conformation and its binding affinity.

[0029] Compared with the prior art, this application has the following advantages:

[0030] This application obtains the molecular map corresponding to the ligand to be docked and the metalloprotein map corresponding to the metalloprotein, and inputs both into a molecular docking model. The molecular docking model includes: a dual-scale metalloprotein map encoding module, a dual-scale molecular map encoding module, a metalloprotein-ligand interaction map construction module, a coordinating atom prediction module, a binding conformation prediction module, and a binding affinity assessment module. The dual-scale metalloprotein map encoding module processes the metalloprotein map to obtain node feature representations, and the dual-scale molecular map encoding module processes the molecular map to obtain node feature representations. The metalloprotein-ligand interaction map construction module processes the molecular nodes and metalloprotein nodes to obtain a metalloprotein-ligand interaction map. The coordinating atom prediction module obtains the predicted ligand coordinating atoms. The binding conformation prediction module processes the metalloprotein-ligand interaction map and the predicted ligand coordinating atoms to obtain the predicted molecular binding conformation between the ligand and the metalloprotein. The binding affinity assessment module scores the predicted molecular binding conformation to obtain the predicted binding affinity of the predicted molecular binding conformation to the metalloprotein.

[0031] This application enables metalloprotein molecule docking, accurately and rapidly predicting the binding mode and binding affinity between drug small molecules and metalloproteins, solving the problems of existing technologies being unable to accurately predict the binding mode between small molecules and metalloproteins and unable to accurately reconstruct the metal coordination environment in the binding pocket.

[0032] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description

[0033] Figure 1 A flowchart illustrating a method for determining the metalloprotein-ligand binding conformation and binding affinity provided in this application embodiment;

[0034] Figure 2 A schematic diagram illustrating a metalloDock process for performing metalloprotein molecule docking, provided as an embodiment of this application;

[0035] Figure 3 A schematic diagram illustrating the metal coordination environment sensing capability of MetalloDock as provided in an embodiment of this application;

[0036] Figure 4 A schematic diagram illustrating the ability of MetalloDock to reconstruct metal coordination geometry, provided in an embodiment of this application;

[0037] Figure 5 A schematic diagram illustrating the virtual screening capability of MetalloDock on a metalloprotein virtual screening dataset, provided as an embodiment of this application;

[0038] Figure 6 A schematic diagram of the virtual screening results of a PSMA target using MetalloDock, provided as an embodiment of this application;

[0039] Figure 7 A schematic diagram of a device for determining the binding conformation and binding affinity of a metalloprotein-ligand, provided in an embodiment of this application;

[0040] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0041] To make the purpose, description and advantages of this application clearer, the technical solutions in the examples of this application will be described in detail below with reference to the accompanying drawings.

[0042] It should be understood that the embodiments described in the examples of this application are only some examples of the technical solution, and not all embodiments. The terminology used in the embodiments and claims of this application is only used to describe specific implementation methods and is not intended to limit the scope of protection of this application. Unless the context clearly indicates otherwise, the singular forms "a," "the," and "the" in the embodiments and claims of this application also include their plural forms.

[0043] It should be noted that relational terms such as "first" and "second" in this document are used only to distinguish different objects or operational steps and do not indicate any actual relationship or sequence between them. Furthermore, the terms "comprising," "including," and any synonymous variations indicate a non-exclusive inclusion relationship. Therefore, an article, method, terminal, or process that comprises several elements may include, in addition to the listed elements, other elements not expressly listed or well-known in the art. Unless specifically limited, the statement "comprising one..." does not exclude the inclusion of other identical or similar elements.

[0044] Or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or terminal that includes a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal. Without further limitation, an element defined by the phrase "including one..." does not exclude the presence of other identical elements in the process, method, article, or terminal that includes said element.

[0045] This application provides a method for determining the metalloprotein-ligand binding conformation and its binding affinity, the method comprising:

[0046] Obtain the metalloprotein map representation of the target metalloprotein and the molecular map representation of the ligand to be docked;

[0047] The metalloprotein map and the molecular map are input into the molecular docking model, which includes: a dual-scale metalloprotein map encoding module, a dual-scale molecular map encoding module, a metalloprotein-ligand interaction map construction module, a coordination atom prediction module, a binding conformation prediction module, and a binding affinity assessment module.

[0048] The metalloprotein map is feature extracted by the dual-scale metalloprotein map encoding module to obtain the node feature representation of the metalloprotein map;

[0049] The molecular graph is subjected to feature extraction by the dual-scale molecular graph encoding module to obtain the node feature representation of the molecular graph;

[0050] Based on the node feature representations of the metalloprotein graph and the molecular graph, the predicted ligand coordinating atoms are obtained through the coordinating atom prediction module.

[0051] The node feature representations of the metalloprotein graph and the molecular graph are processed by the metalloprotein-ligand interaction graph construction module to obtain the metalloprotein-ligand interaction graph.

[0052] The predicted molecular binding conformation between the ligand and the metalloprotein is obtained by calculating the metalloprotein-ligand interaction diagram and the predicted ligand coordinating atoms using the binding conformation prediction module.

[0053] The binding affinity assessment module is used to evaluate the predicted molecule binding conformation to obtain the predicted binding affinity of the ligand to the metalloprotein.

[0054] Furthermore, the step of processing the metalloprotein map through the dual-scale metalloprotein map encoding module to obtain the node feature representation of the metalloprotein map includes:

[0055] The multilayer perceptron is invoked to perform dimension mapping on the features of residue nodes (including metal ions) and edge features within the metalloprotein graph, resulting in dimensionally unified residue node features and dimensionally unified residue edge features.

[0056] The geometric vector perceptron is invoked to process the dimensionally unified residue node features, edge features, and sequence information encoding features of the metalloprotein to obtain an initial residue-level node feature representation;

[0057] The multilayer perceptron is invoked to perform dimension mapping on the atomic node features within the metalloprotein map, resulting in atomic node features with unified dimensions.

[0058] A multi-head attention graph neural network is invoked to process the atomic node features and edge features of the unified dimension to obtain atomic-level node feature representations;

[0059] The atomic-level node feature representations are aggregated according to the correspondence between atoms and residues, then spliced ​​with the initial residue-level node feature representations, and the graph normalization module is called to obtain the normalized metalloprotein node feature representations.

[0060] The normalized metalloprotein node feature representation is fused using a multilayer perceptron to obtain the node feature representation of the metalloprotein map.

[0061] Furthermore, the step of extracting features from the molecular graph using the dual-scale molecular graph encoding module to obtain the node feature representation of the molecular graph includes:

[0062] The multilayer perceptron is invoked to perform dimension mapping processing on the atomic node features and edge features within the molecular graph, resulting in atomic node features and edge features with unified dimensions.

[0063] The multi-head attention graph neural network is invoked to process the atomic node features and edge features of the unified dimension to obtain the initial atomic-level node feature representation;

[0064] The multilayer perceptron is invoked to perform dimension mapping on the fragment node features within the molecular graph, resulting in fragment node features with unified dimensions.

[0065] The multi-head attention graph neural network is invoked to process the segment node features and edge features of the unified dimension to obtain segment-level node feature representations;

[0066] The fragment-level node feature representation is concatenated with the initial atomic-level node feature representation, and the graph normalization module is called to obtain the normalized ligand node feature representation.

[0067] The normalized ligand node feature representation is fused using a multilayer perceptron to obtain the node feature representation of the molecular graph.

[0068] Furthermore, the node feature representation based on the metalloprotein map and the node feature representation of the molecular map, through the coordinating atom prediction module, predicts the ligand coordinating atoms participating in metal coordination, including:

[0069] The feature representations of metal ions in the metalloprotein diagram and the node feature representations of potential coordinating atoms in the molecular diagram are spliced ​​together to obtain the metal-ligand splicing feature.

[0070] The metal-ligand splicing features are processed by a multilayer perceptron, and the node with the highest predicted probability is selected as the ligand coordinating atom based on the output probability distribution.

[0071] Further, the process of processing the node feature representations of the metalloprotein graph and the molecular graph through the metalloprotein-ligand interaction graph construction module to obtain the metalloprotein-ligand interaction graph includes:

[0072] The metalloprotein graph node feature representation and the molecular graph node feature are embedded together into the same graph space;

[0073] While preserving the original edges, the system will traverse all combinations of "metalloprotein graph nodes - molecular graph nodes" and introduce cross-molecular edges for them. At the same time, the distance between the nodes connected by the cross-molecular edges will be specially encoded.

[0074] Further, the step of calculating the predicted molecular binding conformation between the ligand and the metalloprotein using the binding conformation prediction module, based on the metalloprotein-ligand interaction diagram and the predicted ligand coordinating atoms, includes:

[0075] The order in which the coordinates of the ligand atoms are generated is determined based on the predicted topological relationships of the ligands and their covalent bonds; or the order in which the coordinates of the ligand atoms are generated is determined based on the user-specified topological relationships of the ligands and their covalent bonds.

[0076] A dynamic mask is constructed based on the atomic generation order to mask the feature representations of ligand nodes that have not yet been generated.

[0077] Based on the metalloprotein-ligand interaction map, the dynamic mask, and the atomic generation sequence, the predicted molecular binding conformation is generated atom-by-atom in a regressive manner.

[0078] Further, the evaluation of the predicted molecule's binding conformation using the binding affinity assessment module to obtain the predicted binding affinity of the ligand to the metalloprotein includes:

[0079] The node feature representations of the metalloprotein graph and the node feature representations of the molecular graph are spliced ​​together to obtain the metalloprotein-ligand splicing feature.

[0080] Based on the metalloprotein-ligand splicing characteristics and the predicted molecular binding conformation, the predicted binding affinity of the molecular binding conformation to the metalloprotein is calculated.

[0081] Optionally, before the metalloprotein map and the molecular map are input into the molecular docking model, the following steps are also included:

[0082] The binding affinity assessment module was trained based on the real three-dimensional conformation of the protein-ligand;

[0083] After the binding affinity assessment module reaches stable convergence, the binding conformation prediction module is further trained using training samples until stable convergence, thus obtaining the parameter weights of the molecular docking model.

[0084] Method Implementation Examples:

[0085] Reference Figure 1 This application illustrates a flowchart of a method for determining the binding conformation and binding affinity of a metalloprotein-ligand, as shown in the embodiments below. Figure 1 As shown, the execution of this method may include the following steps:

[0086] Step 101: Obtain the molecular map of the ligand to be docked and the metalloprotein map of the metalloprotein.

[0087] In this embodiment, a molecular graph corresponding to the ligand to be docked can be obtained. Specifically, the ligand to be docked can be represented as a 2D molecular graph, which may include an atomic-level subgraph with atoms as nodes and a fragment-level subgraph with fragments as nodes. The node features and edge features of the atomic-level subgraph and the fragment-level subgraph are shown in Table 1 below:

[0088] Table 1

[0089]

[0090] In Table 1, `bool2value` represents converting a Boolean variable to a numerical form, where `True` represents 1 and `False` represents 0. In practical applications, the construction and characterization of molecular maps can be accomplished using open-source Python libraries such as RDKit.

[0091] Metalloproteins can be represented as a 3D metalloprotein graph structure, composed of residue-level subgraphs and atomic-level subgraphs. The residue-level subgraph uses amino acid residues (including metal ions) as nodes, with node coordinates derived from the Cα atom coordinates of each residue or the center coordinates of the metal ion. Using the K-nearest neighbor algorithm, each node is connected to its 30 nearest neighbors in Euclidean space, thus constructing the overall residue-level topology. The atomic-level subgraph uses atoms as nodes, connecting them based on the chemical covalent bonds or virtual non-bonded interactions within the residues to reflect local chemical structural features. The specific graph structure definitions are shown in Table 2.

[0092] Table 2

[0093]

[0094] In Table 2, `bool2value` represents converting a Boolean variable to a numerical form, where `True` represents 1 and `False` represents 0. In practical applications, the construction and characterization of metalloprotein maps can be accomplished using open-source Python libraries such as RDKit and MDAnalysis. After obtaining the molecular map corresponding to the ligand to be docked and the metalloprotein map corresponding to the metalloprotein, step 102 is executed.

[0095] Step 102: In this example, the molecular docking model can consist of the following modules: a dual-scale metalloprotein map encoding module, a dual-scale molecular map encoding module, a metalloprotein-ligand interaction map construction module, a coordinating atom prediction module, a binding conformation prediction module, and a binding affinity assessment module. After obtaining the molecular map corresponding to the ligand to be docked and the metalloprotein map corresponding to the target metalloprotein, both can be input into the molecular docking model, and step 103 can be executed.

[0096] Step 103: This step extracts and encodes feature representations for the metalloprotein map and molecular map by calling the corresponding encoding modules. In this embodiment, the atomic-level sub-map of the metalloprotein map uses Graph Transformer (GT) to extract features, and the residue-level sub-map uses Geometric Vector Perceptrons (GVPs) to encode structural and geometric features; both the atomic-level and fragment-level sub-maps of the molecular map use GT for feature extraction to capture local chemical environments and long-range dependencies.

[0097] Furthermore, GT employs a multi-head self-attention mechanism, which can effectively capture the inherent local and global information in molecular structures, thereby enhancing the representation capabilities of graph data.

[0098] Specifically, firstly, node features are transformed through a learnable linear transformation. Sum of edge features Mapped to a unified dimensional space:

[0099]

[0100] in , , All of these are learnable parameters in the linear layer, and d represents the unified dimension.

[0101] Subsequently, the initialized node and edge features are processed through the ground truth (GT) architecture to generate the final representation. In each layer, the GT employs a multi-head attention mechanism to calculate attention weights based on edge information and updates the node and edge features:

[0102]

[0103]

[0104]

[0105] in The parameters represent the learnable linear transformations of the query, key, and value; H represents the number of attention heads, each with a dimension of d / H. This represents the edge features in the attention head; Indicates the adjacent nodes of node i; These are the learnable parameters of the linear layer. Finally, the representational ability of node and edge features is further enhanced through a feedforward neural network (FFN):

[0106]

[0107] The GVP architecture enhances the representation capabilities of geometric structures by integrating scalar and vector information.

[0108] In GVP, the input scalar features Sum vector features The following processing steps are first performed:

[0109]

[0110]

[0111]

[0112]

[0113]

[0114] in , , These are learnable parameters in a linear layer; , The result is the calculation result after L2 norm normalization; Let represent the activation function, and the final output scalar features and vector features are respectively represented by . , express.

[0115] Furthermore, before being input into the GVP layer, the sequence information is first embedded and then concatenated with additional scalar node features.

[0116]

[0117] Subsequently, the GVP layer calculates the message vector between edges based on the concatenation features of nodes and edges. Node features are defined as follows: , , edge features are defined as ,but:

[0118]

[0119]

[0120]

[0121] The node features obtained after GVP update will be regularized using LayerNorm and Dropout layers.

[0122] After extracting the features of each level of sub-map within the metalloprotein map, such as Figure 2 As shown in Figure a, the node features of the atomic-level subgraph are grouped and aggregated according to the residues to obtain an aggregated representation at the residue level. Subsequently, the aggregated features are subjected to graph normalization and then spliced ​​and fused with the corresponding residue-level subgraph node features. Finally, a multilayer perceptron is used to achieve nonlinear mapping and integration of cross-level features, thereby obtaining a unified representation that combines atomic precision and residue semantics.

[0123] After extracting features from subgraphs at different levels within the molecular graph, such as Figure 2 As shown in Figure a, the node features of the fragment-level subgraph are concatenated with the corresponding atomic node features within the atomic-level subgraph, and graph normalization is performed on the concatenated features to align information at the fragment and atomic levels. Subsequently, a multilayer perceptron is used to further achieve the fusion and joint expression of molecular cross-level features, thereby enhancing the multi-scale consistency and structural expression capability of molecular representation.

[0124] After completing the multi-level feature encoding and fusion of the metalloprotein map and molecular map, the model has obtained a comprehensive representation of the spatial structure information of the metalloprotein and the chemical topological features of the ligands, and then step 104 is executed.

[0125] Step 104: This step uses the coordination atom prediction module to identify and predict the ligand atoms most likely to coordinate with the metal ion center in the metalloprotein binding pocket, providing coordination constraint information for the subsequent construction of the binding conformation.

[0126] like Figure 2As shown in section b, the "Donor atom prediction module" allows for the input of both the coded metalloprotein and molecular maps to the coordination atom prediction module. This module first concatenates the metal node feature vectors with the feature vectors of each potential coordination atom node, and then uses a multilayer perceptron to calculate the probability distribution of coordination between each potential coordination atom and the metal center. Subsequently, the atom with the highest coordination probability is selected as the predicted ligand coordinating atom. This atom serves as the starting atom for the autoregressive generation of the ligand conformation in the subsequent binding conformation prediction module, ensuring that the final prediction result conforms to the prior knowledge of metal coordination chemistry.

[0127] Step 105: In this step, the metalloprotein-ligand interaction graph construction module integrates the metalloprotein graph node feature representation and the molecular graph node feature into the same graph structure to generate a metalloprotein-ligand interaction graph.

[0128] Specifically, metalloprotein graph nodes and molecular graph nodes are embedded together in the same graph space. While retaining the original edges, all combinations of "metalloprotein graph nodes-molecular graph nodes" are further traversed and new cross-molecular edges are added to achieve the interaction of topological information between ligands and metalloproteins.

[0129] In the interaction graph construction, nodes in the metalloprotein graph are still connected to their 30 nearest neighbors based on Euclidean distance. Within the molecular graph, to enhance topological integrity, edges are established between all atomic nodes. The characteristics of edges in the interaction graph include the type of chemical bond (e.g., single, double, triple, aromatic, metal-coordinate, and non-bonded interactions) and the geometric distance between nodes. Specifically, the distance between covalently connected nodes or nodes within the same segment in the molecular graph is calculated based on their true 3D coordinates, while the distances between other nodes are calculated based on their true spatial coordinates and multiplied by a scaling factor of 0.1. The distances between nodes in the metalloprotein graph are calculated based on their true spatial coordinates and multiplied by a scaling factor of 0.1. The distance between metalloprotein nodes and molecular nodes is uniformly encoded as -1 to represent potential non-bonded interactions.

[0130] After completing the construction of the metalloprotein-ligand interaction map, proceed to step 106 to predict the binding conformation of the ligands.

[0131] Step 106: This step processes the metalloprotein-ligand interaction map using the binding conformation prediction module to obtain the predicted ligand binding conformation between the ligand to be docked and the metalloprotein. The processing flow is illustrated below. Figure 2 As shown in "Ligand docking module" in section b.

[0132] In this embodiment, the coordinate generation order of ligand atoms is first determined based on the covalent bond topology of the ligands predicted by the coordination atom prediction module. Then, a dynamic mask is constructed according to the atom generation order to mask the feature representations of ligand nodes that have not yet been generated, thus avoiding the introduction of prior information about ungenerated atoms during conformation generation. Based on this, the binding conformation prediction module prioritizes generating predicted ligand coordination atom coordinates, starting with the metal ion, strictly adhering to the physical priors of metal coordination geometry. Subsequently, according to the generation order of compound atoms, the model uses an autoregressive approach to progressively construct the binding conformation of the compound. In the autoregressive process, each generated atom serves as the parent atom of the next atom. The coordinates of the atom to be generated are first initialized around the parent atom and then calculated by E(n) Equivariant GraphNeural Networks (EGNN).

[0133]

[0134] in Represents a metalloprotein diagram. Indicates the first The partial ligand molecule map generated in the first step, while the initial features and positions are respectively generated by... and express.

[0135] E(n) equivariant graph neural networks, as a classic architecture, are particularly suitable for processing geometric data and simulating physical systems due to their equivariance advantage. In MetalloDock, an EGNN layer with an integrated self-attention mechanism is used to predict ligand binding conformations through autoregression. Initial scalar embeddings of proteins and compounds are then used. First, initialization is performed using graph normalization, and then edge features are... Then, initialization is performed through a learnable linear transformation:

[0136]

[0137]

[0138] in ; , and These are learnable parameters. The underlying module for updating ligand coordinates consists of eight stacked EGNN layers, where the first seven layers only update node features, and the eighth EGNN layer updates atom coordinates. For compound atoms whose coordinates have not yet been predicted, the relevant data is masked to prevent information leakage. The message passing process of the l-th EGNN layer is as follows:

[0139]

[0140]

[0141]

[0142]

[0143]

[0144]

[0145]

[0146] As the basic unit of residual connections in EGNN layers, Gate_Block can effectively transmit information between layers within the module while preserving key features of the previous layer:

[0147]

[0148]

[0149] in and ; and These are the learnable parameters of a linear layer.

[0150] Coords_Update_Block is used to update the positions of the ligand atoms to be docked, ensuring the accuracy of the spatial representation:

[0151]

[0152]

[0153]

[0154] in and ; These are learnable parameters.

[0155] After the prediction of the ligand binding conformation is completed, the conformation rationality can be optimized by post-processing methods. There are two main strategies: (1) Correcting only the ligand conformation: This strategy focuses on correcting unreasonable bond lengths and bond angles in the ligand binding conformation. Its small computational cost makes it suitable for tasks such as large-scale virtual screening. In specific implementation, the molecular force field in RDkit can be used to optimize the ligand conformation in a limited number of steps (such as 10 steps); (2) Optimizing the ligand binding conformation in the protein pocket environment: This strategy minimizes the energy of the metalloprotein-ligand complex under the premise of fixing the protein conformation. This allows for more refined optimization of the ligand binding conformation and exploration of the interaction between the ligand and the metalloprotein. In specific implementation, the force field such as AMBER14SB in OpenMM can be used to optimize the ligand binding conformation.

[0156] After calling the binding conformation prediction module to obtain the predicted ligand binding conformation, step 107 is executed.

[0157] Step 107: This step processes the predicted ligand binding conformation using the binding affinity assessment module to obtain the binding affinity score of the ligand in metalloprotein binding. The processing flow is illustrated below. Figure 2 As shown in the “Scoring module” in section b.

[0158] Based on node feature representations of metalloprotein maps, node feature representations of molecular maps, and predicted ligand binding conformations, the Mixture Density Network (MDN) module can accurately estimate ligand binding affinity based on statistical potential.

[0159]

[0160]

[0161] in and Let p, c, and s represent node features and coordinates, respectively. p, c, and s correspond to metalloprotein nodes, ligand nodes, and scalar features, respectively. These represent the mean, standard deviation, mixing coefficient, and node distance, respectively. This represents the predictive affinity.

[0162] The MDN network can model the distance distribution between each pair of metalloprotein-ligand nodes by outputting a corresponding Gaussian mixture distribution based on the node feature representations of the metalloprotein graph, the node feature representations of the molecular graph, and the coordinates of the predicted ligand binding conformation.

[0163]

[0164]

[0165]

[0166]

[0167]

[0168] in , It represents the nodal features of metalloproteins and ligands. These are the learnable parameters of the linear layer, including the mean. Standard deviation and mixing coefficient This is an essential parameter for MDN networks. The predicted binding affinity score, based on the statistical potential calculated from the metalloprotein-ligand node pairs, quantifies the plausibility of the compound's binding conformation. In virtual screening, a higher score generally indicates a greater likelihood of the compound binding to the target.

[0169] The advantages of the metalloprotein molecule docking model of this application are described below in conjunction with several aspects of testing:

[0170] I. Prediction accuracy of ligand binding conformation

[0171] This application compares the prediction accuracy of MetalloDock, general-purpose molecular docking models, and molecular docking models developed for metalloproteins in terms of docking conformation. As shown in Table 3, MetalloDock achieved the best docking accuracy and conformational plausibility in both test sets with different partitions, and also achieved the best performance in bimetallic coordination systems.

[0172] Table 3:

[0173]

[0174] RMSD represents the error between the predicted conformation and the actual structure, and RMSD < 2Å is usually used as the standard for judging whether the docking conformation is predicted successfully; RMSD < 2Å & PB-valid represents the ratio of docking results with RMSD < 2Å and docking results that pass the conformation rationality test in the PoseBusters tool; Mean RMSD represents the average RMSD of all prediction results; Median RMSD represents the median RMSD of all prediction results; Time-split means that metalloprotein-ligand complexes released after 2022 are used as the test set; Plinder-consistent split means that metalloprotein-ligand complexes selected according to the split criteria of the Plinder V1 dataset are used as the test set, thereby reducing the similarity between the test set and the training set and ensuring the representativeness of the test set.

[0175] II. Metal Coordination Environment Sensing Capability

[0176] This application compares the performance of MetalloDock, general-purpose molecular docking models, and molecular docking models developed specifically for metalloproteins in terms of their ability to sense metal coordination environments. The docking results of each model were analyzed based on the Time-split metalloprotein dataset. The results are as follows: Figure 3 As shown in section a, MetalloDock achieves the best results in reproducing the ligand coordination atom environment around the metal center, accurately identifying the number of missing coordination atoms and the preferred type of coordination atoms in the metal center. Figure 3 Part b further analyzes the prediction preferences and potential biases of each model for different types of coordinating atoms in the docking results. The results show that MetalloDock performs the most accurately in capturing the distribution of coordinating atom types and can effectively reduce systematic biases for specific atom types. These results fully demonstrate that MetalloDock can effectively characterize and capture the complex features of metal coordination environments.

[0177] III. Recurrence rate of metal coordination geometry

[0178] This application compares the performance of MetalloDock, general-purpose molecular docking models, and molecular docking models developed for metalloproteins in terms of their ability to reproduce metal coordination geometry. The docking results of each model were analyzed based on the Time-split metalloprotein dataset. Figure 4 The deviations between the predictions of each model and the coordination geometry parameters in the crystal structure were compared. Figure 4Part a in the figure shows that the MetalloDock coordination distance prediction achieved the smallest average error. Figure 4 Part b in the data shows that its performance in coordination angle prediction is second only to AlphaFold3. Figure 4 Part c of the study evaluated whether the docking predictions of each model could reproduce the coordination geometry in the crystal structure. The results showed that MetalloDock achieved the highest reproduction rate. These results demonstrate that MetalloDock can reproduce the coordination geometry of the metal center with high accuracy, ensuring the accuracy of the binding conformation prediction while accurately modeling the spatial constraints of the metal coordination bonds.

[0179] IV. Testing the performance of virtual screening on a metalloprotein virtual screening dataset.

[0180] In virtual screening tasks, the ability to effectively enrich active molecules is crucial for molecular docking models. This requires the molecular docking model to accurately predict the docking configuration of compounds and accurately assess their binding affinity. This application tested the model on a self-constructed virtual screening dataset of metalloproteins, the composition of which is shown in Table 4.

[0181]

[0182] Data filtering results are as follows Figure 5 As shown, the experimental results indicate that the model can effectively enrich active molecules, and its virtual screening capability surpasses that of a number of traditional molecular docking software and deep learning-based molecular docking models.

[0183] V. Virtual screening and experimental verification of PSMA.

[0184] To investigate the performance of MetalloDock in real-world drug discovery applications, a virtual screening for PSMA targets was performed using MetalloDock on the Specs database (200,168 compounds). For example... Figure 6 As shown in Figure a, the Specs database was first filtered for drug-like properties, followed by virtual screening using MetalloDock. We performed structural clustering on the top 1% of the virtual screening results and selected 27 compounds for purchase and testing. Biological experiments showed that two compounds from MetalloDock achieved sub-micromolar activity against the PSMA target, at 219 nM and 375 nM, respectively. The relevant experimental results are shown in Figure a. Figure 6 b and Figure 6 As shown in c. Figure 6 d and Figure 6The paper presents MetalloDock's predicted binding conformations for these two active molecules, showing that both molecules form metal coordination with the metal center of PSMA, demonstrating the validity of the interaction mode.

[0185] In summary, this application provides a method for determining the metalloprotein-ligand binding conformation and its binding affinity. By obtaining the metalloprotein map representation of the target metalloprotein and the molecular map representation of the ligand to be docked, the metalloprotein map and molecular map are input into a molecular docking model. The molecular docking model includes: a dual-scale metalloprotein map encoding module, a dual-scale molecular map encoding module, a metalloprotein-ligand interaction map construction module, a coordinating atom prediction module, a binding conformation prediction module, and a binding affinity assessment module. The dual-scale metalloprotein map encoding module extracts features from the metalloprotein map to obtain node feature representations, and the dual-scale molecular map encoding module extracts features from the molecular map to obtain node feature representations. Based on the node feature representations of the metalloprotein map and the molecular map, the coordinating atom prediction module obtains the predicted ligand coordinating atoms. The metalloprotein-ligand interaction map construction module processes the node feature representations of the metalloprotein map and the molecular map to obtain the metalloprotein-ligand interaction map. By combining the conformation prediction module with the calculation of the metalloprotein-ligand interaction diagram and the predicted ligand coordinating atoms, the predicted molecular binding conformation between the ligand and the metalloprotein is obtained. The binding affinity assessment module then evaluates the predicted molecular binding conformation to obtain the predicted binding affinity of the ligand to the metalloprotein. The embodiments of this application can effectively improve the prediction accuracy and efficiency of the binding conformation and binding affinity of ligands in metalloproteins.

[0186] Reference Figure 7 This illustration shows a schematic diagram of a device for determining the binding conformation and binding affinity of a metalloprotein-ligand according to an embodiment of this application. Figure 7 As shown, the apparatus 700 for determining the molecular binding conformation and binding strength may include the following modules:

[0187] The metalloprotein-molecular map acquisition module 710 is used to acquire the molecular map corresponding to the ligand to be docked and the metalloprotein map corresponding to the metalloprotein, and input the molecular map and the metalloprotein map into the molecular docking model. The molecular docking model includes: a dual-scale metalloprotein map encoding module, a dual-scale molecular map encoding module, a metalloprotein-ligand interaction map construction module, a coordinating atom prediction module, a binding conformation prediction module, and a binding affinity assessment module.

[0188] The node feature acquisition module 720 is used to call the dual-scale molecular graph encoding module to extract features from the molecular graph to obtain the node feature representation of the molecular graph, and to call the dual-scale metalloprotein graph encoding module to extract features from the metalloprotein graph to obtain the protein node feature representation of the metalloprotein graph.

[0189] The coordination atom prediction module 730 is used to call the coordination atom prediction module to predict the ligand coordination atoms that form metal bonds with the metal based on the node feature representation of the metalloprotein map and the node feature representation of the molecular map.

[0190] The metalloprotein-ligand interaction graph acquisition module 740 is used to call the metalloprotein-ligand interaction graph construction module to process the metalloprotein node feature representation and the molecular node feature representation to obtain the metalloprotein-ligand interaction graph.

[0191] The conformation acquisition module 750 is used to call the conformation prediction module to calculate the metalloprotein-ligand interaction diagram and the predicted ligand coordinating atoms to obtain the predicted molecular binding conformation between the ligand and the metalloprotein.

[0192] The binding affinity acquisition module 760 is used to call the binding affinity assessment module to evaluate the predicted molecule binding conformation and obtain the predicted binding affinity of the ligand to the metalloprotein.

[0193] The node feature acquisition module includes:

[0194] The metalloprotein residue node feature representation unit is used to perform dimension mapping processing on the residue node features and edge features within the metalloprotein graph to obtain dimensionally unified residue node features and dimensionally unified residue edge features. The dimensionally unified residue node features, edge features, and sequence information encoding features of the metalloprotein are then processed to obtain an initial residue-level node feature representation.

[0195] The metalloprotein atomic node feature representation unit is used to perform dimensional mapping processing on the atomic node features within the metalloprotein graph to obtain dimensionally unified atomic node features, and to process the dimensionally unified atomic node features and edge features to obtain atomic-level node feature representations.

[0196] The metalloprotein node feature fusion unit is used to aggregate the atomic-level node feature representation according to the correspondence between atoms and residues, then splice it with the initial residue-level node feature representation, call the graph normalization module to obtain the normalized metalloprotein node feature representation, and then perform feature fusion on the normalized metalloprotein node feature representation to obtain the node feature representation of the metalloprotein graph.

[0197] The molecular atomic node feature representation unit is used to perform dimension mapping processing on the atomic node features and edge features within the molecular graph to obtain dimensionally unified atomic node features and edge features, and to process the dimensionally unified atomic node features and edge features to obtain an initial atomic-level node feature representation.

[0198] The molecular fragment node feature representation unit is used to perform dimension mapping processing on the fragment node features within the molecular graph to obtain dimensionally unified fragment node features, and to process the dimensionally unified fragment node features and edge features to obtain fragment-level node feature representations.

[0199] The molecular node feature fusion unit is used to perform feature splicing between the fragment-level node feature representation and the initial atomic-level node feature representation, and call the graph normalization module to obtain the normalized ligand node feature representation, and then perform feature fusion on the normalized ligand node feature representation to obtain the node feature representation of the molecular graph.

[0200] The coordinating atom prediction module includes:

[0201] The metal-potential coordinating atom feature splicing unit is used to splice the feature representation of metal ions in the metalloprotein diagram with the node feature representation of potential coordinating atoms in the molecular diagram to obtain metal-ligand splicing features.

[0202] The coordination atom prediction unit is used to process the metal-ligand splicing features and select the node with the highest predicted probability as the ligand coordination atom according to the output probability distribution.

[0203] The binding conformation acquisition module calculates the metalloprotein-ligand interaction diagram and the predicted ligand coordinating atoms to obtain the predicted molecular binding conformation between the ligand and the metalloprotein, including:

[0204] The atomic coordinate order generation unit is used to determine the coordinate generation order of ligand atoms based on the predicted ligand coordinating atoms (or user-specified ligand coordinating atoms) and the covalent bond topology of the ligands.

[0205] A dynamic mask generation unit is used to construct a dynamic mask according to the atomic generation order to mask the feature representation of ligand nodes that have not yet been generated.

[0206] The autoregressive conformation generation unit is used to integrate the metalloprotein-ligand interaction map, the dynamic mask, and the atomic generation sequence to generate a predicted molecular binding conformation through autoregressive atom-by-atom reasoning.

[0207] The binding affinity assessment module includes:

[0208] The node feature splicing unit is used to splice the node feature representations of the metalloprotein graph and the node feature representations of the molecular graph to obtain metalloprotein-ligand splicing features.

[0209] An affinity prediction unit is used to integrate the metalloprotein-ligand splicing features and the predicted molecular binding conformation to calculate the predicted binding affinity of the molecular binding conformation to the metalloprotein.

[0210] Optionally, before the metalloprotein map and the molecular map are input into the molecular docking model, the following steps are also included:

[0211] Combined with an affinity pre-training module, it is used to train the binding affinity assessment module based on the real protein-ligand three-dimensional conformation;

[0212] After the binding affinity assessment module reaches stable convergence, the binding conformation prediction module is further trained using training samples until stable convergence, thus obtaining the parameter weights of the molecular docking model.

[0213] The apparatus for predicting metalloprotein-ligand binding conformation and binding affinity provided in this application obtains a metalloprotein map representation of the target metalloprotein and a molecular map representation of the ligand to be docked. The metalloprotein map and molecular map are then input into a molecular docking model, which includes a dual-scale metalloprotein map encoding module, a dual-scale molecular map encoding module, a metalloprotein-ligand interaction map construction module, a coordinating atom prediction module, a binding conformation prediction module, and a binding affinity assessment module. The dual-scale metalloprotein map encoding module extracts features from the metalloprotein map to obtain node feature representations, and the dual-scale molecular map encoding module extracts features from the molecular map to obtain node feature representations. Based on the node feature representations of the metalloprotein map and the molecular map, the coordinating atom prediction module obtains the predicted ligand coordinating atoms. The metalloprotein-ligand interaction map construction module processes the node feature representations of the metalloprotein map and the molecular map to obtain a metalloprotein-ligand interaction map. By combining the conformation prediction module with the calculation of the metalloprotein-ligand interaction diagram and the predicted ligand coordinating atoms, the predicted molecular binding conformation between the ligand and the metalloprotein is obtained. The binding affinity assessment module then evaluates the predicted molecular binding conformation to obtain the predicted binding affinity of the ligand to the metalloprotein. The embodiments of this application can effectively improve the prediction accuracy and efficiency of the binding conformation and binding affinity of ligands in metalloproteins.

[0214] Embodiments of this application disclose a computer device. It includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the program is executed on the processor, it can be used to implement the aforementioned method for predicting metalloprotein-ligand binding conformation and binding affinity.

[0215] Figure 8 A schematic diagram of the structure of an electronic device 800 according to an embodiment of the present invention is given. For example... Figure 8 As shown, the electronic device 800 may include a central processing unit (CPU) 801. The CPU can execute relevant processing operations according to instructions stored in read-only memory (ROM) 802, or instructions loaded from memory unit 808 into random access memory (RAM) 803. RAM 803 may also store other program modules and data resources required by the electronic device 800 during operation. The CPU, ROM, and RAM can be communicated via system bus 804. In addition, input / output (I / O) interface 805 is also connected to system bus 804.

[0216] In this electronic device 800, various external or internal functional modules can be coupled to the system bus via I / O interface 805. For example, it may include input units 806, such as a keyboard, mouse, touch device, or audio acquisition device; output units 807, such as a display screen, audio playback device, etc.; storage units 808, such as magnetic storage media or optical disc storage media; and communication units 809, such as a network adapter, wireless transceiver module, modem, etc. The communication unit 809 enables the electronic device 800 to exchange data with external terminals or servers via the Internet or other communication networks.

[0217] The various method steps or processing flows mentioned in this specification can be executed by processor 801. For example, the methods of any of the above embodiments can be implemented as program instructions stored on a computer-readable medium (e.g., storage unit 808). In some embodiments, the program can be transmitted, installed, or updated to electronic device 800 via ROM or communication unit 809. When the program is loaded into RAM and invoked by the CPU, all or part of the processing actions in the aforementioned methods can be executed.

[0218] This application provides a computer-readable storage medium storing a computer program that, when executed by a processor, can implement the aforementioned method for predicting the metalloprotein-ligand binding conformation and binding affinity.

[0219] In the relevant embodiments of this application, the computer-readable medium used to store program instructions can take many forms, and its carrier is not limited to USB flash drive, mobile storage device, read-only memory (ROM), magnetic recording medium or optical medium, etc. Any medium capable of storing computer executable code is applicable.

[0220] It should be noted that those skilled in the art will understand that the functional structures and algorithm steps described herein can be implemented through hardware circuits, software instructions, or a combination of hardware and software. To avoid unnecessary repetition, this specification mainly describes these components and operations in an abstract manner according to their functions. As for whether the actual implementation is undertaken by hardware or executed by software, this can be determined by those skilled in the art based on specific application requirements, system constraints, and design trade-offs. This choice will not affect the applicability and effectiveness of the technical solution of this application.

[0221] Furthermore, the apparatus and method proposed in this application are not limited to the division method used in the specification. The foregoing embodiments are merely illustrative of the structure and are used to explain the technical principles, and are not intended to limit the actual structure. In specific implementations, the functional units in the apparatus can be recombined, adjusted, or simplified according to the system architecture. For example, some units can be physically independent or integrated into the same processing module; multiple units can also be further split or merged to improve system performance or simplify deployment.

[0222] Similarly, the order of steps described in this application is not necessarily fixed and can be adjusted, combined, or omitted according to different needs of business processes, execution efficiency, or hardware logic. In software implementation scenarios, if certain functional units exist in the form of program modules and can be distributed independently as a product, these modules can be encapsulated and stored in a computer-readable medium. When the program is loaded into the processor, the processor can execute instructions to complete all or part of the processing steps described in this application.

[0223] In summary, although several embodiments are provided in the specification for reference, these embodiments are not intended to limit the scope of protection of this application. Those skilled in the art will recognize that various equivalent modifications and alternatives can be made to the described solutions without departing from the technical concept of this application. All such modifications should be considered to be included within the scope of protection claimed in this application, and the final scope of protection is determined by the submitted claims.

[0224] Although this application has described the technical solutions in conjunction with preferred embodiments, those skilled in the art, upon understanding the basic concept of the invention, may still make various modifications, substitutions, or extensions to these embodiments. Therefore, the appended claims are intended to cover all equivalent adjustments and improvements made based on the core ideas of this application, and are not limited to the embodiments illustrated in the specification.

[0225] This application relates to a method, apparatus, electronic device, and computer-readable storage medium for predicting metalloprotein-ligand binding conformation and binding affinity based on deep learning, which has been described in detail through specific examples. The above embodiments are intended to help understand the principles and working mechanism of this technical solution, and are not intended to limit this application. For those skilled in the art, different variations or extensions can be made in specific operating procedures, system structures, or application scenarios without departing from the basic concept of this application. Therefore, the content of this specification should not be considered as a limitation on the scope of protection of this application; the actual scope of protection should be determined by the content defined in the claims.

Claims

1. A method for determining the metalloprotein-ligand binding conformation and its binding affinity, characterized in that... The method includes: Obtain the metalloprotein map representation of the target metalloprotein and the molecular map representation of the ligand to be docked; The metalloprotein map and the molecular map are input into the molecular docking model, which includes: a dual-scale metalloprotein map encoding module, a dual-scale molecular map encoding module, a metalloprotein-ligand interaction map construction module, a coordination atom prediction module, a binding conformation prediction module, and a binding affinity assessment module. The metalloprotein map is feature extracted by the dual-scale metalloprotein map encoding module to obtain the node feature representation of the metalloprotein map; The molecular graph is subjected to feature extraction by the dual-scale molecular graph encoding module to obtain the node feature representation of the molecular graph; Based on the node feature representations of the metalloprotein graph and the molecular graph, the predicted ligand coordinating atoms are obtained through the coordinating atom prediction module. The node feature representations of the metalloprotein graph and the molecular graph are processed by the metalloprotein-ligand interaction graph construction module to obtain the metalloprotein-ligand interaction graph. The predicted molecular binding conformation between the ligand and the metalloprotein is obtained by calculating the metalloprotein-ligand interaction diagram and the predicted ligand coordinating atoms using the binding conformation prediction module. The binding affinity assessment module is used to evaluate the predicted molecule binding conformation to obtain the predicted binding affinity of the ligand to the metalloprotein.

2. The method according to claim 1, characterized in that: The metalloprotein map is processed by the dual-scale metalloprotein map encoding module to obtain the node feature representation of the metalloprotein map, specifically including: The multilayer perceptron is invoked to perform dimension mapping on the residue node features and edge features within the metalloprotein graph, resulting in residue node features and residue edge features with unified dimensions. The geometric vector perceptron is invoked to process the dimensionally unified residue node features, edge features, and sequence information encoding features of the metalloprotein to obtain an initial residue-level node feature representation; The multilayer perceptron is invoked to perform dimension mapping on the atomic node features within the metalloprotein map, resulting in atomic node features with unified dimensions. A multi-head attention graph neural network is invoked to process the atomic node features and edge features of the unified dimension to obtain atomic-level node feature representations; The atomic-level node feature representations are aggregated according to the correspondence between atoms and residues, then spliced ​​with the initial residue-level node feature representations, and the graph normalization module is called to obtain the normalized metalloprotein node feature representations. The normalized metalloprotein node feature representation is fused using a multilayer perceptron to obtain the node feature representation of the metalloprotein map.

3. The method according to claim 1 or 2, characterized in that: The molecular graph is feature-extracted using the dual-scale molecular graph encoding module to obtain node feature representations of the molecular graph, specifically including: The multilayer perceptron is invoked to perform dimension mapping processing on the atomic node features and edge features within the molecular graph, resulting in atomic node features and edge features with unified dimensions. The multi-head attention graph neural network is invoked to process the atomic node features and edge features of the unified dimension to obtain the initial atomic-level node feature representation; The multilayer perceptron is invoked to perform dimension mapping on the fragment node features within the molecular graph, resulting in fragment node features with unified dimensions. The multi-head attention graph neural network is invoked to process the segment node features and edge features of the unified dimension to obtain segment-level node feature representations; The fragment-level node feature representation is concatenated with the initial atomic-level node feature representation, and the graph normalization module is called to obtain the normalized ligand node feature representation. The normalized ligand node feature representation is fused using a multilayer perceptron to obtain the node feature representation of the molecular graph.

4. The method according to claim 3, characterized in that: Based on the node feature representations of the metalloprotein map and the molecular map, the coordinating atom prediction module predicts the ligand coordinating atoms involved in metal coordination, specifically including: The feature representations of metal ions in the metalloprotein diagram and the node feature representations of potential coordinating atoms in the molecular diagram are spliced ​​together to obtain the metal-ligand splicing feature. The metal-ligand splicing features are processed by a multilayer perceptron, and the node with the highest predicted probability is selected as the ligand coordinating atom based on the output probability distribution.

5. The method according to claim 1, characterized in that: The metalloprotein-ligand interaction graph construction module processes the node feature representations of the metalloprotein graph and the molecular graph to obtain a metalloprotein-ligand interaction graph, specifically including: The metalloprotein graph node feature representation and the molecular graph node feature are embedded together into the same graph space; While preserving the original edges, the system traverses all combinations of "metalloprotein graph nodes – molecular graph nodes" and introduces cross-molecular edges for them, while encoding the distance between the nodes connected by the cross-molecular edges.

6. The method according to claim 5, characterized in that: The binding conformation prediction module calculates the metalloprotein-ligand interaction diagram and the predicted ligand coordinating atoms to obtain the predicted molecular binding conformation between the ligand and the metalloprotein, specifically including: The order in which the coordinates of the ligand atoms are generated is determined based on the predicted topological relationships of the ligands and their covalent bonds; or the order in which the coordinates of the ligand atoms are generated is determined based on the user-specified topological relationships of the ligands and their covalent bonds. A dynamic mask is constructed based on the atomic generation order to mask the feature representations of ligand nodes that have not yet been generated. Based on the metalloprotein-ligand interaction map, the dynamic mask, and the atom generation sequence, the predicted molecular binding conformation is generated atom-by-atom inference in an autoregressive manner.

7. The method according to claim 6, characterized in that: The binding affinity assessment module evaluates the predicted molecule binding conformation to obtain the predicted binding affinity of the ligand to the metalloprotein, specifically including: The node feature representations of the metalloprotein graph and the node feature representations of the molecular graph are spliced ​​together to obtain the metalloprotein-ligand splicing feature. Based on the metalloprotein-ligand splicing characteristics and the predicted molecular binding conformation, the predicted binding affinity of the molecular binding conformation to the metalloprotein is calculated.

8. A device for determining the metalloprotein-ligand binding conformation and its binding affinity, characterized in that, include: The metalloprotein-molecular map acquisition module is used to acquire the molecular map corresponding to the ligand to be docked and the metalloprotein map corresponding to the metalloprotein, and input the molecular map and the metalloprotein map into the molecular docking model. The molecular docking model includes: a dual-scale metalloprotein map encoding module, a dual-scale molecular map encoding module, a metalloprotein-ligand interaction map construction module, a coordination atom prediction module, a binding conformation prediction module, and a binding affinity assessment module. The node feature acquisition module is used to call the dual-scale molecular graph encoding module to extract features from the molecular graph to obtain the node feature representation of the molecular graph, and to call the dual-scale metalloprotein graph encoding module to extract features from the metalloprotein graph to obtain the protein node feature representation of the metalloprotein graph. The coordination atom prediction module is used to call the coordination atom prediction module to predict the ligand coordination atoms that form metal bonds with the metal based on the node feature representation of the metalloprotein map and the node feature representation of the molecular map. The metalloprotein-ligand interaction graph acquisition module is used to call the metalloprotein-ligand interaction graph construction module to process the metalloprotein node feature representation and the molecular node feature representation to obtain the metalloprotein-ligand interaction graph; The conformation acquisition module is used to call the binding conformation prediction module to calculate the metalloprotein-ligand interaction diagram and the predicted ligand coordinating atoms to obtain the predicted molecular binding conformation between the ligand and the metalloprotein. The affinity acquisition module is used to call the affinity assessment module to evaluate the predicted molecule binding conformation and obtain the predicted binding affinity of the ligand to the metalloprotein.

9. A computer device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, can implement the method according to any one of claims 1-8.

10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program that, when executed by a processor, can implement the method described in any one of claims 1-8.