A method for predicting protein metal ion binding sites based on complete subgraph retrieval and statistical potential function
By constructing a statistical potential function and a complete subgraph retrieval algorithm, the binding sites of various metal ions in proteins are identified, solving the problem of inaccurate identification in existing technologies. This achieves efficient and accurate multi-nuclear site prediction with physical interpretability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUAZHONG UNIV OF SCI & TECH
- Filing Date
- 2026-03-18
- Publication Date
- 2026-06-19
Smart Images

Figure CN122245411A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computational biology, structural bioinformatics, and protein function prediction, and more specifically, to a method for predicting protein metal ion binding sites based on complete subgraph retrieval and statistical potential functions. Background Technology
[0002] Metal ions play crucial roles in protein structural stability, catalytic activity, and signal regulation. It is estimated that approximately 30-40% of proteins require metal ions as cofactors to maintain their biological functions. Therefore, accurately identifying metal-binding sites in proteins is of great significance for understanding protein functional mechanisms, designing metalloenzymes, and identifying drug targets.
[0003] Existing methods for predicting metal binding sites can be broadly categorized as follows: First, sequence-based methods predict metal binding sites by identifying conserved amino acid patterns, but struggle to capture the cooperative coordination relationships of distant residues in space. Second, template-matching methods rely on known metal-binding structural fragments and have limited generalization ability to novel coordination patterns. Third, geometrically based methods use only distance or angle thresholds between donor atoms for screening, lacking quantitative assessment of interaction strength and tending to produce high false positive rates. Fourth, machine learning or deep learning-based methods rely on large-scale training datasets, have limited generalization ability to different metal types, and lack physical interpretability. Furthermore, while quantum mechanical / molecular mechanics methods offer high accuracy, their computational cost is extremely high, making them unsuitable for large-scale routine predictions.
[0004] Furthermore, for complex coordination scenarios such as multinucleated metal clusters or bridging coordination structures, existing methods often employ simple spatial clustering strategies, making it difficult to distinguish the boundaries or shared ligand relationships between different coordination clusters. This may lead to the accidental deletion of biologically significant real sites.
[0005] Therefore, there is an urgent need for a prediction method that has a physical basis, is applicable to multiple metal types, is computationally efficient, highly interpretable, and can accurately identify multi-core and complex bridging metal sites. Summary of the Invention
[0006] This invention provides an efficient, accurate, and physically interpretable method for predicting protein metal ion binding sites. Sample data is extracted from a metalloprotein structure database, and metal ions are categorized based on their coordination preferences and ionic radius characteristics. The distance distribution between each type of metal ion and the protein donor atom is statistically analyzed. A statistical potential function is constructed based on the inverse Boltzmann relation to extract the donor atom set of the target protein. An undirected graph structure is constructed based on the spatial distance between donor atoms, and a complete subgraph retrieval algorithm is used to identify candidate coordination clusters. A local grid search is performed in the neighborhood of the geometric center of the candidate clusters, and energy optimization is performed using the statistical potential function. Finally, redundant predictions are removed based on energy thresholds and spatial distances, and the three-dimensional coordinates of the metal ion in the target protein and the residue information coordinated with the metal ion are output. This method achieves or exceeds the levels of existing mainstream methods in terms of accuracy, recall, and F1 score. It also possesses strong multi-nucleus site identification capabilities, and the prediction results are physically interpretable, directly traceable to atomic-level interactions, facilitating mechanistic analysis. This solves the technical problem that, under the condition of knowing the three-dimensional structure of a protein, the prediction of the binding sites of multiple metal ions is not accurate and slow, and cannot be predicted accurately and with physical interpretability.
[0007] This invention provides a method for predicting protein metal ion binding sites based on complete subgraph retrieval and statistical potential functions, comprising the following steps: (1) Obtain the three-dimensional structure of the metal ion protein complex, classify the protein complex according to the coordination characteristics of the metal ion, and construct a statistical potential function for each type of protein complex; (2) Extract the donor atom set of the target protein, construct an undirected graph based on the spatial distance between donor atoms, and use a complete subgraph retrieval algorithm to identify all complete subgraphs in the undirected graph. Each complete subgraph is a candidate coordination cluster, and the geometric center of the donor atom in each candidate coordination cluster is used as the metal initial coordinate of the candidate coordination cluster. (3) The candidate coordination clusters are identified by a complete subgraph retrieval algorithm. Local grid search is performed around the geometric center of the candidate coordination clusters. Energy optimization is performed by combining the statistical potential function in step (1). Redundant predictions are removed based on the energy threshold and spatial distance. The three-dimensional coordinates of the metal ions of the target protein and the residue information coordinated with the metal ions are output.
[0008] Preferably, in step (1), the construction of the statistical potential function specifically involves: statistically analyzing the radial distribution of donor atoms within 10 Å around the metal ions of each type of protein complex, and constructing the statistical potential function based on the inverse Boltzmann relation.
[0009] Preferably, in step (1), the statistical potential function is constructed using the following formula:
[0010] in, Boltzmann's constant, Absolute temperature Indicates the type of metal. Indicates the type of protein atoms. Indicates the distance between metal ions and protein atoms; Indicates that the radius is from arrive Inside, type is Protein atoms and metal types Number density of atomic pairs The thickness of the radially statistical spherical shell is represented by 0.1 Å; Indicates a radius of The average number density of corresponding atomic pairs within a reference sphere of 10 Å; This represents the statistical potential of a metal-ion protein complex based on the inverse Boltzmann relation.
[0011] Preferably, in step (1), the classification specifically refers to: Ca 2+ Class, K + Class, Mg 2+ Classes of transition metal ions and transition metal ions.
[0012] Preferably, the transition metal ions include zinc ions, copper ions, cobalt ions, nickel ions, ferric ions, ferrous ions, or manganese ions.
[0013] Preferably, in step (2), after identifying all complete subgraphs in the undirected graph using the complete subgraph retrieval algorithm, clusters with ≥3 nodes are retained as candidate coordination clusters.
[0014] Preferably, in step (2), after identifying all complete subgraphs in the undirected graph using the complete subgraph retrieval algorithm, subgraphs completely contained by larger subgraphs are deleted to reduce redundancy.
[0015] Preferably, in step (3), the metal ion coordination residues are coordination residues within 4 Å of the metal ion.
[0016] In summary, compared with the prior art, the above-described technical solutions conceived by this invention mainly possess the following technical advantages: (1) This invention provides an efficient, accurate and physically interpretable method for predicting protein metal ion binding sites, which is applicable to various metal types and complex coordination scenarios and has good application prospects.
[0017] (2) Evaluation results on multiple public test sets such as Metal3D and TEMSP show that the key indicators such as precision, recall and F1 score of this method reach or exceed the level of existing mainstream methods. For example, in the TEMSP zinc binding site test set, the F1 score of this method reaches 0.967, the recall rate is 0.978 and the precision is 0.957. At the same time, it has a strong ability to identify multinucleus sites. In the test cases such as binucleate zinc peptidase PepV, the prediction error of metal spacing is less than 0.2 Å.
[0018] (3) This method constructs a potential function based on the principles of statistical mechanics, and the prediction results are physically interpretable and can be directly traced back to the interactions at the atomic level, which facilitates mechanism analysis; this method also covers Ca 2+ Mg 2+ Na + K + Alkali / alkaline earth metals and Zn 2+ Mn 2+ Fe 2+ / Fe 3+ Transition metals, with good versatility, can simultaneously output the three-dimensional coordinates of metal ions and information on coordination residues, and can be directly used for downstream structural analysis and functional studies. Attached Figure Description
[0019] Figure 1 : A schematic diagram of the workflow of the MetalKB method. (a) Statistical potential function construction stage: Extract donor atom coordination information from the metalloprotein database and construct a statistical potential function based on the distance distribution; (b) Prediction stage: Based on the donor atom map, candidate clusters are identified through complete subgraph retrieval, and the final site is output after energy optimization and redundancy removal.
[0020] Figure 2 Overview of the prediction phase of the MetalKB method. It shows four main steps: input structure, candidate site detection, energy scoring, and output metal-binding sites.
[0021] Figure 3 : Schematic diagram of candidate site sampling based on complete subgraph retrieval in the MetalKB method. (a) Protein structure is converted into donor atom graph, and edges are established when the interatomic spacing is within the statistical coordination range; (b, c) Two typical coordination modes of carboxyl oxygen atoms.
[0022] Figure 4 Examples of multinuclear metal-binding site prediction. (a) Prediction results for the binuclear zinc peptidase PepV; (b) Prediction results for the trinuclear zinc cluster in human H3K9 methyltransferase. Gold spheres represent experimentally determined locations, and red spheres represent locations predicted by this method. Detailed Implementation
[0023] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.
[0024] like Figure 1 As shown, the method of this invention mainly consists of two stages: a statistical potential function construction stage and a prediction stage. In the statistical potential function construction stage, sample data is extracted from a high-resolution metalloprotein structure database. Based on the coordination preference and ionic radius characteristics of metal ions, they are classified into different categories, and the distance distribution between each type of metal ion and the protein donor atom is statistically analyzed. A statistical potential function is constructed based on the inverse Boltzmann relation, and combined with the van der Waals potential function to form a mixed potential model, which is used to quantitatively evaluate the interaction energy between metal ions and protein atoms.
[0025] like Figure 2 As shown, the prediction stage of this invention can be summarized into four core steps: first, the three-dimensional structure of the target protein is input; then, candidate site detection is performed, and possible coordination clusters are identified through a complete subgraph retrieval algorithm; next, the candidate sites are scored using a constructed statistical potential function; finally, after redundancy removal, high-confidence metal-binding sites are output. This process provides a clear framework for subsequent detailed sampling and optimization.
[0026] In the specific implementation of the prediction phase, such as Figure 3 As shown in 'a', this invention extracts the donor atom set of the target protein, constructs an undirected graph structure based on the spatial distance between donor atoms, and establishes connecting edges when the distance between two atoms is within a statistically determined reasonable coordination range. A maximal cluster detection algorithm (i.e., complete subgraph retrieval) is used to identify complete subgraphs containing at least three donor atoms as candidate coordination clusters, and subgraphs completely contained by larger subgraphs are deleted to reduce redundancy. Figure 3 As shown in b and c, this invention can effectively distinguish different coordination modes of carboxyl oxygen atoms (such as monodentate and bidentate coordination) by introducing virtual nodes, thus avoiding the limitations of traditional grid sampling methods.
[0027] This invention proposes an energy optimization and redundancy removal process. For each candidate cluster, the geometric center of the donor atom is calculated as the initial coordinates of the metal ion, and local grid energy optimization is performed in its neighborhood space, selecting the position with the lowest energy as the final predicted coordinates. All prediction results are sorted by energy, and redundancy is removed based on the spatial distance between metal ions, retaining sites with lower energy. Finally, the three-dimensional coordinates of the predicted metal ion and the information of the residues coordinated with it are output.
[0028] This invention proposes a universal prediction framework applicable to multiple metal types. Through metal classification and the construction of specific potential functions, this method is applicable to various metal types, including transition metals, alkali metals, and alkaline earth metals, and can effectively identify multinucleated metal clusters and bridging coordination modes. Figure 4 This paper presents a typical example of the method in predicting multinucleated metal sites. The gold sphere represents the experimentally determined location, and the red sphere represents the predicted location by the method. It can be seen that the predicted coordinates are in high agreement with the experimental values.
[0029] Example 1 The specific implementation of this invention (MetalKB) is as follows: I. Construction of Statistical Potential Function High-resolution metalloprotein structures (resolution better than 2.5 Å) were obtained from the MESPEUS database and redundancy was removed using a 30% sequence homology threshold. Metals were classified into four classes based on their coordination characteristics (e.g., Ca). 2+ K + Class, Mg 2+ (Class, transition metal class), and construct potential functions independently for each class.
[0030] The radial distribution of donor atoms within a 10 Å radius around each type of metal ion was statistically analyzed, with a bin thickness of 0.1 Å. A knowledge statistical potential was constructed based on the inverse Boltzmann relation and fused with the Lennard-Jones 12-6 van der Waals potential to obtain a mixed potential function.
[0031] The formula for calculating the statistical potential function is as follows:
[0032] in, Boltzmann's constant, Absolute temperature Set it to 1 to simplify the calculation. Indicates that the radius is from arrive Inside, type is Protein atoms and metal types Number density of atomic pairs This indicates that in the reference sphere (with radius ), The average number density of corresponding atomic pairs within 10 Å was used as a comparison benchmark.
[0033] The specific calculation formulas for the two items are as follows:
[0034]
[0035] in, This represents the total number of samples, which is the number of all metal ion binding sites in the dataset. For the first The distance among the sites is place Number of atomic pairs For all at this site The total number of atomic pairs The volume of the reference sphere. Shell thickness. Set it to 0.1 Å.
[0036] To further improve the stability and physical plausibility of the potential function, we initially extracted... Smoothing was performed to obtain the mixed statistical potential. The specific method is as follows:
[0037] in, The original potential energy function is... Lennard-Jones 12-6 Van der Waals.
[0038] II. Sampling of candidate metal sites With Zn 2+ For example, input the target protein PDB file and extract Cys-SG, His-ND1 / NE2, Asp-OD1 / OD2, and Glu-OE1 / OE2 as the donor atom set. Construct an undirected graph based on the interatomic distance: if the distance between two atoms is within a reasonable coordination range (Zn 2+ If the distance is 2.4-5.2 Å, then edges are established. A complete subgraph retrieval algorithm (such as the Bron-Kerbosch maximal clique enumeration algorithm) is used to identify all complete subgraphs in the graph. Cliques with ≥3 nodes are retained as candidate coordination clusters, and subclusters completely contained by larger cliques are deleted. The geometric center of the donor atom is calculated as the initial coordinates of the metal.
[0039] III. Energy Optimization and Redundancy Removal In the local space around the initial coordinates (e.g., 2.5 Å) 3 A grid search is performed with an appropriate step size (e.g., 0.25 Å). The energy of each point is calculated using the constructed statistical potential function, and the point with the lowest energy is taken as the final predicted coordinate. All predicted sites are sorted in ascending order of energy. If the distance between two predicted metals is less than a preset threshold (e.g., 2.5 Å), the one with the lower energy is retained. Candidate sites are further screened according to an energy cutoff threshold (e.g., -1.7), and predictions with energies higher than the threshold are removed to reduce false positives. The final output is the three-dimensional coordinates of the metal ions that meet the energy conditions and the coordinating residues within a distance of 4 Å. The energy of each point is calculated using the atomic pair potential function. The summation is as follows:
[0040] in, The distance between the metal ion and the protein atom is 6 Å, which is taken as the range in this study. For the aforementioned distance is Metal ions With atom type The mixed statistical potential function of the interactions between protein atoms.
[0041] IV. Case Verification In the binuclear zinc peptidase PepV (PDB ID: 1LFW) assay, MetalKB successfully identified two zinc ions and the bridging ligand Asp119, with an average deviation of 0.18 Å between predicted and experimental coordinates. In the human H3K9 methyltransferase (PDB ID: 2R3A) assay, it accurately predicted the trinuclear zinc cluster in the pre-SET region and the mononuclear zinc site in the post-SET region, with coordination patterns highly consistent with the crystal structure.
[0042] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for predicting protein metal ion binding sites based on complete subgraph retrieval and statistical potential function, characterized in that, Includes the following steps: (1) Obtain the three-dimensional structure of the metal ion protein complex, classify the protein complex according to the coordination characteristics of the metal ion, and construct a statistical potential function for each type of protein complex; (2) Extract the donor atom set of the target protein, construct an undirected graph based on the spatial distance between donor atoms, and use a complete subgraph retrieval algorithm to identify all complete subgraphs in the undirected graph. Each complete subgraph is a candidate coordination cluster, and the geometric center of the donor atom in each candidate coordination cluster is used as the metal initial coordinate of the candidate coordination cluster. (3) The candidate coordination clusters are identified by a complete subgraph retrieval algorithm. Local grid search is performed around the geometric center of the candidate coordination clusters. Energy optimization is performed by combining the statistical potential function in step (1). Redundant predictions are removed based on the energy threshold and spatial distance. The three-dimensional coordinates of the metal ions of the target protein and the residue information coordinated with the metal ions are output.
2. The method for predicting protein metal ion binding sites based on complete subgraph retrieval and statistical potential function as described in claim 1, characterized in that, In step (1), the construction of the statistical potential function specifically involves: statistically analyzing the radial distribution of donor atoms within 10 Å around the metal ions of each type of protein complex, and constructing the statistical potential based on the inverse Boltzmann relation.
3. The method for predicting protein metal ion binding sites based on complete subgraph retrieval and statistical potential function as described in claim 2, characterized in that, In step (1), the statistical potential function is constructed using the following formula: in, Boltzmann's constant, Absolute temperature Indicates the type of metal. Indicates the type of protein atoms. Indicates the distance between metal ions and protein atoms; Indicates that the radius is from arrive Inside, type is Protein atoms and metal types Number density of atomic pairs The thickness of the radially statistical spherical shell is represented by 0.1 Å; Indicates a radius of The average number density of corresponding atomic pairs within a reference sphere of 10 Å; This represents the statistical potential of a metal-ion protein complex based on the inverse Boltzmann relation.
4. The method for predicting protein metal ion binding sites based on complete subgraph retrieval and statistical potential function as described in claim 1, characterized in that, In step (1), the classification specifically refers to: Ca 2+ Class, K + Class, Mg 2+ Classes of transition metal ions and transition metal ions.
5. The method for predicting protein metal ion binding sites based on complete subgraph retrieval and statistical potential function as described in claim 4, characterized in that, The transition metal ions include zinc ions, copper ions, cobalt ions, nickel ions, ferric ions, ferrous ions, or manganese ions.
6. The method for predicting protein metal ion binding sites based on complete subgraph retrieval and statistical potential function as described in claim 1, characterized in that, In step (2), after identifying all complete subgraphs in the undirected graph using the complete subgraph retrieval algorithm, the cliques with ≥3 nodes are retained as candidate coordination clusters.
7. The method for predicting protein metal ion binding sites based on complete subgraph retrieval and statistical potential function as described in claim 1, characterized in that, In step (2), after identifying all complete subgraphs in the undirected graph using the complete subgraph retrieval algorithm, subgraphs completely contained by larger subgraphs are deleted to reduce redundancy.
8. The method for predicting protein metal ion binding sites based on complete subgraph retrieval and statistical potential function as described in claim 1, characterized in that, In step (3), the metal ion is coordinated by residues within 4 Å of the metal ion.