Method and system for automatic identification of protein bimetallic sites based on three-dimensional spatial features and candidate set construction, and readable storage medium
Patent Information
- Application Number
- CN202610638362.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-11
- Publication Date
- 2026-08-04
AI Technical Summary
若缺乏针对同一残基金属簇与高密度金属邻域的抑制机制,筛选流程将被团簇结构主导,难以在全库尺度上稳定运行,也难以为后续的先验统计与机器学习建模提供干净、可控的数据基础
[0043] In summary, the present invention has the following advantages: the method provided by the present invention can automatically and traceably construct bimetallic candidate sites and provide interpretable structural information on large-scale real analysis and computational prediction protein structure data, and is applicable to scenarios such as high-throughput mining of bimetallic enzyme sites, functional annotation and protein design screening of metalloproteins, and rational design of industrial biocatalysts.
Smart Images

Figure CN122511344A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of structural bioinformatics, computational structural biology, and automated mining of biological big data. Specifically, it relates to a method and system for the automatic identification of bimetallic sites, candidate pair generation, pseudo-candidate suppression filtering, and structured characterization of bridging / coordination environments for large-scale protein three-dimensional structure data (such as mmCIF, PDB, and their compressed formats). This invention aims to provide standardized underlying data support for downstream artificial intelligence scoring models, rational metalloenzyme design, precise drug target discovery, and industrial biocatalyst modification. Background Technology
[0002] Bimetallic proteins mediate complex and industrially valuable catalytic reactions in nature, such as nitrogen fixation, carbon fixation, oxygen reduction, and antibiotic degradation. Their core bimetallic sites are widely found in hydrolases, oxidoreductases, and metalloproteinases. These sites achieve synergistic polarization and efficient catalysis of the substrate through the close spatial proximity and electronic coupling of two metal ions. Accurately identifying and annotating real bimetallic catalytic centers from massive amounts of protein structural data has become a core challenge for green chemistry and biomedicine.
[0003] Existing methods for identifying and annotating metallic sites mainly rely on manual annotation of structural databases, literature review, or rule-based selection based on a limited number of features. For monometallic sites, these methods can meet functional annotation requirements to some extent. However, for bimetallic properties that are highly dependent on microenvironment synergy, manual annotation or simple distance thresholds are often insufficient to fully cover the diversity of real structures, let alone support large-scale, reproducible, automated selection.
[0004] The structural characteristics of bimetallic sites are reflected not only in the metal-metal distance but also in the bridging mode and local coordination environment. For example, two metals may form a cooperative bridge through different donor atoms of the same residue, or an atomic-level bridge through water molecules or other small molecules. Furthermore, the coordination shell composition, donor atom type, and geometric constraints of the two metals jointly determine the true chemical semantics of the site. Traditional methods that only output results showing that the two metals are close to each other often fail to further distinguish between true bimetallic cooperative sites and accidentally adjacent independent metals, thus reducing the reliability and interpretability of the screening results.
[0005] On the other hand, automated screening for large-scale structural databases encounters significant spurious candidate problems. Common inorganic clusters, metallic mineral cores, or multimetallic cofactors (such as metal clusters, iron-sulfur cluster-related structural units, etc.) in structures generate a large number of metal pair combinations that satisfy the distance window, leading to a significant increase in the number of candidates and misclassifying non-target structures as bimetallic sites. Without a mechanism to inhibit metal clusters and high-density metal neighborhoods of the same residue, the screening process will be dominated by cluster structures, making it difficult to operate stably at the whole-database scale and providing a clean and controllable data foundation for subsequent prior statistics and machine learning modeling.
[0006] Furthermore, existing toolchains often scatter processes such as candidate generation, bridging / coordination environment extraction, filtering rules, and structured output across different scripts or software, lacking a unified traceable identification system and standardized output format. This makes it difficult to reproduce the results and integrate them with subsequent scoring models. When faced with different input formats such as mmCIF and PDB, the differences in atomic indexing, residue identification, and structure analysis further amplify the difficulty of data governance and engineering deployment.
[0007] Therefore, it is necessary to propose a closed-loop method and system capable of automatically performing metal screening, bimetallic candidate generation, pseudo-candidate suppression filtering, bridging and coordination environment extraction, and structured output on ultra-large-scale protein three-dimensional structure data. While ensuring computational efficiency and scalability, this method should improve the structural semantic expression ability and screening reliability of bimetallic site candidates, providing core technical support for functional annotation of metalloenzymes, discovery of bimetallic active sites and protein design and screening, and biomanufacturing of high-value-added chemicals. Summary of the Invention
[0008] To achieve the aforementioned objectives, this invention proposes an automatic identification and candidate construction method, system, and readable storage medium for bimetallic sites in large-scale protein three-dimensional structure data. This technical solution revolves around a closed-loop process encompassing structural data acquisition and quality control, metal atom extraction, bimetallic candidate enumeration and pseudo-candidate filtering, bridging and coordination environment extraction, and structured result output. Through multi-level filtering and structural semantic extraction mechanisms, it achieves automated and scalable identification of bimetallic candidate sites.
[0009] The automatic identification and candidate set construction method for protein bimetallic sites based on three-dimensional spatial features provided by this invention is achieved through the following technical solutions: A method for automatic identification and candidate set construction of protein bimetallic sites based on three-dimensional spatial features includes the following steps: S1. Structural data acquisition: Acquire the three-dimensional structural data of the protein to be processed, and read the atomic element types and three-dimensional coordinate information from the atomic table of the three-dimensional structural data; S2. Structural quality control: At least one of the following is used to perform structural quality control on the protein three-dimensional structure data obtained in S1: decompression validity check, atom table field integrity check, structural parsing validity check, and multi-model structure exclusion. S3. Metal Atom Extraction: Based on a preset set of metal elements, a set of metal atoms is obtained by filtering from the atomic table of the three-dimensional structure data, and the structure identifier, atom index, chain identifier, residue name, residue number and three-dimensional coordinates are recorded for each metal atom; S4. Bimetallic Candidate Enumeration: Within the same structure, pairwise combine the sets of metal atoms obtained in S3 to calculate the arbitrary metal-metal distance d. mm And generate bimetallic candidate pairs under the condition of satisfying the preset distance window; S5. Multi-level pseudo-candidate suppression filtering: Pseudo-candidate suppression filtering is performed on the bimetallic candidate pairs in S4. The pseudo-candidate suppression filtering includes at least one of the following: blacklist filtering, same residue metal cluster filtering, and high-density metal neighborhood filtering, to obtain a filtered set of bimetallic candidate pairs. S6. Structured Result Output: For each bimetallic candidate pair after S5 multi-level pseudo-candidate suppression filtering, identify the bridging information and coordination environment information related to the two metals within a preset spatial neighborhood, generate a structured description of the candidate site and output it.
[0010] Preferably, the protein three-dimensional structure data includes at least one of mmCIF format, PDB format, and compressed mmCIF format.
[0011] Preferably, the set of metal atoms in S3 is obtained by screening a preset set of metal elements, which includes at least transition metal elements; and further supports the exclusion of metal atoms based on an element blacklist and / or a residue name blacklist.
[0012] Preferably, the preset distance window condition in S4 is: when the metal-to-metal distance d mm When the conditions between the first threshold and the second threshold are met, the bimetallic candidate pair is generated, wherein the first threshold is the lower limit threshold of the metal-metal distance, the second threshold is the upper limit threshold of the metal-metal distance, and the first threshold is less than the second threshold.
[0013] More preferably, the first threshold is 2.0 Å; the second threshold is 7.0 Å.
[0014] Preferably, the blacklist filtering in S5 is as follows: when the element type of any metal in the candidate pair belongs to the preset element blacklist, and / or the residue name of any metal belongs to the preset residue name blacklist, the candidate pair is removed; wherein, the preset element blacklist is used to exclude non-target metal elements that are prone to forming ion clusters, high-density metal environments, or inorganic cluster interference.
[0015] More preferably, the preset element blacklist includes alkali metal elements and / or alkaline earth metal elements.
[0016] Preferably, the same-residue metal cluster filtering in S5 is as follows: the number of metal atoms in a residue is calculated using the residue identifier bond as the statistical unit; when the number of metal atoms is not less than the threshold K, the residue is determined to be a metal cluster residue; when the two metal atoms of a candidate pair belong to the same metal cluster residue, the candidate pair is eliminated.
[0017] Preferably, the same-residue metal cluster filtering only includes residues whose names do not belong to the standard amino acid set in the metal atom count, so as to reduce the risk of false rejection of the protein's normal coordination environment.
[0018] Preferably, the high-density metal neighborhood filtering in S5 involves: calculating the number of metal neighbors within a preset radius R for each metal atom in the structure as the neighborhood degree; and removing the candidate pair when the neighborhood degree of either endpoint is not less than the threshold D; wherein the preset radius R is the spatial neighborhood radius used to count the number of metal neighbors, and the threshold D is the neighborhood degree threshold used to determine the high-density metal neighborhood.
[0019] More preferably, the preset radius R is 4.0 Å and the threshold D is 8.
[0020] Preferably, the pseudo-candidate suppression filtering further includes the following method: using the midpoint of the two metal coordinates of the candidate pair as the query point, and statistically analyzing the data within a preset radius R. m The number of other metal atoms besides the two metals in the candidate pair is considered; if the number of other metal atoms is greater than a threshold M, the candidate pair is discarded; wherein, R m M represents the spatial query radius of the local metallic environment at the midpoint, where M is the threshold number of other metal atoms allowed to exist within the spatial query radius.
[0021] More preferably, the R m The value is 4.0 Å, and M is 1.
[0022] Preferably, the pseudo-candidate suppression filtering further includes the following method: taking each metal of the candidate pair as the center, statistically analyzing its position within a preset density radius R. d The number of metals within the range, and when the number of any metal exceeds a threshold N, the candidate pair is discarded; wherein, R dThe statistical radius is the local metal density, and N is the threshold number of metal atoms allowed to exist within the statistical radius.
[0023] More preferably, the R d The value is 10.0 Å, and N is 10.
[0024] Preferably, the bridging information in S6 includes: for the two metals M1 and M2 of a candidate pair, traversing the candidate atom α in the local atomic table of the candidate site, when the candidate atom α simultaneously satisfies dist(α, M1) ≤ r c And dist(α, M2) ≤ r c When α is determined to be a bridging atom, α is a candidate atom participating in the bridging determination, and r is... c The spatial neighborhood radius identified by the bridge connection.
[0025] More preferably, the r c It is 2.8 Å.
[0026] Preferably, the bridging atoms are classified by source or chemical type, and the classification includes at least: bridging formed by amino acid side chain donors and / or bridging formed by water molecules.
[0027] Preferably, the coordination environment information in S6 includes: collecting data located at r c The first and second sets of coordinating atoms within the radius are sorted according to their distance from the corresponding metal, and then truncated to retain the first K. c Coordinating atoms are used to generate structured coordination shell information; wherein, the K c The threshold for the maximum number of coordinating atoms that each metal can retain in a structured coordination shell.
[0028] More preferably, the K c It is 32.
[0029] Preferably, the donor atoms in the coordination shell are statistically analyzed, and the donor atoms include at least one of the following: side-chain oxygen donor, histidine nitrogen donor, cysteine sulfur donor, main-chain carbonyl oxygen donor, and water molecule oxygen donor.
[0030] Preferably, the number of protein source donors is obtained based on the composition statistics, and when the number of protein source donors for any candidate metal is lower than a threshold T... p The candidate pair is removed at that time; wherein, the T p This is the lower limit threshold for the number of protein source donors used to determine whether the local coordination environment of a candidate site is reasonable.
[0031] More preferably, the T p The value is 2.
[0032] Preferably, in step S6, the coordinating atoms are mapped to a set of residues based on the two metal coordination shells and the intersection is calculated. When the intersection is not empty, it is determined that there is a residue bridge; otherwise, it is determined that there is no residue bridge.
[0033] Preferably, when residue bridging exists, the residue bridging is further subdivided into at least one of carboxylic acid residue bridging, histidine residue bridging, cysteine residue bridging, main chain carbonyl oxygen bridging, and other residue bridging; and when there are ≥2 categories of residue bridging, the final bridging category is output according to a preset priority.
[0034] Preferably, when it is determined that there is no residue bridging, the candidate sites are further subdivided based on the metal-metal distance and the sparsity of the coordination shells on both sides. The subdivision is divided into at least one of tightly coupled candidates, environmentally coupled candidates, weakly coupled candidates, long-distance independent candidates, and noise candidates.
[0035] Preferably, the subclass division is segmented using at least three distance thresholds, including a tight coupling threshold, a coupling upper bound threshold, and a long distance threshold.
[0036] More preferably, the tight coupling threshold is 3.10 Å, the upper coupling threshold is 4.20 Å, and the long-distance threshold is 4.60 Å.
[0037] Preferably, the sparsity of the coordination shell is determined by the minimum value min(n1) of the number of atoms in the first coordination shell and the number of atoms in the second coordination shell. 1, n2), and based on min(n) 1, The subclass is determined by the joint constraint of n2 and metal-metal distance.
[0038] Preferably, the structured candidate site description in S6 is output by combining table fields and nested structure data. The nested structure data includes at least a list of bridging atoms and a list of two metal coordination shells, and each record includes an atom index, residue information and distance information.
[0039] Preferably, the atomic index information in the structured candidate site description preferentially uses atomic sequence numbers under PDB input and atomic table row indexes or their mapping indexes under mmCIF input to ensure traceability and consistency of candidate sites across structural formats; when the atomic sequence number is missing, an incremental index is used as a fallback index.
[0040] Preferably, parallel batch processing is performed on multi-structure inputs at the structural granularity, and a list of successful processing, a list of failed processing, and statistical information on the reasons for failure are output.
[0041] The system provided by this invention, which utilizes a method for automatic identification and candidate set construction of protein bimetallic sites based on three-dimensional spatial features, is achieved through the following technical solution: A system built using a method for automatic identification and candidate set construction of protein bimetallic sites based on three-dimensional spatial features includes a structure data acquisition module, a structure quality control module, a metal atom extraction module, a bimetallic candidate enumeration module, a pseudo-candidate suppression and filtering module, and a structured result output module. The structural data acquisition module is used to execute the automatic identification and candidate set construction method S1 for protein bimetallic sites based on three-dimensional spatial features, acquire the three-dimensional structural data of the protein to be processed, and read the atomic table information from the three-dimensional structural data. The structure quality control module is used to execute the method S2 for automatic identification and candidate set construction of protein bimetallic sites based on three-dimensional spatial features, and to perform structure quality control on the protein three-dimensional structure data. The metal atom extraction module is used to execute the automatic identification and candidate set construction method S3 for protein bimetallic sites based on three-dimensional spatial features, and to screen the set of metal atoms from the atomic table and establish corresponding index information. The bimetallic candidate enumeration module is used to execute the automatic identification and candidate set construction method S4 of protein bimetallic sites based on three-dimensional spatial features, generate bimetallic candidate pairs and calculate metal-metal distances. The pseudo-candidate suppression and filtering module is used to execute the automatic identification and candidate set construction method S5 for protein bimetallic sites based on three-dimensional spatial features, and to eliminate pseudo-candidate pairs caused by inorganic clusters or high-density metal environments. The structured result output module is used to execute the method S6 for automatic identification and candidate set construction of protein bimetallic sites based on three-dimensional spatial features, identify bridging information and coordination environment information, and generate and output structured candidate site descriptions.
[0042] A readable storage medium storing a computer program, which, when executed by a processor, implements a method for preparing an automatic identification and candidate set construction method for protein bimetallic sites based on three-dimensional spatial features.
[0043] In summary, the present invention has the following advantages: the method provided by the present invention can automatically and traceably construct bimetallic candidate sites and provide interpretable structural information on large-scale real analysis and computational prediction protein structure data, and is applicable to scenarios such as high-throughput mining of bimetallic enzyme sites, functional annotation and protein design screening of metalloproteins, and rational design of industrial biocatalysts. Attached Figure Description
[0044] Figure 1This is the overall flowchart of the automatic identification of bimetallic sites in this invention.
[0045] Figure 2 This is a flowchart of the pseudo-candidate suppression filtering module of the present invention.
[0046] Figure 3 This is a flowchart of the bridging identification and residue-free bridging subclass classification process of the present invention.
[0047] Figure 4 This is a schematic diagram showing the distribution of metal-metal distances corresponding to different residue bridging categories.
[0048] Figure 5 This diagram illustrates the subclassing of candidates without residue bridging and the relationship between metal-metal distances. Detailed Implementation
[0049] To further understand the inventiveness and technical advancements of this invention, the preferred embodiments of this invention will be discussed in detail below with reference to examples and comparative examples.
[0050] A method for automatic identification and candidate set construction of protein bimetallic sites based on three-dimensional spatial features includes the following steps: S1. Structural Data Acquisition: Acquire the three-dimensional structural data of the protein to be processed. The three-dimensional structural data of the protein includes at least one of the compressed formats of mmCIF, PDB, and mmCIF. Read the atomic element types and three-dimensional coordinate information from the atomic table of the three-dimensional structural data. S2. Structural quality control: At least one of the following is used to perform structural quality control on the protein three-dimensional structure data obtained in S1: decompression validity check, atom table field integrity check, structural parsing validity check, and multi-model structure exclusion. S3. Metal Atom Extraction: Based on a pre-defined set of metal elements, a set of metal atoms is obtained by filtering from the atomic table of the three-dimensional structure data, and the structure identifier, atom index, chain identifier, residue name, residue number and three-dimensional coordinates are recorded for each metal atom; The set of metal atoms obtained in S3 is obtained by screening a preset set of metal elements, which must contain at least transition metal elements; and further supports the exclusion of metal atoms based on an element blacklist and / or a residue name blacklist. S4. Bimetallic candidate enumeration: Within the same structure, pairwise combinations of the metal atom sets obtained in S3 are performed to calculate arbitrary metal-metal distances dmm, and bimetallic candidate pairs are generated under the condition of satisfying the preset distance window. The preset distance window condition is: when the metal-to-metal distance d mmBimetallic candidate pairs are generated when the conditions between the first threshold and the second threshold are met, wherein the first threshold is the lower limit threshold of the metal-metal distance, the second threshold is the upper limit threshold of the metal-metal distance, the first threshold is less than the second threshold, preferably, the first threshold is 2.0 Å; the second threshold is 7.0 Å; S5. Multi-level pseudo-candidate suppression filtering: Pseudo-candidate suppression filtering is performed on the bimetallic candidate pairs in S4. The pseudo-candidate suppression filtering includes at least one of the following: blacklist filtering, same residue metal cluster filtering, and high-density metal neighborhood filtering, to obtain a filtered set of bimetallic candidate pairs. Blacklist filtering in S5: When the element type of any metal in a candidate pair belongs to a preset element blacklist, and / or the residue name of any metal belongs to a preset residue name blacklist, the candidate pair is removed; wherein, the preset element blacklist is used to exclude non-target metal elements that are prone to forming ion clusters, high-density metal environments or inorganic cluster interference; preferably, the preset element blacklist includes alkali metal elements and / or alkaline earth metal elements. Same-residue metal cluster filtering in S5: The number of metal atoms in a residue is calculated using the residue identifier bond as the statistical unit. When the number of metal atoms is not less than the threshold K, the residue is identified as a metal cluster residue. When the two metal atoms of a candidate pair belong to the same metal cluster residue, the candidate pair is removed. Same-residue metal cluster filtering only includes residues whose residue names do not belong to the standard amino acid set in the metal atom count, so as to reduce the risk of false removal of the protein's normal coordination environment. High-density metal neighborhood filtering in S5: For each metal atom in the structure, the number of its metal neighbors within a preset radius R is calculated as the neighborhood degree. When the neighborhood degree of either endpoint of a candidate pair is not less than the threshold D, the candidate pair is eliminated. Here, the preset radius R is the spatial neighborhood radius used to count the number of metal neighbors, and the threshold D is the neighborhood degree threshold used to determine the high-density metal neighborhood. Preferably, the preset radius R is 4.0 Å, and the threshold D is 8. False candidate suppression filtering also includes the following method: using the midpoint of the two metal coordinates of the candidate pair as the query point, and statistically analyzing the results within a preset radius R. m The number of other metal atoms besides the two metals in the candidate pair is considered; if the number of other metal atoms exceeds a threshold M, the candidate pair is discarded. Where R... m R is the spatial query radius of the local metallic environment at the midpoint, M is the threshold for the number of other metal atoms allowed to exist within the spatial query radius, and preferably, R m The value is 4.0 Å, and M is 1; The pseudo-candidate suppression filtering also includes the following method: taking each metal of the candidate pair as the center, counting the number of metals within a preset density radius Rd, and removing the candidate pair when the number of any metal exceeds a threshold N; where Rd is the statistical radius of the local metal density, and N is the threshold for the number of metal atoms allowed to exist within the statistical radius; preferably, Rd is 10.0 Å, and N is 10; S6. Structured Result Output: For each bimetallic candidate pair after S5 multi-level pseudo-candidate suppression filtering, identify the bridging information and coordination environment information related to the two metals within the preset spatial neighborhood, generate a structured description of the candidate site and output it. The bridging information includes: for two metal pairs M1 and M2, traversing the candidate atom α in the local atomic table of the candidate site, when candidate atom α simultaneously satisfies dist(α, M1) ≤ r c And dist(α, M2) ≤ r c When α is determined to be a bridging atom, r is considered a candidate atom in the bridging determination. c For the spatial neighborhood radius of the bridge connection, preferably, r c It is 2.8 Å; The coordination environment information includes: collecting the first and second sets of coordination atoms located within the rc radius, sorting them according to their distance from the corresponding metal, and truncating and retaining the first Kc coordination atoms to generate structured coordination shell information. Kc is the threshold for the maximum number of coordination atoms retained by each metal in the structured coordination shell. Preferably, Kc is 32. The composition statistics of the donor atoms in the coordination shell are performed. The donor atoms include at least one of the following: side chain oxygen donor, histidine nitrogen donor, cysteine sulfur donor, main chain carbonyl oxygen donor, and water molecule oxygen donor. The bridging atoms are classified by their source or chemical type, and the classification includes at least: bridging formed by amino acid side chain donors and / or bridging formed by water molecules; The number of protein source donors is obtained based on compositional statistics. When the number of protein source donors for any candidate metal is lower than the threshold T... p The candidate pair is removed at that time; where T p As a threshold for the minimum number of protein source donors used to determine whether the local coordination environment of a candidate site is reasonable, preferably, T p It is 2; Based on the two metal coordination shells, the coordinating atoms are mapped to the set of residues and the intersection is calculated. When the intersection is not empty, it is determined that there is residue bridging; otherwise, it is determined that there is no residue bridging. When residue bridging exists, it is further subdivided into at least one of carboxylic acid residue bridging, histidine residue bridging, cysteine residue bridging, main chain carbonyl oxygen bridging, and other residue bridging; and when there are ≥2 types of residue bridging, the final bridging category is output according to the preset priority. When it is determined that there is no residue bridging, the candidate sites are further subdivided based on the metal-metal distance and the sparsity of the coordination shells on both sides. The sub-classification is divided into at least one of tightly coupled candidates, environmentally coupled candidates, weakly coupled candidates, distant independent candidates, and noise candidates. The sparsity of the coordination shell is determined by the minimum value (n1) of the number of atoms in the first coordination shell and the number of atoms in the second coordination shell. 1, n2), and based on min(n) 1, n2) The subclass is determined by the joint constraint of the metal-metal distance; The subclass division is segmented using at least three distance thresholds, including a tight coupling threshold, a coupling upper bound threshold, and a far distance threshold. Preferably, the tight coupling threshold is 3.10 Å, the coupling upper bound threshold is 4.20 Å, and the far distance threshold is 4.60 Å. The structured candidate site description is output by combining tabular fields and nested structured data. The nested structured data includes at least a list of bridging atoms and a list of two metal coordination shells, and each record includes an atom index, residue information and distance information. In the structured candidate site description, the atomic index information prioritizes the atomic sequence number under PDB input and uses the atomic table row index or its mapping index under mmCIF input to ensure traceability and consistency of candidate sites across structural formats; when the atomic sequence number is missing, an incremental index is used as a fallback index. Perform parallel batch processing on multi-structure inputs at the structure granularity, and output a list of successful processing, a list of failed processing, and statistical information on the reasons for failure.
[0051] A system built using a method for automatic identification and candidate set construction of protein bimetallic sites based on three-dimensional spatial features includes a structural data acquisition module, a structural quality control module, a metal atom extraction module, a bimetallic candidate enumeration module, a pseudo-candidate suppression and filtering module, and a structured result output module.
[0052] The structural data acquisition module is used to execute the automatic identification and candidate set construction method S1 of protein bimetallic sites based on three-dimensional spatial features, acquire the three-dimensional structural data of the protein to be processed, and read the atomic table information from the three-dimensional structural data; The structure quality control module is used to execute the method S2 for automatic identification and candidate set construction of protein bimetallic sites based on three-dimensional spatial features, and to perform structure quality control on the protein three-dimensional structure data; The metal atom extraction module is used to execute the automatic identification and candidate set construction method S3 for protein bimetallic sites based on three-dimensional spatial features, which filters the set of metal atoms from the atomic table and establishes corresponding index information; The bimetallic candidate enumeration module is used to execute the automatic identification and candidate set construction method S4 based on three-dimensional spatial features of protein bimetallic sites, generate bimetallic candidate pairs and calculate metal-metal distances; The pseudo-candidate suppression and filtering module is used to execute the automatic identification and candidate set construction method S5 for protein bimetallic sites based on three-dimensional spatial features, and to eliminate pseudo-candidate pairs caused by inorganic clusters or high-density metal environments. The structured results output module is used to execute the automatic identification and candidate set construction method S6 for protein bimetallic sites based on three-dimensional spatial features, identify bridging information and coordination environment information, and generate and output structured candidate site descriptions.
[0053] A readable storage medium storing a computer program that, when executed by a processor, implements a method for automatic identification and candidate set construction of protein bimetallic sites based on three-dimensional spatial features.
[0054] Example: A method for automatic identification and candidate set construction of protein bimetallic sites based on three-dimensional spatial features, including the following steps: S1. Structured Data Acquisition [Data Acquisition Script (s0_rcsb_new_mmcif_dataset.py)]: This script batch-retrieves PDB IDs matching the query criteria from the RCSBSearch API, removes existing IDs, randomly samples N IDs, and downloads the corresponding .cif.gz files using a multi-threaded approach. The execution flow of the data acquisition script includes the following steps: S1.1, read query_json (query JSON exported from RCSB web), force request_options.return_all_hits=True and remove paginate to prevent conflicts; S1.2, POST to the RCSB Search API, parse result_set.identifier to obtain the full set of PDB IDs; S1.3, read in exclude_ids, remove duplicates, randomly sample the remaining IDs and write out_ids; S1.4 If the process can be run online, use ThreadPoolExecutor to download the .cif.gz structure in a multi-threaded manner, supporting retries, temporary .part files, and failure reason tables; S2. Structural Quality Control [Quality Control (QC) and Valid Structure List (s1_qc_mmcif_gz_mp.py) and Multi-Model Structure Filtering in mmCIF.gz]: Quality Control (QC) and Valid Structure List for mmCIF.gz (s1_qc_mmcif_gz_mp.py): Performs QC on all .cif.gz files in the directory, filters out a list of usable structure IDs, and outputs qc_report, valid_structures, and failed_structures. The main workflow employs two-level strategies: quick_check and parse_check_decompress_then_biopython. The specific steps for quality control and valid structure list generation are as follows: S2.1, Quick check: Read the first max_bytes after gzip decompression, and exclude empty files, HTML error pages, and missing _atom_site structures; S2.2, Parsing check: For the fast_ok file, first completely decompress it to a temporary .cif file, and then call the parser to perform structure parsing verification; S2.3, Multi-process parallel processing: Use multiprocessing.Pool for batch processing and output qc_report.tsv (including fast_ok / parse_ok / reason / time consumption) and a list of valid / failed; Multi-model structure filtering (s1_filter_multimodel.py): Further excludes multi-model structures from the valid / failed list passed by QC, outputting a single_model valid list. The multi-model structure filtering detection logic is based on the number of unique values in _atom_site.pdbx_PDB_model_num, and does not rely on complex parsers. The specific steps are as follows: S2.4, for each .cif.gz: scan the loop_ block and locate the loop containing the _atom_site.* column; if the _atom_site.pdbx_PDB_model_num column does not exist, it is considered a single model; S2.5, if the column exists, count the number of unique values; if the number is greater than 1, it is determined to be a multi-model and excluded; S2.6 outputs three files: nmodels.tsv, single_model.valid_structures.txt, and multi_model.structures.txt; S3. Metal Atom Extraction [Extracting metal atoms from mmCIF atomic tables (s2_extract_metal_atoms_mp.py) + Extracting metal atoms from design PDB (s2_extract_metal_atoms_pdb_mp.py)]; Extracting metal atoms from the mmCIF atom table (s2_extract_metal_atoms_mp.py: Extracts metal atoms from the mmCIF _atom_site and generates a unified metal_atoms.tsv file for subsequent bimetal enumeration; its execution flow is as follows:) S3.1, Read mmCIF: Load _atom_site using MMCIF2Dict; if any required fields (type_symbol, coordinates, group_PDB) are missing, mark it as a failure; S3.2, Field fallback: chain_id / res_id / res_name / atom_name take auth_* first, label_* as a fallback; S3.3, line-by-line scanning of atom_site: elements belonging to METAL_ELEMENTS are retained; supports atomic-level element exclusion and residue name exclusion (used to remove interference from alkali metals, FeS4, etc.). S3.4, Multi-process output: Structure level expansion to atomic level record, number of metals in each structure, and failure reason record; Extracting metal atoms from a design PDB (s2_extract_metal_atoms_pdb_mp.py): When the input is not an RCSB mmCIF but a self-designed / generated PDB structure, it provides a metal_atoms_tsv that is fully compatible with S3. The script explicitly uses the PDB serial number for the atom_index first, otherwise increments as a fallback. The specific execution flow is as follows: S3.5, Traverse the PDB list and parse its structure.
[0055] S3.6, scan metal elements one residue / atom at a time (with the same set of loose METAL_ELEMENTS).
[0056] S3.7, the output columns are the same as those in S2, including the following property parameters: structure_id, atom_index, element, group_PDB, chain_id, res_name, res_id, atom_name, x, y, z; S4. Bimetallic Candidate Enumeration and Three-Type Filtering (s3_enumerate_dinuclear_mp.py): Grouping metal_atoms.tsv by structure, enumerating all metal pairs, performing filtering and distance window selection, and outputting dinuclear_candidates_*.tsv. Cluster residue and density degree are key pseudo-candidate suppression mechanisms, and the specific execution flow is as follows: S4.1, Parsing Input: The robust TSV parser skips bad lines and records badlines to avoid IndexError.
[0057] S4.2, Cluster Residue Filtering: Counts the number of metals in residue_key=(chain_id,res_id,res_name), and those ≥ cluster_k are considered as organic_cluster_residue; the default is restrict_non_aa=True, which only counts non-standard AA to avoid false positives.
[0058] S4.3, dense degree filtering: For each metal i, calculate deg(i), where deg(i) represents the number of metal neighbors within a preset radius R, where the preset radius R is 4.0 Å; if a candidate pair (i,j) satisfies max(deg(i),deg(j)) ≥D, then exclude the candidate pair, where the threshold D is 8; S4.4, Blacklist Filtering and Distance Window: First filter by element / residue blacklist, then retain candidate pairs by metal-metal distance window; wherein, when the metal-metal distance d_mm meets the condition between the first threshold and the second threshold, the candidate pair is retained, the first threshold is the lower limit threshold of metal-metal distance 2.0 Å, and the second threshold is the upper limit threshold of metal-metal distance 7.0 Å.
[0059] S4.5, Output: Records the statistics of kept_pairs and each filter class (blacklist / cluster / dense / distance). S5. The structured results output includes the extraction of bridging atoms and coordination shells of candidate sites (bridge_and_coord_mp.py), statistics and plotting of residue bridging categories based on coordination shells (analysis_residue_bridge_hetero.py), and further subdivision of no_res_bridge buckets (no_res_bridge_subclass.py). Bridge and coordination shell extraction for candidate sites (bridge_and_coord_mp.py): For each candidate pair in dinuclear_candidates.tsv, the system retrieves a list of bridging atoms that are close to both metals (bridge_atoms_json) and a list of coordination shells for each metal (coord1_json / coord2_json), along with metal cluster and density statistics, truncation information, and donor counts. The final output is dinuclear_sites_*.tsv. The specific execution flow is as follows: S5.1, collect the set of metal coordinates metals_xyz within the structure in advance for subsequent cluster / density statistics; S5.2, Metal Cluster Filtering: Using the midpoint between two metals as the center, calculate the preset radius R. m The number of other metal atoms besides the two metals in the candidate pair, the preset radius R m The value is 4.0 Å; when the number of other metal atoms is greater than a threshold M, the candidate pair is discarded, where the threshold M is 1; further, taking each metal in the candidate pair as the center, its value within a preset density radius R is statistically analyzed. d The amount of metal within, the preset density radius R d The threshold is 10.0 Å; when the quantity of any metal exceeds a threshold N, the candidate pair is discarded, where the threshold N is 10. S5.3, Bridged Atom Recognition: Traverse the candidate atom α in the local atom table of candidate sites. When candidate atom α simultaneously satisfies dist(α,M1)≤r c And dist(α,M2)≤r c When r is identified as a bridging atom, it is determined to be a bridging atom. c The value is 2.8 Å, and the bridge-res_set of bridging residues is statistically analyzed; S5.4, Coordination Shell Extraction: Collect the coordinate shells located at r c The first and second sets of coordinating atoms within the radius are sorted according to their distance from the corresponding metal, and then truncated to retain the first K. c One coordinating atom, the K c 32 (to control JSON size); S5.5. Output fields: Write n_bridge_atom / n_bridge_res, coord_total, donor count, etc., and write JSON fields to TSV; furthermore, when the number of protein source donors for any candidate metal is lower than the threshold T p The candidate pair is removed at that time, and the T p It is 2; Residue bridging category statistics and plotting based on coordination shell (analysis_residue_bridge_hetero.py): This reads dinuclear_sites_*.tsv, parses coord1_json / coord2_json, calculates the intersection of coordinating residues on both sides (common), and then determines the residue-bridge, further subdividing it into six categories: carbox, his, cys, bbO, other, and no_res. The classification priority is: carboxylate > his > cys > bbO > other > none. The specific execution flow is as follows: S5.6 parses the safe_load structure into JSON.
[0060] S5.7, build_res_to_atoms: Maps coordination atoms to residue_key.
[0061] S5.8, common_res = set(m1.keys())&set(m2.keys()), and calculate various bridging flags. S5.9, output dinuclear_sites_residue_bridge_hetero.tsv, and draw d_mm histograms by res_bridge_category, as well as comparison hiss_res_bridge / cys_res_bridge histograms; The `no_res_bridge` bucket is further subdivided (`no_res_bridge_subclass.py`): For candidates of `no_res_bridge` (where the intersection of the two coordinating residues is empty), subclassing is performed based on features such as `d_mm` and coordination shell sparsity to distinguish potentially real coupling pockets from distant independent metals / noise. The script clearly defines the criteria and output file for `no_res_bridge`, and the specific execution flow is as follows: S5.10 recalculates common_res (based on the residue identity of coord JSON) to determine no_res_bridge.
[0062] S5.11, if bridge_atoms_json exists but common_res is not, it is determined to be a truncation / labeling anomaly class such as water_atom_bridge or carbox_atom_bridge_anomaly; otherwise, it generates subclasses such as env_coupled_pocket / far_apart_independent by segmenting according to sparsity (n_coord1+n_coord2, min(n_coord1,n_coord2)) and d_mm threshold, where the tight coupling threshold is 3.10 Å, the upper bound of coupling threshold is 4.20 Å, and the far-distance threshold is 4.60 Å.
[0063] S5.12, Output: dinuclear_sites_no_res_bridge_subclass.tsv and summary.txt (containing tight_cut / coupled_hi / far_cut and the count of each bucket). In a preferred embodiment, the present invention first obtains protein three-dimensional structure data in batches from a public structure database using the script s0_rcsb_new_mmcif_dataset.py, and performs deduplication and random sampling to construct a set of structures to be processed. Subsequently, the downloaded mmCIF files are subjected to quality control using s1_qc_mmcif_gz_mp.py, including file integrity checks, atom table field detection, and structure resolution validity verification. Multi-model structures are excluded using s1_filter_multimodel.py, thereby obtaining a single-model and structurally complete effective dataset.
[0064] In the metal atom extraction stage, this invention automatically filters metal atoms from the atomic table using s2_extract_metal_atoms_mp.py (for mmCIF structures) or s2_extract_metal_atoms_pdb_mp.py (for PDB structures) to generate a metal atom set file in a uniform format. Subsequently, s3_enumerate_dinuclear_mp.py is used to combine metal atoms within the same structure in pairs, calculate metal-metal distances, and generate bimetallic candidate pairs (e.g., ...) within a preset distance window. Figure 1 (As shown). A multi-level pseudo-candidate suppression mechanism is introduced at this stage, including blacklist filtering based on element or residue name, metal cluster filtering based on the number of metal residues, and a high-density filtering strategy based on local metal neighborhood density, to effectively eliminate interfering candidates generated by inorganic clusters or high metal density environments (such as...). Figure 2 (As shown).
[0065] For the selected bimetallic candidate pairs, this invention uses the bridge_and_coord_mp.py script to automatically identify bridging atoms within a preset spatial neighborhood and constructs coordination shell information for the two metals respectively. At the same time, it counts the number of bridging atoms, the number of bridging residues, and the donor composition, thereby forming a structured candidate site description that includes bridging and coordination environment.
[0066] Based on this, the intersection of coordination shell residues was analyzed using analysis_residue_bridge_hetero.py, and candidate sites were divided into different residue bridging categories; further, no_res_bridge_subclass.py was used to perform semantic subclassing of candidate sites without residue bridging based on metal-metal distance and coordination shell sparsity (e.g., Figure 3 (As shown).
[0067] By executing the above modules in sequence, this invention achieves an automated closed-loop processing from raw structural data to structured description of bimetallic candidate sites. It can run stably on large-scale protein structural data and output standardized result files for subsequent statistical analysis, prior modeling, or the construction of machine learning scoring models.
[0068] Example 1: Candidate Construction and Classification Statistics for a Large-Scale Database. In this example, to verify the stability and statistical discrimination ability of the method of the present invention on large-scale structural data, 50,000 protein structure entries containing metal elements were randomly selected from a public protein structure database as input data. The above structures were then subjected to steps including metal atom extraction, bimetallic candidate enumeration, multi-level pseudo-candidate suppression filtering, bridging and coordination shell extraction, and structured output to obtain a structured set of bimetallic candidate sites. After progressively filtering and structural analysis of the data through the above process, 3,132 bimetallic candidate sites were identified from the 50,000 input structures, and their bridging patterns and coordination environments were further statistically analyzed.
[0069] After candidate site construction, based on the intersection analysis of the two metal coordination shell residue sets, candidate sites were divided into two main categories: those with residue bridging and those without. For candidates with residue bridging, further classification was performed according to the chemical type of the bridging residues, resulting in categories such as carboxylic acid residue bridging, histidine residue bridging, cysteine residue bridging, main chain carbonyl oxygen bridging, and other residue bridging (e.g., ...). Figure 4(As shown). Statistical results show that different bridging categories exhibit significant differences in the distribution of metal-metal distances d_mm. For example, carboxylic acid bridging categories are mainly concentrated in the medium distance range, while some sulfur coordination bridging categories show a more compact distance distribution. This statistical stratification result indicates that the method of this invention can not only identify bimetallic candidate sites, but also reveal the differences in geometric features corresponding to different bridging modes at the structural statistical level.
[0070] For candidate sites determined to lack residue bridging, this invention further performs semantic subclassification based on metal-metal distance and coordination shell sparsity. Statistical results show that candidates without residue bridging can be stably classified into several semantically clear categories, including env_coupled_pocket, weakly_coupled_sparse, weakly_coupled_wide, and far_apart_independent, etc. (e.g.) Figure 5 (As shown). The counts and proportions of each subclass are given by an automatically generated summary file, and exhibit a stable distribution at a structure scale of 50,000. This result demonstrates that the method of this invention can achieve structural semantic hierarchies for bimetallic candidates on large-scale data, rather than simply relying on distance filtering.
[0071] In summary, this embodiment verifies the scalability, stability, and structural statistical discrimination ability of the method of the present invention in large-scale databases, proving that it can effectively construct a set of bimetallic candidate sites with structural semantic information.
[0072] Example 2: Interpretation verification of no_res_bridge subclass threshold division. In this example, to verify the rationality and interpretability of the no-res_bridge candidate subclass division rule, the 679 no_res_bridge type candidate sites obtained in Example 1 were further analyzed by distance and coordination shell joint analysis.
[0073] Specifically, this invention sets three distance segmentation thresholds: a tight coupling threshold (tight_cut = 3.1 Å), a coupling upper bound threshold (coupled_hi = 4.2 Å), and a far-distance threshold (far_cut = 4.6 Å). Based on the range of metal-metal distances, the candidates are initially divided into tightly coupled, moderately coupled, and far-distance ranges. Simultaneously, by combining the minimum number of coordinating atoms in the two coordinating shells (min(n_coord1, n_coord2)), the sparsity of the structure is constrained, thereby avoiding simple division based solely on distance.
[0074] Statistical analysis shows that when d_mm < 3.1 Å and the coordination shell is relatively intact, the candidates mostly exhibit tightly coupled structures without explicit residue bridging; when 3.1 Å ≤ d_mm < 4.2 Å and the coordination environment is relatively intact, the candidates mostly fall into the environment coupling pocket category; when d_mm > 4.6 Å and the coordination shell is sparse, most belong to the distant independent metal category. In the distance-sparseness two-dimensional space, each subclass exhibits a relatively clear distribution boundary, indicating that the adopted segmentation threshold and joint constraints have good structural semantic discrimination ability.
[0075] Therefore, this embodiment verifies that the joint rule of segmented distance threshold and coordination sparsity adopted in this invention can effectively distinguish bimetallic candidate types with different structural semantics, and has clear physical interpretability and statistical stability.
[0076] Technical Effects: As can be seen from the above embodiments, the method of the present invention can operate stably on large-scale protein three-dimensional structure data and form a complete automated processing flow from the original structure to the structured description of bimetallic candidate sites. At a scale of 50,000 random structures, the present invention not only successfully constructed a bimetallic candidate set, but also effectively reduced the interference caused by inorganic clusters and high metal density environments through a multi-level pseudo-candidate suppression mechanism, ensuring the structural purity and analyzability of the candidate set. Simultaneously, through bridging atom recognition and coordination shell extraction, the structural semantics of candidate sites are enhanced, so that the output results are no longer limited to simple metal-metal distance screening, but include bridging patterns, donor composition, and local environment information.
[0077] Furthermore, statistical results of residue bridging categories show that different bridging types exhibit significant differences in metal-metal distance distribution, verifying that the present invention has the ability to distinguish different bimetallic cooperative modes at the structural level. For candidates without residue bridging, through the joint constraint of distance segmentation and coordination shell sparsity, they can be stably classified into multiple structural semantic subclasses, and exhibit a clear hierarchical distribution in the distance-sparseness space, proving that the set threshold has physical interpretability and statistical stability.
[0078] It should be noted that this specific embodiment is merely an explanation of the technical solution of the present invention and is not intended to limit the present invention. After reading this specification, those skilled in the art can make modifications to this embodiment without contributing any inventive step, but as long as they are within the scope of the claims of the present invention, they are protected by patent law.
Claims
1. A method for automatic identification and candidate set construction of protein bimetallic sites based on three-dimensional spatial features, characterized in that: Includes the following steps: S1. Structural data acquisition: Acquire the three-dimensional structural data of the protein to be processed, and read the atomic element types and three-dimensional coordinate information from the atomic table of the three-dimensional structural data; S2. Structural quality control: At least one of the following is used to perform structural quality control on the protein three-dimensional structure data obtained in S1: decompression validity check, atom table field integrity check, structural parsing validity check, and multi-model structure exclusion. S3. Metal Atom Extraction: Based on a preset set of metal elements, a set of metal atoms is obtained by filtering from the atomic table of the three-dimensional structure data, and a structure identifier, atom index, chain identifier, residue name, residue number and three-dimensional coordinates are recorded for each metal atom; S4. Bimetallic candidate enumeration: Within the same structure, pairwise combine the sets of metal atoms obtained in S3 and calculate the arbitrary metal-metal distance d. mm And generate bimetallic candidate pairs under the condition of satisfying the preset distance window; S5. Multi-level pseudo-candidate suppression filtering: Pseudo-candidate suppression filtering is performed on the bimetallic candidate pairs in S4. The pseudo-candidate suppression filtering includes at least one of the following: blacklist filtering, same residue metal cluster filtering, and high-density metal neighborhood filtering, to obtain a filtered set of bimetallic candidate pairs. S6. Structured Result Output: For each bimetallic candidate pair after S5 multi-level pseudo-candidate suppression filtering, identify the bridging information and coordination environment information related to the two metals within a preset spatial neighborhood, generate a structured description of the candidate site and output it.
2. The method for automatic identification and candidate set construction of protein bimetallic sites based on three-dimensional spatial features according to claim 1, characterized in that: The protein three-dimensional structure data includes at least one of the compressed formats of mmCIF, PDB, and mmCIF.
3. The method for automatic identification and candidate set construction of protein bimetallic sites based on three-dimensional spatial features according to claim 1, characterized in that: The set of metal atoms in S3 is obtained by screening a preset set of metal elements, which includes at least transition metal elements; and further supports the exclusion of metal atoms based on an element blacklist and / or a residue name blacklist.
4. The method for automatic identification and candidate set construction of protein bimetallic sites based on three-dimensional spatial features according to claim 1, characterized in that: The preset distance window condition mentioned in S4 is: when the metal-metal distance d mm When the conditions between the first threshold and the second threshold are met, the bimetallic candidate pair is generated, wherein the first threshold is the lower limit threshold of the metal-metal distance, the second threshold is the upper limit threshold of the metal-metal distance, the first threshold is less than the second threshold, and the first threshold is 2.0 Å; The second threshold is 7.0 Å.
5. The method for automatic identification and candidate set construction of protein bimetallic sites based on three-dimensional spatial features according to claim 1, characterized in that: The blacklist filtering in S5 is as follows: when the element type of any metal in the candidate pair belongs to the preset element blacklist, and / or the residue name of any metal belongs to the preset residue name blacklist, the candidate pair is removed; wherein, the preset element blacklist is used to exclude non-target metal elements that are prone to forming ion clusters, high-density metal environments or inorganic cluster interference; the preset element blacklist includes alkali metal elements and / or alkaline earth metal elements.
6. The method for automatic identification and candidate set construction of protein bimetallic sites based on three-dimensional spatial features according to claim 1, characterized in that: The same-residue metal cluster filtering in S5 is as follows: the number of metal atoms in a residue is calculated using the residue identifier bond as the statistical unit. When the number of metal atoms is not less than the threshold K, the residue is determined to be a metal cluster residue. When the two metal atoms of a candidate pair belong to the same metal cluster residue, the candidate pair is removed.
7. The method for automatic identification and candidate set construction of protein bimetallic sites based on three-dimensional spatial features according to claim 6, characterized in that: The same-residue metal cluster filtering only includes residues whose names do not belong to the standard amino acid set in the metal atom count, in order to reduce the risk of false rejection of the protein’s normal coordination environment.
8. The method for automatic identification and candidate set construction of protein bimetallic sites based on three-dimensional spatial features according to claim 1, characterized in that: The high-density metal neighborhood filtering in S5 involves calculating the number of metal neighbors within a preset radius R for each metal atom in the structure as its neighborhood degree. When the neighborhood degree of either endpoint of a candidate pair is not less than a threshold D, the candidate pair is eliminated. Here, the preset radius R is the spatial neighborhood radius used to count the number of metal neighbors, and the threshold D is the neighborhood degree threshold used to determine high-density metal neighborhoods. The preset radius R is 4.0 Å, and the threshold D is 8.
9. The method for automatic identification and candidate set construction of protein bimetallic sites based on three-dimensional spatial features according to claim 1, characterized in that: The pseudo-candidate suppression filtering also includes the following method: using the midpoint of the two metal coordinates of the candidate pair as the query point, and statistically analyzing the data within a preset radius R. m The number of other metal atoms besides the two metals in the candidate pair is considered; if the number of other metal atoms is greater than a threshold M, the candidate pair is discarded; wherein, R m R is the spatial query radius of the local metallic environment at the midpoint, where M is the threshold number of other metal atoms allowed to exist within the spatial query radius; m The value is 4.0 Å, and M is 1.
10. The method for automatic identification and candidate set construction of protein bimetallic sites based on three-dimensional spatial features according to claim 9, characterized in that: The pseudo-candidate suppression filtering also includes the following method: taking each metal in the candidate pair as the center, statistically analyzing its position within a preset density radius R. d The number of metals within the range, and when the number of any metal exceeds a threshold N, the candidate pair is discarded; wherein, R d R is the statistical radius of the local metal density, where N is the threshold number of metal atoms allowed within the statistical radius; d The value is 10.0 Å, and N is 10.
11. The method for automatic identification and candidate set construction of protein bimetallic sites based on three-dimensional spatial features according to claim 1, characterized in that: The bridging information in S6 includes: for two metals M1 and M2 in a candidate pair, traversing the candidate atom α in the local atomic table of the candidate site, when the candidate atom α simultaneously satisfies dist(α, M1) ≤ r c And dist(α, M2) ≤ r c When α is determined to be a bridging atom, α is a candidate atom participating in the bridging determination, and r is... c The spatial neighborhood radius for bridge identification; the r c It is 2.8 Å.
12. The method for automatic identification and candidate set construction of protein bimetallic sites based on three-dimensional spatial features according to claim 11, characterized in that: The bridging atoms are classified by their source or chemical type, including at least bridging formed by amino acid side chain donors and / or bridging formed by water molecules.
13. The method for automatic identification and candidate set construction of protein bimetallic sites based on three-dimensional spatial features according to claim 1, characterized in that: The coordination environment information in S6 includes: collecting data located at r... c The first and second sets of coordinating atoms within the radius are sorted according to their distance from the corresponding metal, and then truncated to retain the first K. c Coordinating atoms are used to generate structured coordination shell information; wherein, the K c The threshold for the maximum number of coordinating atoms retained in the structured coordination shell for each metal; the K c It is 32.
14. The method for automatic identification and candidate set construction of protein bimetallic sites based on three-dimensional spatial features according to claim 13, characterized in that: The composition of the donor atoms in the coordination shell is statistically analyzed, and the donor atoms include at least one of the following: side chain oxygen donor, histidine nitrogen donor, cysteine sulfur donor, main chain carbonyl oxygen donor, and water molecule oxygen donor.
15. The method for automatic identification and candidate set construction of protein bimetallic sites based on three-dimensional spatial features according to claim 14, characterized in that: The number of protein source donors is obtained based on the composition statistics. When the number of protein source donors for any candidate metal is lower than the threshold T... p The candidate pair is removed at that time; wherein, the T p The threshold value for the number of protein source donors used to determine whether the local coordination environment of a candidate site is reasonable; the T p The value is 2.
16. The method for automatic identification and candidate set construction of protein bimetallic sites based on three-dimensional spatial features according to claim 1, characterized in that: In S6, the coordinating atoms are mapped to a set of residues based on the two metal coordination shells and the intersection is calculated. When the intersection is not empty, it is determined that there is a residue bridge; otherwise, it is determined that there is no residue bridge.
17. The method for automatic identification and candidate set construction of protein bimetallic sites based on three-dimensional spatial features according to claim 16, characterized in that: When residue bridging exists, the residue bridging is further subdivided into at least one of carboxylic acid residue bridging, histidine residue bridging, cysteine residue bridging, main chain carbonyl oxygen bridging, and other residue bridging; and when there are ≥2 types of residue bridging, the final bridging category is output according to a preset priority.
18. The method for automatic identification and candidate set construction of protein bimetallic sites based on three-dimensional spatial features according to claim 16, characterized in that: When it is determined that there is no residue bridging, the candidate sites are further subdivided based on the metal-metal distance and the sparsity of the coordination shells on both sides. The subdivision is divided into at least one of tightly coupled candidates, environmentally coupled candidates, weakly coupled candidates, long-distance independent candidates, and noise candidates.
19. The method for automatic identification and candidate set construction of protein bimetallic sites based on three-dimensional spatial features according to claim 18, characterized in that: The subclass division is segmented using at least three distance thresholds, including a tight coupling threshold, a coupling upper bound threshold, and a far distance threshold, wherein the tight coupling threshold is 3.10 Å, the coupling upper bound threshold is 4.20 Å, and the far distance threshold is 4.60 Å.
20. The method for automatic identification and candidate set construction of protein bimetallic sites based on three-dimensional spatial features according to claim 18, characterized in that: The sparsity of the coordination shell is determined by the minimum value min(n1) of the number of atoms in the first coordination shell and the number of atoms in the second coordination shell. 1, n2), and based on min(n) 1, The subclass is determined by the joint constraint of n2 and metal-metal distance.
21. The method for automatic identification and candidate set construction of protein bimetallic sites based on three-dimensional spatial features according to claim 1, characterized in that: The structured candidate site description in S6 is output by combining table fields and nested structure data. The nested structure data includes at least a list of bridging atoms and a list of two metal coordination shells, and each record includes an atom index, residue information and distance information.
22. The method for automatic identification and candidate set construction of protein bimetallic sites based on three-dimensional spatial features according to claim 21, characterized in that: The atomic index information in the structured candidate site description prioritizes atomic sequence numbers under PDB input and atomic table row indexes or their mapping indexes under mmCIF input to ensure traceability and consistency of candidate sites across structural formats; when the atomic sequence number is missing, an incremental index is used as a fallback index.
23. The method for automatic identification and candidate set construction of protein bimetallic sites based on three-dimensional spatial features according to claim 1, characterized in that: Perform parallel batch processing on multi-structure inputs at the structure granularity, and output a list of successful processing, a list of failed processing, and statistical information on the reasons for failure.
24. A system constructed using the automatic identification and candidate set construction method for protein bimetallic sites based on three-dimensional spatial features according to any one of claims 1-23, characterized in that: It includes a structural data acquisition module, a structural quality control module, a metal atom extraction module, a bimetallic candidate enumeration module, a pseudo-candidate suppression and filtering module, and a structured result output module. The structural data acquisition module is used to execute the automatic identification and candidate set construction method S1 of protein bimetallic sites based on three-dimensional spatial features according to any one of claims 1-23, to acquire the three-dimensional structural data of the protein to be processed, and to read the atomic table information from the three-dimensional structural data; The structure quality control module is used to execute the automatic identification and candidate set construction method S2 of protein bimetallic sites based on three-dimensional spatial features as described in any one of claims 1-23, and to perform structure quality control on the protein three-dimensional structure data; The metal atom extraction module is used to execute the automatic identification and candidate set construction method S3 of protein bimetallic sites based on three-dimensional spatial features according to any one of claims 1-23, to screen the set of metal atoms from the atomic table and establish corresponding index information; The bimetallic candidate enumeration module is used to execute the automatic identification and candidate set construction method S4 of protein bimetallic sites based on three-dimensional spatial features according to any one of claims 1-23, to generate bimetallic candidate pairs and calculate metal-metal distances; The pseudo-candidate suppression and filtering module is used to execute the automatic identification and candidate set construction method S5 of protein bimetallic sites based on three-dimensional spatial features according to any one of claims 1-23, and to remove pseudo-candidate pairs caused by inorganic clusters or high-density metal environments. The structured result output module is used to execute the automatic identification and candidate set construction method S6 of protein bimetallic sites based on three-dimensional spatial features according to any one of claims 1-23, identify bridging information and coordination environment information, and generate and output structured candidate site descriptions.
25. A readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the method for automatic identification and candidate set construction of protein bimetallic sites based on three-dimensional spatial features as described in any one of claims 1 to 23.