Method and system for constructing multi-layer prior of bimetallic site based on large-scale structure statistics, and readable storage medium
By employing a multi-level prior construction method based on large-scale structural statistics, the problem of insufficient joint modeling capability in bimetallic site statistical analysis is solved, and a stable and scalable multi-level statistical modeling framework is established to support structural analysis for de novo protein design and green biomanufacturing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- UNIV OF SCI & TECH OF CHINA
- Filing Date
- 2026-05-11
- Publication Date
- 2026-07-31
AI Technical Summary
Existing statistical analysis of bimetallic sites has shortcomings in joint modeling capabilities, structural partitioning control, hierarchical feature coverage, and prior unified coding, making it difficult to form a stable, scalable, and generalizable statistical modeling framework.
Through a multi-level prior construction method based on large-scale structural statistics, including structure-level data partitioning, continuous-discrete hybrid probability modeling, and unified structured coding, a global and metal-pair-specific multi-level statistical distribution system is established, forming an scalable multi-level prior file.
It achieves stable and scalable statistical modeling of bimetallic sites, supporting structural analysis and site quality assessment in fields such as de novo protein design and green biomanufacturing, and provides a reliable prior model.
Smart Images

Figure CN122493931A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of structural bioinformatics and computational structural modeling technology, specifically relating to a method, system, and readable storage medium for constructing multilayer priors for bimetallic sites based on large-scale structural statistics. The invention aims to establish a spatial topological and physicochemical constraint model for bimetallic centers, providing underlying physicochemical rules and prior knowledge constraints for structure-based artificial intelligence scoring systems, de novo protein design, and directed enzyme evolution. Background Technology
[0002] Metal ions play a variety of crucial roles in protein systems, including catalysis, electron transfer, molecular recognition, and structural stability. Among numerous metalloprotein structures, bimetallic sites are important structural units with synergistic effects, widely found in hydrolases, oxidoreductases, and various metal-dependent enzyme systems. Compared to monometallic sites, bimetallic sites achieve complex functions such as multi-electron transfer, bridging intermediate stability, and reaction pathway regulation through the spatial proximity and electronic coupling between two metal ions (e.g., the rapid scavenging of superoxide anions by Cu / Zn superoxide dismutase). Their geometry, intermetallic distance, bridging mode, and coordination environment all significantly influence their function. Therefore, systematic structural statistical analysis and feature modeling of bimetallic sites are of great significance for understanding their structural patterns and constructing stable statistical expression systems.
[0003] With the continuous expansion of structural biology databases, the sample size of bimetallic sites has increased significantly, providing conditions for summarizing statistical regularities from real structural data. However, existing bimetallic site analyses largely rely on empirical rules, often simplifying the screening process to independent judgments of a few indicators, such as limiting only the range of intermetallic distances or screening based solely on the type of bridging residues and the lowest coordination number. This approach decomposes multidimensional structural constraints into independent local conditions, making it difficult to describe the cooperative distribution relationship between intermetallic distances, bridging patterns, coordination numbers, and geometric configurations. Consequently, its statistical expressive power is limited, and it is difficult to provide a complete site configuration spectrum.
[0004] In large-scale structural data environments, structural repetition and homology similarity can significantly impact statistical reliability. Databases commonly contain multiple conformations of the same protein, approximate structures under different resolution conditions, or homologous protein structures. If modeling and evaluation are not isolated and controlled at the structural level, patterns appearing in the training set may be repeated in the evaluation set in an approximate form, causing distribution bias and weakening the generalization ability on new structures. Therefore, it is necessary to introduce a data partitioning mechanism based on structure to reduce the leakage risk caused by implicit repetition.
[0005] Furthermore, the patterns of bimetallic sites are not only reflected in basic geometric scales but also in coordination configurations and local geometric constraints. Simply statistically analyzing basic features such as inter-metal distances and bridging types is often insufficient to cover higher-order patterns such as coordination geometry, planarity, key angle distribution, and configuration deviation. It also fails to reflect the coupling relationship between these features and basic characteristics within the same framework, thus hindering the formation of a complete structural prior representation system. In addition, existing statistical results are mostly presented as scattered thresholds or independent parameter files, lacking a unified structured coding method to organize distribution models of different levels, different metal pairs, and different feature sets. This makes it difficult to expand, integrate, and maintain a versioned prior system. As sample size and feature dimensions further increase, the lack of unified coding and integration management will become a major obstacle to the construction and reuse of large-scale prior systems.
[0006] Therefore, it is necessary to construct a multi-level statistical prior construction method for bimetallic sites based on a large-scale structural database. By using structural hierarchical partitioning and a continuous-discrete hybrid probability modeling mechanism, a multi-level statistical distribution system for both global and metal pairs can be established and encoded and output in a unified structured format, thereby forming a stable, scalable and generalizable statistical modeling framework for bimetallic sites. Summary of the Invention
[0007] To address the shortcomings of existing bimetallic site statistical analysis in terms of joint modeling capabilities, structural-level partitioning control, hierarchical feature coverage, and unified prior coding, this invention proposes a multi-layer prior construction method, system, and readable storage medium for bimetallic sites based on large-scale structural statistics.
[0008] The present invention provides a multilayer prior construction method for bimetallic sites based on large-scale structural statistics, which is achieved through the following technical solution: A multi-level prior construction method for bimetallic sites based on large-scale structural statistics includes the following steps: S1, establish a standardized Level 1 candidate dataset; S2, Summary of the size statistics and distribution of metal-pair particles; S3, perform structural-level partitioning of the training and test sets: perform structural-level partitioning of the Level 1 bimetallic candidate dataset; S4.1, based on the structural-level partitioning results, exports the full candidate data of Level 1 as training and test sets, and simultaneously exports different labeled subsets (such as positive samples, hard negative samples, etc.) for hierarchical statistical modeling and comparative analysis. S4.2, Phase1 Prior: Establish a continuous-discrete mixed probability distribution model on the training set for the basic geometry and basic features of the bridging / donor layer Phase1, forming a global prior and a metal-pair-specific prior, and output the Phase1 prior file (v1) in a structured format. S4.3, Phase2 Prior: Based on the standard row data of Phase1, further extract the coordination geometry and configuration-related features of the coordination geometry configuration layer Phase2 to form a second-layer feature table, providing input for fitting the higher-order geometric statistical Phase2 prior (v2); establish a second-layer continuous-discrete mixed probability distribution model on the higher-order configuration features of the coordination geometry configuration layer Phase2 on the training set, forming a global prior and a metal-pair-specific prior, and output the Phase2 prior file (v2) in a structured format. S4.4 unifies and merges the two-layer priors v1 (Phase1) and v2 (Phase2) into a unified code, generating an extensible multi-layer prior file v2.5, making different layers traceable, extensible and reusable under the same structured framework; The method further includes: S5. Prior scoring the test subset based on the multi-layer prior file, outputting the total prior score and sub-item contribution of the candidate sites, and statistically summarizing or visually verifying the prior score distribution of different label subsets to evaluate the discriminative power and stability of the multi-layer prior. Specifically, S5.1. Use the constructed prior file to score the candidate data, output the total score and the contribution of each item, to verify the ability of the multi-level statistical prior to distinguish different label strata, and to provide a basis for subsequent statistical analysis and visualization. S5.2. Statistically summarize and visualize the prior scoring results to show the distribution differences and hierarchical trends of different label subsets in the prior scores, which is used to illustrate the effectiveness and stability of the prior modeling. S5.3. The above scripts are chained together in batch processing to form a reproducible experimental process, realizing the automated execution from Level 1 data preparation, structure-level partitioning, v1 prior fitting, Phase 2 feature extraction, v2 prior fitting to v2.5 fusion and scoring verification.
[0009] Preferably, step S1. Establishing a standardized Level 1 candidate dataset: obtaining structural data from a protein structure database or obtaining a bimetallic candidate data table by parsing structural data from a protein structure database, and constructing a standardized Level 1 bimetallic candidate dataset; Preferably, the Level 1 bimetallic candidate dataset construction includes a hierarchical labeling system, which includes at least explicit positive samples, implicit positive samples, difficult negative samples, and other negative samples. Based on the hierarchical labeling system, the system outputs the count summary of each label subset and its distribution statistics by metal pair.
[0010] Preferably, the Level 1 bimetallic candidate dataset includes at least structural identifiers, the types of the two metal elements, and information on the distance between the metals; Preferably, the Level 1 bimetallic candidate dataset further includes one or more of the following: bridging category, bridging subclass, coordination number, coordinating atom type, and coordinating residue type.
[0011] Preferably, the scale statistics and distribution summary of the S2. metal-pair granularity provide a data foundation for the subsequent construction of global priors and metal-pair-specific priors, and are used to identify long-tailed metal pairs and modelable subsets.
[0012] More preferably, the execution process of S2. Scale statistics and distribution summary of metal-pair granularity includes the following steps: S2.1, Read the standardized candidate data table (Level 1 full data or a specified hierarchical subset), parse the metal element types and combine them into metal-pair bonds; S2.2 performs counting statistics using metal-pair as the key, outputs the overall distribution, head and long-tail lists, and can further summarize the counting information on the label dimension; S2.3 Export the statistical results as a tabular file for subsequent data partitioning, prior fitting, and quality checks.
[0013] Preferably, step S3. Dividing the training set and test set at the structural level: dividing the Level 1 bimetallic candidate dataset at the structural level based on the structural identifiers or structural cluster identifiers in the Level 1 bimetallic candidate dataset.
[0014] Preferably, the structural-level partitioning is performed by sampling and isolating structures or structural clusters as the smallest unit, so that the training subset and the test subset are independent of each other at the structural level, thereby reducing the statistical bias caused by structural similarity and implicit repetition.
[0015] Preferably, step S3 involves dividing the training and testing sets at the structural level to ensure that the same structure does not simultaneously enter the modeling and verification process, thereby reducing the risk of statistical bias and leakage caused by implicit duplication. The execution process includes the following steps: S3.1, Read the full Level 1 data and extract the structure identifier (structure_id) and its corresponding group identifier; S3.2, Sampling and partitioning are performed using structure as the smallest unit to generate the structure set for train / test; S3.3 outputs a list of train / test structures and corresponding segmentation result files, which can be used for filtering when exporting the dataset and fitting the prior.
[0016] Preferably, in step S4.1, based on the structural-level partitioning results, the full candidate data of Level 1 is exported as a training set and a test set, and different labeled subsets (such as positive samples, hard negative samples, etc.) are exported simultaneously for hierarchical statistical modeling and comparative analysis. The execution process includes the following steps: S4.1.1, Read the full Level1 data and the train / test structure list, and filter by structure affiliation to obtain level1_train and level1_test; S4.1.2, stratify and segment the data according to the label system, and export multiple subset files (e.g., explicit positive samples, implicit positive samples, difficult negative samples, other negative samples). S4.1.3 synchronously outputs a summary of label counts and a label distribution table by metal-pair, providing support for subsequent prior fitting and result interpretation.
[0017] Preferably, in S4.2, the Phase 1 prior stage mainly covers basic statistical laws such as inter-metal distance, bridging type, and coordination number. The specific execution flow of the Phase 1 prior stage is as follows: S4.2.1, Read the level1_train data, select continuous and discrete features for Phase1 modeling, and define the binning strategy for the discrete class space and continuous variables.
[0018] S4.2.2 performs histogram statistics and smoothing on continuous features, performs class frequency statistics and normalizes discrete features to obtain probability distributions; at the same time, it constructs two-level distributions: global (GLOBAL) and by metal pair (BY_METAL_PAIR).
[0019] S4.2.3 encodes the distribution parameters, binning information, and metadata into a compressed structured file output (such as json.gz), and allows prior scoring on the test set to verify the distribution's stability and discriminative ability.
[0020] Preferably, the specific execution flow of S4.3, Phase 2 prior is as follows: S4.3.1 reads the full or specified subset of Level 1 data to locate the local coordination environment (donor atom set, geometric neighborhood, etc.) corresponding to each bimetallic candidate site. S4.3.2 Calculate and output higher-order configuration features, such as coordination geometry type labels, configuration deviation index, key angle statistics, planarity correlation index, etc., and align them with the original candidate rows; S4.3.3, export the full table after Phase2 feature enhancement (e.g., phase2_geom.tsv) for the v2 prior fitting script to read directly; S4.3.4, Read the Phase2 feature table and filter it according to the existing structural level division results to obtain phase2_train, ensuring the independence of the fitted data; S4.3.5, perform binning statistics and category statistics on the selected continuous and discrete features of Phase2 respectively, construct two-level probability distributions of GLOBAL and BY_METAL_PAIR and perform smoothing processing; S4.3.6 outputs v2 prior files (structured compressed format) and retains feature directories, binning information and metadata for easy subsequent fusion and version management; Preferably, in step S4.4, the two prior layers v1 (Phase 1) and v2 (Phase 2) are uniformly organized and fused into a single encoding to generate an extensible multi-layer prior file v2.5. This ensures that different layers are distributed within the same structured framework, making them traceable, extensible, and reusable. The specific execution flow is as follows: S4.4.1, Read the v1 and v2 prior files, and parse their distribution blocks, bin definitions, feature directories and metal pair index structures; S4.4.2 Organize the two-layer prior according to a unified data structure, clarify the hierarchical relationship between GLOBAL and BY_METAL_PAIR, and write the metadata required for fusion (version, feature list, binning consistency check information, etc.). S4.4.3 outputs the merged v2.5 prior file (json.gz), forming a unified carrier for multi-layered priors.
[0021] Preferably, both the Phase1 prior and the Phase2 prior simultaneously construct a global prior distribution and a metal-pair-specific prior distribution, wherein the global prior distribution is used for backtracking of long-tailed metal pairs, and the metal-pair-specific prior distribution is used for fine-grained expression of metal pairs with sufficient samples.
[0022] Preferably, in the metal-pair-specific prior distribution, when the sample size of the target metal pair is lower than a preset threshold, the global prior distribution is used as a backoff distribution to estimate the probability of the metal pair. The preset threshold is a configurable parameter and is written into the version metadata.
[0023] Preferably, in the Phase1 prior, the continuous-discrete mixed probability distribution model is constructed using a binned histogram method, wherein the binning boundary and the number of bins are written into the Phase1 prior as binning information.
[0024] More preferably, the binning histogram statistics employ smoothing to avoid zero-probability issues, and the smoothing includes Laplace smoothing or additive smoothing.
[0025] Preferably, or in the Phase 1 prior, the class probability distribution is obtained by class frequency statistics and normalization on the continuous-discrete mixed probability distribution model, wherein the discrete features include at least bridging classes, bridging subclasses, minimum coordination number or combinations thereof.
[0026] Preferably, the second-layer feature table includes at least one or more of the following: coordination geometry type label, configuration deviation index, key angle statistical index, and planarity index.
[0027] Preferably, in the Phase2 prior, when establishing the second-layer continuous-discrete mixed probability distribution model for the second-layer feature table, the continuous features are statistically analyzed using binned histograms and then smoothed, while the discrete features are statistically analyzed using category frequencies and then normalized to obtain the category probability distribution.
[0028] Preferably, the unified organization and fusion coding in S4 includes: parsing the distribution blocks, binning definitions and feature directories of the Phase1 prior file (v1) and the Phase2 prior file (v2), performing consistency checks on the binning information of the two priors, and writing the two priors into the same multi-level prior file v2.5 according to a preset hierarchical structure, so that different levels are traceable, scalable and reusable under the same structured framework.
[0029] Preferably, the multi-layer prior file v2.5 is organized using a distribution block structure, which includes at least a Phase1 distribution block and a Phase2 distribution block, and each distribution block contains a global distribution sub-block and a metal-pair distribution sub-block respectively; the hierarchical index is used to indicate the mapping relationship between distribution blocks, feature directories and binning information.
[0030] Preferably, the multi-level prior file includes at least a hierarchical index, a feature directory, binning information, probability distribution parameters, and version metadata.
[0031] Preferably, the multi-layer prior file v2.5 is output in a structured compression format, which includes JSON compression format or JSON-GZ format.
[0032] Preferably, the version metadata includes at least one or more of the following: build time, data size identifier, feature list, binning configuration, and prior version number.
[0033] Preferably, in step S5.1, the candidate data is scored using the constructed prior file, and the total score and component contributions are output. This is used to verify the ability of the multi-level statistical prior to distinguish between different label strata and to provide a foundation for subsequent statistical analysis and visualization. The specific execution process is as follows: S5.1.1 reads the candidate data table and v2.5 prior file, loads the global and metal-pair distribution parameters, and establishes a lookup table mapping from features to distributions.
[0034] S5.1.2 Calculate the logarithmic probability or normalized probability contribution of each feature for each candidate site, and combine them into a total score according to hierarchical rules, while retaining the sub-item scoring fields.
[0035] S5.1.3 outputs a data table with scoring results (e.g., scored_level1_test.tsv) for statistical comparison and visualization analysis.
[0036] Preferably, step S5.2 involves statistically summarizing and visually validating the prior scoring results to demonstrate the distribution differences and hierarchical trends of different label subsets in the prior scores, thereby illustrating the effectiveness and stability of the prior modeling. The main execution flow is as follows: S5.2.1, Read the scored_level1_test data, group the prior scores by label category, and calculate the mean, median, quantiles and other indicators; S5.2.2 generates distribution plots such as box plots, violin plots, and histogram overlays, and outputs a summary table of key statistics; S5.2.3 Export the charts and statistical files to the specified directory to provide direct evidence for the effect of the implementation example.
[0037] Preferably, in step S5.3, the above scripts are chained together in a batch processing manner to form a reproducible experimental process, realizing automated execution from Level 1 data preparation, structure-level partitioning, v1 prior fitting, Phase 2 feature extraction, v2 prior fitting to v2.5 fusion and scoring verification. The specific execution process is as follows: S5.3.1 sets the data path and scale parameters, and sequentially calls the statistics, partitioning, exporting and fitting scripts to automatically generate v1 related outputs; S5.3.2, Based on the results of v1, call the Phase2 feature extraction script and filter the training set to complete the v2 prior fitting; S5.3.3 executes the prior fusion and scoring verification script, generating v2.5 prior files and accompanying distribution maps and statistical tables.
[0038] The present invention provides a method, system, and readable storage medium for constructing multilayer priors of bimetallic sites based on large-scale structural statistics, which is achieved through the following technical solutions: A system is constructed using a multi-layer prior algorithm for bimetallic sites based on large-scale structural statistics. It includes a data acquisition and normalization module, a metal pair statistics module, a structural partitioning module, a first-level data export module, a first-layer prior fitting module, a second-layer feature extraction module, a second-layer prior fitting module, a multi-layer prior fusion encoding module, and a scoring and verification module. The data acquisition and standardization construction module is used to execute S1 in the multi-level prior construction method for bimetallic sites based on large-scale structural statistics, acquire structural data or candidate data tables, and construct a standardized first-level bimetallic candidate dataset. The metal pair statistics module is used to execute S2 in the multi-level prior construction method for bimetallic sites based on large-scale structural statistics, and to perform counting statistics and distribution summarization of the first-level bimetallic candidate dataset by metal pair. The structure-level partitioning module is used to execute S3 in the bimetallic site multilayer prior construction method based on large-scale structural statistics, and to generate training structure sets and test structure sets based on structure identifiers or structure cluster identifiers. The first-level data export module is used in S4 of the bimetallic site multilayer prior construction method based on large-scale structural statistics to derive the training subset, the test subset, and the hierarchical label subset according to the training structure set and the test structure set. The first-layer prior fitting module is used to execute S5 in the multi-layer prior construction method for bimetallic sites based on large-scale structural statistics, and to establish a continuous-discrete mixed probability distribution model for the Phase1 feature on the training subset to obtain the Phase1 prior. The second-layer feature extraction module is used to execute S6 in the multi-layer prior construction method for bimetallic sites based on large-scale structural statistics, extracting coordination geometry and configuration-related features to form the Phase2 feature table. The second-layer prior fitting module is used to execute S7 in the multi-layer prior construction method for bimetallic sites based on large-scale structural statistics, and to establish a continuous-discrete mixed probability distribution model on the Phase2 training subset to obtain the Phase2 prior. The multi-level prior fusion coding module is used to execute S8 in the bimetallic site multi-level prior construction method based on large-scale structural statistics. It performs unified organization and fusion coding of Phase1 prior and Phase2 prior, and outputs multi-level prior files.
[0039] The scoring and verification module is used to execute S9 in the multi-level prior construction method for bimetallic sites based on large-scale structural statistics. Based on the multi-level prior file, the module performs prior scoring on the test subset, outputs the total prior score and the contribution of each item, and generates statistical summary results or distribution plots.
[0040] A computer-readable storage medium storing a computer program, which, when executed by a processor, implements a preparation method for a multilayer prior construction method for bimetallic sites based on large-scale structural statistics.
[0041] In summary, the present invention has the following advantages: 1. The method of this invention breaks through the limitations of a single statistical threshold, improves the statistical expression capability through joint distribution modeling, and forms an scalable multi-layer statistical modeling system through hierarchical prior fusion.
[0042] 2. The method of this invention can construct stable and generalizable statistical priors for bimetallic sites on large-scale structural data, providing a basic modeling framework for structural analysis and site quality assessment in cutting-edge fields such as de novo protein design and green biomanufacturing.
[0043] 3. The method of this invention can be used for summarizing the structural regularity of bimetallic sites, continuously updating the database and evaluating its quality. It can serve as a basic module for related structural analysis, candidate screening and machine learning modeling, and supports structural library updates and cross-metal system migration. Attached Figure Description
[0044] Figure 1 This is the overall flowchart of the multi-level prior construction of the present invention (including basic geometric statistics Phase 1 and higher-order coordination configuration statistics Phase 2).
[0045] Figure 2 This is a detailed diagram of the mixed probability modeling of underlying features in Phase 1 of this invention.
[0046] Figure 3 This is a schematic diagram of the multi-layered prior document structure of the present invention.
[0047] Figure 4 The histogram distribution of the Phase1 prior score across different label subsets is shown.
[0048] Figure 5 Comparison of boxplot distributions of the Phase1 prior score across different label subsets.
[0049] Figure 6 Comparison of violin plot distributions of Phase1 prior scores across different label subsets.
[0050] Figure 7Graphs showing the mean, median, and P10-P90 intervals of the Phase1 prior score across different label subsets. Detailed Implementation
[0051] To further understand the inventiveness and technical advancements of this invention, the preferred embodiments of this invention will be discussed in detail below with reference to examples and comparative examples.
[0052] Example: A method for constructing multilayer priors for bimetallic sites based on large-scale structural statistics, comprising the following steps: S1. Establish a standardized Level 1 candidate dataset: Obtain structural data from a protein structure database or parse the structural data from a protein structure database to obtain a bimetallic candidate data table, and construct a standardized Level 1 bimetallic candidate dataset. The Level 1 bimetal candidate dataset is constructed by including a hierarchical labeling system, which includes at least explicit positive samples, implicit positive samples, hard negative samples, and other negative samples. Based on the hierarchical labeling system, the dataset outputs the count summary of each label subset and its distribution statistics by metal pair. The Level 1 bimetal candidate dataset contains at least the structural identifier, the types of the two metal elements, and the distance information between the metals. The Level 1 bimetal candidate dataset also further includes one or more of the following: bridging category, bridging subclass, coordination number, coordinating atom type, and coordinating residue type. S2, the scale statistics and distribution summary of metal-pair granularity, provides a data foundation for the subsequent construction of global priors and metal-pair-specific priors, and is used to identify long-tailed metal pairs and modelable subsets. The execution process includes the following steps: S2.1, Read the standardized candidate data table (Level 1 full data or a specified hierarchical subset), parse the metal element types and combine them into metal-pair bonds; S2.2 performs counting statistics using metal-pair as the key, outputs the overall distribution, head and long-tail lists, and can further summarize the counting information on the label dimension; S2.3 Export the statistical results as a tabular file for subsequent data partitioning, prior fitting, and quality checks; S3, perform structural-level partitioning of the training and test sets: perform structural-level partitioning of the Level 1 bimetallic candidate dataset; The training and test sets are divided at the structural level: The Level 1 bimetallic candidate dataset is divided at the structural level based on the structural identifiers or structural cluster identifiers in the Level 1 bimetallic candidate dataset. Structural-level partitioning uses structures or structural clusters as the smallest unit for sampling and isolation, making the training subset and the test subset independent of each other at the structural level, thereby reducing statistical bias caused by structural similarity and implicit repetition. Specifically, S3 divides the training and test sets at the structural level to ensure that the same structure does not enter the modeling and validation process simultaneously, thereby reducing the risk of statistical bias and leakage caused by implicit duplication. The specific execution process includes the following steps: S3.1, Read the full Level 1 data and extract the structure identifier (structure_id) and its corresponding group identifier; S3.2, Sampling and partitioning are performed using structure as the smallest unit to generate the structure set for train / test; S3.3 outputs a list of train / test structures and corresponding segmentation result files for filtering when exporting the dataset and fitting the prior. S4.1, based on the structural-level partitioning results, exports the full Level 1 candidate data as training and test sets, and simultaneously exports different labeled subsets (such as positive samples, hard negative samples, etc.) for hierarchical statistical modeling and comparative analysis. The specific execution process includes the following steps: S4.1.1, Read the full Level1 data and the train / test structure list, and filter by structure affiliation to obtain level1_train and level1_test; S4.1.2, stratify and segment the data according to the label system, and export multiple subset files (e.g., explicit positive samples, implicit positive samples, difficult negative samples, other negative samples). S4.1.3 synchronously outputs a summary of label counts and a label distribution table by metal-pair, providing support for subsequent prior fitting and result interpretation. S4.2, Phase1 Prior: Establish a continuous-discrete mixed probability distribution model on the training set for the basic geometry and basic features of the bridging / donor layer Phase1, forming a global prior and a metal-pair-specific prior, and output the Phase1 prior file (v1) in a structured format. Phase 1 prior simultaneously constructs a global prior distribution and a metal-pair-specific prior distribution. The global prior distribution is used for backtracking long-tailed metal pairs, while the metal-pair-specific prior distribution is used for fine-grained expression of metal pairs with sufficient samples. In the metal-pair-specific prior distribution, when the sample size of the target metal pair is lower than a preset threshold, the global prior distribution is used as the backtracking distribution to estimate the probability of the metal pair. The preset threshold is a configurable parameter and is written into the version metadata. In the Phase 1 prior, the continuous-discrete mixed probability distribution model is constructed using a binned histogram method. The binning boundary and the number of bins are written into the Phase 1 prior as binning information. The binned histogram statistics are smoothed to avoid the zero probability problem. The smoothing process includes Laplace smoothing or additive smoothing. Alternatively, in the Phase 1 prior, the class probability distribution is obtained by using class frequency statistics and normalization on the continuous-discrete mixed probability distribution model. The discrete features include at least the bridging class, the bridging subclass, the minimum coordination number, or a combination thereof. Phase 1, the prior stage, mainly covers basic statistical laws such as inter-metal distance, bridging type, and coordination number. The specific execution process includes the following steps: S4.2.1, Read the level1_train data, select continuous and discrete features for Phase1 modeling, and define the binning strategy for the discrete class space and continuous variables.
[0053] S4.2.2 performs histogram statistics and smoothing on continuous features, performs class frequency statistics and normalizes discrete features to obtain probability distributions; at the same time, it constructs two-level distributions: global (GLOBAL) and by metal pair (BY_METAL_PAIR).
[0054] S4.2.3 encodes the distribution parameters, binning information and metadata into a compressed structured file output (such as json.gz), and can perform prior scoring on the test set to verify the distribution stability and discriminative ability; S4.3, Phase2 Prior: Based on the standard row data of Phase1, further extract the coordination geometry and configuration-related features of the Phase2 coordination geometry configuration layer to form a second-layer feature table. The second-layer feature table includes at least one or more of the following: coordination geometry type label, configuration deviation index, key angle statistical index, and planarity index, providing input for fitting the higher-order geometric statistical Phase2 prior (v2). On the training set, establish a second-layer continuous-discrete mixed probability distribution model for the higher-order configuration features of the Phase2 coordination geometry configuration layer to form a global prior and a metal-pair-specific prior, and output the Phase2 prior file (v2) in a structured format. In the Phase 2 prior, when establishing the second-layer continuous-discrete hybrid probability distribution model for the second-layer feature table, the continuous features are statistically analyzed using binned histograms and smoothed, while the discrete features are statistically analyzed using class frequency and normalized to obtain the class probability distribution. Phase 2 prior simultaneously constructs a global prior distribution and a metal-pair-specific prior distribution. The global prior distribution is used for backtracking long-tailed metal pairs, while the metal-pair-specific prior distribution is used for fine-grained expression of metal pairs with sufficient samples. In the metal-pair-specific prior distribution, when the sample size of the target metal pair is lower than a preset threshold, the global prior distribution is used as the backtracking distribution to estimate the probability of the metal pair. The preset threshold is a configurable parameter and is written into the version metadata. The specific execution process for Phase 2 prior is as follows: S4.3.1 reads the full or specified subset of Level 1 data to locate the local coordination environment (donor atom set, geometric neighborhood, etc.) corresponding to each bimetallic candidate site. S4.3.2 Calculate and output higher-order configuration features, such as coordination geometry type labels, configuration deviation index, key angle statistics, planarity correlation index, etc., and align them with the original candidate rows; S4.3.3, export the full table after Phase2 feature enhancement (e.g., phase2_geom.tsv) for the v2 prior fitting script to read directly; S4.3.4, Read the Phase2 feature table and filter it according to the existing structural level division results to obtain phase2_train, ensuring the independence of the fitted data; S4.3.5, perform binning statistics and category statistics on the selected continuous and discrete features of Phase2 respectively, construct two-level probability distributions of GLOBAL and BY_METAL_PAIR and perform smoothing processing; S4.3.6 outputs v2 prior files (structured compressed format) and retains feature directories, binning information and metadata for easy subsequent fusion and version management; S4.4 unifies and merges the two-layer priors v1 (Phase1) and v2 (Phase2) into a unified code, generating an extensible multi-layer prior file v2.5, making different layers traceable, extensible and reusable under the same structured framework; The unified organization and fusion coding includes: parsing the distribution blocks, binning definitions and feature directories of the Phase1 prior file (v1) and the Phase2 prior file (v2), performing consistency checks on the binning information of the two priors, and writing the two priors into the same multi-level prior file v2.5 according to the preset hierarchical structure, so that the different levels are distributed under the same structured framework, making them traceable, scalable and reusable. The multi-layer prior file v2.5 is organized using a distributed block structure, which includes at least a Phase1 distributed block and a Phase2 distributed block, and each distributed block contains a global distributed sub-block and a metal-pair distributed sub-block; the hierarchical index is used to indicate the mapping relationship between the distributed blocks, feature directories and binning information. The multi-level prior file v2.5 includes at least a hierarchical index, feature directory, binning information, probability distribution parameters, and version metadata; The multi-layer prior file v2.5 is output in a structured compression format, which includes JSON compression format or JSON-GZ format; Version metadata should include at least one or more of the following: build time, data size identifier, feature list, binning configuration, and prior version number; The specific execution process is as follows: S4.4.1, Read the v1 and v2 prior files, and parse their distribution blocks, bin definitions, feature directories and metal pair index structures; S4.4.2 Organize the two-layer prior according to a unified data structure, clarify the hierarchical relationship between GLOBAL and BY_METAL_PAIR, and write the metadata required for fusion (version, feature list, binning consistency check information, etc.). S4.4.3 outputs the merged v2.5 prior file (json.gz), forming a unified carrier for multi-layered priors.
[0055] The multi-level prior construction method for bimetallic sites based on large-scale structural statistics also includes the following steps: S5. Based on the multi-layer prior file, perform prior scoring on the test subset, output the total prior score and sub-item contribution of the candidate sites, and perform statistical summary or visualization verification on the prior score distribution of different label subsets to evaluate the discriminative power and stability of the multi-layer prior. S5.1. Use the constructed prior file to perform prior scoring on the candidate data, output the total score and the contribution of each item, to verify the ability of the multi-level statistical prior to distinguish different label strata, and to provide a foundation for subsequent statistical analysis and visualization. The specific execution process is as follows: S5.1.1 reads the candidate data table and v2.5 prior file, loads the global and metal-pair distribution parameters, and establishes a lookup table mapping from features to distributions.
[0056] S5.1.2 Calculate the logarithmic probability or normalized probability contribution of each feature for each candidate site, and combine them into a total score according to hierarchical rules, while retaining the sub-item scoring fields.
[0057] S5.1.3 outputs a data table with scoring results (e.g., scored_level1_test.tsv) for statistical comparison and visualization analysis; S5.2. Statistically summarize and visualize the prior scoring results to show the distribution differences and stratification trends of different label subsets in the prior scores, which is used to illustrate the effectiveness and stability of the prior modeling. The specific execution process is as follows: S5.2.1, Read the scored_level1_test data, group the prior scores by label category, and calculate the mean, median, quantiles and other indicators; S5.2.2 generates distribution plots such as box plots, violin plots, and histogram overlays, and outputs a summary table of key statistics; S5.2.3, Export the charts and statistical files to the specified directory to provide direct evidence for the effect of the implementation example; S5.3. The above scripts are chained together in batch processing to form a reproducible experimental workflow, realizing automated execution from Level 1 data preparation, structure-level partitioning, v1 prior fitting, Phase 2 feature extraction, v2 prior fitting to v2.5 fusion and scoring verification. The specific execution flow is as follows: S5.3.1 sets the data path and scale parameters, and sequentially calls the statistics, partitioning, exporting and fitting scripts to automatically generate v1 related outputs; S5.3.2, Based on the results of v1, call the Phase2 feature extraction script and filter the training set to complete the v2 prior fitting; S5.3.3 executes the prior fusion and scoring verification script, generating v2.5 prior files and accompanying distribution maps and statistical tables.
[0058] A system built using a multi-layer prior construction method for bimetallic sites based on large-scale structural statistics includes a data acquisition and standardization construction module, a metal pair statistics module, a structural partitioning module, a first-level data export module, a first-layer prior fitting module, a second-layer feature extraction module, a second-layer prior fitting module, a multi-layer prior fusion coding module, and a scoring verification module. The data acquisition and standardization construction module is used to execute S1 in the multi-level prior construction method for bimetallic sites based on large-scale structural statistics, acquire structural data or candidate data tables, and construct a standardized first-level bimetallic candidate dataset. The metal pair statistics module is used to execute S2 in the multi-level prior construction method for bimetallic sites based on large-scale structural statistics, and to perform counting statistics and distribution summarization of the first-level bimetallic candidate dataset by metal pair. The structure-level partitioning module is used to execute S3 in the bimetallic site multilayer prior construction method based on large-scale structural statistics, and to generate training structure sets and test structure sets based on structure identifiers or structure cluster identifiers. The first-level data export module is used in S4 of the bimetallic site multilayer prior construction method based on large-scale structural statistics to derive the training subset, the test subset, and the hierarchical label subset according to the training structure set and the test structure set. The first-layer prior fitting module is used to execute S5 in the multi-layer prior construction method for bimetallic sites based on large-scale structural statistics, and to establish a continuous-discrete mixed probability distribution model for the Phase1 feature on the training subset to obtain the Phase1 prior. The second-layer feature extraction module is used to execute S6 in the multi-layer prior construction method for bimetallic sites based on large-scale structural statistics, extracting coordination geometry and configuration-related features to form the Phase2 feature table. The second-layer prior fitting module is used to execute S7 in the multi-layer prior construction method for bimetallic sites based on large-scale structural statistics, and to establish a continuous-discrete mixed probability distribution model on the Phase2 training subset to obtain the Phase2 prior. The multi-level prior fusion coding module is used to execute S8 in the bimetallic site multi-level prior construction method based on large-scale structural statistics, to uniformly organize and fuse the Phase1 prior and Phase2 prior, and output the multi-level prior file. The scoring and verification module is used to execute S9 in the multi-level prior construction method for bimetallic sites based on large-scale structural statistics. Based on the multi-level prior file, the module performs prior scoring on the test subset, outputs the total prior score and the contribution of each item, and generates statistical summary results or distribution plots.
[0059] Preferably, both the first-layer prior fitting module and the second-layer prior fitting module are configured to simultaneously construct a global prior distribution and a specific prior distribution divided by metal pairs.
[0060] A readable storage medium storing a computer program, which, when executed by a processor, implements a preparation method for a multilayer prior construction method for bimetallic sites based on large-scale structural statistics.
[0061] Example 1: v1 prior construction and interpretable scoring for a 50,000 structure.
[0062] A multi-level prior construction method for bimetallic sites based on large-scale structural statistics was developed. Using a dataset of 50,000 structures as input, including the `dinuclear_sites_50000.tsv` file and bridging annotation tables, a bimetallic candidate data table and its bridging category / subclass annotation data were constructed, forming a standardized candidate set for Level 1 modeling. In this dataset, after enumerating and filtering bimetallic sites on all structures, 1344 protein structures containing bimetallic candidate sites were obtained, corresponding to 3132 bimetallic candidate sites. Subsequently, the training and test sets were divided using structures as the smallest unit. The training set contained 1105 protein structures, corresponding to 2525 bimetallic candidate sites, involving 22 different metal-pair types; the test set contained 239 protein structures, corresponding to 607 bimetallic candidate sites, involving 6 metal-pair types. In this data distribution, ZN-ZN metal pairs were the predominant type, containing 843 and 205 candidate sites in the training and test sets, respectively.
[0063] The process of this embodiment includes metal-pair statistics, structural partitioning, Level 1 training / test set export, Phase 1 mixed probability distribution fitting, and prior scoring and distribution validation on the test set. In one feasible implementation, the script v1_1_metal_pair_stats.py performs metal-pair granular statistics and summaries on the Level 1 candidate data; v1_2_split_train_test_by_structure.py generates training and testing structure sets with structure as the smallest unit; v1_3_export_level1_dataset.py exports level1_train.tsv, level1_test.tsv, and hierarchical label subsets, and generates label counts and summaries of label distribution by metal-pair; subsequently, v1_4_fit_dinuclear_priors.py fits the Phase 1 prior distribution on level1_train.tsv and outputs priors_v1.json.gz; during this fitting process, a global prior distribution and a metal-pair-specific prior distribution are constructed simultaneously, and the minimum sample size threshold for constructing the metal-pair-specific prior is set by the parameter --min_pair_count. In this embodiment, the threshold is set to 50. That is, when the number of candidate sites corresponding to a certain metal-pair in the training set is less than 50, a separate prior distribution for that metal-pair is not constructed; instead, the global prior distribution is used as a backoff distribution for probability estimation. When the sample size is greater than or equal to 50, a separate prior distribution for that metal-pair is constructed. To verify the effect of the prior modeling, a scored_level1_test.tsv file is generated on the test set.
[0064] Histogram overlay (e.g.) Figure 4 As shown), box plot (such as) Figure 5 (as shown) and violin diagram (as shown) Figure 6 As shown in the figure, the score distribution of explicit and implicit positive samples is generally higher than that of negative samples, and the difficult negative samples are further shifted to the lower score range compared to other negative samples; the mean / median and p10-p90 interval statistics also show a consistent trend (e.g. Figure 7 As shown in the figure, this demonstrates that the constructed Phase1 mixed probability prior can form a stable hierarchical discrimination effect on the test set.
[0065] The complete batch processing flow in this embodiment can be implemented by serializing job_v1_database_deal.sh. The relevant outputs include priors_v1.json.gz, scored_level1_test.tsv, and distribution plots and statistical summary files, which can serve as supporting materials for the Phase 1 prior construction and interpretable verification of this invention.
[0066] Example 2: Phase 2 geometric feature extraction, v2 prior fitting and v2.5 fusion.
[0067] Based on the Level 1 standard row data and structural partitioning results of Example 1, Phase 2 coordination geometric configuration features are further constructed and fitted with the second layer statistical prior, which is finally fused with the v1 prior to form the v2.5 multilayer prior.
[0068] This embodiment first performs Phase 2 feature extraction on all Level 1 candidates, generating level1_all_labeled.phase2_geom_50000.tsv using v2_1_extract_phase2_features.py; in a feasible implementation, this step can be batch-processed using the job script job_b6_1_geom_phase2_50000.sh. Then, the training structure set obtained in Embodiment 1 is reused, and only a subset of the training data is used to fit the v2 prior. The two-level mixed probability distribution of Phase 2 features (GLOBAL and BY_METAL_PAIR) is constructed using v2_2_fit_dinuclear_priors_v2.py, and priors_v2.json.gz is output; during this fitting process, the minimum sample size threshold for constructing the metal-pair-specific prior is also set using the parameter --min_pair_count. In this embodiment, the threshold is set to 50. That is, when the number of candidate sites for a metal-pair in the Phase2 training subset is less than 50, a separate Phase2-specific prior distribution for that metal-pair is not constructed; instead, the global Phase2 prior distribution is used as the backoff distribution. When the sample size is greater than or equal to 50, a Phase2-specific prior distribution for that metal-pair is constructed. This process can be batch-processed using job_b6_2_fit_priors_v2.sh. After completing the two-layer prior fitting, priors_v1.json.gz and priors_v2.json.gz are uniformly organized and fused using v2p5_merge_priors.py to generate priors_v2p5.json.gz, thus forming a unified carrier for the multi-layer priors. To further verify the interpretability of the fused priors, the prior total score and component contributions can be output for the test set or hierarchical subsets using v2p5_score_with_priors.py. This can be used to diagnose the contribution structure of each feature layer to the total score, and the hierarchical differences of different label subsets can be explained by combining the distribution plot and statistical table.
[0069] This embodiment produces level1_all_labeled.phase2_geom_50000.tsv, priors_v2.json.gz, and priors_v2p5.json.gz, and provides interpretable verification results of the fused priors.
[0070] The technical effects of this invention are as follows: As can be seen from the above embodiments, this invention can complete the entire process verification on large-scale structural datasets, from Level 1 candidate data construction, structural partitioning, two-layer feature extraction to multi-layer prior fitting and fusion encoding, forming a reproducible and scalable multi-layer statistical prior system for bimetallic sites. Through the fitting of the Phase 1 basic layer prior and the scoring results on the test set, it can be observed that different label subsets exhibit stable hierarchical distribution characteristics in the total prior score. This indicates that the constructed mixed probability distribution can effectively characterize the joint statistical laws of inter-metal distance, bridging category, and coordination correlation features, and forms distinguishable statistical differences between difficult negative samples and positive samples, thereby improving the effectiveness and interpretability of the prior expression.
[0071] After feature extraction at the Phase 2 configuration layer and fitting of the v2 prior, this invention further incorporates higher-order laws such as coordination geometry type, configuration deviation degree, and angle / planarity into the same statistical framework, achieving a hierarchical modeling extension from the base layer to the configuration layer. A unified v2.5 prior file is obtained through the fusion encoding of the v1 and v2 priors. Prior parameters, binning information, feature directories, and version metadata are centrally organized in a structured compressed format, giving the prior system good maintainability and scalability, and enabling iterative updates for new metal pairs and new feature layers.
[0072] It should be noted that this specific embodiment is merely an explanation of the technical solution of the present invention and is not intended to limit the present invention. After reading this specification, those skilled in the art can make modifications to this embodiment without contributing any inventive step, but as long as they are within the scope of the claims of the present invention, they are protected by patent law.
Claims
1. A multi-level prior construction method for bimetallic sites based on large-scale structural statistics, characterized in that: Includes the following steps: S1, establish a standardized Level 1 candidate dataset; S2, Summary of the size statistics and distribution of metal-pair particles; S3, perform structural-level partitioning of the training and test sets: perform structural-level partitioning of the Level 1 bimetallic candidate dataset; S4, based on the structural partitioning results, exports the full candidate data of Level 1 into training and testing sets, and simultaneously exports training and testing subsets for hierarchical statistical modeling and comparative analysis. S5, establish a continuous-discrete mixed probability distribution model for the Phase1 features on the training subset, wherein the Phase1 features include at least continuous features of inter-metal distance, bridging category and / or coordination-related discrete features, to obtain the Phase1 prior; S6: Based on the Level1 bimetallic candidate dataset, extract the coordination geometry and configuration-related features of Phase2 to form the Phase2 feature table, and obtain the Phase2 training subset according to the structural level partition in S3. S7. On the Phase2 training subset, establish a continuous-discrete mixed probability distribution model for the Phase2 features to obtain the Phase2 prior. S8, the Phase1 prior and Phase2 prior are uniformly organized and fused into an encoding to output a multi-level prior file. The multi-level prior file contains at least a hierarchical index, feature directory, binning information, probability distribution parameters and version metadata. S9. Based on the multi-layer prior file, perform prior scoring on the test subset and output statistical summary results or visualization results.
2. The method according to claim 1, characterized in that: S1, establish a standardized Level 1 candidate dataset: obtain structural data from a protein structure database or obtain a bimetallic candidate data table by parsing structural data from a protein structure database, and construct a standardized Level 1 bimetallic candidate dataset; the Level 1 bimetallic candidate dataset includes at least structural identifiers, types of two metal elements and inter-metal distance information, and further includes one or more of the following: bridging category, bridging subclass, coordination number, coordinating atom type and coordinating residue type.
3. The method according to claim 2, characterized in that: The Level 1 bimetal candidate dataset is constructed by a hierarchical labeling system, which includes at least explicit positive samples, implicit positive samples, difficult negative samples, and other negative samples. Based on the hierarchical labeling system, the system outputs the count summary of each label subset and its distribution statistics by metal pair.
4. The method according to claim 1, characterized in that: The execution process of S2. Scale statistics and distribution summary of metal-pair granularity includes the following steps: First, read the standardized Level 1 candidate dataset, parse the metal element types and combine them into metal-pair keys; then, perform counting statistics with metal-pair as keys, output the overall distribution, head and long tail lists, and can further summarize the counting information on the label dimension; finally, export the statistical results as a table file for subsequent data partitioning, prior fitting and quality checks.
5. The method according to claim 1, characterized in that: The structural-level partitioning described in S3 uses structures or structural clusters as the smallest unit for sampling and isolation, and generates a training structure set, a test structure set, and corresponding training and test subsets based on the partitioning results.
6. The method according to claim 1, characterized in that: The construction of the continuous-discrete mixed probability distribution model adopts the binning histogram method and / or the category frequency statistics and normalization to obtain the category probability distribution. When the construction of the continuous-discrete mixed probability distribution model adopts the binning histogram method, the binning histogram statistics adopt smoothing to avoid the zero probability problem. The smoothing includes Laplace smoothing or additive smoothing. Then the binning boundary and the number of bins are written as binning information into the Phase1 prior. When the construction of the continuous-discrete mixed probability distribution model adopts the method of obtaining the category probability distribution by frequency statistics and normalization, the discrete features in the continuous-discrete mixed probability distribution model include at least one of bridging categories, bridging subclasses, and minimum coordination number.
7. The method according to claim 1, characterized in that: Both the Phase1 prior in S5 and the Phase2 prior in S7 simultaneously construct a global prior distribution and a specific prior distribution partitioned by metal pairs. The global prior distribution is used for backtracking long-tailed metal pairs, while the specific prior distribution partitioned by metal pairs is used for fine-grained expression of metal pairs with sufficient samples.
8. The method according to claim 7, characterized in that: In the dedicated prior distribution divided by metal pairs, when the sample size of the target metal pair is lower than a preset threshold, the global prior distribution is used as a backoff distribution to estimate the probability of the metal pair. The preset threshold is a configurable parameter, and the configurable parameter is 50. It is used to distinguish between metal pairs with sufficient samples and long-tailed metal pairs, and writes the version metadata of the multi-layer prior file.
9. The method according to claim 1, characterized in that: The Phase2 feature table in S6 includes at least one or more of the following: coordination geometry type label, configuration deviation index, key angle statistical index, and planarity index.
10. The method according to claim 1, characterized in that: When establishing a continuous-discrete mixed probability distribution model for the Phase2 features in S7, the continuous features are statistically analyzed using binned histograms and then smoothed, while the discrete features are statistically analyzed using class frequency and then normalized to obtain the class probability distribution.
11. The method according to claim 1, characterized in that: The unified organization and fusion coding in S8 includes: parsing the distribution blocks, binning definitions and feature directories of Phase1 prior and Phase2 prior, performing consistency checks on the binning information of the two priors, and writing the two priors into the same multi-level prior file according to a preset hierarchical structure.
12. The method according to claim 11, characterized in that: The multi-layered prior files are output in a structured compression format, which includes JSON compression format or JSON-GZ format.
13. The method according to claim 11, characterized in that: The version metadata includes at least one of the following: build time, data size identifier, feature list, binning configuration, and prior version number.
14. The method according to claim 1 or 11, characterized in that: The multi-layered prior file is organized using a distributed block structure, which includes at least a Phase1 distributed block and a Phase2 distributed block, and each distributed block contains a global distributed sub-block and a metal-pair distributed sub-block respectively; the hierarchical index is used to indicate the mapping relationship between the distributed blocks, feature directories and binning information.
15. The method according to claim 1, characterized in that: S9, which involves prior scoring the test subset based on the multi-layer prior file, specifically includes: reading the test subset and the multi-layer prior file, loading the global prior distribution and the specific prior distribution divided by metal pairs; selecting the corresponding specific prior distribution or global backoff distribution according to the metal pair type of the candidate site; calculating the probability contribution or log probability contribution of the candidate site under the Phase1 feature and Phase2 feature respectively; summarizing the probability contribution or log probability contribution of each feature into the prior total score of the candidate site according to a preset combination rule, and outputting the prior total score and the sub-item contribution corresponding to each feature; grouping and statistically analyzing the prior total score according to the label category, generating one or more statistical results among the mean, median, and quantile, and / or generating one or more visualization results among the box plot, violin plot, and histogram overlay plot, to verify the discriminative power and stability of the multi-layer prior.
16. A system constructed using the multilayer prior construction method for bimetallic sites based on large-scale structural statistics as described in any one of claims 1-15, characterized in that: It includes a data acquisition and standardization construction module, a metal pair statistics module, a structure-level partitioning module, a first-level data export module, a first-layer prior fitting module, a second-layer feature extraction module, a second-layer prior fitting module, a multi-layer prior fusion coding module, and a scoring verification module; The data acquisition and standardization construction module is used to execute S1 in the method for constructing a multi-level prior of bimetallic sites based on large-scale structural statistics as described in any one of claims 1-15, to acquire structural data or candidate data tables, and to construct a standardized first-level bimetallic candidate dataset. The metal pair statistics module is used to execute S2 in the method for constructing a multi-level prior of bimetallic sites based on large-scale structural statistics as described in any one of claims 1-15, and to perform counting statistics and distribution summarization of the first-level bimetallic candidate dataset by metal pair. The structure-level partitioning module is used to execute S3 in the bimetallic site multilayer prior construction method based on large-scale structural statistics as described in any one of claims 1-15, and to generate a training structure set and a test structure set based on the structure identifier or structure cluster identifier. The first-level data export module is used to execute S4 in the method for constructing a multilayer prior of bimetallic sites based on large-scale structural statistics as described in any one of claims 1-15, and to derive a training subset, a test subset and a hierarchical label subset according to the training structure set and the test structure set. The first-layer prior fitting module is used to execute S5 in the method for constructing a multi-layer prior of bimetallic sites based on large-scale structural statistics as described in any one of claims 1-15, and to establish a continuous-discrete mixed probability distribution model for the Phase1 feature on the training subset to obtain the Phase1 prior. The second layer feature extraction module is used to execute S6 in the method for constructing a multilayer prior of bimetallic sites based on large-scale structural statistics as described in any one of claims 1-15, and extract coordination geometry and configuration-related features to form the Phase2 feature table. The second-layer prior fitting module is used to execute S7 in the method for constructing a multilayer prior of bimetallic sites based on large-scale structural statistics as described in any one of claims 1-15, to establish a continuous-discrete mixed probability distribution model on the Phase2 training subset to obtain the Phase2 prior; The multi-layer prior fusion encoding module is used to execute S8 in the method for constructing a multi-layer prior of bimetallic sites based on large-scale structural statistics as described in any one of claims 1-15, to uniformly organize and fuse the Phase1 prior and Phase2 prior, and output the multi-layer prior file. The scoring and verification module is used to execute S9 in the method for constructing a multi-layer prior of bimetallic sites based on large-scale structural statistics as described in any one of claims 1-15, to score the test subset based on the multi-layer prior file, output the total prior score and the contribution of each item, and generate a statistical summary result or distribution map.
17. The system according to claim 16, characterized in that: Both the first-layer prior fitting module and the second-layer prior fitting module are configured to simultaneously construct a global prior distribution and a specific prior distribution based on metal pairs.
18. A readable storage medium having a computer program stored thereon, the computer program, when executed by a processor, implementing the method for constructing multilayer priors of bimetallic sites based on large-scale structural statistics as described in any one of claims 1-15.