Multifunctional catalyst performance prediction method based on high-throughput calculation and machine learning

By combining high-throughput computing and machine learning, constructing a dataset, and using the interpretable model SISSO to establish descriptor formulas, the problem of low efficiency in screening multifunctional catalysts in existing technologies is solved, enabling rapid prediction and optimization of catalyst performance.

CN122024909APending Publication Date: 2026-05-12NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NORTHWESTERN POLYTECHNICAL UNIV
Filing Date
2026-03-28
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately screen for multifunctional catalysts that can meet the needs of various applications, and machine learning methods lack clear interpretability in complex environments, resulting in low catalyst development efficiency.

Method used

By combining high-throughput computing and machine learning, and through the construction of datasets, multi-objective Bayesian optimization, and the interpretable machine learning model SISSO, a descriptor formula is established to achieve rapid analysis and optimization of catalyst performance.

Benefits of technology

It significantly improves the efficiency of catalyst research and development, reduces the consumption of material resources and reliance on labor-intensive experiments, and provides clear mathematical expressions to guide the development and improvement of multifunctional catalysts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122024909A_ABST
    Figure CN122024909A_ABST
Patent Text Reader

Abstract

The invention discloses a multifunctional catalyst performance prediction method based on high-throughput calculation and machine learning, and relates to the technical field of catalytic material prediction. The method comprises the following steps: acquiring basic structure information of a target catalyst, correspondingly calculating catalytic performance data and related characteristic data of the target catalyst, and pairing the basic structure information and the related characteristic data to establish a data set; importing the data set into a plurality of machine learning models for training, and performing hyper-parameter tuning by using multi-target Bayesian optimization; based on the screened optimal model, sorting and discriminating key features influencing the performance of the catalyst by using an SHAP value, and importing the key features into an interpretable machine learning model SISSO for multi-task training; based on an interpretable machine learning model SISSO and through a symbol regression method, obtaining an explicit mathematical relationship between the feature combination and the catalytic performance as a descriptor formula for application; the multi-aspect catalytic performance of the catalyst can be rapidly predicted, and the research and development efficiency of the catalyst is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of catalytic material prediction technology, and particularly relates to a method for predicting the performance of multifunctional catalysts based on high-throughput computing and machine learning. Background Technology

[0002] Catalysis is the cornerstone of modern chemical industry. With the aid of catalysts, raw materials can be gradually transformed into high-value-added products. Different catalysts exhibit significant differences in catalytic activity, which directly determines the kinetic rate and the depth of the reaction. Suitable catalysts can reduce the conditions required for chemical reactions to occur, thus allowing product production and energy conversion to take place under relatively relaxed conditions. For example, room-temperature sodium-sulfur batteries (RT-SSBs) are highly efficient due to their high theoretical energy density (1274 Wh / kg). -1 Na₂S₄, with its abundant natural resources and low cost, is considered the most promising large-scale energy storage technology. However, its practical application still faces many challenges, mainly including the "shuttle effect" caused by the dissolution of polysulfides and the high energy barriers accompanying the liquid-solid (Na₂S₄→Na₂S₂) and solid-solid (Na₂S₂→Na₂S) phase transitions during discharge, resulting in slow overall reaction kinetics. To address these issues, the introduction of catalysts has proven to be an effective strategy. However, the development of traditional catalysts mainly relies on an experience-driven "trial and error" approach, i.e., screening materials through repeated synthesis, characterization, and performance testing. This approach is not only time-consuming and resource-intensive but also highly dependent on the researcher's personal experience, resulting in slow progress in the rational design of catalysts, often limited by the chemist's intuitive understanding.

[0003] In recent years, with the increasing maturity of theoretical calculation methods and the significant improvement in computing power, researchers have been able to predict catalytic performance and explore reaction mechanisms using methods such as first-principles calculations. However, the complexity of actual reaction conditions places higher demands on catalysts, and single-function catalysts are no longer sufficient to meet practical needs. In actual reaction systems, simply pursuing high reaction rates is often insufficient to meet application requirements; catalysts also need to possess multiple functional properties to adapt to complex reaction environments. For example, for easily soluble intermediate products during the reaction process, the catalyst needs to have specific adsorption capabilities to confine them at the reaction interface, prevent the loss of active substances, and maintain the continuous progress of the reaction. Therefore, for specific reactions, it is necessary to comprehensively consider their characteristics and search for multifunctional catalysts, making the development of "multifunctional catalysts" an inevitable choice.

[0004] These catalysts require the synergistic processing of continuous intermediate adsorption and transformation processes, leading to highly dynamic shifts in the rate-controlling step within the reaction network, making it difficult to accurately define the reaction mechanism. This complexity makes it challenging to accurately screen catalysts that can meet diverse application requirements using only first-principles calculations. Therefore, there is an urgent need to develop a new method capable of accurately predicting multifunctional catalysts.

[0005] Against this backdrop, machine learning methods, with their ability to automatically learn from data and optimize decisions, have gradually become an important tool for accelerating catalyst design. Thanks to the continuous evolution of algorithmic models and the significant enhancement of computational processing power, current machine learning algorithms can efficiently extract patterns from massive amounts of experimental data and theoretical calculations, and predict unknown outcomes through constructed data models. However, existing machine learning algorithms (e.g., random forests, neural networks, etc.) are typically black-box models. Although these models can achieve high prediction accuracy through training on massive amounts of data, their internal representation mechanisms are extremely complex and highly abstract, making it difficult to establish an explicit physical mapping relationship between the model's input variables and output results. In other words, these black-box models often lack clear interpretability; users cannot intuitively understand the key feature factors affecting the prediction results and their specific weight contributions, resulting in an inability to deeply reveal the underlying patterns behind the data at the mechanistic level. This, to some extent, limits their application and promotion in fields requiring highly reliable interpretation.

[0006] Therefore, the present invention aims to provide an intuitive descriptive extraction method that combines uninterpretable machine learning models with interpretable machine learning models, so as to achieve rapid analysis and optimization of candidate catalysts. Summary of the Invention

[0007] The purpose of this invention is to provide a multifunctional catalyst performance prediction method based on high-throughput computing and machine learning, in order to solve the problems mentioned in the background art, such as the difficulty of accurately screening catalysts that can meet the needs of multiple applications by relying solely on first-principles calculations, and the difficulty of machine learning methods in predicting catalyst performance by understanding the deep correlations behind the data in complex environments.

[0008] To achieve the above objectives, the present invention employs the following technical solution:

[0009] In its first aspect, this invention proposes a method for predicting the performance of multifunctional catalysts based on high-throughput computing and machine learning, comprising the following steps: S1. Dataset Construction: Obtain the basic structural information of the target catalyst from public databases, and calculate the catalytic performance data and related characteristic data of the target catalyst accordingly. Pair the catalytic performance data and related characteristic data to establish a dataset. S2. Machine Learning Modeling and Evaluation: Divide the dataset into training and testing sets, and import them into multiple machine learning models for training. Use multi-objective Bayesian optimization to tune the hyperparameters of multiple machine learning models, search for the best parameter combination, and evaluate the model fitting effect. S3. Interpretable Machine Learning: Based on the selected best model, the key features affecting catalyst performance are identified by sorting the SHAP values. After deduplication, these features are imported into the interpretable machine learning model SISSO for multi-task training. S4. Descriptor Construction and Application: Based on the interpretable machine learning model SISSO, the explicit mathematical relationship between feature combinations and catalytic performance is searched in the feature space as a descriptor formula; multiple key performance characteristics of the target catalyst are predicted simultaneously using a single descriptor formula; a performance-feature variation curve is established based on the descriptor formula; the feature corresponding to the optimal solution is found; and potential high-performance multifunctional catalysts are found accordingly.

[0010] Preferably, the calculation of the catalytic performance data and related characteristic data of the target catalyst in S1 is as follows: Calculate the catalytic performance data of the target catalyst: obtain the basic structural information of the target catalyst from public databases and construct a catalyst surface model; use first principles to obtain the adsorption configuration, find the optimal adsorption configuration and obtain the adsorption energy, and search for reaction pathways to obtain the reaction energy barrier; Calculate relevant characteristic data of the target catalyst: Extract the catalytic active center characteristics of the target catalyst from public databases. The catalytic active center characteristics include physical characteristics. After cleaning by the program, the calculated electronic structure characteristics are introduced and finally correlated with the multiple performance indicators of the catalyst obtained by theoretical calculation.

[0011] Furthermore, a program is used to assist in finding the optimal adsorption configuration, i.e. the most stable configuration. This program can batch traverse different surfaces, automatically build and submit computational tasks for multiple adsorbent species, and complete file copying, structure generation and job submission in one go, significantly improving the computational efficiency of adsorption configuration screening.

[0012] Furthermore, reaction pathways are searched and analyzed to determine the structures of each intermediate and transition state. The configurations of catalyst adsorption intermediates and transition states are optimized, and the energy of catalyst adsorption intermediates and the energy barriers for transition between intermediates are calculated using the climbing image elastic band method.

[0013] Furthermore, the catalytic performance data are output features predicted from input features collected using first-principles calculations, including: discharge reaction energy barrier, adsorption energy of reactants, and reactant decomposition reaction energy barrier.

[0014] Furthermore, the physical characteristics include electronegativity ( atomic mass () atomic radius () ), Fermi level ( ), first ionization energy ( ), electron affinity ( ), valence electron number ( ); The electronic structure features include the outermost orbital spin-up of the atom (… ) and downward ( The number of electrons filled, the center of the outermost orbital ( ), bond length ( ).

[0015] Preferably, the dataset is established in S1 as follows: For the preprocessing of relevant feature data, a Python-based program is used to organize features from different databases and introduce Pearson correlation coefficient to analyze the correlation between features. Features with highly redundant information are filtered out to simplify the input and avoid collinearity. Then, standardization is used to transform the original data into dimensionless, order-of-magnitude standardized values. The Pearson correlation coefficient method measures the degree of linear correlation between two data sets, with a value between -1 and 1. A strong linear correlation indicates highly redundant information, requiring the removal of these feature columns to simplify the input features and avoid collinearity. Furthermore, this method can also identify feature columns highly correlated with catalytic performance in advance.

[0016] A dataset was created by pairing catalytic performance data with processed relevant feature data.

[0017] Preferably, the machine learning modeling in S2 is as follows: The machine learning model employs at least one of random forest, gradient boosting, adaptive boosting, autocorrelation determination regression, and support vector regression. Furthermore, during the training process of the machine learning model, if the number of training samples is lower than a preset minimum threshold, only one sample is selected as the test set. If the number of training samples is higher than a preset maximum threshold, the training samples are divided into a training set, a test set, and a validation set, with a division ratio of 80%, 16%, and 4%, respectively. If the number of training samples is higher than the preset minimum threshold but lower than the preset maximum threshold, the training samples are divided into a training set and a test set, with a division ratio of 80% and 20%, respectively.

[0018] Preferably, when using multi-objective Bayesian optimization to tune the hyperparameters of the model, the optimization index should be selected from at least two of the following: maximizing prediction accuracy, maximizing robustness, maximizing sparsity, minimizing prediction uncertainty, and minimizing model complexity.

[0019] Furthermore, when using multi-objective Bayesian optimization to tune the hyperparameters of the model, the optimization process is constrained by the practical electrochemical meaning to avoid obtaining hyperparameter combinations that only perform well statistically but lack rationality at the electrochemical level. For example, the relationship between the catalyst's adsorption energy for reactants and the discharge energy barrier is constructed to constrain the process, and a reasonable range for the catalyst's adsorption energy for reactants is set according to the Sabatier principle.

[0020] Specifically, physicochemical principles are incorporated as constraints into the optimization process to ensure that the model predictions conform to the basic laws of electrochemical reactions. The physicochemical constraints introduced are either the Sabatier principle or the Brønsted-Evans-Polanyi (BEP) relationship.

[0021] Preferably, the model accuracy is evaluated using the coefficient of determination (R²), root mean square error (RMSE), mean absolute error (MAE), and maximum absolute error (MaxAE).

[0022] Preferably, the key features affecting catalyst performance in S3 are as follows: Based on the SHAP value, the top ten features with the greatest impact on each catalytic performance were extracted from the optimal machine learning model corresponding to that catalytic performance.

[0023] Preferably, the interpretable machine learning model in S3 is a deterministic independence filter and sparse operator (SISSO).

[0024] Preferably, in S3, the interpretable machine learning model SISSO is used for multi-task training, as follows: The top ten features that have the greatest impact on each catalytic performance are summarized. After removing duplicates, the remaining features are used as input data and the corresponding catalytic performance is used as output data to construct a new dataset. This dataset is then input into the interpretable machine learning model SISSO for multi-task training. The model is evaluated using the leave-P method for cross-validation, and the optimal model is selected to predict the catalytic performance of the catalyst.

[0025] Preferably, the descriptor construction and application in S4 are as follows: When training with the interpretable machine learning model SISSO, the parameters of the interpretable machine learning model SISSO include: number of multi-task learning (ntask), number of features (nsf), set of mathematical operators (ops), dimension (desc_dim), feature complexity (fcomplexity), and feature unit set (funit), which are used to determine whether the selected descriptor has an explicit mathematical expression and conforms to the structure-activity relationship of physical and chemical intuition. The descriptor formula is evaluated, and the optimal mathematical expression is selected as the final descriptor.

[0026] Preferably, by substituting the values ​​of the desired features into the mathematical expression of the descriptor, multiple catalytic properties of the catalyst can be predicted simultaneously, including: the discharge reaction energy barrier, the adsorption energy of the reactants, and the reactant decomposition reaction energy barrier.

[0027] Preferably, in step S4, the search is for potential high-performance multifunctional catalysts, specifically as follows: An activity trend graph is constructed based on the descriptor formula to find the physical characteristics corresponding to the local or global highest points of the curve, and then to find the corresponding materials in reality.

[0028] To facilitate the implementation of catalyst screening methods, this invention provides, in a second aspect, a multifunctional catalyst performance prediction system based on high-throughput computing and machine learning, comprising: The model building unit is used to build catalyst models and acquire feature data. It uses the adsorption energy, reaction energy barrier and other data obtained from the calculation simulation as output data, and combines the screened atomic features as input data to train machine learning association models for different catalytic needs. Furthermore, the specific steps of constructing a catalyst model and obtaining feature inputs through the model building unit in this invention include: obtaining the basic structure through a publicly available material structure library, and constructing a specific catalyst model using software; performing quantum chemical calculations based on a first-principles software package to determine the most stable adsorption configuration and obtain the adsorption energy, and searching for reaction pathways to obtain the reaction energy barrier; integrating multiple publicly available databases to obtain atomic features, using a program to remove redundant and null features, and using correlation coefficients to screen out highly linearly correlated features, thereby using these as feature inputs for machine learning.

[0029] The validation and evaluation unit divides the data into training and test sets, uses K-fold cross-validation to optimize and evaluate the model, and selects the machine learning model with the best prediction performance and the feature importance ranking. Furthermore, the specific steps for the verification and evaluation unit to optimize and evaluate the model in this invention are as follows: the data on catalytic performance and related characteristics are divided into a training set and a test set, wherein the amount of data in the training set accounts for 70-80% of the total amount of data, and the K-fold cross-validation method is used to divide the training set; A machine learning model was established for various catalytic requirements of multifunctional catalysts, and the adsorption energy and reaction energy barrier obtained from the calculation simulation were used as output data. After the training set met the convergence criteria, the model was used to predict the catalytic performance in the test set and compared with the calculation results until the convergence condition was met. Furthermore, the model with the best predictive catalytic performance was selected by evaluation indicators such as the coefficient of determination and root mean square error, and the importance of features was ranked.

[0030] The descriptor generation unit synthesizes features ranked highest in importance from the optimal model, removes duplicates, and inputs them into the interpretable machine learning model SISSO to obtain mathematical expression models of the relationship between features and catalytic performance under different dimensions and feature complexities. The accuracy of the fitted formula is evaluated using leave-one-out cross-validation or leave-P cross-validation, and the model with the most accurate predictions is selected as the final descriptor. This descriptor is then used to predict the catalytic performance of other catalysts.

[0031] Furthermore, the specific steps for the descriptor generation unit to obtain descriptors in this invention include: integrating the features ranked first in importance from the optimal machine learning models for different catalytic requirements, constructing a dataset with catalytic performance after removing duplicates; inputting the dataset into the interpretable machine learning model SISSO to obtain models between specific features and catalytic performance under different dimensions and feature complexities; determining whether to use leave-one-out cross-validation or leave-P cross-validation based on the dataset size, and selecting the most accurate model as the final descriptor by combining the model and validation results.

[0032] Compared with the prior art, the beneficial effects of the present invention are: (1) The method in this invention combines mechanism analysis and batch processing to determine the structure of each intermediate and transition state and calculate the energy and energy barrier. Then, features are extracted from the database and the feature set is screened using the Pearson correlation coefficient. The calculated catalytic performance and features are then input into machine learning models such as random forests for training to obtain multiple preset catalytic performance outputs. The optimal model is selected based on the evaluation index. Finally, the optimal machine learning model is used to rank the importance of features and select the most important features. Deterministic independence screening and sparse operators are used to construct an interpretable model that conforms to physicochemical intuition. The optimal model is selected through cross-validation for performance prediction, thereby establishing a descriptor for rapidly predicting multifunctional catalyst molecules and realizing rapid prediction of various aspects of catalyst catalytic performance, which greatly improves the efficiency of catalyst research and development.

[0033] (2) The method in this invention establishes a complete workflow that includes catalyst molecule preparation, mechanism analysis, energy calculation of intermediate molecules and transition state molecules, data preprocessing, machine learning algorithm selection, model optimization, model evaluation, feature secondary processing, interpretable machine learning, model data analysis, and descriptors with clear mathematical expressions. This significantly reduces the consumption of large amounts of material resources and reliance on labor-intensive experimental evaluation. The conclusions drawn can effectively guide the development, synthesis and improvement of multifunctional catalyst molecules. Attached Figure Description

[0034] Figure 1 This is a flowchart of the multifunctional catalyst performance prediction method based on high-throughput computing and machine learning in Embodiment 1 of the present invention; Figure 2 This is a flowchart of the multifunctional catalyst performance prediction method in Example 2 of the present invention. Detailed Implementation

[0035] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0036] Example 1: The black-box nature of traditional machine learning makes it difficult to uncover the deep relationships behind data, thus hindering its ability to fully trust decision-making in complex environments. This invention proposes a multifunctional catalyst performance prediction method based on high-throughput computing and machine learning, such as... Figure 1 As shown, it includes the following steps: S1: Construct a catalyst model.

[0037] First, the basic structural information of the target catalyst is obtained from publicly available material structure databases (such as Materials Project, PubChem, Crystallography Open Database, etc.). Then, the obtained basic structure is processed using material modeling software (such as Materials Studio, VESTA, etc.), and a specific catalyst surface model is constructed by adding or deleting atoms, cutting the surface along specific crystal plane indices, and constructing heterojunctions.

[0038] S2: Quantum chemical calculations and data acquisition.

[0039] After the model was built, quantum chemical calculations were performed using first-principles calculation software packages based on density functional theory (such as VASP, CP2K, Gaussian, etc.). During the calculation, the most stable structure of the interaction between reactant molecules, intermediate products and surface model was determined by calculating the surface Gibbs free energy under specific reaction temperature and pressure conditions.

[0040] To achieve high-throughput computing, an automated Python script was developed and implemented. This script is used for batch construction and job submission of catalyst adsorption models for reactants or intermediates. After calculation, it can automatically extract energy data. Then, by comparing the energy of different adsorption configurations, the most stable adsorption configuration is determined, and the adsorption energy of the catalyst for reactants or intermediates is calculated accordingly. In addition, molecular dynamics simulations or the climbing elastic band method (CI-NEB) are used to search for reaction pathways between each adsorption steady state and obtain the corresponding energy barriers and other information. These key results are then compiled and used as input data for subsequent machine learning models.

[0041] S3: Acquire and integrate atomic feature data to build a feature database.

[0042] Intrinsic characteristic information of each atom constituting the catalytic active center of the catalyst is obtained from multiple publicly available elemental databases (e.g., Magpie, Mendeleev, Villars, etc.). This information can also be supplemented by electronic structure information of the material obtained through structure optimization. The aforementioned atomic characteristics include, but are not limited to: electronegativity (…). atomic mass () atomic radius () ), Fermi level ( ), first ionization energy ( ), electron affinity ( ), valence electron number ( Physicochemical parameters such as the outermost orbital spin-up of the atom obtained through calculation ( ) ) and downward ( The number of electrons filled, the center of the outermost orbital ( )wait.

[0043] After acquiring the raw feature data, a data processing program is used to clean and integrate the multi-source feature database, which serves as the standardized feature input for the machine learning model. The program integrates the feature database using a Python-based data analysis script to analyze feature lists from different databases. First, data cleaning is performed: identical feature columns appearing repeatedly in different databases are retained only once to remove duplicates; simultaneously, data integrity is checked, removing feature columns with a large number of null or missing values. Next, feature filtering is performed to reduce data dimensionality. The Pearson correlation coefficient between each feature column is calculated, and feature pairs with an absolute correlation coefficient greater than a preset threshold (e.g., 0.9) are removed, thus eliminating highly linearly correlated redundant feature columns, simplifying the input feature space for machine learning, and improving model training efficiency.

[0044] S4: Splitting the dataset and setting up cross-validation.

[0045] The catalytic performance data of the catalyst collected in step S2 (such as adsorption energy, reaction energy barrier, etc., as labels) are paired with the relevant feature data (as features) obtained in step S3, and divided into training set and test set. In order to ensure that the model has good generalization ability and prevent overfitting, the training set data accounts for 70% to 80% of the total data, and the remaining 20% ​​to 30% is used as the test set.

[0046] During the model training phase, the training set is further divided using the K-fold cross-validation method. Specifically, the training set data is randomly divided into K subsets (e.g., K=5 or K=10). Each time, one subset is selected as the validation set, and the remaining K-1 subsets are used as the training set for model training and validation. This process is repeated K times, and the average of the K validation results is taken as the evaluation metric for model performance, thereby ensuring the robustness of the model parameters.

[0047] S5: Construct a machine learning prediction model for a multifunctional catalyst.

[0048] To address the diverse catalytic requirements of multifunctional catalysts (such as adsorption strength of intermediate products and reaction energy barriers for catalytic conversion), corresponding machine learning models are established. Using the characteristic properties of each atom in the catalyst model (features filtered through step S3) as input data, and the adsorption energy and reaction energy barriers calculated in step S2 as output data, multiple different types of machine learning algorithm models are constructed, such as neural networks, random forests, gradient boosting trees, or support vector machines. These models are trained using training set data to obtain a model of the catalyst's specific catalytic performance.

[0049] During training, the change in the loss function is monitored. After the training set meets the preset convergence criteria (e.g., the loss function decreases below a threshold or a preset number of iterations is reached), the trained model is used to predict the catalytic performance of catalysts in the test set, and the prediction results are compared with the catalytic performance labels calculated in step S2. If the model's performance on the test set does not meet the preset convergence conditions (e.g., the prediction error is greater than a preset threshold), the hyperparameters of the machine learning model are adjusted, and the model is further optimized using the training set until the model's prediction of the catalyst catalytic performance in the test set meets the convergence conditions and has high accuracy.

[0050] S6: Model evaluation and feature importance ranking.

[0051] After obtaining multiple trained machine learning models, their predictive performance is quantitatively evaluated using defined evaluation metrics. These metrics include, but are not limited to, the coefficient of determination (R²), root mean square error (RMSE), mean absolute error (MAE), and maximum absolute error (MaxAE). By comparing the evaluation metrics of each model, the model with the best predictive catalytic performance is selected as the benchmark model.

[0052] Furthermore, the feature importance analysis function built into the optimal machine learning model (such as Gini importance of random forest or gradient-based feature importance) is used to analyze the input features, calculate the contribution of each feature to the prediction result of catalytic performance, and rank all features according to their importance to select the key features that contribute significantly to the model prediction.

[0053] S7: Extract descriptors based on interpretable machine learning.

[0054] A comprehensive analysis is conducted on the optimal machine learning models fitted for different catalytic requirements obtained in step S6. Based on the SHAP values, the top ten features of each model are extracted according to their importance. The top ten features of all models are then summarized, and after removing duplicates, the remaining features are used as input data, with the corresponding catalytic performance as output data, to construct a new dataset.

[0055] The dataset is input into the interpretable machine learning model SISSO. The SISSO program can search in a huge feature space in single-task or multi-task mode, and obtain explicit mathematical models (i.e. mathematical expressions) between specific feature combinations and single or multiple catalytic performances under different dimensions (number of feature terms) and different feature complexities through symbolic regression.

[0056] To ensure the accuracy and generalization ability of the fitted formula, the evaluation method was determined based on the total number of samples in the newly constructed dataset: leave-one-out cross-validation was used when the dataset size was small; leave-P cross-validation was used when the dataset size was large. Finally, considering the model's complexity, physical interpretability, and cross-validation results, the model with the most accurate prediction was selected as the final descriptor for predicting the catalytic performance of this type of catalyst.

[0057] S8: Predicting the performance of new catalysts using descriptors.

[0058] Once the descriptors with clear physical meaning and mathematical form are obtained, they can be applied to the rapid screening of novel catalyst materials. For other candidate catalyst materials that have not undergone first-principles calculations, only their atomic characteristic parameters need to be obtained and substituted into the mathematical expression of the descriptor obtained in step S7 to quickly calculate the predicted catalytic performance of the catalyst, thus eliminating the need for expensive quantum chemical calculations and greatly improving the efficiency of catalyst development. Simultaneously, a performance-feature curve can be established based on the descriptor formula to find the feature corresponding to the optimal solution, and based on this, potential high-performance multifunctional catalysts can be identified.

[0059] Example 2: This embodiment uses the process of screening axially coordinated transition metal single-atom catalysts for sodium-sulfur batteries as an example. Figure 2 As shown, the steps for screening catalytic materials are as follows: S1: Construct a catalyst model.

[0060] The graphite structure file was opened in Materials Studio software. Excess atomic layers were removed, leaving only a single layer of atoms to construct the graphene substrate. Subsequently, some carbon atoms were removed from this graphene substrate and M-N4 structural units were inserted to form metal-nitrogen coordination active centers. The metal center M consisted of 27 transition metal atoms from 3d, 4d, and 5d, and was combined with six axial ligands (F, Cl, Br, I, OH, NH2) to construct a total of 162 different axially coordinated single-atom catalyst surface models.

[0061] S2: Quantum chemical calculations and data acquisition.

[0062] After the model was built, the theoretical calculation software VASP was used for structural optimization and energy calculation. To efficiently handle large batches of models, an automated script based on Python was developed and run to achieve the following workflow: automatically and batch-build models of a series of reactants and intermediates such as S8, Na2S8, Na2S6, Na2S4, Na2S2, Na2S, Na, and S, and submit them to VASP calculation jobs. After the calculation, the script automatically extracts energy data, determines the most stable configuration by comparing the energies of different adsorption configurations, and calculates the adsorption energy of the catalyst for each species accordingly. Furthermore, the reaction energy barrier of the critical step Na2S decomposition process is calculated using the climbing elastic band method (CI-NEB). Finally, three key catalytic performance data are extracted and organized: the discharge reaction energy barrier, the adsorption energy of Na2S6 (used to evaluate the ability to suppress the shuttle effect), and the Na2S decomposition reaction energy barrier.

[0063] S3: Acquire and integrate atomic feature data.

[0064] Intrinsic physicochemical characteristics of the active center of the catalyst (i.e., the metal atom M and axial ligands in the M-N4 structure) were extracted from three public elemental databases: Magpie, Mendelev, and Villars. These characteristics include, but are not limited to, electronegativity, atomic mass, atomic radius, Fermi level, first ionization energy, electron affinity, and number of valence electrons. A data processing script written in Python was used to clean and integrate the features from multiple databases: first, duplicate feature columns between different databases were removed; second, features with too many missing values ​​were checked and deleted; finally, feature filtering was performed, and the Pearson correlation coefficients between features were calculated. Redundant features with an absolute correlation coefficient greater than 0.9 were removed, resulting in a simplified feature set, which served as input data for subsequent machine learning.

[0065] S4: Splitting the dataset and setting up cross-validation.

[0066] The three key catalytic performance data (labels) obtained in step S2 are paired with the corresponding feature data obtained in step S3 to form a complete dataset. The data is then divided into a training set (70%~80% of the total data) and a test set (20%~30%).

[0067] During the model training phase, K-fold cross-validation is used to further divide the training set, and the average value of the validation results is taken as a preliminary evaluation of the model performance to ensure the robustness of the model parameters and prevent overfitting.

[0068] S5: Build machine learning prediction models.

[0069] For three key catalytic performance parameters (discharge reaction energy barrier, adsorption energy for Na2S6, and decomposition reaction energy barrier for Na2S), corresponding machine learning prediction models were established. Using the atomic features selected in step S3 as input and the corresponding performance data calculated in step S2 as output, five algorithms, including random forest, gradient boosting tree, and adaptive boosting, were used for model training. The model was trained using the training set data, and the change in the loss function was monitored. When the training set reached a preset convergence criterion (e.g., the decrease in the loss function was below a threshold or the maximum number of iterations was reached), the model's prediction performance was evaluated using the test set. If the prediction error on the test set exceeded a preset threshold, the model's hyperparameters were adjusted and retrained until the model achieved a high prediction accuracy on the test set.

[0070] S6: Model evaluation and feature importance ranking.

[0071] The coefficient of determination (R²), root mean square error (RMSE), mean absolute error (MAE), and maximum absolute error (MaxAE) were used as evaluation metrics to quantitatively evaluate and compare multiple machine learning models trained for each catalytic performance in step S5. Based on the evaluation results, the model with the best predictive performance was selected as the random forest model.

[0072] Furthermore, the feature importance analysis function built into the random forest model was used to calculate the contribution of each input feature to the prediction results and to rank the features by importance. For the three key catalytic performances, the top ten most important features were extracted, resulting in a total of 30 features.

[0073] S7: Extract descriptors based on interpretable machine learning.

[0074] The top ten most important features extracted from the three random forest models in step S6 are summarized, merged, and duplicates are removed to obtain 16 unique key features. Using these 16 features as input and their corresponding catalytic performance as output, a new dataset is constructed.

[0075] The dataset was input into the interpretable machine learning model SISSO. Symbolic regression was used to search for explicit mathematical relationships (i.e., descriptor formulas) between feature combinations and catalytic performance within a vast feature space. Based on the dataset sample size, leave-P cross-validation was employed to evaluate the accuracy of the generated candidate descriptor models. Finally, considering the model's prediction accuracy, complexity, and physical interpretability, the optimal mathematical expression for each key catalytic performance was selected as the final descriptor.

[0076] S8: Predicting the performance of new catalysts using descriptors.

[0077] Once the descriptors with clear mathematical forms and physical meanings are obtained, they can be used for rapid screening of novel catalysts. For transition metal single-atom catalysts modified with arbitrary axial ligands, it is only necessary to obtain a few key characteristic parameters corresponding to their active centers and substitute them into the descriptor formula obtained in step S7 to directly predict the three key indicators affecting the performance of sodium-sulfur batteries: the discharge reaction barrier, the adsorption energy for Na2S6, and the decomposition barrier of Na2S. Furthermore, by comprehensively examining the trends of these three properties with respect to the characteristics in the descriptors, it is possible to selectively screen for potentially high-performance multifunctional catalysts with specific combinations of characteristics.

[0078] The above description is only for the purpose of helping to understand the method and core essence of the present invention, but the scope of protection of the present invention is not limited thereto. For those skilled in the art, any equivalent substitutions or modifications made to the technical solution and inventive concept disclosed in the present invention within the scope of the technology disclosed in the present invention should be covered within the scope of protection of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A method for predicting the performance of multifunctional catalysts based on high-throughput computing and machine learning, characterized in that, Includes the following steps: S1. Dataset Construction: Obtain the basic structural information of the target catalyst from public databases, and calculate the catalytic performance data and related characteristic data of the target catalyst accordingly. Pair the catalytic performance data and related characteristic data to establish a dataset. S2. Machine Learning Modeling and Evaluation: Divide the dataset into training and testing sets, and import them into multiple machine learning models for training. Use multi-objective Bayesian optimization to tune the hyperparameters of multiple machine learning models, search for the best parameter combination, and evaluate the model fitting effect. S3. Interpretable Machine Learning: Based on the selected best model, the key features affecting catalyst performance are identified by sorting the SHAP values. After deduplication, the features are imported into the interpretable machine learning model SISSO for multi-task training. S4. Descriptor Construction and Application: Based on the interpretable machine learning model SISSO, the explicit mathematical relationship between feature combinations and catalytic performance is searched in the feature space as a descriptor formula. A single descriptor formula is used to simultaneously predict multiple key properties of a target catalyst. Based on the descriptor formula, a curve showing the performance changes with the characteristics is established to find the characteristics corresponding to the optimal solution, and based on this, potential high-performance multifunctional catalysts are sought.

2. The method for predicting the performance of multifunctional catalysts based on high-throughput computing and machine learning according to claim 1, characterized in that, The catalytic performance data and related characteristic data of the target catalyst are calculated in S1 as follows: Calculate the catalytic performance data of the target catalyst: obtain the basic structural information of the target catalyst from public databases and construct a catalyst surface model; use first principles to obtain the adsorption configuration, find the optimal adsorption configuration and obtain the adsorption energy, and search for reaction pathways to obtain the reaction energy barrier; Calculate relevant characteristic data of the target catalyst: Extract the catalytic active center characteristics of the target catalyst from public databases. The catalytic active center characteristics include physical characteristics, and the electronic structure characteristics are obtained through calculation.

3. The method for predicting the performance of multifunctional catalysts based on high-throughput computing and machine learning according to claim 2, characterized in that, The physical characteristics include electronegativity, atomic mass, atomic radius, Fermi level, first ionization energy, electron affinity, and number of valence electrons; The electronic structure features include the number of electrons filling the outermost spin-up and spin-down orbitals of the atom, the band center of the outermost orbitals, and the bond length.

4. The method for predicting the performance of multifunctional catalysts based on high-throughput computing and machine learning according to claim 2, characterized in that, The dataset is established in S1 as follows: The relevant feature data is preprocessed, and the correlation between features is analyzed based on the relevant feature data and the Pearson correlation coefficient is introduced to filter out features with highly redundant information. Then, the original data is transformed into dimensionless and order-of-magnitude standardized values ​​using standardization processing. A dataset was created by pairing catalytic performance data with processed relevant feature data.

5. The method for predicting the performance of multifunctional catalysts based on high-throughput computing and machine learning according to claim 1, characterized in that, The machine learning modeling in S2 is as follows: The machine learning model employs at least one of random forest, gradient boosting, adaptive boosting, autocorrelation determination regression, and support vector regression. When using multi-objective Bayesian optimization to tune the hyperparameters of a model, the optimization metric should be at least two of the following: maximizing prediction accuracy, maximizing robustness, maximizing sparsity, minimizing prediction uncertainty, and minimizing model complexity.

6. The method for predicting the performance of multifunctional catalysts based on high-throughput computing and machine learning according to claim 5, characterized in that, When using multi-objective Bayesian optimization to tune the hyperparameters of the model, physicochemical principles are incorporated as constraints into the optimization process to ensure that the model predictions conform to the basic laws of electrochemical reactions. The physicochemical principle constraints introduced are either the Sabatier principle or the Brønsted-Evans-Polanyi relation.

7. The method for predicting the performance of multifunctional catalysts based on high-throughput computing and machine learning according to claim 1, characterized in that, The key features affecting catalyst performance in S3 are as follows: Based on the SHAP value, the top ten features with the greatest impact on each catalytic performance were extracted from the optimal machine learning model corresponding to that catalytic performance.

8. The method for predicting the performance of multifunctional catalysts based on high-throughput computing and machine learning according to claim 7, characterized in that, The interpretable machine learning model SISSO in S3 is trained on multiple tasks, as follows: The top ten features that have the greatest impact on each catalytic performance are summarized. After removing duplicates, the remaining features are used as input data and the corresponding catalytic performance is used as output data to construct a new dataset. The dataset is then fed into the interpretable machine learning model SISSO for multi-task training.

9. The method for predicting the performance of multifunctional catalysts based on high-throughput computing and machine learning according to claim 1, characterized in that, The construction and application of descriptors in S4 are as follows: When training with the interpretable machine learning model SISSO, the parameters of the interpretable machine learning model SISSO include the number of multi-task learning elements, the number of features, the set of mathematical operators, the dimension, the feature complexity, and the feature dimension set, which are used to determine whether the selected descriptor formula has an explicit mathematical expression and conforms to the structure-property relationship of physical and chemical intuition. The descriptor formula is evaluated, and the optimal mathematical expression is selected as the final descriptor. An activity trend graph is constructed based on the descriptor formula to find the physical characteristics corresponding to the local or global highest points of the curve, and then to find the corresponding materials in reality.

10. A multifunctional catalyst performance prediction system based on high-throughput computing and machine learning applied to the method of any one of claims 1-9, characterized in that, include: The model building unit is used to build catalyst models and acquire feature data. It uses the catalytic performance data obtained from computational simulation as output data and combines the selected feature data as input data to train machine learning association models for different catalytic needs. The validation and evaluation unit uses K-fold cross-validation to optimize and evaluate the model, and selects the machine learning model with the best prediction performance and the feature importance ranking. The descriptor generation unit is used to synthesize the features with the highest importance in the optimal model, remove duplicates, and input them into the interpretable machine learning model SISSO to obtain mathematical expression models between features and catalytic performance under different dimensions and feature complexities. The accuracy of the mathematical expression models is evaluated, and the model with the most accurate prediction is selected as the final descriptor. This descriptor is then used to predict the catalytic performance of other catalysts.