Multi-algorithm fused toluene catalytic combustion catalyst rapid screening method
By constructing a multi-dimensional feature system and using multiple algorithm optimization methods, high-performance toluene catalytic combustion catalysts were screened, solving the problems of long time consumption and high resource consumption in existing technologies, and achieving rapid and accurate catalyst screening.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- XINJIANG UNIVERSITY
- Filing Date
- 2026-02-02
- Publication Date
- 2026-05-12
AI Technical Summary
Existing methods for screening toluene catalytic combustion catalysts are time-consuming and resource-intensive, and existing models lack accuracy and stability in prediction, making it difficult to quickly respond to the demand for high-performance catalysts.
By constructing a multi-dimensional feature system, selecting 11 supervised regression models, combining multiple algorithms for training and optimization, screening the optimal model, and using key feature thresholds and physical constraints to quickly screen high-performance catalysts.
It significantly improves the accuracy and efficiency of catalyst screening, reduces resource waste, shortens the R&D cycle, and provides an efficient catalyst screening pathway.
Smart Images

Figure CN122024930A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of catalytic materials, specifically to a rapid screening method for toluene catalytic combustion catalysts using a multi-algorithm fusion approach. Background Technology
[0002] Toluene is a common volatile organic pollutant in industrial production and daily life, widely originating from emissions from industries such as chemical processing, coating, and printing. Direct emission without effective treatment not only pollutes the atmosphere but may also harm human health. Therefore, achieving efficient toluene purification has become an important research direction in the environmental protection field. Catalytic combustion, with its advantages of high purification efficiency, low energy consumption, and no secondary pollution, has become a key technology for treating toluene waste gas. The catalyst, as the core component of this technology, directly determines the efficiency and stability of toluene catalytic combustion. Among these, single-metal-metal oxide catalysts have received widespread attention in toluene catalytic combustion research due to their relatively low preparation cost and highly tunable catalytic activity. With the continuous deepening of related research, a large amount of literature on single-metal-metal oxide catalysts has been accumulated, covering various aspects such as catalyst composition, preparation parameters, testing conditions, and toluene conversion rate. This provides a foundation for data-driven catalyst screening. How to establish efficient and accurate screening methods based on this massive amount of data has become an important requirement for further promoting the development of toluene catalytic combustion technology.
[0003] Traditional methods for screening toluene catalytic combustion catalysts are largely based on trial and error. Researchers need to repeatedly prepare catalyst samples with different compositions and parameters based on experience, and test their toluene conversion performance through numerous experiments. This process not only consumes a large amount of reagent, equipment, and manpower resources, but also has a long development cycle, making it difficult to quickly respond to the demand for high-performance catalysts in practical applications. At the same time, existing screening methods often do not comprehensively consider the factors affecting catalyst performance, usually focusing only on the composition of the catalyst or a single preparation parameter, ignoring the comprehensive effect of multi-dimensional characteristics such as elemental physicochemical properties, valence electron structure, thermodynamic stability, and catalyst microstructure on catalytic performance. This results in an inability to systematically capture the intrinsic relationship between various factors and toluene conversion rate. In addition, some model-based screening methods often use single regression algorithms, lacking systematic optimization of model parameters and comparative verification of the performance of multiple models. This makes the accuracy and stability of model predictions insufficient, making it difficult to reliably guide the screening of high-performance catalysts, further restricting the improvement of catalyst development efficiency. Summary of the Invention
[0004] The purpose of this invention is to overcome the shortcomings of existing technologies and provide a rapid screening method for toluene catalytic combustion catalysts using a multi-algorithm fusion approach. This method involves collecting literature to construct a standardized dataset, then building a multi-dimensional feature system. Eleven supervised regression models are selected for training and optimization. The optimal model is selected based on indicators such as mean absolute error. Key features and correlation patterns are then explored. Finally, the feature parameters of the catalyst to be screened are input into the optimal model to predict the conversion rate. Combined with physical constraints and key feature thresholds, the method determines whether the catalyst is a high-performance candidate catalyst, thus improving screening efficiency and accuracy.
[0005] To solve the above-mentioned technical problems, this invention provides the following technical solution: a rapid screening method for toluene catalytic combustion catalysts using multi-algorithm fusion, the specific steps of which are as follows: S100, Dataset Construction: Collect literature containing single-metal-metal oxide catalysts, extract the composition, preparation parameters, test conditions and toluene conversion data of single-metal-metal oxide catalysts, and form a standardized dataset through structuring, standardization and missing value imputation. S200, Construction of Multidimensional Feature System: Based on a standardized dataset, a multidimensional input feature set is constructed, the target variable is determined to be toluene conversion rate, and a feature-target variable dataset is formed; S300, multi-algorithm model training and optimization: taking the feature-target variable dataset as input, 11 supervised regression models are selected, the training set and test set are divided in an 8:2 ratio, the numerical features are standardized, the hyperparameter search range is designed for each model, and the parameters of all models are tuned by grid search combined with 10-fold cross-validation. S400, Optimal Model Selection and Feature Mining: The average absolute error and root mean square error of cross-validation are used as indicators to compare the training set, test set and cross-validation performance of 11 models to select the optimal model; based on the optimal model, the key features affecting the catalytic combustion performance of toluene are mined and the correlation between each key feature and the toluene conversion rate is clarified. S500, rapid catalyst screening: After organizing and standardizing the characteristic parameters of the catalyst to be screened according to the characteristic system, the optimal model is input to predict its toluene catalytic combustion conversion rate. Combined with the physical constraints of the catalytic reaction and the key feature thresholds, the performance level of the model prediction results is judged. When the predicted value of the toluene conversion rate of the catalyst to be screened is ≥75% and the key features meet the preset threshold range, it is judged as a high-performance candidate catalyst.
[0006] Furthermore, in the construction of the S100 dataset, literature on single-metal-metal oxide catalysts was collected through the Web of Science Core Collection database and review literature in the field of toluene catalytic oxidation. The search keywords were toluene, metal oxide, and catalytic oxidation. The collected data included composition, preparation parameters, test conditions, and toluene conversion rate. The composition covered the type of metal oxide support, the type of loaded metal, and the molar fraction of the loaded metal. The preparation parameters covered the catalyst synthesis method, calcination temperature, calcination time, and calcination atmosphere. The test conditions covered the reaction temperature, catalyst mass, reaction space velocity, and toluene content. The toluene conversion rate was the specific percentage conversion value under the corresponding test conditions.
[0007] Furthermore, in the construction of the S200 multidimensional feature system, based on a standardized dataset and following the principles of accessibility, uniqueness of characterization, and physical completeness, a multidimensional input feature set is constructed through the integration of multi-source tools and databases. Elemental physicochemical features are extracted using the Matminer and Magpie toolkits, covering size-related physical properties and electronic attributes; multi-component systems are processed using mass-weighted averaging. Valence electron structure features are constructed using Pymatgen's ElectronicStructure module, including orbital filling features, electronic energy level statistical features, and encoding features. Thermodynamic stability features are obtained from the FToxid database, including standard enthalpy of formation, standard Gibbs free energy of formation, decomposition temperature, lattice energy, phase transition temperature, and oxygen vacancy formation energy for metals and metal oxides. Experimental preparation and testing parameter features are manually extracted from the standardized dataset; categorical variables are encoded using one-hot encoding, covering material proportioning, post-processing parameters, synthesis methods, metal oxide preparation methods, and related parameters for testing conditions. Catalyst structural features are extracted from literature characterization data, including metal oxide space group symbols, metal existence forms, specific surface area, average pore size, and total pore volume.
[0008] Furthermore, in the construction of the S200 multidimensional feature system, the basis for determining the toluene conversion rate as the target variable is that this parameter is the core indicator for evaluating the catalytic combustion performance of toluene, and the standardized dataset already includes the specific conversion percentage values under the corresponding test conditions, possessing data integrity and usability. The formation process of the feature-target variable dataset is as follows: multidimensional input features are associated with the target variable according to the sample dimensions. The elemental physicochemical features, valence electron structure features, thermodynamic stability features, experimental preparation and test parameter features, and catalyst structure features of each sample strictly correspond to the toluene conversion rate under the same test conditions. The unified data format is numerical. The integrity of the associated dataset is verified, and abnormal samples that cannot be matched with the target variable are removed, finally forming a structured feature-target variable dataset. The multidimensional input feature set includes multiple dimensions from elemental physicochemical features, valence electron structure features, thermodynamic stability features, experimental preparation and test parameter features, and catalyst structure features.
[0009] Furthermore, in the training and optimization of the S300 multi-algorithm model, the 11 supervised regression models selected are specifically random forest regression, decision tree regression, extreme random tree regression, gradient boosting regression, minimum absolute shrinkage and selection operator regression, partial least squares regression, support vector regression, ridge regression, elastic net regression, extreme gradient boosting regression, and Bayesian ridge regression.
[0010] Furthermore, in the S300 multi-algorithm model training and optimization, the specific steps for parameter tuning using grid search combined with 10-fold cross-validation are as follows: First, the divided training set is evenly divided into 10 mutually exclusive subsets according to the 10-fold cross-validation rule, with 9 subsets serving as training subsets and 1 subset serving as validation subset. Ten rounds of training and validation are performed alternately to cover all subsets. Then, for each of the 11 supervised regression models, based on their respective algorithmic characteristics and catalytic performance prediction scenario requirements, reasonable hyperparameter search ranges are designed, covering key dimensions such as model complexity, regularization strength, and kernel function parameters. Subsequently, grid search is initiated, and the model is trained one by one in the 10-fold cross-validation process according to preset hyperparameter combinations, simultaneously recording the cross-validation mean absolute error and root mean square error corresponding to each set of hyperparameters. After all hyperparameter combinations have been traversed, the hyperparameter combination with the smallest cross-validation mean absolute error and root mean square error is selected as the optimal parameters. Finally, the model is retrained on the complete training set using these optimal parameters.
[0011] Furthermore, in the S400 optimal model selection, the mathematical expression for the mean absolute error (MAE) of cross-validation is: The root mean square error (RMSE) is expressed mathematically as follows: ,in, These represent the number of samples in the training set, test set, or cross-validation subset, respectively. This represents the true value of toluene conversion rate. The values are model predictions. When comparing, the MAE and RMSE of the 11 supervised regression models are calculated in the training set, test set, and cross-validation respectively. The model with the smallest cross-validation MAE and the smallest RMSE is selected as the optimal model.
[0012] Furthermore, in the S400 optimal model selection, feature mining specifically adopts a combination of feature importance ranking, partial dependency graphs, and the SHapley Additive exPlanations interpretation model. The selection criteria for key features are those that rank in the top 20% of feature importance scores and have a significant monotonic or nonlinear correlation with toluene conversion rate as verified by partial dependency graphs. The correlation between each key feature and toluene conversion rate is quantified and characterized by SHAP values, clarifying the contribution direction and intensity of a single feature to toluene conversion rate and the dynamic trend of the contribution as the feature value changes.
[0013] Furthermore, in the rapid screening of the S500 catalyst, the optimal model is a fusion model of the screened optimal supervised regression model and the physical constraints of the catalytic reaction; the physical constraints include the catalyst thermodynamic stability boundary, the reasonable range of oxygen vacancy formation energy, and the reaction kinetic rate limit.
[0014] Furthermore, in the rapid screening of the S500 catalyst, the key characteristic thresholds are determined based on the SHAP value analysis results and the catalytic reaction mechanism, specifically including: oxygen vacancy formation energy of 2-4 eV, catalyst specific surface area ≥50 m² / g, and reaction space velocity ≤20000. The calcination temperature is 300-600℃, the molar fraction of the supported metal is 0.5%-3%, the standard Gibbs free energy of formation of the metal oxide support is ≤-200 kJ / mol, and the average pore size of the catalyst is 2-20 nm.
[0015] Beneficial effects Compared with existing technologies, this rapid screening method for toluene catalytic combustion catalysts based on multi-algorithm fusion has the following advantages: I. This invention integrates multiple tools and databases to construct an input feature set covering key information across multiple dimensions. The system incorporates core elements such as material properties, preparation and testing parameters, and structural features to ensure the comprehensiveness and completeness of the feature system. Simultaneously, multiple supervised regression models are selected for training and optimization. Scientific dataset partitioning and parameter tuning strategies are employed, and the optimal model is selected through comprehensive comparison of multiple indicators. This significantly improves the accuracy and stability of toluene conversion rate prediction. This multi-dimensional feature fusion and multi-algorithm optimization and verification mode can deeply capture the intrinsic relationship between various key factors and catalytic performance, avoiding the limitations of single features or models. It provides solid technical support for subsequent catalyst performance evaluation and significantly improves the scientific rigor and reliability of the screening process.
[0016] Second, this invention establishes an efficient catalyst screening mechanism by combining the optimal model with the physical constraints of the catalytic reaction and integrating the correlation patterns and threshold standards derived from feature mining. The screening process does not require complex experimental verification; performance level can be quickly determined simply by inputting standardized feature parameters, effectively shortening the catalyst development cycle. At the same time, based on scientific feature analysis, the influence of key factors is clarified, enabling precise identification of high-performance candidate catalysts and reducing the waste of resources caused by blind experiments. This screening method, which balances efficiency and accuracy, provides a new path for the development of toluene catalytic combustion catalysts, helping related fields to rapidly advance the screening and application of high-performance catalysts, and has significant practical value and application prospects.
[0017] Other advantages, objectives and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination or study, or may be learned from the practice of the invention. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.
[0019] Figure 1 A flowchart of a rapid screening method for toluene catalytic combustion catalysts using a multi-algorithm fusion approach; Figure 2 This is a schematic diagram showing the connection relationship of each step in a rapid screening method for toluene catalytic combustion catalysts using a multi-algorithm fusion approach. Detailed Implementation
[0020] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.
[0021] Example 1: Research and screening of novel single-metal-metal oxide toluene catalytic combustion catalysts in university laboratories.
[0022] A university environmental catalysis laboratory plans to develop a novel, highly efficient toluene catalytic combustion catalyst. Targeting a single metal-metal oxide system, the laboratory needs to rapidly screen high-performance candidate catalysts that meet the requirements of industrial waste gas treatment, shorten the development cycle, and reduce development costs. Specific steps are as follows: Figure 1 As shown.
[0023] S100, Dataset Construction: Using the Web of Science Core Collection database and authoritative review articles in the field of toluene catalytic oxidation, and employing the keywords "toluene," "metal oxide," and "catalytic oxidation," a total of 320 articles related to single-metal-metal oxide catalysts published in the past 10 years were systematically collected. This ensured coverage of research results on different metal oxide supports, loaded metal types, and preparation processes, providing rich and comprehensive basic data for subsequent model training and avoiding insufficient model generalization ability due to limited data. Key data were extracted from the literature, including component composition, preparation parameters, test conditions, and toluene conversion rate. The extracted data were structured and formatted, and the mean interpolation method was used to handle the small number of missing loaded metal molar fractions and specific surface area data to avoid deviations in subsequent feature system construction due to missing data, ensuring the completeness and consistency of the dataset. Finally, a standardized dataset containing 286 valid samples was formed, such as... Figure 2 As shown.
[0024] S200, Construction of a Multidimensional Feature System: Based on standardized datasets and following the principles of accessibility, unique representation, and physical completeness, a multi-dimensional input feature set is constructed by integrating multiple source tools and databases. This ensures that the features can comprehensively cover the composition, structure, preparation, and operation-related attributes of the catalyst, providing sufficient information for the model to learn accurately. The physicochemical characteristics of elements were extracted using the Matminer and Magpie toolkits, including atomic radius, ionic radius, electronegativity, melting point, boiling point, and other size-related, physical property-related, and electronic property-related characteristics of metals and metal oxides. For multi-component systems, a mass-weighted average method was used to ensure that the overall characteristics of multi-component catalysts could be accurately quantified, avoiding the inability of single-component characteristics to represent the overall performance of the catalyst. Valence electron structure characteristics were constructed using Pymatgen's Electronic Structure module, including orbital filling states, electronic energy level distribution statistics, and corresponding coding features. These features reveal the electronic interactions between the metal and the support, which are key factors affecting the activity of catalyst active sites. Thermodynamic stability characteristics were obtained based on the FToxid database, including the standard enthalpy of formation, standard Gibbs free energy of formation, decomposition temperature, lattice energy, phase transition temperature, and oxygen vacancy formation energy of metals and metal oxides. These parameters are directly related to the stability and anti-attenuation ability of the catalyst during use. Experimental preparation and testing parameters were manually extracted, and one-hot encoding was used for categorical variables such as synthesis methods and metal oxide preparation methods. The coding encompasses parameters related to material ratios, post-processing parameters, various synthesis methods, reaction temperatures, catalyst mass, and other test conditions, ensuring the model can learn the influence of preparation and operating conditions on catalytic performance. Catalyst structural features are extracted from literature characterization data, including metal oxide space group symbols, metal presence forms, specific surface area, average pore size, and total pore volume. These structural parameters directly affect the mass transfer efficiency between reactants and products within the catalyst and the accessibility of active sites. Toluene conversion is determined as the target variable. This parameter is a core indicator for evaluating toluene catalytic combustion performance, and the standardized dataset includes complete conversion percentage values under the corresponding test conditions, directly reflecting the actual catalytic effect of the catalyst and providing a clear optimization direction for model training. Multidimensional input features are associated with target variables according to sample dimensions to ensure that the elemental physicochemical characteristics, valence electron structure characteristics, thermodynamic stability characteristics, experimental preparation and testing parameter characteristics, and catalyst structure characteristics of each sample strictly correspond to the toluene conversion rate under the same test conditions. This eliminates the problem of feature-performance mismatch caused by differences in test conditions. All data formats are unified to numerical type. The integrity of the associated dataset is checked, and 34 abnormal samples whose features do not match the target variable are removed. Finally, a structured feature-target variable dataset containing 252 samples is formed.
[0025] S300, multi-algorithm model training optimization: Using the constructed feature-target variable dataset as input, 11 supervised regression models were selected: Random Forest Regression, Decision Tree Regression, Extreme Random Tree Regression, Gradient Boosting Regression, Minimum Absolute Shrinkage and Selection Operator Regression, Partial Least Squares Regression, Support Vector Regression, Ridge Regression, Elastic Net Regression, Extreme Gradient Boosting Regression, and Bayesian Ridge Regression. These models, covering different algorithmic principles, can explore the correlation between features and toluene conversion rate from multiple perspectives, improving the reliability of subsequent optimal model selection and avoiding prediction bias caused by the limitations of a single model. The dataset was divided into training and test sets in an 8:2 ratio. This ensured that the training set had sufficient sample size to support the model in fully learning the correlation between features and performance, while the test set effectively verified the model's generalization ability on unseen samples, preventing overfitting. Numerical features in the training set were standardized to eliminate the influence of dimensional differences and prevent the model from overemphasizing certain features due to different feature numerical ranges, ensuring that each feature plays a balanced role in model training. For each model, considering its algorithmic characteristics and the specific requirements of its catalytic performance prediction scenarios, a reasonable hyperparameter search range is designed. Hyperparameters cover key dimensions such as model complexity, regularization strength, and kernel function parameters. For example, the number of decision trees in random forest regression ranges from 100 to 500, and the maximum depth ranges from 5 to 20. Support vector regression kernel types include linear kernels, multinomial kernels, and radial basis kernels, along with their corresponding parameter ranges. This ensures that each model can find its optimal operating state within a suitable parameter space. Parameter tuning is performed using grid search combined with 10-fold cross-validation. The training set is first divided into 10 mutually exclusive subsets according to the 10-fold cross-validation rule. Nine subsets are selected as the training subset and one subset as the validation subset in each round of training and validation to cover all subsets, reducing parameter tuning bias caused by random sample partitioning and improving the reliability of the optimal parameters. Simultaneously, the mean absolute error (MAE) and root mean square error (RMSE) of cross-validation for each set of hyperparameters are recorded. These two error metrics accurately measure the deviation between the model's predicted values and the actual values, providing a quantitative basis for selecting the optimal parameters. After all hyperparameter combinations have been iterated, the hyperparameter combination with the smallest cross-validation MAE and RMSE is selected as the optimal parameters for each model. Finally, each model is retrained on the complete training set using the optimal parameters to obtain 11 optimized regression models, ensuring that each model is in the best predictive state.
[0026] S400, Optimal Model Selection and Feature Mining: Calculate the MAE and RMSE of the 11 optimized models on the training set, test set, and cross-validation respectively. The mathematical expression for the mean absolute error (MAE) of cross-validation is: The root mean square error (RMSE) is expressed mathematically as follows: ,in, These represent the number of samples in the training set, test set, or cross-validation subset, respectively. This represents the true value of toluene conversion rate. The predicted values are used for model comparison. The MAE and RMSE of 11 supervised regression models were calculated in the training, test, and cross-validation sets. The model with the smallest cross-validation MAE and RMSE was selected as the optimal model. Comparison revealed that the extreme gradient boosting regression model had the smallest MAE and RMSE in cross-validation, exhibiting the best overall performance. Therefore, it was selected as the final prediction model. This model can more accurately establish the mapping relationship between features and toluene conversion rate, providing a reliable tool for subsequent catalyst performance prediction and reducing the interference of prediction errors on the selection results. Based on this optimal model, feature mining was performed using a combination of feature importance ranking, partial dependency graphs, and the SHapley Additive exPlanations interpretation model. This combination of methods comprehensively and accurately identifies features that play a key role in catalytic performance, avoiding misjudgment or omission of features due to a single method, and ensuring that the mined key features have practical guiding significance. The key feature screening criteria were set at the top 20% of feature importance scores, with a significant monotonic or nonlinear correlation to toluene conversion verified by partial dependency plots. Seven key features were ultimately selected: oxygen vacancy formation energy, catalyst specific surface area, space velocity, calcination temperature, molar fraction of loaded metal, standard Gibbs free energy of formation of the metal oxide support, and average pore size of the catalyst. These features directly guide the design and optimization of subsequent catalysts, helping researchers focus on adjusting core parameters. The correlation between each key feature and toluene conversion was quantified using SHAP values. It was found that the oxygen vacancy formation energy has the strongest promoting effect on conversion in the 2-4 eV range; the larger the specific surface area and the lower the space velocity, the higher the overall conversion rate; and the conversion rate remains at a relatively high level within the calcination temperature range of 300-600℃. These trends provide a clear basis for judging the performance of the catalysts to be screened, avoiding blind testing.
[0027] S500, rapid catalyst screening: The catalyst to be screened is a TiO2-supported Pd single-metal-metal oxide catalyst prepared in the laboratory. Its characteristic parameters are summarized as follows: the metal oxide support is TiO2, the standard Gibbs free energy of formation is -245 kJ / mol; the supported metal is Pd, with a molar fraction of 1.8%; the preparation method is sol-gel method, calcination temperature is 480℃, calcination time is 3h, and the calcination atmosphere is air; the test conditions are reaction temperature 280℃, catalyst mass 0.15g, reaction space velocity 16000h-1, and toluene content 1200ppm; the catalyst structure is space group symbol P42 / mnm, the metal exists in the elemental form, specific surface area is 68m² / g, average pore size is 12nm, total pore volume is 0.22cm³ / g; oxygen vacancy formation energy is 3.1eV. Completely summarizing these parameters ensures comprehensive characteristic information for input into the model, avoids prediction bias due to missing parameters, and ensures that the prediction results accurately reflect the catalyst performance. After organizing and standardizing the above feature parameters according to the constructed multidimensional feature system, they are input into the saved optimal extreme gradient boosting regression model to predict that the toluene catalytic combustion conversion rate is 86%. This prediction value can quickly and preliminarily determine the performance level of the catalyst without waiting for time-consuming experimental tests, thus shortening the screening cycle. Based on the physical constraints of the catalytic reaction and the key characteristic thresholds identified, the performance level of this catalyst was determined as follows: the predicted conversion rate is 86% ≥ 75%, and the oxygen vacancy formation energy is 3.1 eV (within the range of 2-4 eV), the specific surface area is 68 m² / g (≥ 50 m² / g), the reaction space velocity is 16000 h⁻¹ (≤ 20000 h⁻¹), the calcination temperature is 480℃ (within the range of 300-600℃), the molar fraction of loaded metal is 1.8% (within the range of 0.5%-3%), the standard Gibbs free energy of formation on the support is -245 kJ / mol (≤ -200 kJ / mol), and the average pore size is 12 nm. All key characteristics meet the preset threshold ranges. Therefore, this catalyst is determined to be a high-performance candidate catalyst, providing a clear conclusion for further in-depth research and application in the laboratory and reducing ineffective R&D investment.
[0028] In summary, in the screening and development of novel single-metal-metal oxide toluene catalytic combustion catalysts in university laboratories, a standardized dataset of 286 samples was first constructed using Web of Science and relevant review literature. Then, following established principles, a multi-dimensional feature system was built to form a feature-target variable dataset of 252 samples. Eleven supervised regression models were trained and optimized using grid search combined with 10-fold cross-validation. Extreme gradient boosting regression was selected as the optimal model, and key features were extracted. The predicted conversion rate of the TiO2-supported Pd catalyst was 86%, and based on constraints and thresholds, it was determined to be a high-performance candidate, effectively shortening the development cycle and providing reliable support for further in-depth research.
[0029] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A rapid screening method for toluene catalytic combustion catalysts using a multi-algorithm fusion approach, characterized in that, The specific steps of this method are as follows: S100, Dataset Construction: Collect literature containing single-metal-metal oxide catalysts, extract the composition, preparation parameters, test conditions and toluene conversion data of single-metal-metal oxide catalysts, and form a standardized dataset through structuring, standardization and missing value imputation. S200, Construction of Multidimensional Feature System: Based on a standardized dataset, a multidimensional input feature set is constructed, the target variable is determined to be toluene conversion rate, and a feature-target variable dataset is formed; S300, multi-algorithm model training and optimization: taking the feature-target variable dataset as input, 11 supervised regression models are selected, the training set and test set are divided in an 8:2 ratio, the numerical features are standardized, the hyperparameter search range is designed for each model, and the parameters of all models are tuned by grid search combined with 10-fold cross-validation. S400, Optimal Model Selection and Feature Mining: The average absolute error and root mean square error of cross-validation are used as indicators to compare the training set, test set and cross-validation performance of 11 models to select the optimal model; based on the optimal model, the key features affecting the catalytic combustion performance of toluene are mined and the correlation between each key feature and the toluene conversion rate is clarified. S500, rapid catalyst screening: After organizing and standardizing the characteristic parameters of the catalyst to be screened according to the characteristic system, the optimal model is input to predict its toluene catalytic combustion conversion rate. Combined with the physical constraints of the catalytic reaction and the key feature thresholds, the performance level of the model prediction results is judged. When the predicted value of the toluene conversion rate of the catalyst to be screened is ≥75% and the key features meet the preset threshold range, it is judged as a high-performance candidate catalyst.
2. The rapid screening method for toluene catalytic combustion catalysts based on multi-algorithm fusion according to claim 1, characterized in that, In constructing the S100 dataset, literature on single-metal-metal oxide catalysts was collected using the Web of Science Core Collection database and review literature in the field of toluene catalytic oxidation. The search keywords were toluene, metal oxide, and catalytic oxidation. The collected data included composition, preparation parameters, test conditions, and toluene conversion rate. The composition covered the type of metal oxide support, the type of loaded metal, and the molar fraction of the loaded metal. The preparation parameters covered the catalyst synthesis method, calcination temperature, calcination time, and calcination atmosphere. The test conditions covered the reaction temperature, catalyst mass, reaction space velocity, and toluene content. The toluene conversion rate was the specific percentage conversion value under the corresponding test conditions.
3. The rapid screening method for toluene catalytic combustion catalysts based on multi-algorithm fusion according to claim 1, characterized in that, In the construction of the S200 multidimensional feature system, based on a standardized dataset and following the principles of accessibility, uniqueness of representation, and physical completeness, a multidimensional input feature set is constructed through the integration of multi-source tools and databases. Elemental physicochemical features are extracted using the Matminer and Magpie toolkits, covering size-related physical properties and electronic attributes; multi-component systems are processed using mass-weighted averaging. Valence electronic structure features are constructed using Pymatgen's ElectronicStructure module, including orbital filling features, electronic energy level statistical features, and encoding features. Thermodynamic stability features are obtained from the FToxid database, including standard enthalpy of formation, standard Gibbs free energy of formation, decomposition temperature, lattice energy, phase transition temperature, and oxygen vacancy formation energy for metals and metal oxides. Experimental preparation and testing parameter features are manually extracted from the standardized dataset; categorical variables are encoded using one-hot encoding, covering parameters related to material proportioning, post-processing, synthesis methods, metal oxide preparation methods, and testing conditions. The structural characteristics of the catalyst were extracted from literature characterization data, including the metal oxide space group symbol, the metal presence form, the specific surface area, the average pore size, and the total pore volume.
4. The rapid screening method for toluene catalytic combustion catalysts using multi-algorithm fusion according to claim 1, characterized in that, In the construction of the S200 multidimensional feature system, the target variable is determined to be the toluene conversion rate because this parameter is the core indicator for evaluating the catalytic combustion performance of toluene, and the standardized dataset already includes the specific conversion percentage values under the corresponding test conditions, possessing data integrity and usability. The formation process of the feature-target variable dataset is as follows: multidimensional input features are associated with the target variable according to the sample dimensions. The elemental physicochemical features, valence electron structure features, thermodynamic stability features, experimental preparation and test parameter features, and catalyst structure features of each sample strictly correspond to the toluene conversion rate under the same test conditions. The unified data format is numerical. The integrity of the associated dataset is verified, and abnormal samples that cannot be matched with the target variable are removed, finally forming a structured feature-target variable dataset. The multidimensional input feature set includes multiple dimensions from elemental physicochemical features, valence electron structure features, thermodynamic stability features, experimental preparation and test parameter features, and catalyst structure features.
5. The rapid screening method for toluene catalytic combustion catalysts using multi-algorithm fusion according to claim 1, characterized in that, In the training and optimization of the S300 multi-algorithm model, the 11 supervised regression models selected are: random forest regression, decision tree regression, extreme random tree regression, gradient boosting regression, minimum absolute shrinkage and selection operator regression, partial least squares regression, support vector regression, ridge regression, elastic net regression, extreme gradient boosting regression, and Bayesian ridge regression.
6. The rapid screening method for toluene catalytic combustion catalysts using multi-algorithm fusion according to claim 1, characterized in that, In the S300 multi-algorithm model training and optimization, the specific steps for parameter tuning using grid search combined with 10-fold cross-validation are as follows: First, the divided training set is evenly divided into 10 mutually exclusive subsets according to the 10-fold cross-validation rule, with 9 subsets serving as training subsets and 1 subset serving as validation subset. Ten rounds of training and validation are performed alternately to cover all subsets. Next, reasonable hyperparameter search ranges are designed for the algorithmic characteristics and catalytic performance prediction scenario requirements of the 11 supervised regression models. The hyperparameters cover key dimensions such as model complexity, regularization strength, and kernel function parameters. Then, grid search is initiated, and the model is trained one by one in the 10-fold cross-validation process according to the preset hyperparameter combinations, while simultaneously recording the cross-validation mean absolute error and root mean square error for each set of hyperparameters. After all hyperparameter combinations have been traversed, the hyperparameter combination with the smallest cross-validation mean absolute error and root mean square error is selected as the optimal parameters. Finally, the model is retrained on the complete training set using these optimal parameters.
7. The rapid screening method for toluene catalytic combustion catalysts based on multi-algorithm fusion according to claim 1, characterized in that, In the S400 optimal model selection, the mathematical expression for the mean absolute error (MAE) of cross-validation is: The root mean square error (RMSE) is expressed mathematically as follows: ,in, These represent the number of samples in the training set, test set, or cross-validation subset, respectively. This represents the true value of toluene conversion rate. The values are model predictions. When comparing, the MAE and RMSE of the 11 supervised regression models are calculated in the training set, test set, and cross-validation respectively. The model with the smallest cross-validation MAE and the smallest RMSE is selected as the optimal model.
8. The rapid screening method for toluene catalytic combustion catalysts based on multi-algorithm fusion according to claim 1, characterized in that, In the selection of the optimal S400 model, feature mining specifically adopts a combination of feature importance ranking, partial dependency graphs, and the SHapley Additive exPlanations interpretation model. The selection criteria for key features are those that rank in the top 20% of feature importance scores and have a significant monotonic or nonlinear correlation with toluene conversion rate as verified by partial dependency graphs. The correlation between each key feature and toluene conversion rate is quantified and characterized by SHAP values, clarifying the contribution direction and intensity of a single feature to toluene conversion rate and the dynamic trend of the contribution as the feature value changes.
9. The rapid screening method for toluene catalytic combustion catalysts based on multi-algorithm fusion according to claim 1, characterized in that, In the rapid screening of the S500 catalyst, the optimal model is the fusion model of the screened optimal supervised regression model and the physical constraints of the catalytic reaction; the physical constraints include the catalyst thermodynamic stability boundary, the reasonable range of oxygen vacancy formation energy, and the reaction kinetic rate limit.
10. The rapid screening method for toluene catalytic combustion catalysts using multi-algorithm fusion according to claim 1, characterized in that, In the rapid screening of the S500 catalyst, the key characteristic thresholds were determined based on SHAP value analysis results and catalytic reaction mechanism, specifically including: oxygen vacancy formation energy of 2-4 eV, catalyst specific surface area ≥50 m² / g, and reaction space velocity ≤20000. The calcination temperature is 300-600℃, the molar fraction of the supported metal is 0.5%-3%, the standard Gibbs free energy of formation of the metal oxide support is ≤-200 kJ / mol, and the average pore size of the catalyst is 2-20 nm.