Machine learning algorithm for screening of metal elements in diatomic catalysts throughout the whole cycle based on product yield prediction and application thereof

By constructing interpretable descriptors using the FSIFSC algorithm and machine learning models, the limitations of traditional catalyst screening methods are overcome, enabling efficient screening of metal elements throughout the entire lifecycle and direct prediction of product yields, significantly improving catalyst performance and efficiency.

CN122266546APending Publication Date: 2026-06-23DALIAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
DALIAN UNIV OF TECH
Filing Date
2026-03-23
Publication Date
2026-06-23

Smart Images

  • Figure CN122266546A_ABST
    Figure CN122266546A_ABST
Patent Text Reader

Abstract

The application discloses a machine learning algorithm for screening of metal elements in a diatomic catalyst full cycle based on product yield prediction and application thereof, and the algorithm comprises the following steps: (1) collecting attribute features of metal elements in the diatomic catalyst as a feature space 1; (2) preliminarily screening the attribute features according to comprehensive scores of feature importance to obtain a feature space 2; (3) performing feature operation processing on the feature space 2 to generate a multivariate descriptor space matrix; (4) performing dimension reduction processing on the multivariate descriptor space matrix by combining a gradient boosting regression model with a recursive feature elimination algorithm; (5) further screening to obtain an interpretable descriptor by using a least absolute shrinkage and selection operator; and (6) applying the interpretable descriptor to a random forest regression model to screen a catalyst for a propane dehydrogenation reaction to prepare propylene. The application can solve the problem that a traditional descriptor is difficult to directly predict a catalytic product yield.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of catalyst screening technology. Specifically, it relates to a machine learning algorithm for full-cycle metal element screening of diatomic catalysts based on product yield prediction and its application. Background Technology

[0002] Catalysts are core functional materials used in chemical engineering, materials science, and energy fields to regulate reaction processes and improve conversion efficiency; their performance directly determines process economics and product quality. However, traditional catalyst research and screening systems have long faced the following common industry bottlenecks:

[0003] (1) Limitations of the screening scope: Catalytic reactions involve the effects of multiple types of elements in the s, p, d, and f blocks of the periodic table. However, existing screening methods rely on single or a few features (such as atomic electronegativity and bond energy), which can only cover d-block metal elements and cannot achieve screening of metal elements in the entire periodic table, resulting in the omission of a large number of potentially high-efficiency catalytic components. (2) Insufficient accuracy of performance prediction: Traditional descriptors focus on the qualitative judgment of reaction selectivity and are difficult to directly and quantitatively predict core performance indicators (such as product yield and reaction rate). At the same time, they lack accurate characterization of the electronic structure of elements (such as orbital occupancy and electron cloud distribution), and cannot distinguish the differences in catalytic behavior of different types of metal elements (such as alkali metals AMs, transition metals TMs, alkaline earth metals AEMs, lanthanides LNMs, post-transition metals PTMs, and other metals OMs), which limits the targeted design of high-performance catalysts. (3) Lack of interpretability of multi-feature models; Although the application of machine learning in catalyst design has improved prediction efficiency, the models constructed with multiple feature parameters often exhibit "black box" characteristics - the relationship between features and catalytic performance cannot be clearly defined, making it difficult to verify the screening results through experiments and unable to provide an interpretable theoretical basis for subsequent catalyst structure optimization. (4) Contradiction between screening efficiency and practicality: Existing methods have high computational complexity under multiple feature dimensions, making it difficult to balance between "high-throughput screening" and "result verifiability".

[0004] Therefore, there is an urgent need in this field for a descriptor method that covers all metal elements in the entire lifecycle, has both quantitative performance prediction capabilities and model interpretability, in order to achieve efficient and accurate screening of catalysts, while establishing a clear logical correlation between features and performance, and providing theoretical support for the targeted design of catalytic materials. Summary of the Invention

[0005] Therefore, the technical problem to be solved by the present invention is to provide a machine learning algorithm for full-cycle metal element screening of diatomic catalysts based on product yield prediction and its application, so as to solve the problem that traditional descriptors can only predict catalyst selectivity but cannot directly predict product yield.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0007] A machine learning algorithm for full-cycle metal element screening of diatomic catalysts based on product yield prediction includes the following steps:

[0008] Step (1): Collect the intrinsic and derived properties of metal elements in the diatomic catalyst that are closely related to catalytic activity to form the feature space 1; the intrinsic properties of metal elements include atomic number, total number of frontier orbital electrons, first ionization energy, atomic radius, total number of valence electrons, period number, atomic mass, electronegativity and activation energy; the derived properties of metal elements include frontier orbital occupancy.

[0009] The formula for calculating the frontier orbital occupancy of metallic elements is:

[0010] (1);

[0011] In equation (1), Metallic elements The occupancy rate of the front-line tracks, Metallic elements The total number of front-line orbital electrons, Metallic elements The total number of valence electrons;

[0012] Step (2): Based on the FSIFSC algorithm, calculate the comprehensive importance score of each attribute feature in feature space 1 according to the feature importance score and the comprehensive SHAP value; perform preliminary screening of the attribute features in the feature dataset according to the comprehensive importance score results to obtain feature space 2; the FSIFSC algorithm (Feature Sparsification and Interpretable Feature Space Construction) is used to reduce the dimensionality of huge high-dimensional feature spaces and screen to obtain the final interpretable, low-dimensional feature space that can be further selected using methods such as GBR. The "FSIFSC algorithm" in this invention is a general term for an algorithm that integrates feature importance scoring, SHAP value calculation, and subsequent single feature operation, label operation, and combination operation.

[0013] Step (3): Based on the FSIFSC algorithm, perform single feature operations and label each initial feature in feature space 2, and then perform label operations and combination operations to construct a multivariate descriptor space matrix.

[0014] Step (4): Use the Gradient Boosting Regression (GBR) model to evaluate the importance of descriptors, and use the Recursive Feature Elimination (RFE) algorithm to reduce the dimensionality of the multivariate descriptor space matrix to obtain the dimensionality-reduced descriptor space matrix. Since the number of descriptors in the constructed multivariate descriptor space matrix is ​​huge (hundreds of thousands or more), it is impossible to directly use the Least Absolute Shrinkage and Selection Operator (LASSO) to identify the optimal descriptor. This invention uses a combination of GBR and RFE to reduce the dimensionality of the multivariate descriptor space matrix, which can significantly reduce the number of descriptors entering the "LASSO" algorithm while retaining the key variables that contribute the most to the product yield prediction, thereby improving the computational efficiency.

[0015] Step (5): Use the minimum absolute shrinkage and selection operators to further filter the optimal interpretable descriptors from the dimension-reduced descriptor space matrix;

[0016] Step (6): Apply the optimal interpretable descriptor to the random forest regression model and construct the diatomic catalyst screening model through model training; use the diatomic catalyst screening model to screen target diatomic catalysts and predict product yield.

[0017] Because different metals have varying abilities to adsorb reactants and desorb products, the catalytic performance of diatomic catalysts is highly dependent on the specific properties of the metal. Traditional methods primarily rely on the "d-band theory" to predict the catalytic performance of transition metals (TMs). However, the "d-band theory" can only describe the electronic contribution of d orbitals. The electronic contributions of alkali metals (AMs), alkaline earth metals (AEMs), lanthanides (LNMs), post-transition metals (PTMs), and other metals (OMs) are not limited to d orbitals, making the "d-band theory" unable to predict the catalytic performance of metals in other regions. This invention discovers that frontier orbital occupancy reflects the number of electrons in the metal donor or acceptor. Therefore, introducing frontier orbital occupancy allows for the prediction of s-block, p-block, d-block, and f-block metals, and reflects a characteristic of catalytic differences between metals. Therefore, choosing frontier orbital occupancy as a frontier orbital characteristic of metal elements for constructing interpretable descriptors can effectively distinguish the differences in catalytic behavior between metal elements in different regions, thereby accurately predicting the catalytic performance of metals with different electronic structures. This invention uses frontier orbital occupancy as an important feature and combines it with the FSIFSC algorithm to construct an interpretable descriptor that can predict the catalytic performance of metal elements throughout the entire cycle and directly predict the catalytic yield of diatomic catalysts.

[0018] In the above-mentioned machine learning algorithm for screening metal elements in the full cycle of diatomic catalysts based on product yield prediction, in step (1), the metal elements in the diatomic catalyst are located in the s-block, p-block, d-block, or f-block, that is: the metal elements of the diatomic catalyst in this invention cover the s-block, p-block, d-block, and f-block, and can be any element in the full cycle of metal elements.

[0019] In the above machine learning algorithm for full-cycle metal element screening of diatomic catalysts based on product yield prediction, step (2) involves the following feature importance scoring method: using catalytic activity, selectivity, and product yield as prediction targets, random forest feature importance analysis is used to score the importance of each attribute feature; the formula for calculating the importance score of each attribute feature is as follows:

[0020] (2);

[0021] In equation (2), Score the importance of each attribute feature; The feature importance score is calculated using catalytic activity as the prediction target; The feature importance score is calculated using selectivity as the prediction objective; The feature importance score is calculated using product yield as the prediction target;

[0022] Overall SHAP value The calculation formula is as follows:

[0023] (3);

[0024] (4);

[0025] In equation (3), The combined SHAP value for all attribute features; The SHAP value is calculated using catalytic activity as the prediction target. The SHAP value is calculated with selectivity as the prediction objective. The SHAP value is calculated using product yield as the prediction target.

[0026] In equation (4), It can be 1, 2, or 3; For the entire feature set; for The middle does not contain features A subset of; For subset The number of features; For only subsets The model prediction value of the feature; The total number of features;

[0027] Comprehensive score of the importance of each attribute feature The value is calculated as follows:

[0028] (5).

[0029] The machine learning algorithm for full-cycle metal element screening of diatomic catalysts based on product yield prediction, as described above, includes the following specific calculation steps in step (3):

[0030] Step (3-1), Single Feature Operation: Perform exponentiation, multiplication, and division operations on each initial feature in feature space 2, and add the results of the single feature operations to the feature space of feature space 2 to obtain feature space 3; the method of exponentiation is: perform nth power operation and nth root operation on each initial feature to obtain two exponentiation features, where n is greater than or equal to 1; both multiplication and division operations are performed using a factor of m; m is greater than or equal to 1; that is:

[0031] Exponentiation is used to perform calculations of raising to the power of n and taking the nth root, where n is a variable. For example, for a radius rM, when n=2, exponentiation yields... and Then, the results of these two exponentiation operations are added to the feature space; the new feature space contains the initial features, the square term, and the square root term; the calculation method of multiplication and division operations is similar to that of exponentiation operations, multiplying or dividing by m, where m is also a variable that can be changed;

[0032] Step (3-2): Label all physical quantity features in feature space 3 according to the original physical quantity type. That is, label the physical quantity features generated by single feature operation with the same initial feature in feature space 3 with the same label, and label the physical quantity features generated by single feature operation with different initial features with different labels to distinguish different physical quantity features.

[0033] Step (3-3), Same-label operation: Perform addition, subtraction and absolute value operation, multiplication and division operations on all physical quantity features with the same label in feature space 3 respectively, and add the results of the same-label operation to feature space 3 to obtain feature space 4; the same-label operation is to merge the same physical quantity features in feature space to enhance the local information expression ability.

[0034] Step (3-4), Different label operations: Perform multiplication and division operations on any two physical quantity features with different labels in feature space 3 respectively, and add the results of the different label operations to feature space 3 to obtain feature space 5; the different label operations are to integrate the different physical quantity features in the feature space in order to construct a complete multivariate descriptor space matrix;

[0035] Steps (3-5), Combination operation: Combine the physical quantity features in feature space 4 and feature space 5 through multiplication and division operations to generate a new feature space; integrate feature space 4, feature space 5 and the new feature space to form feature space 6, which is the multivariate descriptor space matrix, and the physical quantity features in feature space 6 are the descriptors.

[0036] In the above machine learning algorithm for full-cycle metal element screening of diatomic catalysts based on product yield prediction, in step (4), GBR is first used to evaluate the importance of each feature. The feature importance score obtained from GBR helps to identify the feature that contributes the most to the target prediction. The iterative process of recursive feature elimination gradually removes features and builds a model for the remaining features, which eliminates irrelevant or redundant features and achieves the purpose of dimensionality reduction. The specific method of dimensionality reduction using recursive feature elimination is as follows:

[0037] Step (4-1), coarse screening: Using the accuracy of product yield prediction as an indicator, a lightweight gradient boosting regression model (LightGBM) is adopted. The score value of the contribution of each descriptor in the multivariate descriptor space matrix to the target prediction is calculated sequentially through univariate evaluation. Based on the calculation results of the score value of each descriptor, 80-90% of the descriptors with lower score values ​​in the multivariate descriptor space matrix are removed, and the remaining descriptors with higher score values ​​are retained to obtain the coarse screening descriptor space matrix.

[0038] Step (4-2), intermediate screening: Based on the recursive feature elimination algorithm combined with the gradient boosting regression model, the model is constructed and trained using the descriptors in the coarse screening descriptor space matrix. The gradient boosting regression model is used to calculate the score value of the descriptor's contribution to the target prediction. 10-15% of the descriptors with lower scores are eliminated in a single batch from the coarse screening descriptor space matrix, and the remaining descriptors with higher scores are retained. The intermediate screening elimination step is repeated iteratively until 1000-1500 descriptors remain in the coarse screening descriptor space matrix, thus obtaining the intermediate screening descriptor space matrix.

[0039] Step (4-3), Fine Screening: Based on the recursive feature elimination algorithm combined with the gradient boosting regression model, the model is constructed and trained using the descriptors in the mid-screen descriptor space matrix. The gradient boosting regression model is used to calculate the score value of the descriptor's contribution to the target prediction. 3-5% of the descriptors with lower scores are eliminated in a single batch from the mid-screen descriptor space matrix, and the remaining descriptors with higher scores are retained. The fine screening elimination step is repeated iteratively until 40-80 descriptors remain in the mid-screen descriptor space matrix, resulting in a dimensionality-reduced descriptor space matrix.

[0040] In the aforementioned machine learning algorithm for full-cycle metal element screening of diatomic catalysts based on product yield prediction, in step (4-1), the number of lightweight gradient boosting model trees is 10-20, the maximum tree depth is 2-3, the feature sampling rate is 0.01-0.05, the sample sampling rate is 0.5-0.7, and the learning rate is 0.3-0.5; in step (4-2), the number of gradient boosting model trees is 20-30, the maximum tree depth is 3-4, and the feature sampling rate is 0.05-0.1. The sample sampling rate is 0.7-0.8, and the learning rate is 0.1-0.2; the number of descriptors in the mid-screen descriptor space matrix is ​​1200; in step (4-3), the number of gradient boosting model trees is 50-100, the maximum tree depth is 4-5, the feature sampling rate is 0.1-0.2, the sample sampling rate is 0.8-0.9, the learning rate is 0.05-0.1, and L1 / L2 regularization of 0.1-1.0 is added; the number of descriptors in the dimensionality reduction descriptor space matrix is ​​50.

[0041] This invention employs a hierarchical screening process of GBR-RFE (coarse screening-medium screening-fine screening) when performing dimensionality reduction screening of descriptors in a multidimensional descriptor space matrix. It also performs targeted GBR selection and refined setting of parameters at each stage. This collaborative design of "adapting the RFE algorithm in stages and dynamically adjusting the elimination step size" not only solves the problems of low screening efficiency and high resource consumption of ultra-high dimensional descriptors, but also significantly improves the screening accuracy of descriptors. This is beneficial for screening descriptors that can both predict the catalytic performance of metal elements throughout the entire cycle and directly predict the catalytic yield.

[0042] In the above machine learning algorithm for full-cycle metal element screening of diatomic catalysts based on product yield prediction, in step (5), the minimum absolute shrinkage and selection operator is used to process the descriptors in the dimension-reduced descriptor space matrix, and the final feature selection is performed by reducing the coefficients of non-critical features to zero.

[0043] In step (6), the diatomic catalyst is used to catalyze the dehydrogenation reaction. During model training: the catalytic activity, selectivity, and product yield of the diatomic catalyst are the objectives, and the descriptor value of the optimal interpretable descriptor is used as input to train the random forest regression model. Guided by the predictive performance of catalytic activity, selectivity, and product yield, the hyperparameters are searched using the grid search automatic parameter search method to determine the hyperparameter depth of the random forest regression model. The five-fold cross-validation method is used to evaluate the model performance, and the root mean square error and coefficient of determination of the product yield are used as the evaluation indicators of the model performance.

[0044] Density functional theory was used to calculate the reaction feed conversion barrier, product dehydrogenation barrier, and product desorption energy of the diatomic catalyst. The reaction feed conversion barrier was used as the catalytic activity label, and the difference between the product dehydrogenation barrier and the product desorption energy was used as the selectivity label. The product yield label was obtained by multiplying the catalytic activity label value and the selectivity label value after homogenization.

[0045] The method for screening target diatomic catalysts and predicting product yields using a model is as follows: input the descriptor values ​​of the diatomic catalysts to be screened into the trained diatomic catalyst screening model, and the diatomic catalyst screening model outputs the product yields corresponding to the diatomic catalysts to be screened; draw a volcano plot based on the output product yields and determine the final required diatomic catalysts;

[0046] Root mean square error and coefficient of determination The calculation formula is:

[0047] (7);

[0048] (8);

[0049] In equations (7) and (8), For sample size; The actual value; These are the model's predicted values; The mean of the true values; The sum of squared residuals represents the total amount of prediction error in the model. The total sum of squares reflects the overall variability of the sample data;

[0050] Model performance evaluation method: root mean square error smaller and A larger value indicates better predictive performance of the model; the root mean square error of each fold in five-fold cross-validation. and coefficient of determination The smaller the fluctuation, the more stable the model's generalization ability.

[0051] This invention constructs machine learning variables and objectives by calculating the descriptor values ​​and performance indicators of diatomic catalysts. First, it determines the best-performing model. This invention compares various machine learning models such as GBR, RF, and SVR, finding that the Random Forest Regression model achieves the highest accuracy. Furthermore, it uses a grid search method to automatically find hyperparameters and determine the optimal hyperparameter depth. The resulting model achieves a score of 0.95 under five-fold cross-validation. This confirms the high accuracy of machine learning work.

[0052] An application of a machine learning algorithm for full-cycle metal element screening of diatomic catalysts based on product yield prediction is presented. This algorithm is used to screen diatomic catalysts for the propane dehydrogenation to propylene reaction and predict the propylene yield. The calculation formula for the optimal interpretable descriptor is as follows:

[0053] (6);

[0054] In equation (6), For interpretable descriptor values; Metal atomic radius; Metal atomic radius of metals and metal These are two different metal sites in a diatomic catalyst; Metal The occupancy rate of the front-line tracks; Metal Electronegativity.

[0055] The optimal interpretable descriptor selected by this invention has the following characteristics: (1) U: Universal, meaning applicable to all periodic table metal elements; (2) S: Selective & Simultaneous, capable of simultaneously predicting catalytic activity and catalytic selectivity; (3) F: Frontier orbital-based, constructed based on frontier orbital theory. Therefore, this invention defines the constructed optimal interpretable descriptor as φ. USF .

[0056] The application of the machine learning algorithm for full-cycle metal element screening of diatomic catalysts based on product yield prediction identified IrGa@NC as the diatomic catalyst for the propane dehydrogenation to propylene reaction. During the screening of diatomic catalysts, the molar ratio of the two metal elements was predetermined to be 1:1, and the mass ratio of the metal elements to the nitrogen-doped carbon material was kept constant.

[0057] The application of the above-mentioned machine learning algorithm for full-cycle metal element screening of diatomic catalysts based on product yield prediction, and the preparation method of IrGa@NC are as follows:

[0058] Step A: Mix zinc nitrate methanol solution with a molar concentration of 0.1-0.15 mol / L and 2-methylimidazolium methanol solution with a molar concentration of 0.3-0.5 mol / L in equal volume ratio, stir at room temperature for 12-16 h; then centrifuge to collect the solid precipitate and vacuum dry.

[0059] Step B: The solid product obtained by vacuum drying is heated to 900-1050℃ for 1-3 hours under an inert atmosphere at a heating rate of 5℃ / min to obtain nitrogen-doped carbon material.

[0060] Step C: Disperse nitrogen-doped carbon material in anhydrous ethanol to obtain a black suspension; add equimolar amounts of Ga(NO3)3·9H2O and H2IrCl6·xH2O to anhydrous ethanol to prepare a mixed salt solution; the mass concentration of nitrogen-doped carbon material in the black suspension is 5-10 g / L; the molar concentration of Ga(NO3)3·9H2O in the mixed salt solution is 5-15 mol / L;

[0061] Step D: Add the mixed salt solution dropwise to the black suspension, heat to 70-90℃ and stir continuously for 4-8 hours; after the stirring reaction is completed, allow it to cool naturally to room temperature; wash the solid product obtained by centrifugation and then vacuum dry it to obtain a black powder;

[0062] Step E: Place the black powder in an argon atmosphere and heat it to 700-900℃ at a heating rate of 5℃ / min and calcine for 1-2 hours. After calcination, allow it to cool naturally to room temperature to obtain the final product.

[0063] The technical solution of the present invention achieves the following beneficial technical effects:

[0064] 1. This invention relates to a machine learning algorithm for full-cycle metal element screening of diatomic catalysts based on product yield prediction. It utilizes a general, broad-spectrum descriptor generated by the Feature Sparse and Interpretable Feature Space Construction (FSIFSC) algorithm. This descriptor employs physically meaningful metal atom properties, enabling comprehensive screening of all metal elements (including s, p, d, and f blocks) in the propane dehydrogenation to propylene reaction, and directly predicting the propylene (C3H6) yield. By introducing features such as frontier orbital occupancy, this descriptor can effectively distinguish the influence of alkali metals, transition metals, post-transition metals, and other metal elements on catalytic performance, solving the problem that traditional descriptors can only predict catalyst selectivity but cannot directly predict catalytic yield.

[0065] 2. The FSIFSC algorithm proposed in this invention combines multivariate arithmetic operations in feature engineering with feature importance evaluation methods in machine learning models, ensuring that the descriptor still possesses good sparsity and interpretability in high-dimensional feature spaces. The interpretable descriptor (φ) constructed in this invention... USF A combination of high-throughput computing and machine learning was used to screen various diatomic catalyst combinations containing s-block, p-block, d-block, and f-block elements. The propane conversion energy barrier ΔG for each combination was calculated. activity The difference between the energy barrier for propylene desorption and further dehydrogenation represents ΔG. selectivityBy combining the volcano-like relationships predicted by descriptors to determine the optimal catalyst region, an IrGa@NC diatomic catalyst with excellent performance potential was ultimately selected. The random forest regression model trained on descriptors showed excellent accuracy, confirming the precision of the machine learning work. This invention provides theoretical support and practical tools for the efficient design of catalysts for propane dehydrogenation to propylene.

[0066] 3. The optimal interpretability descriptor φ constructed using this invention USF The screened catalyst IrGa@NC achieves a propylene yield of 51.10% at 580°C, which is 3 to 5 times higher than existing catalysts and significantly superior to traditional transition metal-based catalysts. Density functional theory (DFT) calculations verify that the Ir and Ga atoms in the propane dehydrogenation to propylene catalyst IrGa@NC screened using the method of this invention are stably anchored in a diatomic form within the nitrogen-doped carbon matrix, and their highest occupied molecular orbitals (HOMOs) are appropriately positioned, effectively regulating the adsorption and desorption intensities of reactants and products. This conforms to the Sabatier principle and provides theoretical support for achieving synergistic optimization of high selectivity and high activity.

[0067] 4. The interpretable descriptor constructed by the machine learning algorithm for full-cycle metal element screening of diatomic catalysts based on product yield prediction in this invention is applicable to all metal elements (including s-block, p-block, d-block, and f-block), and is particularly suitable for performance prediction of diatomic catalysts (DACs) composed of transition metals and post-transition metals. Compared with traditional DFT calculations, this invention achieves a speedup of 232,228 times, and is suitable for high-throughput material design and rapid prediction of catalytic performance. Attached Figure Description

[0068] Figure 1 A schematic diagram of the combined operation process of input features based on the FSIFSC algorithm in this embodiment of the invention;

[0069] Figure 2 A comparison of the prediction accuracy of different models for propane conversion in the embodiments of the present invention; in the figure: (a) DTR model; (b) ETR model; (c) KNR model; (d) KRR model; (e) GBM model; (f) GBR model; (g) SVR model; (h) XGBR model; (i) GPR model;

[0070] Figure 3 A comparison of the prediction accuracy of different models for propylene selectivity in the embodiments of the present invention; in the figure: (a) DTR model; (b) ETR model; (c) KNR model; (d) KRR model; (e) GBM model; (f) GBR model; (g) SVR model; (h) XGBR model; (i) GPR model;

[0071] Figure 4 A comparison of the prediction accuracy of different models for propylene yield in the embodiments of the present invention; in the figure: (a) DTR model; (b) ETR model; (c) KNR model; (d) KRR model; (e) GBM model; (f) GBR model; (g) SVR model; (h) XGBR model; (i) GPR model;

[0072] Figure 5 Comparison of root mean square error (RMSE) and coefficient of determination (R²) of different models in the embodiments of the present invention; (a) propane conversion, (b) propylene selectivity, (c) propylene yield;

[0073] Figure 6 A comparison of the catalytic activity (a) and selectivity predictions (b) of 23 diatomic catalysts screened using the random forest model and their DFT calculation results in this embodiment of the invention.

[0074] Figure 7 A schematic diagram of the reaction network for the propane dehydrogenation to propylene reaction on a diatomic catalyst in an embodiment of the present invention;

[0075] Figure 8 DFT models of diatomic active sites with four different coordination structures (Qv1, Qv2, Qv3 and Qv4) in this embodiment of the invention;

[0076] Figure 9 The relationship between the propane conversion energy barrier and 10 randomly different DACs in the propane dehydrogenation to propylene reaction in this embodiment of the invention and the average value of the four coordination structures is shown in the figure.

[0077] Figure 10 The present invention presents the energy barrier for propylene dehydrogenation and the energy difference for propylene desorption in the propane dehydrogenation to propylene reaction, as well as their relationship with the Qv1 coordination structure (a), Qv2 coordination structure (b), Qv3 coordination structure (c) and Qv4 coordination structure (d) of 10 randomly different DACs, and their relationship with the average value of the four coordination structures.

[0078] Figure 11 XPS spectra of the catalyst synthesized in this embodiment of the invention before use; in the figure: (a) C 1s, (b) N 1s, (c) Ir 4f, (d) Ga 2p;

[0079] Figure 12 FT-EXAFS fitting curves of Ir foil (a) and IrO2 (b) in embodiments of the present invention;

[0080] Figure 13 FT-EXAFS fitting curves of Ga foil (a) and Ga2O3 (b) in embodiments of the present invention;

[0081] Figure 14 AC-HAADF-STEM image (a) of the catalyst synthesized in this embodiment after stability testing and discarding, and intensity distribution diagrams of corresponding regions A1(b) and A2(c) in (a);

[0082] Figure 15 XRD diffraction pattern of the catalyst synthesized in the embodiments of the present invention after stability testing;

[0083] Figure 16 XPS spectra of the catalyst synthesized in this embodiment of the invention after stability testing; in the figure: (a) C1s, (b) N 1s, (c) Ir 4f, (d) Ga 2p;

[0084] Figure 17 A comparison of the efficiency of the DFT method and the ML method in this embodiment of the invention is shown in the figure.

[0085] Figure 18 The graph shows the prediction accuracy of the random forest regression model for propane conversion (a), propylene selectivity (b), and propylene yield (c) in this embodiment of the invention. Detailed Implementation

[0086] This embodiment uses the screening of diatomic catalysts for the propane dehydrogenation to propylene reaction as an example to provide a detailed explanation of the machine learning algorithm for full-cycle metal element screening of diatomic catalysts based on product yield prediction.

[0087] 1. Construction of interpretable descriptors

[0088] The method for constructing the optimal interpretable descriptor in this embodiment includes the following steps:

[0089] Step (1), as follows Figure 1 As shown, the intrinsic and derived properties of the metal elements in the diatomic catalyst for propane dehydrogenation to propylene reaction are collected to form characteristic space 1; the metal elements in the diatomic catalyst are located in the s-block, p-block, d-block, or f-block; the intrinsic properties of the metal elements include atomic number. Total number of front-line orbital electrons First ionization energy atomic radius Total number of valence electrons n, number of periods atomic mass and electronegativity Derived properties of metallic elements include frontier orbital occupancy. ;

[0090] Metal elements The formula for calculating the frontline track occupancy is:

[0091] (1);

[0092] In equation (1), Metallic elements The occupancy rate of the front-line tracks, Metallic elements The total number of front-line orbital electrons, Metallic elements The total number of valence electrons;

[0093] Step (2): Based on the FSIFSC algorithm, calculate the comprehensive importance score of each attribute feature in feature space 1 according to the feature importance score and the comprehensive SHAP value; perform preliminary screening of the attribute features in the feature dataset according to the comprehensive importance score results to obtain feature space 2;

[0094] The feature importance scoring method is as follows: using catalytic activity (propane conversion), selectivity (propylene selectivity), and product yield (propylene yield) as prediction targets, random forest feature importance analysis is used to score the importance of each attribute feature; the formula for calculating the importance score of each attribute feature is as follows:

[0095] (2);

[0096] In equation (2), Score the importance of each attribute feature; The feature importance score is calculated using catalytic activity as the prediction target; The feature importance score is calculated using selectivity as the prediction objective; The feature importance score is calculated using product yield as the prediction target;

[0097] Overall SHAP value The calculation formula is as follows:

[0098] (3);

[0099] (4);

[0100] In equation (3), The combined SHAP value for all attribute features; The SHAP value is calculated using catalytic activity as the prediction target. The SHAP value is calculated with selectivity as the prediction objective. The SHAP value is calculated using product yield as the prediction target.

[0101] In equation (4), It can be 1, 2, or 3; For SHAP value, For the entire feature set; for The middle does not contain features A subset of; For subset The number of features; For only subsets The model prediction value of the feature; The total number of features;

[0102] Comprehensive score of the importance of each attribute feature The value is calculated as follows:

[0103] (5).

[0104] In this embodiment, after the initial screening in step (2), the frontline track occupancy is selected. atomic radius and electronegativity The feature space consists of three important attribute features.

[0105] Step (3): Based on the FSIFSC algorithm, perform single feature operations and label the three initial features in feature space 2, and then perform label operations and combination operations to construct the multivariate descriptor space matrix; specifically, the following steps are included:

[0106] Step (3-1), Single Feature Operation: Perform exponentiation, multiplication, and division operations on each initial feature in feature space 2, and add the results of the single feature operations to the feature space of feature space 2 to obtain feature space 3; the method of exponentiation is: perform nth power operation and nth root operation on each initial feature to obtain two exponentiation features, where n is greater than or equal to 1; both multiplication and division operations are performed using a factor of m; m is greater than or equal to 1; that is:

[0107] Exponentiation is used to perform calculations of raising to the power of n and taking the nth root, where n is a variable. For example, for a radius rM, when n=2, exponentiation yields... and Then, the results of these two exponentiation operations are added to the feature space; the new feature space contains the initial features, the square term, and the square root term; the calculation method of multiplication and division operations is similar to that of exponentiation operations, multiplying or dividing by m, where m is also a variable that can be changed;

[0108] Step (3-2): Label all physical quantity features in feature space 3 according to the original physical quantity type. That is, label the physical quantity features generated by single feature operation with the same initial feature in feature space 3 with the same label, and label the physical quantity features generated by single feature operation with different initial features with different labels to distinguish different physical quantity features.

[0109] Step (3-3), Same-label operation: Perform addition, subtraction and absolute value operation, multiplication and division operation on all physical quantity features with the same label in feature space 3 respectively, and add the same-label operation results to feature space 3 to obtain feature space 4;

[0110] Step (3-4), Different label operations: Perform multiplication and division operations on the physical quantity features of any two different labels in feature space 3, and add the results of the different label operations to feature space 3 to obtain feature space 5;

[0111] Steps (3-5), Combination operation: Combine the physical quantity features in feature space 4 and feature space 5 through multiplication and division operations to generate a new feature space; integrate feature space 4, feature space 5 and the new feature space to form a high-dimensional feature space 6, which is the multivariate descriptor space matrix.

[0112] This demonstrates that with only three initial features, the FSIFSC algorithm can expand the feature space dimension to the millions.

[0113] Step (4): Descriptor importance is evaluated using the Gradient Boosting Regression (GBR) model, and the dimensionality of the multivariate descriptor space matrix is ​​reduced using the Recursive Feature Elimination (RFE) algorithm to obtain the dimensionality-reduced descriptor space matrix. The specific method is as follows:

[0114] Step (4-1), Coarse Screening: Using the accuracy of product yield prediction as an indicator, a lightweight gradient boosting regression model is adopted. Through univariate evaluation, the score value of the contribution of each descriptor in the multivariate descriptor space matrix to the target prediction is calculated sequentially. Based on the calculation results of each descriptor score value, 90% of the descriptors with lower scores in the multivariate descriptor space matrix are removed, and the remaining descriptors with higher scores are retained to obtain the coarse screening descriptor space matrix. The number of trees in the lightweight gradient boosting model is 10-20, the maximum tree depth is 2-3, the feature sampling rate is 0.01-0.05, the sample sampling rate is 0.5-0.7, and the learning rate is 0.3-0.5.

[0115] Step (4-2), Intermediate Screening: Based on the recursive feature elimination algorithm combined with the gradient boosting regression model, a model is constructed and trained using descriptors in the coarse-screened descriptor space matrix. The gradient boosting regression model is used to calculate the score value of the descriptor's contribution to the target prediction. 15% of the descriptors with lower scores are eliminated in a single batch from the coarse-screened descriptor space matrix, and the remaining descriptors with higher scores are retained. The intermediate screening elimination step is repeated iteratively until 1200 descriptors remain in the coarse-screened descriptor space matrix, thus obtaining the intermediate-screened descriptor space matrix. The number of gradient boosting model trees is 20-30, the maximum tree depth is 3-4, the feature sampling rate is 0.05-0.1, the sample sampling rate is 0.7-0.8, and the learning rate is 0.1-0.2.

[0116] Step (4-3), Fine screening: Based on the recursive feature elimination algorithm combined with the gradient boosting regression model, the model is constructed and trained using the descriptors in the mid-screen descriptor space matrix. The gradient boosting regression model is used to calculate the score value of the descriptor's contribution to the target prediction. 5% of the descriptors with lower scores are eliminated in a single batch from the mid-screen descriptor space matrix, and the remaining descriptors with higher scores are retained. The fine screening elimination step is repeated iteratively until 50 descriptors remain in the mid-screen descriptor space matrix, and the dimension-reduced descriptor space matrix is ​​obtained.

[0117] Step (5): The descriptors in the dimension-reduced descriptor space matrix are processed using the minimum absolute shrinkage and selection operator. Variable selection is performed by reducing the coefficients of non-critical features to zero, thereby selecting the optimal interpretable descriptor.

[0118] The calculation formula for the optimal interpretable descriptor for screening diatomic catalysts in the propane dehydrogenation to propylene reaction constructed in this embodiment is as follows:

[0119] (6);

[0120] In equation (6), For interpretable descriptor values; Metal atomic radius; Metal atomic radius of metals and metal These are two different metal sites in a diatomic catalyst; Metal The occupancy rate of the front-line tracks; Metal Electronegativity.

[0121] 2. Selection of regression models

[0122] To determine the optimal regression model, this embodiment uses propane conversion, propylene selectivity, and propylene yield as evaluation indicators, and compares the prediction accuracy and performance evaluation scores of nine machine learning models: RF, DTR (Decision Tree Regressor), ETR (ExtraTrees Regressor), KNR (K-Nearest Neighbors Regressor), KRR (Kernel Ridge Regression), GBM (Gradient Boosting Machine), GBR (Gradient Boosting Regressor), SVR (Support Vector Regression), XGBR (XGBoost Regression), and GPR (Gaussian Process Regression).

[0123] The tests performed on the catalytic performance included propane conversion ( ), propylene selectivity ( ) and propylene yield ( ).in:

[0124] (9);

[0125] (10);

[0126] (11);

[0127] In equations (9) and (10), This refers to the amount of C3H8 molecules flowing into the reactor.

[0128] This refers to the amount of C3H8 molecules flowing out of the reactor. This represents the amount of C3H6 molecules flowing out of the reactor.

[0129] like Figures 2 to 4 as well as Figure 18 The results compare the prediction accuracy of different models for propane conversion, propylene selectivity, and propylene yield. Figures 2 to 4 as well as Figure 18 The selectivity values ​​calculated using density functional theory were compared with the predicted values ​​of each model to characterize the model's prediction accuracy. Figures 2 to 4 as well as Figure 18 It can be seen that the RF (Random Forest) model has the best prediction accuracy for propane conversion, propylene selectivity, and propylene yield.

[0130] like Figure 5 To evaluate the predictive performance of different models on propane conversion, propylene selectivity, and propylene yield, RMSE and R0 were determined. 2 The comparison results. From Figure 5 As can be seen, RF outperforms other regression models in predicting all three objectives.

[0131] To ensure the stability of the prediction accuracy of the random forest model, this embodiment also randomly selected 23 catalysts for further evaluation, such as... Figure 6 As shown (ML refers to machine learning, DFT refers to density functional theory calculation), from Figure 6 As can be seen, using Random Forest as the regression model, the average error between its prediction results and the DFT calculation results is less than 0.1, proving that the Random Forest regression model has high prediction accuracy. Therefore, this embodiment selects the Random Forest model (RF) as the regression model.

[0132] 3. Screening of diatomic catalysts for propane dehydrogenation to propylene reaction

[0133] The interpretable descriptors constructed in this embodiment are applied to a random forest regression model to train and predict propylene yield, thereby screening diatomic catalysts for the propane dehydrogenation to propylene reaction. Density functional theory is used to calculate the propane conversion barrier, propylene dehydrogenation barrier, and propylene desorption energy of diatomic catalysts with different metal element combinations. The propane conversion barrier is then used as the catalytic activity label, and the difference between the propylene dehydrogenation barrier and the propylene desorption energy is used as the selectivity label. The catalytic activity label value and the selectivity label value are normalized and multiplied to obtain the propylene yield label. The descriptor value of the optimal interpretable descriptor for the diatomic catalyst is used as input to train the random forest regression model. Five-fold cross-validation is used to train and validate the constructed random forest regression model, with the root mean square error (RMSE) and coefficient of determination (CDO) of propylene yield used as evaluation indicators. A grid search method is used to automatically find hyperparameters, determine the hyperparameter depth of the model, and obtain the optimal prediction model. The root mean square error (RMSE) is used to determine the hyperparameter depth. and coefficient of determination The calculation formula is:

[0134] (7);

[0135] (8);

[0136] In equations (7) and (8), For sample size; The actual value; These are the model's predicted values; The mean of the true values; The sum of squared residuals represents the total amount of prediction error in the model. The total sum of squares reflects the overall variability of the sample data.

[0137] Figure 7 This diagram illustrates the reaction network of a diatomic catalyst in the propane dehydrogenation to propylene reaction. Starting with propane adsorption dehydrogenation, the conversion rate is measured using the first dehydrogenation step (TS1) energy barrier, while the selectivity of propylene is measured by the difference between the TS3 energy barrier (the dehydrogenation step) and the dissociation energy barrier. Currently, four configurations of diatomic catalysts are generally accepted. Figure 8 There are four different load carriers: Qv1, Qv2, Qv3, and Qv4. This embodiment selects the final configuration from two aspects: First, the reaction energy barriers of different configurations with the same metal combination are calculated, and it is found that the energy on the Qv2 structure is relatively close to the average value of the four configurations (see...). Figure 9 and Figure 10 Secondly, literature confirms that Qv2 is the structure with the lowest energy; therefore, Qv2 was selected as the support for DFT calculations in this embodiment. In this embodiment, when screening diatomic catalysts, the molar ratio of the two metal elements was predetermined to be 1:1, and the mass ratio of the metal elements to the nitrogen-doped carbon material was kept constant.

[0138] The optimal prediction model trained in this embodiment was used to screen catalysts for propane dehydrogenation to propylene, and the final selected catalyst for propane dehydrogenation to propylene was IrGa@NC.

[0139] 4. Preparation of IrGa@NC catalyst

[0140] In this embodiment, the IrGa@NC catalyst was synthesized by impregnation-calcination method, and the structure of the IrGa@NC catalyst was characterized by X-ray diffraction (XRD), aberration-corrected high-angle annular dark-field scanning transmission electron microscopy (AC-HAADF-STEM), X-ray photoelectron spectroscopy (XPS), and X-ray absorption fine structure (XAFS).

[0141] (1) Preparation method of IrGa@NC catalyst

[0142] First, 1.495 g of Zn(NO3)2·6H2O was dissolved in 40 mL of methanol to prepare a zinc nitrate methanol solution, and 1.026 g of 2-methylimidazole was dissolved in 40 mL of methanol to prepare a 2-methylimidazole methanol solution. The zinc nitrate methanol solution and the 2-methylimidazole methanol solution were mixed and stirred thoroughly at room temperature, and then stirred for 15 h at room temperature. The solid precipitate was then collected by centrifugation, washed with methanol, and vacuum dried (60 °C, 12 h) to obtain a solid product. The solid product was placed in a tube furnace and pyrolyzed at high temperature under an argon atmosphere (heated to 1000 °C for 2 h at a heating rate of 5 °C / min) to obtain nitrogen-doped carbon material (NC). 0.2 g of NC was added to 30 mL of anhydrous ethanol and sonicated to form a homogeneous black suspension. Subsequently, equimolar amounts (0.01 mol each) of two metal salts (Ga...) were added... (NO3)3·9H2O and H2IrCl6·xH2O) were dissolved in 1 mL of anhydrous ethanol and stirred to obtain a mixed salt solution. The mixed salt solution was then added dropwise to a black suspension and stirred continuously at 80 °C for 6 h. After the reaction was completed, the mixture was allowed to cool naturally to room temperature. The solid product obtained by centrifugation was washed and then vacuum dried (60 °C, 12 h) to obtain a black powder. Finally, the black powder was placed in an argon atmosphere and heated to 800 °C at a heating rate of 5 °C / min and calcined for 1 h. After calcination, the mixture was allowed to cool naturally to room temperature to obtain the diatomic catalyst IrGa@NC.

[0143] (2) Characterization analysis of the diatomic catalyst IrGa@NC

[0144] from Figure 11 It can be seen that Ir and Ga atoms are uniformly dispersed in the nitrogen-doped carbon support in a diatomic form, and no metal nanoparticles were found. Figure 12 The Fourier transform extended X-ray absorption fine structure (FT-EXAFS) fitting curves of Ir foil (metallic Ir) and IrO2 (oxidized Ir) at the Ir L3-edge are shown. These reference samples were used for comparison with the coordination environment of Ir atoms in the IrGa@NC catalyst. This figure illustrates that Ir exists in the oxidized form of the IrGa@NC catalyst prepared in this embodiment, exhibiting atomic-level dispersion characteristics and primarily coordinating with nitrogen (Ir–N) rather than forming Ir nanoparticles. Figure 13 The FT-EXAFS fitting curves of Ga foil (metallic Ga) and Ga₂O₃ (oxidized Ga) at the Ga K-edge are shown. This figure is consistent with... Figure 12 Together, they serve as a reference for the coordination structure of bimetallic atoms in IrGa@NC.

[0145] The catalytic performance of IrGa@NC was experimentally tested. The results showed that under the conditions of propane dehydrogenation to propylene reaction at 580°C, IrGa@NC exhibited excellent catalytic performance, with a propane conversion rate of 56.6%, a propylene selectivity of 94.0%, and a propylene yield of 51.1%, which is 3 to 5 times higher than that of existing reported catalysts. Moreover, its performance remained stable in the 10-hour stability test, with no obvious deactivation.

[0146] Figure 14 The image shows the aberration-corrected high-angle annular dark-field scanning transmission electron microscopy (AC-HAADF-STEM) image and intensity distribution of the IrGa@NC catalyst after a 10-hour PDH reaction stability test. This figure verifies the relationship between the structural reliability of IrGa@NC and its excellent catalytic performance under actual reaction conditions.

[0147] Figure 15 X-ray diffraction (XRD) patterns of the IrGa@NC samples before and after stability testing are presented. After stability testing, no obvious crystal diffraction peaks of metallic Ir or Ga are found in the patterns; only the characteristic peaks of the carbon support are retained. This indicates that even after a long reaction period, Ir and Ga remain highly dispersed in atomic form on the support, without significant aggregation or precipitation. This is consistent with the observations of AC-HAADF-STEM, jointly demonstrating that the structure of IrGa@NC exhibits good thermal stability during the reaction process.

[0148] 5. Efficiency Verification

[0149] To conduct a comparative study, this embodiment uses central processing unit hours (CPU-h) as an indicator to evaluate two methods—one based on density functional theory (DFT) and the other based on a random forest regression model—in predicting the free energy change of catalytic activity (ΔG). activity ) and selective free energy change (ΔG) selectivity The computational cost at that time.

[0150] In density functional theory calculations, the optimization of the geometric structure of diatomic catalysts (DACs), the calculation of the transition state (propane conversion barrier) of the first-step dehydrogenation reaction of propane, the calculation of the difference between the propylene dehydrogenation barrier and the propylene desorption energy, and the calculation of the Gibbs free energy were all completed on a 128-core supercomputer using VASP software, with an average time of approximately one hour. Figure 17 As shown, on average, a single ΔG activity and ΔG selectivity The calculation takes about 30 CPU hours, which is equivalent to about 3870 CPU hours on a single-core CPU.

[0151] In contrast, in the machine learning prediction described above in this embodiment, the entire process—from model training and testing to predicting the selectivity, conversion rate, and product yield—can be completed in approximately 60 seconds on a single-core CPU. In other words, predicting the catalytic activity of transition metal diatomic catalysts using machine learning methods is more than 232,228 times faster than density functional theory calculations.

[0152] The test results of this embodiment fully verify the effectiveness and practicality of the descriptor constructed by the machine learning algorithm for full-cycle metal element screening of diatomic catalysts based on product yield prediction in guiding the design of high-performance diatomic catalysts. Compared with the prior art, the inventiveness of this invention can be summarized in the following table:

[0153] Table 1

[0154]

Claims

1. A machine learning algorithm for full-cycle metal element screening of diatomic catalysts based on product yield prediction, characterized in that, Includes the following steps: Step (1): Collect the intrinsic and derived properties of the metal elements in the diatomic catalyst to form the feature space 1; the intrinsic properties of the metal elements include atomic number, total number of frontier orbital electrons, first ionization energy, atomic radius, total number of valence electrons, period number, atomic mass, electronegativity and activation energy; the derived properties of the metal elements include frontier orbital occupancy. The formula for calculating the frontier orbital occupancy of metallic elements is: (1); In equation (1), Metallic elements The occupancy rate of the front-line tracks, Metallic elements The total number of front-line orbital electrons, Metallic elements The total number of valence electrons; Step (2): Based on the FSIFSC algorithm, calculate the comprehensive importance score of each attribute feature in feature space 1 according to the feature importance score and the comprehensive SHAP value; perform preliminary screening of the attribute features in the feature dataset according to the comprehensive importance score results to obtain feature space 2; Step (3): Based on the FSIFSC algorithm, perform single feature operations and label each initial feature in feature space 2, and then perform label operations and combination operations to construct a multivariate descriptor space matrix. Step (4): Use the gradient boosting regression model to evaluate the importance of descriptors, and use the recursive feature elimination algorithm to reduce the dimensionality of the multivariate descriptor space matrix to obtain the dimensionality-reduced descriptor space matrix. Step (5): Use the minimum absolute shrinkage and selection operators to further filter the optimal interpretable descriptors from the dimension-reduced descriptor space matrix; Step (6): Apply the optimal interpretable descriptor to the random forest regression model and construct the diatomic catalyst screening model through model training; use the diatomic catalyst screening model to screen target diatomic catalysts and predict product yield.

2. The machine learning algorithm for full-cycle metal element screening of diatomic catalysts based on product yield prediction according to claim 1, characterized in that, In step (1), the metal element in the diatomic catalyst is located in the s-block, p-block, d-block, or f-block.

3. The machine learning algorithm for full-cycle metal element screening of diatomic catalysts based on product yield prediction according to claim 1, characterized in that, In step (2), the scoring method for feature importance is as follows: using catalytic activity, selectivity, and product yield as prediction targets, random forest feature importance analysis is used to score the importance of each attribute feature; the formula for calculating the importance score of each attribute feature is: (2); In equation (2), Score the importance of each attribute feature; The feature importance score is calculated using catalytic activity as the prediction target; The feature importance score is calculated using selectivity as the prediction objective; The feature importance score is calculated using product yield as the prediction target; Overall SHAP value The calculation formula is as follows: (3); (4); In equation (3), The combined SHAP value for all attribute features; The SHAP value is calculated using catalytic activity as the prediction target. The SHAP value is calculated with selectivity as the prediction objective. The SHAP value is calculated using product yield as the prediction target. In equation (4), It can be 1, 2, or 3; For the entire feature set; for The middle does not contain features A subset of; For subset The number of features; For only subsets The model prediction value of the feature; The total number of features; Comprehensive score of the importance of each attribute feature The value is calculated as follows: (5)。 4. The machine learning algorithm for full-cycle metal element screening of diatomic catalysts based on product yield prediction according to claim 1, characterized in that, The specific calculations in step (3) include the following steps: Step (3-1), Single Feature Operation: Perform exponentiation, multiplication, and division operations on each initial feature in feature space 2, and add the results of the single feature operations to the feature space of feature space 2 to obtain feature space 3; The method of exponentiation is as follows: perform nth power operation and nth root operation on each initial feature to obtain two exponentiation operation features, where n is greater than or equal to 1; Multiplication and division operations are performed using m factors; m is greater than or equal to 1; Step (3-2): Label all physical quantity features in feature space 3 according to the original physical quantity type. That is, label the physical quantity features generated by single feature operation with the same initial feature in feature space 3 with the same label, and label the physical quantity features generated by single feature operation with different initial features with different labels to distinguish different physical quantity features. Step (3-3), Same-label operation: Perform addition, subtraction and absolute value operation, multiplication and division operation on all physical quantity features with the same label in feature space 3 respectively, and add the same-label operation results to feature space 3 to obtain feature space 4; Step (3-4), Different label operations: Perform multiplication and division operations on the physical quantity features of any two different labels in feature space 3, and add the results of the different label operations to feature space 3 to obtain feature space 5; Step (3-5), Combination operation: Combine the physical quantity features in feature space 4 and feature space 5 through multiplication and division operations to generate a new feature space; Feature space 4, feature space 5 and the new feature space are integrated to form feature space 6, which is the multivariate descriptor space matrix.

5. The machine learning algorithm for full-cycle metal element screening of diatomic catalysts based on product yield prediction according to claim 4, characterized in that, The specific method for step (4) is as follows: Step (4-1), coarse screening: Using the accuracy of product yield prediction as an indicator, a lightweight gradient boosting regression model is adopted. The score value of the contribution of each descriptor in the multivariate descriptor space matrix to the target prediction is calculated sequentially through univariate evaluation. Based on the calculation results of the score value of each descriptor, 80-90% of the descriptors with lower score values ​​in the multivariate descriptor space matrix are removed, and the remaining descriptors with higher score values ​​are retained to obtain the coarse screening descriptor space matrix. Step (4-2), intermediate screening: Based on the recursive feature elimination algorithm combined with the gradient boosting regression model, the model is constructed and trained using the descriptors in the coarse screening descriptor space matrix. The gradient boosting regression model is used to calculate the score value of the descriptor's contribution to the target prediction. 10-15% of the descriptors with lower scores are eliminated in a single batch from the coarse screening descriptor space matrix, and the remaining descriptors with higher scores are retained. The intermediate screening elimination step is repeated iteratively until 1000-1500 descriptors remain in the coarse screening descriptor space matrix, thus obtaining the intermediate screening descriptor space matrix. Step (4-3), Fine Screening: Based on the recursive feature elimination algorithm combined with the gradient boosting regression model, the model is constructed and trained using the descriptors in the mid-screen descriptor space matrix. The gradient boosting regression model is used to calculate the score value of the descriptor's contribution to the target prediction. 3-5% of the descriptors with lower scores are eliminated in a single batch from the mid-screen descriptor space matrix, and the remaining descriptors with higher scores are retained. The fine screening elimination step is repeated iteratively until 40-80 descriptors remain in the mid-screen descriptor space matrix, resulting in a dimensionality-reduced descriptor space matrix.

6. The machine learning algorithm for full-cycle metal element screening of diatomic catalysts based on product yield prediction according to claim 5, characterized in that, In step (4-1), the number of lightweight gradient boosting model trees is 10-20, the maximum tree depth is 2-3, the feature sampling rate is 0.01-0.05, the sample sampling rate is 0.5-0.7, and the learning rate is 0.3-0.

5. In step (4-2), the number of gradient boosting model trees is 20-30, the maximum tree depth is 3-4, the feature sampling rate is 0.05-0.1, the sample sampling rate is 0.7-0.8, and the learning rate is 0.1-0.

2. The number of descriptors in the middle-screen descriptor space matrix is ​​1200. In step (4-3), the number of gradient boosting model trees is 50-100, the maximum tree depth is 4-5, the feature sampling rate is 0.1-0.2, the sample sampling rate is 0.8-0.9, the learning rate is 0.05-0.1, and L1 / L2 regularization of 0.1-1.0 is added. The number of descriptors in the dimensionality reduction descriptor space matrix is ​​50.

7. The machine learning algorithm for full-cycle metal element screening of diatomic catalysts based on product yield prediction according to claim 1, characterized in that, In step (5), the descriptors in the dimension-reduced descriptor space matrix are processed using the minimum absolute shrinkage and selection operator, and the final feature selection is performed by reducing the coefficients of non-key features to zero. In step (6), the diatomic catalyst is used to catalyze the dehydrogenation reaction. During model training: the catalytic activity, selectivity, and product yield of the diatomic catalyst are the objectives, and the descriptor value of the optimal interpretable descriptor is used as input to train the random forest regression model. Guided by the predictive performance of catalytic activity, selectivity, and product yield, the hyperparameters are searched using the grid search automatic parameter search method to determine the hyperparameter depth of the random forest regression model. The five-fold cross-validation method is used to evaluate the model performance, and the root mean square error and coefficient of determination of the product yield are used as the evaluation indicators of the model performance. Density functional theory was used to calculate the reaction feed conversion barrier, product dehydrogenation barrier, and product desorption energy of the diatomic catalyst. The reaction feed conversion barrier was used as the catalytic activity label, and the difference between the product dehydrogenation barrier and the product desorption energy was used as the selectivity label. The product yield label was obtained by multiplying the catalytic activity label value and the selectivity label value after homogenization. The method for screening target diatomic catalysts and predicting product yields using a model is as follows: input the descriptor values ​​of the diatomic catalysts to be screened into the trained diatomic catalyst screening model, and the diatomic catalyst screening model outputs the product yields corresponding to the diatomic catalysts to be screened; draw a volcano plot based on the output product yields and determine the final required diatomic catalysts.

8. An application of a machine learning algorithm for full-cycle metal element screening of diatomic catalysts based on product yield prediction, characterized in that, The machine learning algorithm for full-cycle metal element screening of diatomic catalysts based on product yield prediction as described in any one of claims 1-7 was used to screen out diatomic catalysts for propane dehydrogenation to propylene reaction and predict propylene yield. The formula for calculating the optimal interpretable descriptor is as follows: (6); In equation (6), For interpretable descriptor values; Metal atomic radius; Metal atomic radius of metals and metal These are two different metal sites in a diatomic catalyst; Metal The occupancy rate of the front-line tracks; Metal Electronegativity.

9. The application of the machine learning algorithm for full-cycle metal element screening of diatomic catalysts based on product yield prediction as described in claim 8, characterized in that, IrGa@NC was selected as the diatomic catalyst for the propane dehydrogenation to propylene reaction.

10. The application of the machine learning algorithm for full-cycle metal element screening of diatomic catalysts based on product yield prediction according to claim 9, characterized in that, The preparation method of IrGa@NC is as follows: Step A: Mix zinc nitrate methanol solution with a molar concentration of 0.1-0.15 mol / L and 2-methylimidazolium methanol solution with a molar concentration of 0.3-0.5 mol / L in equal volume ratio, and stir at room temperature for 12-16 h. The solid precipitate was then collected by centrifugation and vacuum dried. Step B: The solid product obtained by vacuum drying is heated to 900-1050℃ for 1-3 hours under an inert atmosphere at a heating rate of 5℃ / min to obtain nitrogen-doped carbon material. Step C: Disperse nitrogen-doped carbon material in anhydrous ethanol to obtain a black suspension; add equimolar amounts of Ga(NO3)3·9H2O and H2IrCl6·xH2O to anhydrous ethanol to prepare a mixed salt solution; the mass concentration of nitrogen-doped carbon material in the black suspension is 5-10 g / L; the molar concentration of Ga(NO3)3·9H2O in the mixed salt solution is 5-15 mol / L; Step D: Add the mixed salt solution dropwise to the black suspension, heat to 70-90℃ and stir continuously for 4-8 hours; after the stirring reaction is completed, allow it to cool naturally to room temperature; wash the solid product obtained by centrifugation and then vacuum dry it to obtain a black powder; Step E: Place the black powder in an argon atmosphere and heat it to 700-900℃ at a heating rate of 5℃ / min and calcine for 1-2 hours. After calcination, allow it to cool naturally to room temperature to obtain the final product.