A descriptor construction and optimization method and system for a fischer-tropsch synthesis iron-based supported catalyst
By employing high-throughput testing and a hierarchical weighted screening mechanism, the problem of immature descriptors for iron-based supported catalysts in Fischer-Tropsch synthesis was solved. Key descriptors were automatically selected, shortening the catalyst development cycle and improving the model's discriminative ability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ANHUI UNIV
- Filing Date
- 2026-03-09
- Publication Date
- 2026-06-05
AI Technical Summary
The lack of a systematic descriptor classification and multi-stage robust screening mechanism in existing technologies makes it difficult to reliably identify a subset of key descriptors with statistical significance and low redundancy for the performance of iron-based supported catalysts in Fischer-Tropsch synthesis from a large number of candidate physical descriptors in high-throughput, small-sample data-driven scenarios.
By acquiring the formulation and structural information of iron-based supported catalyst samples, high-throughput testing was conducted. Combined with data cleaning and standardization, the samples were divided into four levels and weighted according to the levels. A lightweight prediction model was constructed using grouped K-fold partitioning and multi-stage screening to select a set of key descriptors.
It enables the automatic selection of key descriptors that are sensitive to the thermal performance of catalysts and have low redundancy from multi-source candidate descriptors, shortening the catalyst development cycle, reducing the resource investment in blind trial and error experiments, and improving the model's ability to distinguish different host phase/support combinations.
Smart Images

Figure CN122157878A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of catalyst performance prediction technology, and in particular to a method and system for constructing and optimizing descriptors for iron-based supported catalysts in Fischer-Tropsch synthesis. Background Technology
[0002] Fischer-Tropsch synthesis (FTS) is an important industrial pathway that uses syngas (CO / H2) as a raw material to convert it into liquid fuels and chemicals under the action of a catalyst, enabling the clean and high-value utilization of various carbon sources such as coal, natural gas, and biomass. Currently, FTS has significant application value in the production of synthetic fuels, sustainable jet fuel, waxes, and α-olefins. Iron-based catalysts, due to their wide availability of raw materials, relatively low cost, certain sulfur resistance, and the ability to exhibit water-gas shift (WGS) activity, are suitable for use in syngas systems with low H2 / CO ratios and occupy an important position in both industrial and basic research. To further improve the utilization efficiency and specific surface area of active sites, supported systems are often used in industrial practice and academic research. Iron or iron oxide precursors are dispersed on supports such as SiO2, Al2O3, TiO2, and carbon materials, and the reduction / carbonization behavior, surface electronic structure, and product selectivity are controlled by adding promoters such as K, Cu, Mn, and Zn.
[0003] In recent years, the continuous development of high-throughput experimental and automated characterization platforms has made it possible to systematically acquire large-scale catalytic formulation-process-performance data in a short period of time. Through parallel preparation, rapid activation, automated sampling analysis, and data recording, large-scale reaction databases can be formed in a short period of time, laying the data foundation for machine learning-assisted catalyst design. However, high-throughput data also faces several typical challenges: the data sources are diverse, covering text, numerical, categorical variables, and missing values, and there are quality problems such as inconsistent units, chaotic formats, noise, and outliers; the sample distribution is often unbalanced, with significant differences in coverage of different material systems, ratio ranges, and operating conditions, which can easily lead to systematic bias in the model; catalytic systems usually have strong nonlinear and high-order interaction characteristics, and it is difficult to fully reveal the key driving factors affecting performance by simply relying on linear statistical correlation methods; in addition, the performance descriptors of iron-based supported catalysts in Fischer-Tropsch synthesis reaction systems are still immature, and there is a lack of a set of candidate descriptors that have been systematically validated, which brings great difficulties to data-driven modeling.
[0004] Chinese patent CN111128311A discloses a method and system for screening catalytic materials based on high-throughput experiments and high-throughput calculations. This approach uses catalytic performance data obtained from high-throughput experiments and adsorption energies calculated based on first-principles calculations as a common foundation. Through machine learning, it sequentially establishes a structural feature-adsorption energy correlation model, an adsorption energy-catalytic performance correlation model, and a correction model for theoretical and experimental catalytic performance values, thus forming a complete catalyst structure-activity relationship model. The model is continuously corrected using rolling updates of experimental data, achieving an iterative closed loop of "theoretical screening—high-performance material preparation and characterization—theoretical screening." This approach is significant in accelerating catalyst screening and reducing R&D costs. However, the descriptor construction in this approach highly depends on the adsorption energies obtained from quantum chemical calculations. When faced with purely experimentally driven datasets of "small sample size to medium dimensionality," it lacks a systematic candidate descriptor classification, weighting, and multi-stage robust screening mechanism. This makes it difficult to automatically identify the most explanatory and low-redundancy subset of key descriptors from a large number of physicochemical property descriptors under statistical significance constraints. Summary of the Invention
[0005] In view of this, the present invention provides a method and system for constructing and optimizing descriptors for iron-based supported catalysts in Fischer-Tropsch synthesis, in order to solve the problem in the prior art that there is a lack of a systematic descriptor classification, hierarchical weighting and multi-stage robust screening mechanism for iron-based supported catalysts in Fischer-Tropsch synthesis, which makes it difficult to reliably identify a subset of key descriptors with statistical significance and low redundancy from a large number of candidate physical descriptors in high-throughput, small-sample data-driven scenarios.
[0006] The technical solution of this invention is implemented as follows: On one hand, this invention provides a method for constructing and optimizing descriptors for iron-based supported catalysts used in Fischer-Tropsch synthesis, comprising: S1. Obtain the formulation and structural information of the iron-based supported catalyst sample, and obtain the thermal imaging temperature under Fischer-Tropsch synthesis conditions through high-throughput testing, and determine the thermal imaging temperature as the target variable; at the same time, collect candidate descriptors of the main phase and the support from material databases and literature sources to form an initial dataset. S2. The initial dataset is preprocessed using data cleaning and standardization strategies to obtain a normalized descriptor dataset; S3. On the normalized descriptor dataset, the candidate descriptors are divided into four levels based on the Fischer-Tropsch synthesis reaction mechanism, and initial weights are assigned to the descriptors according to the level for weighted constraints, resulting in a hierarchical weighted descriptor set. S4. Based on the preset mixing rules, construct composite descriptors for the monomer property descriptors in the layered weighted descriptor set according to the actual mass ratio of the main phase to the carrier, and obtain an extended descriptor set. S5. Grouping is performed using the pair_id identifier between the main phase and the support as the grouping criterion. Within each training fold, the extended descriptor set is subjected to progressive screening of univariate statistical screening, multivariate stable selection, and importance consistency verification to obtain the key descriptor set. The key descriptor set is used as the optimized descriptor system for Fischer-Tropsch synthesis iron-based supported catalysts to predict the thermal imaging temperature of the catalyst sample to be confirmed under Fischer-Tropsch synthesis conditions.
[0007] Based on the above technical solutions, preferably, the candidate descriptors include composition descriptors, structure descriptors, and property descriptors; the composition descriptor includes the type of main phase material, the type of support material, and the mass ratio of the main phase to the support; the structure descriptor includes the electronegativity range, minimum atomic number, crystal system, and space group of the main phase and the support; the property descriptors include the band gap, thermal conductivity, shear modulus, Debye temperature, enthalpy change, entropy change, density, melting point range, and average boiling point of the main phase and the support.
[0008] Based on the above technical solutions, preferably, in step S2, the data cleaning and standardization strategy includes: removing outliers from high-throughput thermal imaging temperature data and removing samples that significantly deviate from the normal range; filling missing values for numerical descriptors with the median or mean, and classifying missing values for categorical descriptors into independent categories; standardizing the format of descriptor field names, removing invisible characters and redundant spaces, and unifying the naming method; performing dimensionless processing on numerical descriptors using the z-score standardization method; and merging rare categories with a frequency of less than 5 times in categorical descriptors into independent rare class labels, and then converting them into numerical form using a Bayesian smoothing target encoding method based on grouped K-folds.
[0009] Based on the above technical solutions, preferably, step S3 specifically includes: Based on the Fischer-Tropsch synthesis mechanism, the candidate descriptors in the normalized descriptor dataset are divided into four levels from bottom to top: L0 (formulation and identity), L1 (electronic scale and bonding), L2 (lattice mechanics and thermal transport), and L3 (phase stability and phase transition indicator). A hierarchical weight vector is then constructed. Weighted scaling is applied to the descriptors at each level after intra-group standardization. For a sample with index i and a numeric descriptor with index j, if its level index is... l Then its weighted descriptor value ,satisfy: in The pre-weighted descriptor value, The z-score is standardized by a transformation computed on the training data; and the hierarchical weights satisfy... , , ; The descriptors corresponding to each level include: L0 level includes the main phase material type, support material type, main phase to support mass ratio, and main phase to support pairing identifier; L1 level includes the electronegativity range, minimum atomic number, and band gap of the main phase and support; L2 level includes the thermal conductivity, shear modulus, Debye temperature, enthalpy change, entropy change, and density of the main phase and support; L3 level includes the melting point range, average boiling point, crystal system, and space group of the main phase and support.
[0010] Based on the above technical solutions, preferably, in step S4, the mixing rule includes an arithmetic weighted mixing rule and a harmonic weighted mixing rule, based on the actual mass ratio of the main phase to the carrier. For weights; for material property descriptors subject to arithmetic weighted mixing rules, the composite descriptor value is... For physical property descriptors applicable to harmonic weighted mixing rules, the composite descriptor value is... The composite descriptor includes at least thermal conductivity, shear modulus, Debye temperature, enthalpy change, entropy change, density, melting point range, average boiling point, electronegativity range, minimum atomic number, and band gap.
[0011] Based on the above technical solutions, preferably, the univariate statistical screening in step S5 takes the extended descriptor set as input, is completed within each fold training set, and includes at least the following rule: the missing rate is greater than the missing threshold. Descriptors are analyzed and near-zero variance descriptors are removed; the absolute value of the Pearson correlation coefficient between any two numerical descriptors is greater than the collinearity threshold. Only one of the descriptors with stronger physical interpretability is retained; for numerical descriptors, the variance inflation factor (VIF) is calculated and descriptors with VIF greater than the collinearity threshold are removed. The descriptors were used; a permutation test was used to obtain the p-value for the correlation between the descriptors and the target variable, and BH-FDR correction was applied. After removing the BH-FDR correction, the q-value was greater than the significance threshold. The descriptors that remain after being screened by the above rules constitute the initial set of descriptors.
[0012] Based on the above technical solutions, preferably, in step S5, the multivariate stable selection takes the pre-screened descriptor set as input, performs bootstrapping in each training set grouped by pair_id, repeats B times to obtain subsamples, and runs on each subsample. The constrained feature selection model records whether a descriptor is selected, thereby calculating the selection probability of each descriptor j. , retain satisfaction The descriptor, where And based on the theoretical bound constraint of the expected number of false positives, a probability threshold is selected such that: ; in To determine the expected number of false positives, denoted as the average number of descriptors selected in each subsample, and p is the number of candidate descriptors that enter the stable selection stage. The descriptors retained after stable selection constitute the set of descriptors after stable selection.
[0013] Based on the above technical solution, preferably, in step S5, the importance consistency verification takes the stable selected descriptor set as input, adopts a gradient boosting decision tree model that supports native category processing, and combines it with grouped cross-validation based on pair_id for training. SHAP importance and grouping permutation importance are calculated at each fold, and the descriptors ranked in the top M by importance in each fold are recorded as the Top-M set of that fold; only those descriptors with an importance not less than a certain percentage threshold are retained. The cross-validation compromise belongs to the Top-M set and its The descriptor, where The importance differences between folds were assessed using a paired Wilcoxon test and BH-FDR correction. The descriptors retained after the above verification are the set of key descriptors.
[0014] Based on the above technical solutions, preferably, in step S5, predicting the thermal imaging temperature of the catalyst sample to be confirmed under Fischer-Tropsch synthesis conditions includes: constructing a lightweight prediction model based on a set of key descriptors, performing Yeo-Johnson transformation on the target variable and then training the regression model, and when the candidate models include CatBoost, random forest regression, Bagging regression and XGBoost regression, the optimal model is selected for predicting the thermal imaging temperature of the catalyst sample to be confirmed, with the difference between the prediction performance of the grouped cross-validation and the coefficient of determination of the training set and the test set not exceeding 0.15 as the model selection criterion.
[0015] This invention also provides a system for constructing and optimizing iron-based supported catalyst descriptors for Fischer-Tropsch synthesis, comprising: The data acquisition module is used to perform high-throughput Fischer-Tropsch synthesis catalytic performance tests on iron-based supported catalysts, acquire thermal imaging temperature data as target variables, and collect candidate descriptors of the host phase and support through code extraction, public databases and literature collection to form the original dataset. The data processing module is used to perform data cleaning and standardization preprocessing on the original dataset, including outlier removal, missing value imputation, character unification and field normalization, as well as standardization of numerical descriptors and encoding conversion of categorical descriptors to obtain a normalized descriptor dataset. The descriptor construction module is used to divide the candidate descriptors in the normalized descriptor dataset into four levels and assign initial weights according to the level to perform weighted constraints, based on the Fischer-Tropsch synthesis reaction mechanism, to obtain a hierarchical weighted descriptor set; then, according to the preset mixing rules, the monomer property descriptors in the hierarchical weighted descriptor set are used to construct composite descriptors according to the actual mass ratio of the main phase to the support, to obtain an extended descriptor set. The descriptor pre-screening module is used to perform K-fold grouping based on the pair_id of the main phase and the carrier, and to perform univariate statistical screening on the extended descriptor set within the training fold of each fold to obtain the descriptor set after initial screening. The descriptor selection module is used to sequentially perform multivariate stable selection and importance consistency verification on the initially screened descriptor set within each training fold. Through false positive control based on selection probability and cross-fold SHAP importance consistency screening, the key descriptor set is obtained. The prediction and validation module is used to build a lightweight prediction model based on a set of key descriptors, predict the thermal imaging temperature of the catalyst sample to be confirmed under Fischer-Tropsch synthesis conditions, and verify the prediction results through high-throughput experiments.
[0016] The present invention has the following advantages over the prior art: (1) This invention, driven by high-throughput thermal imaging temperature data, combined with mechanism-oriented hierarchical construction of descriptors, construction of composite descriptors and multi-stage progressive screening process, systematically solves the problems of immature iron-based supported catalyst descriptors, high candidate descriptor dimensionality and unclear correlation with performance. It can automatically screen out a set of key descriptors that are sensitive to the thermal performance of Fischer-Tropsch synthesis catalysts, have low redundancy and are physically interpretable from multi-source candidate descriptors, and build a lightweight prediction model based on this to predict the thermal imaging temperature of new catalyst formulations. At the same time, through high-throughput experimental closed-loop verification, it realizes the complete process from data acquisition to descriptor optimization to performance prediction, which can shorten the development cycle of catalyst formulations and reduce the resource investment of blind trial and error experiments to a certain extent.
[0017] (2) Guided by the Fischer-Tropsch synthesis reaction mechanism, this invention divides candidate descriptors into multiple levels and assigns differentiated initial weights according to the level. At the same time, it constructs composite descriptors of the main phase and the carrier based on the actual mass ratio, so that the descriptor system can explicitly express the ratio of the main phase and the carrier and their physical property coupling at the coding level. Compared with simply splicing single descriptors, the constructed descriptors are more targeted in representing the "loading relationship", which is conducive to improving the model's ability to distinguish different main phase / carrier combinations.
[0018] (3) In the descriptor screening stage, the present invention adopts a K-fold partitioning based on pair_id, and performs a progressive process in each fold training set, including univariate statistical screening, multivariate stable selection based on the bound of selection probability theory, and cross-fold SHAP importance consistency verification. By combining statistical testing with false positive control, the risk of introducing false relevant descriptors due to random fluctuations under small sample conditions is effectively suppressed. The selected key descriptor set has good stability in multiple folds, and the SHAP value can provide a descriptor-level explanation basis for subsequent experimental targeted optimization. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a flowchart of the method for constructing and optimizing the descriptor of the Fischer-Tropsch synthesis iron-based supported catalyst of the present invention; Figure 2 The SHAP value graph shows the predicted performance of the iron-based supported catalyst for Fischer-Tropsch synthesis used in this invention. Figure 3 This is a schematic diagram of the model evaluation for predicting the performance of the iron-based supported catalyst used in the Fischer-Tropsch synthesis in this invention. Detailed Implementation
[0021] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0022] like Figure 1As shown, this invention provides a method for constructing and optimizing a descriptor for an iron-based supported catalyst used in Fischer-Tropsch synthesis, comprising: S1. Obtain the formulation and structural information of the iron-based supported catalyst sample, and obtain the thermal imaging temperature under Fischer-Tropsch synthesis conditions through high-throughput testing, and determine the thermal imaging temperature as the target variable; at the same time, collect candidate descriptors of the main phase and the support from material databases and literature sources to form an initial dataset. S2. The initial dataset is preprocessed using data cleaning and standardization strategies to obtain a normalized descriptor dataset; S3. On the normalized descriptor dataset, the candidate descriptors are divided into four levels based on the Fischer-Tropsch synthesis reaction mechanism, and initial weights are assigned to the descriptors according to the level for weighted constraints, resulting in a hierarchical weighted descriptor set. S4. Based on the preset mixing rules, construct composite descriptors for the monomer property descriptors in the layered weighted descriptor set according to the actual mass ratio of the main phase to the carrier, and obtain an extended descriptor set. S5. Grouping is performed using the pair_id identifier between the main phase and the support as the grouping criterion. Within each training fold, the extended descriptor set is subjected to progressive screening of univariate statistical screening, multivariate stable selection, and importance consistency verification to obtain the key descriptor set. The key descriptor set is used as the optimized descriptor system for Fischer-Tropsch synthesis iron-based supported catalysts to predict the thermal imaging temperature of the catalyst sample to be confirmed under Fischer-Tropsch synthesis conditions.
[0023] In one embodiment of the present invention, step S1 includes: high-throughput preparation of Fischer-Tropsch synthesis reaction for iron-based catalysts. The catalyst preparation process includes operation steps such as weighing, filling, and fixing to ensure that the catalyst particles have a fixed size distribution. The high-throughput Fischer-Tropsch synthesis catalytic performance testing process is as follows: weighing, filling, fixing, gas purging, gas pressurization, programmed temperature rise, reaction end, cooling and depressurization, and device shutdown.
[0024] Data processing was performed on the catalyst during the reaction process. After the programmed temperature rise, data was recorded and exported according to the experimental time. The exported data was divided into low-temperature and high-temperature ranges using software, and integrated with the data exported from the programmed temperature rise. The required data images were then plotted using Origin. Based on the thermal imaging temperature obtained from the high-throughput experiment, the instantaneous temperature at a certain temperature was selected as the target variable. The higher the value, the better the thermal performance of the catalyst, thus achieving a direct connection between the experimental raw signal and the prediction model.
[0025] Candidate descriptors employ a multi-source aggregation strategy: code extraction utilizes the Jupyter editor, linking to databases via API to generate statistical descriptors for the host phase and support phase based on chemical formulas, including statistical descriptors generated by tools such as Matmine and Magpie, such as electronegativity (MagpieData range Electronegativity) and atomic number (MagpieData minimum Number); public databases include databases such as the Materials Project, used to collect information such as bandgap and crystal system; literature collection is used to obtain thermodynamic information of materials, such as physicochemical property descriptors like entropy (S) and enthalpy (H), ensuring that each data point is traceable. The initial dataset contains approximately 33 descriptors (6 categorical and 27 numerical, with n=348 samples), covering three categories: composition descriptors, structure descriptors, and property descriptors. Composition descriptors include the host phase material type (material_A), the support material type (support_material_C), and the mass ratio of the host phase to the support (Active_Components_Ratio). Structure descriptors include the electronegativity range (MagpieDatarange Electronegativity_A / C), minimum atomic number (MagpieData minimum Number_A / C), crystal system (Crystal System_A / C), and space group (Space Group_A / C). Property descriptors include the band gap (Bandgap_A / C), thermal conductivity (k_A / C), shear modulus (G_A / C), and Debye temperature. _A / C, enthalpy change H_A / C, entropy change S_A / C, density_A / C, melting point range_melting_point_range_A / C, and boiling point mean_A / C.
[0026] In one embodiment of the present invention, step S2, the data cleaning and standardization strategy includes the following processing content.
[0027] For outlier removal, manual experience is used to initially screen whether the experimental results of the samples are too high or contain outliers. The thermal imaging temperature data of the Fischer-Tropsch synthesis reaction of iron-based catalysts obtained from high-throughput experiments are observed. If there are samples that deviate significantly from the normal range, they are removed. Otherwise, they are retained as dataset samples for subsequent use in building catalyst performance prediction models. After the dataset is divided, the IQR / MCD detection method is used to further process outliers in the numerical descriptors within the training layer. The test layer is not involved in this processing to prevent data leakage.
[0028] Regarding missing value handling, for a missing value x in a numeric descriptor, let the set of missing locations be... The non-missing set is Use the median to fill in the gaps (for greater robustness): Alternatively, when the data is approximately normal, mean imputation can be used: For missing values of categorical descriptors, they are treated as a separate category: .
[0029] In terms of character normalization, character unification is essentially a rule mapping function. This involves removing invisible characters (such as \u00A0) and redundant spaces from column names, standardizing naming rules (replacing spaces with underscores), and uniformly handling special characters (such as...). , (etc.), standardize the writing style of materials (such as...) / / (e.g., normalize variants first); define the set of missing strings. ("NA", "--", "None", "Untested", etc.): It provides uniform symbol processing for any cell string, including removing thousands separator commas and replacing "−" with "-".
[0030] In terms of standardization and encoding, z-score standardization is used to make numerical descriptors dimensionless. For rare categories with a frequency of less than 5 in categorical descriptors, they are merged into independent rare class labels (__RARE__ labels), and then converted into numerical form using a Bayesian smoothing target encoding method based on grouped K-folds. Fitting is performed within each group fold, and only transformation is performed on the validation and test folds to prevent target leakage. For category columns such as material_A, support_material_C, CrystalSystem, and Space Group, the numerical codes of the categories should not be treated as continuous values. For CatBoost models, the cat_features column name can be directly filled in to enable native processing of category descriptors. For models such as XGBoost that cannot directly recognize category descriptors, they need to be converted to pandas category format and the training set and test set categories must be aligned.
[0031] In one embodiment of the present invention, step S3 specifically includes: according to the Fischer-Tropsch synthesis reaction mechanism, dividing the candidate descriptors in the normalized descriptor dataset into four levels from bottom to top, namely L0 formulation and identity layer, L1 electronic scale and bonding layer, L2 lattice mechanics and thermal transport layer, and L3 phase stability and phase transition indicator layer, and constructing a layered weight vector. Weighted scaling of descriptors at each layer after intra-group standardization is achieved by establishing feature group weights at the front end of the model and using group-wise scaling (multiplying the z-score of each group by the corresponding value). This implements "soft constraints." For sample indices... The sample and descriptor indexes are The numeric descriptor, if its hierarchical index is The weighted descriptor value satisfy: in The pre-weighted descriptor value, The z-score is standardized on the training data; the hierarchical weights satisfy... , , .
[0032] The L0 formulation and identity layer correspond to the fields material_A, support_material_C, Active_Components_Ratio, and pair_id (main phase / support disordered pair identifier). This layer captures the combination relationship between the main phase and the support. The formulation identity determines the catalytic surface and interface properties. The categorical effect of the A / C combination (support polarity, defects, basicity, thermal conductivity) has the highest priority for supported systems. pair_id serves as a single-class feature input to capture A×C interactions. Rare categories are merged into __RARE__ and then processed by intra-group K-fold target encoding / Bayesian smoothing to prevent leakage. Initial weight suggestions are provided. .
[0033] The L1 electronic scale and bonding layer corresponding fields are MagpieData range Electronegativity_A / C, MagpieData minimum Number_A / C, and Band gap_A / C. Electronegativity difference and band gap affect electron transfer and surface charge distribution, thereby regulating the adsorption / activation and selectivity of CO, H, and O species; substitution tests and FDR correction are performed within grouped cross-validation; and highly collinear terms with L2 / L3 ( (VIF>10) Select those with stronger physical interpretability; initial weighting suggestion If the SHAP indicator is strongly correlated with the target, it will be increased to 1.0.
[0034] The corresponding fields for L2 lattice mechanics and thermal transport layers are k_A / C and G_A / C. _A / C, H_A / C, S_A / C, and density_A / C. Thermal conductivity. shear modulus Debye temperature Harmony The reaction zone temperature gradient and phonon scattering are jointly determined, while density affects the pore structure and effective heat capacity. Given the critical role of the strongly exothermic nature of Fischer-Tropsch synthesis in thermal management, the initial weights are appropriately increased, and it is recommended that... .
[0035] The L3 phase stability and phase transition indicator fields are melting_point_range_A / C, boiling_point_mean_A / C, Crystal System_A / C, and Space Group_A / C. Melting / boiling point statistics and crystallographic parameters indicate the phase stability window, phase transition tendency, and the reversibility of redox cycles, indirectly indicating resistance to sintering and carbon deposition. The crystal system / space group retains its original category, using either CatBoost native processing or target encoding. Initial weight suggestions are provided. The entire process uses GroupKFold(pair_id), with early stopping during training and test folds used only for final evaluation to avoid A / C leakage.
[0036] Implementation and stability assurance of hierarchical weights. Hierarchical prior weighting: Feature group weights α=[w0,w1,w2,w3] are established at the front end of the model, and "soft constraints" are achieved through group-wise scaling (e.g., multiplying each group's z-score by w_i). Threshold criteria: Collinearity: |ρ|>0.90, VIF>10, only one is guaranteed; Significance: Permutation test + BH-FDR q<0.10; Confidence learning: Rare items (frequency <5) are merged into __RARE__ to avoid spurious correlations. Group validation: GroupKFold(pair_id) is used throughout the process, with early stopping during training; testing is only used for final evaluation to avoid A / C leakage. This system links "mechanistic hypothesis → hierarchical prior → stable selection → interpretability".
[0037] In one embodiment of the present invention, in step S4, a composite descriptor is constructed from the monomer property descriptors using a mixing rule based on mass ratio. First, notation and preprocessing are performed, and all monomer properties are represented by the suffixes _A and _C (such as k_A, k_C, etc.).
[0038] The fundamental theory of complex mixing rules comes from the general formula of weighted average. ,in and These are the weighting coefficients for the main phase and the carrier, respectively. In practical applications, they are represented by the actual mass ratio of the main phase to the carrier. For the weights, where Main phase quality, To improve carrier quality, the general formula is specialized into two specific hybrid rules.
[0039] For materialized property descriptors to which arithmetic weighted mixing rules apply, when , At that time, the composite descriptor value is: ; For physicochemical property descriptors applicable to the harmonic weighted mixing rule, based on the inverse weighting principle of material conductivity, the composite descriptor value is: ; in The main corresponding monomer physical property value, These represent the physical properties of the carrier corresponding to the monomer. The arithmetic weighted mixing rule applies to extensive and average properties such as density, entropy change, and enthalpy change, while the harmonic weighted mixing rule applies to conductive properties such as thermal conductivity and shear modulus.
[0040] In practice, all monomer properties are represented by the suffixes _A and _C (e.g., k_A, k_C, etc.). For each pair of "main phase A / carrier C", 11 composite physicochemical property descriptors are calculated according to the above mixing rules. The numerical characteristics are converted according to the rules, such as merging k_A and k_C into k, to form a composite descriptor. The composite descriptor includes at least thermal conductivity k, shear modulus G, and Debye temperature. Enthalpy change H, entropy change S, density, melting point range, boiling point mean, electronegativity range, minimum atomic number, and band gap.
[0041] All constructs are implemented solely within the training layer based on individual table lookups to prevent data leakage; missing values are filled with stratification constants or group medians, and missing value indicators are set; thermal conductivity k, shear modulus G, and Debye temperature are related to the strong exothermic characteristics of FTS. Entropy change S and density These are given high priority. After constructing the composite descriptor as described above, the final example feature list is as follows: material_A, support_material_C, Active_Components_Ratio, MagpieData rangeElectronegativity, Band gap, MagpieData minimum Number, Crystal System_A, CrystalSystem_C, melting_point_range, boiling_point_mean, density, 、k、G、H、S.
[0042] In one embodiment of the present invention, in step S5, given the unstable nature of feature selection in scenarios with small samples (n≈348) and medium dimensions (approximately 33 descriptors), a phased, repeatable, and statistically guaranteed progressive screening process is adopted. Specifically, it includes three stages: univariate statistical screening, multivariate stable selection, and importance consistency verification. The pair_id is used as the grouping basis for K-fold partitioning. The entire process uses GroupKFold(pair_id) to prevent the same formulation variant from entering both the training and testing folds simultaneously, which could lead to optimistic bias.
[0043] Univariate statistical screening takes an expanded descriptor set as input and is performed within each fold of the training set, following the rule: the missing data rate is greater than the missing data threshold. Descriptors are descriptors that have near-zero variance, and descriptors with near-zero variance are typically set to... For any two numerical descriptors, calculate the Pearson correlation coefficient: in For a certain descriptor's first Each sample value For the target value, , These are the corresponding means. When the absolute value of the Pearson correlation coefficient between any two numerical descriptors is greater than the collinearity threshold... At that time, only one descriptor with stronger physical interpretability is retained, usually set to Calculate the variance inflation factor (VIF) for the numerical descriptors, and remove descriptors with VIF values greater than the collinearity threshold. The descriptor is usually set The correlation between the descriptors and the target variable was tested using a permutation test to obtain the p-value, followed by BH-FDR correction. The effect size was also reported (for numerical descriptors). Category-type descriptors are used for targets / After removing BH-FDR correction, the q-value is greater than the significance threshold. The descriptor is usually set The descriptors retained after screening according to the above rules constitute the initial set of descriptors.
[0044] Multivariate stable selection takes the initial set of descriptors as input, performs bootstrapping by pair_id within each fold of the training set, and repeats the process. The subsamples are obtained, and the Lasso and mRMR→EN branches of the Elastic Net / L1 constraints are run separately, recording the results for each descriptor. Calculate the probability of selection to determine whether a candidate is selected. , retain satisfaction The descriptor, where And based on the theoretical bound constraint of the expected number of false positives, a probability threshold is selected such that: in To determine the expected number of false positives, This represents the average number of descriptors selected in each subsample. The total number of candidate descriptors that have entered the stable selection phase; based on this, take... To keep the expected false positive rate below 1, and with Descriptors are determined collaboratively through nested cross-validation. The descriptors retained after stable selection constitute the stable selection descriptor set.
[0045] Importance consistency verification takes a stable set of post-selected descriptors as input and trains a CatBoost gradient boosting decision tree model that supports native class processing, combined with GroupKFold cross-validation based on pair_id. At each fold, SHAP importance and group permutation importance are calculated separately; group permutation importance is defined as: in For only the first The data is arranged with random permutations, and the Score can be taken as follows: Or RMSE (note that the smaller the RMSE, the better); sort the importance of each compromise and put it in the top. The descriptor of the name is denoted as this fold. The set is retained only if it is not less than the proportion threshold. The cross-validation trade-off belongs to set and its The descriptor, where The differences in importance between folds were assessed using a paired Wilcoxon test and corrected for by the BH-FDR. The descriptors retained after the above verification are the key descriptor set.
[0046] Combining the three stages with physical interpretability, a robust subset of descriptors of approximately 12-18 is ultimately formed, with a default threshold of missing rate <35%. VIF<10 SHAP consistency After FDR correction In a specific embodiment of the present invention, after the above progressive filtering, 15 descriptors are ultimately retained, as shown in Table 1.
[0047] Table 1. Descriptors describing catalyst performance This scheme controls spurious correlations and overfitting by combining stable selection with statistical tests, and uses pair-level encoding to improve the expressive power of load relationships, obtaining a statistically significant and physically interpretable subset of key descriptors with small sample sizes.
[0048] In one embodiment of the present invention, step S5, predicting the thermal imaging temperature of the catalyst sample to be confirmed under Fischer-Tropsch synthesis conditions, includes: constructing a lightweight prediction model based on a set of key descriptors, performing a Yeo-Johnson transformation on the target variable to improve its distribution skewness, using PowerTransformer(method='yeo-johnson'), with the core formula being: in The original target value, The target value after transformation. The transformation parameters are automatically determined by maximum likelihood estimation.
[0049] The candidate models include four types: CatBoost, RandomForestRegressor, Bagging Regression, and XGBoost Regression. Predictive performance was evaluated using grouped cross-validation and the coefficients of determination on the training and test sets. A difference of no more than 0.15 is used as the model selection criterion; the optimal model is selected for predicting the thermal imaging temperature of the catalyst sample to be confirmed. The core principles of the four models are as follows: CatBoost is essentially a gradient boosting tree, and its core principle is to add a tree in each round to minimize the loss: ,in For the first The overall model of the wheel For the first A new tree, For learning rate, The loss function is RMSE / MAE (RMSE / MAE are commonly used in regression); random forest regression takes the average of predictions from multiple trees: ,in For the first The output of each decision tree The number of trees; Bagging performs bootstrap sampling from the training set: Each base learner is trained on a resampled dataset, and the output is averaged during regression; XGBoost incorporates the loss and model complexity regularization into the objective function: ,in For data loss (such as squared error). For regularization terms (penalty leaf count) Leaf weight ), , For regularization coefficients. CatBoost is friendly to category descriptors and can directly recognize them without encoding; other models cannot directly recognize category descriptors and need to encode category descriptors such as Crystal_System accordingly. When using models such as XGBoost, it is necessary to convert them to pandas category format and ensure that the training set and test set categories are aligned.
[0050] The model evaluation index uses the coefficient of determination. The calculation formula is: in For the first The true value of each sample For predicted values, The mean of the true values. The sample size is the numerator, the residual sum of squares (RSS) is the denominator, and the total sum of squares (TSS) is the denominator; the test set is the sample size. The closer the value is to 1, the better the model's prediction performance. The training set and test set are used as examples. An error difference of no more than 0.15 controls overfitting.
[0051] In a specific embodiment of the present invention, after predicting the thermal imaging temperature of the catalyst sample to be confirmed based on the key descriptor set and the optimal prediction model, the contribution of each descriptor to the model prediction is analyzed using the SHAP value, such as... Figure 2 As shown, the contribution of different descriptors to the model is analyzed based on the SHAP value plot, and the final descriptors are selected. After completing the descriptor engineering, 15 descriptors are finally obtained. The model is trained using this dataset to obtain the test set for the model. Model evaluation. Based on the training results, CatBoost performed best, and the training set... The value is 0.83, in the test set. The value is 0.71, across the four models on both the training and test sets. Since all values are the highest, CatBoost is chosen as the optimal model for predicting catalyst performance. Figure 3 As shown, the model, trained to predict catalyst performance with different ratios, verifies the predictions through high-throughput experiments. The predictions are compared with actual thermal imaging temperature data obtained from the high-throughput experiments to ensure the verifiability and transferability of the descriptor selection results. A closed-loop mutual verification mechanism is established between high-throughput experimental data and the prediction results of the new catalyst model, thus supporting the forward-looking prediction of novel catalyst performance and enabling dynamic iterative updates of the database and performance prediction model. This solution has automated implementation capabilities, significantly shortening the R&D cycle and reducing experimental and resource investment costs.
[0052] This invention also provides a system for constructing and optimizing iron-based supported catalyst descriptors for Fischer-Tropsch synthesis, the system comprising: The data acquisition module is used to perform high-throughput Fischer-Tropsch synthesis catalytic performance tests on iron-based supported catalysts, acquire thermal imaging temperature data as target variables, and collect candidate descriptors of the host phase and support through code extraction, public databases and literature collection to form the original dataset. The data processing module performs data cleaning and standardization preprocessing on the original dataset, including outlier removal, missing value imputation, character unification and field normalization, as well as standardization of numerical descriptors and encoding conversion of categorical descriptors to obtain a normalized descriptor dataset. The data acquisition module includes a high-throughput experimental interface submodule, a code extraction submodule, and an external database and literature entry submodule. The high-throughput experimental interface submodule connects to the data acquisition card of the high-throughput Fischer-Tropsch synthesis apparatus, automatically receiving raw test data such as temperature signals and reaction timestamps from each reaction channel during the programmed heating process. Within the system, data from different channels and temperature zones are associated and stored according to experiment number and catalyst formulation number. Built-in data processing scripts organize the raw thermal imaging signals into thermal imaging temperature data suitable for modeling, and the instantaneous temperature at selected time points or temperature zones is written into the database as the target variable. The code extraction submodule is configured with a script interface based on the Jupyter environment. The system has pre-built APIs for interfacing with materials informatics tools such as Matminer and Magpie. After the user inputs the chemical formulas of the host phase and support in the interface, the code extraction submodule automatically calls external materials databases to batch generate statistical descriptors related to the host phase and support and writes them into the raw data table. The external database and literature entry submodule connects to public databases such as the Materials Project via a network interface, extracting structural and band parameters such as band gap, crystal system, and space group by material name or chemical formula. It also provides a manual input interface for entering thermodynamic and physical property data such as thermal conductivity, Debye temperature, entropy, and enthalpy obtained from literature into the system. These three submodules work together to unify the target variables obtained from high-throughput experiments and the multi-source pooled host phase and support candidate descriptors into a raw dataset stored in a relational database.
[0053] The descriptor construction module is used to divide the candidate descriptors in the normalized descriptor dataset into four levels and assign initial weights according to the level to perform weighted constraints, based on the Fischer-Tropsch synthesis reaction mechanism, to obtain a hierarchical weighted descriptor set; then, according to the preset mixing rules, the monomer property descriptors in the hierarchical weighted descriptor set are used to construct composite descriptors according to the actual mass ratio of the main phase to the support, to obtain an extended descriptor set. The descriptor pre-screening module is used to group the descriptor set into K-fold partitions based on the pair_id of the main phase and the carrier. Within the training fold of each fold, the extended descriptor set is subjected to univariate statistical screening to obtain the initial screening descriptor set. The descriptor selection module is used to perform multivariate stable selection and importance consistency verification on the initial screening descriptor set within the training fold of each fold. Through false positive control based on selection probability and cross-fold SHAP importance consistency screening, the key descriptor set is obtained. The prediction and validation module is used to build a lightweight prediction model based on a set of key descriptors, predict the thermal imaging temperature of the catalyst sample to be confirmed under Fischer-Tropsch synthesis conditions, and verify the prediction results through high-throughput experiments.
[0054] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for constructing and optimizing a descriptor for an iron-based supported catalyst in Fischer-Tropsch synthesis, characterized in that, include: S1. Obtain the formulation and structural information of the iron-based supported catalyst sample, and obtain the thermal imaging temperature under Fischer-Tropsch synthesis conditions through high-throughput testing, and determine the thermal imaging temperature as the target variable; at the same time, collect candidate descriptors of the main phase and the support from material databases and literature sources to form an initial dataset. S2. The initial dataset is preprocessed using data cleaning and standardization strategies to obtain a normalized descriptor dataset; S3. On the normalized descriptor dataset, the candidate descriptors are divided into four levels based on the Fischer-Tropsch synthesis reaction mechanism, and initial weights are assigned to the descriptors according to the level for weighted constraints, resulting in a hierarchical weighted descriptor set. S4. Based on the preset mixing rules, construct composite descriptors for the monomer property descriptors in the layered weighted descriptor set according to the actual mass ratio of the main phase to the carrier, and obtain an extended descriptor set. S5. Grouping is performed using the pair_id identifier between the main phase and the support as the grouping criterion. Within each training fold, the extended descriptor set is subjected to progressive screening of univariate statistical screening, multivariate stable selection, and importance consistency verification to obtain the key descriptor set. The key descriptor set is used as the optimized descriptor system for Fischer-Tropsch synthesis iron-based supported catalysts to predict the thermal imaging temperature of the catalyst sample to be confirmed under Fischer-Tropsch synthesis conditions.
2. The method for constructing and optimizing a Fischer-Tropsch synthesis iron-based supported catalyst descriptor as described in claim 1, characterized in that, The candidate descriptors include composition descriptors, structure descriptors, and property descriptors; the composition descriptors include the type of main phase material, the type of support material, and the mass ratio of the main phase to the support; the structure descriptors include the electronegativity range, minimum atomic number, crystal system, and space group of the main phase and the support; the property descriptors include the band gap, thermal conductivity, shear modulus, Debye temperature, enthalpy change, entropy change, density, melting point range, and average boiling point of the main phase and the support.
3. The method for constructing and optimizing a Fischer-Tropsch synthesis iron-based supported catalyst descriptor as described in claim 1, characterized in that, In step S2, the data cleaning and standardization strategies include: removing outliers from high-throughput thermal imaging temperature data and eliminating samples that significantly deviate from the normal range; imputing missing values for numerical descriptors using the median or mean, and classifying missing values for categorical descriptors into independent categories; standardizing the format of descriptor field names, removing invisible characters and redundant spaces, and unifying the naming convention; using the z-score standardization method to make numerical descriptors dimensionless; and merging rare categories with a frequency of less than 5 in categorical descriptors into independent rare class labels, then converting them into numerical form using a Bayesian smoothing target encoding method based on grouped K-folds.
4. The method for constructing and optimizing a Fischer-Tropsch synthesis iron-based supported catalyst descriptor as described in claim 1, characterized in that, Step S3 specifically includes: Based on the Fischer-Tropsch synthesis reaction mechanism, the candidate descriptors in the normalized descriptor dataset are divided into four levels from bottom to top: L0 formulation and identity layer, L1 electronic scale and bonding layer, L2 lattice mechanics and thermal transport layer, and L3 phase stability and phase transition indicator layer. A hierarchical weight vector is then constructed. Weighted scaling is applied to the descriptors at each level after intra-group standardization. For a sample with index i and a numeric descriptor with index j, if its level index is... l Then its weighted descriptor value ,satisfy: in The pre-weighted descriptor value, The z-score is standardized by a transformation computed on the training data; and the hierarchical weights satisfy... , , ; The descriptors corresponding to each level include: L0 level includes the main phase material type, support material type, main phase to support mass ratio, and main phase to support pairing identifier; L1 level includes the electronegativity range, minimum atomic number, and band gap of the main phase and support; L2 level includes the thermal conductivity, shear modulus, Debye temperature, enthalpy change, entropy change, and density of the main phase and support; L3 level includes the melting point range, average boiling point, crystal system, and space group of the main phase and support.
5. The method for constructing and optimizing a Fischer-Tropsch synthesis iron-based supported catalyst descriptor as described in claim 4, characterized in that, In step S4, the mixing rules include arithmetic weighted mixing rules and harmonic weighted mixing rules, based on the actual mass ratio of the main phase to the carrier. The weight is used for the material property descriptor that applies the arithmetic weighted mixing rule; the composite descriptor value is... For physical property descriptors applicable to harmonic weighted mixing rules, the composite descriptor value is... The composite descriptor includes at least thermal conductivity, shear modulus, Debye temperature, enthalpy change, entropy change, density, melting point range, average boiling point, electronegativity range, minimum atomic number, and band gap.
6. The method for constructing and optimizing a Fischer-Tropsch synthesis iron-based supported catalyst descriptor as described in claim 4, characterized in that, In step S5, the univariate statistical screening takes the extended descriptor set as input and is completed within each fold of the training set, and includes at least the following rule: the missing rate is greater than the missing threshold. Descriptors are analyzed and near-zero variance descriptors are removed; the absolute value of the Pearson correlation coefficient between any two numerical descriptors is greater than the collinearity threshold. Only one of the descriptors with stronger physical interpretability is retained at that time; Calculate the variance inflation factor (VIF) for the numerical descriptors and remove those with VIF greater than the collinearity threshold. The descriptor; The correlation between the descriptor and the target variable was tested using a permutation test to obtain the p-value, followed by BH-FDR correction. After removing the BH-FDR correction, the q-value was greater than the significance threshold. The descriptors that remain after being screened by the above rules constitute the initial set of descriptors.
7. The method for constructing and optimizing a Fischer-Tropsch synthesis iron-based supported catalyst descriptor as described in claim 4, characterized in that, In step S5, multivariate stable selection uses the pre-screened descriptor set as input, performs bootstrapping by pair_id in each fold of the training set, repeats B times to obtain subsamples, and runs on each subsample. The constrained feature selection model records whether a descriptor is selected, thereby calculating the selection probability of each descriptor j. , retain satisfaction The descriptor, where And based on the theoretical bound constraint of the expected number of false positives, a probability threshold is selected such that: ; in To determine the expected number of false positives, denoted as the average number of descriptors selected in each subsample, and p is the number of candidate descriptors that enter the stable selection stage. The descriptors retained after stable selection constitute the set of descriptors after stable selection.
8. The method for constructing and optimizing a Fischer-Tropsch synthesis iron-based supported catalyst descriptor as described in claim 4, characterized in that, In step S5, the importance consistency verification takes the stable selected descriptor set as input, employs a gradient boosting decision tree model that supports native class processing, and combines it with pair cross-validation based on pair_id for training. At each fold, SHAP importance and group permutation importance are calculated, and the top M descriptors in importance ranking at each fold are recorded as the Top-M set for that fold; only those descriptors with an importance not less than a certain percentage threshold are retained. The cross-validation compromise belongs to the Top-M set and its The descriptor, where The importance differences between folds were assessed using a paired Wilcoxon test and BH-FDR correction. The descriptors retained after the above verification are the set of key descriptors.
9. The method for constructing and optimizing a Fischer-Tropsch synthesis iron-based supported catalyst descriptor as described in claim 1, characterized in that, In step S5, predicting the thermal imaging temperature of the catalyst sample to be confirmed under Fischer-Tropsch synthesis conditions includes: constructing a lightweight prediction model based on a set of key descriptors, performing Yeo-Johnson transformation on the target variable and then training the regression model, and selecting the optimal model for predicting the thermal imaging temperature of the catalyst sample to be confirmed when the candidate models include CatBoost, Random Forest Regression, Bagging Regression and XGBoost Regression, with the difference between the prediction performance of the grouped cross-validation and the coefficient of determination of the training set and the test set not exceeding 0.15 as the model selection criterion.
10. A system for constructing and optimizing descriptors for iron-based supported catalysts in Fischer-Tropsch synthesis, characterized in that, include: The data acquisition module is used to perform high-throughput Fischer-Tropsch synthesis catalytic performance testing on iron-based supported catalysts, acquire thermal imaging temperature data as target variables, and collect candidate descriptors of the host phase and support through code extraction, public databases and literature collection to form the original dataset. The data processing module is used to perform data cleaning and standardization preprocessing on the original dataset, including outlier removal, missing value imputation, character unification and field normalization, as well as standardization of numerical descriptors and encoding conversion of categorical descriptors to obtain a normalized descriptor dataset. The descriptor construction module is used to divide the candidate descriptors in the normalized descriptor dataset into four levels and assign initial weights according to the level to perform weighted constraints, based on the Fischer-Tropsch synthesis reaction mechanism, to obtain a hierarchical weighted descriptor set; then, according to the preset mixing rules, the monomer property descriptors in the hierarchical weighted descriptor set are used to construct composite descriptors according to the actual mass ratio of the main phase to the support, to obtain an extended descriptor set. The descriptor pre-screening module is used to perform K-fold grouping based on the pair_id of the main phase and the carrier, and to perform univariate statistical screening on the extended descriptor set within the training fold of each fold to obtain the descriptor set after initial screening. The descriptor selection module is used to sequentially perform multivariate stable selection and importance consistency verification on the descriptor set after initial screening within each training fold. Through false positive control based on selection probability and cross-fold SHAP importance consistency screening, the key descriptor set is obtained. The prediction and validation module is used to build a lightweight prediction model based on a set of key descriptors, predict the thermal imaging temperature of the catalyst sample to be confirmed under Fischer-Tropsch synthesis conditions, and verify the prediction results through high-throughput experiments.
Citation Information
Patent Citations
Catalytic material screening method and catalytic material screening system based on high-throughput experiment and high-throughput calculation
CN111128311A