Fischer-Tropsch synthesis catalyst performance prediction method and system based on high-throughput data
By combining infrared thermal imaging and machine learning, a performance prediction model for Fischer-Tropsch synthesis catalysts was constructed, which solved the problems of long screening cycles and low accuracy in Fischer-Tropsch synthesis catalysts, and achieved rapid and efficient prediction and screening of catalyst performance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-26
- Publication Date
- 2026-04-03
AI Technical Summary
Existing technologies suffer from problems such as long screening cycles, low screening efficiency, and insufficient accuracy of machine learning models in predicting high-performance catalysts.
Temperature field images of the catalyst reaction bed were acquired using infrared thermal imaging technology to obtain structural feature data and thermodynamic property data. A high-throughput database was constructed, and a catalyst performance prediction model was established by combining machine learning models with cross-validation and multi-objective Bayesian optimization strategies. The model was then optimized through a closed-loop mechanism of prediction-experiment-reflux.
It enables rapid prediction and efficient screening of catalyst performance, improves prediction accuracy and model generalization ability, and ensures accurate identification of high-performance catalysts.
Smart Images

Figure CN121789826A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of catalyst performance prediction technology, and in particular to a method and system for predicting the performance of Fischer-Tropsch synthesis catalysts based on high-throughput data. Background Technology
[0002] Fischer-Tropsch synthesis, a core technology in coal and natural gas chemical industries, can convert syngas into liquid hydrocarbon fuels and high-value chemicals, which is of great significance for the efficient utilization of carbon resources. Traditional Fischer-Tropsch synthesis catalyst development mainly relies on researchers' empirical judgment of the reaction mechanism, combined with point-by-point experiments on a small number of formulations to optimize catalyst composition and structure. This process has low experimental throughput, long cycles, and is difficult to fully cover the multi-dimensional combination space of support type, active component ratio, and auxiliary agent type, easily missing potential high-performance formulations. As Fischer-Tropsch synthesis reaction systems become increasingly complex, catalytic activity is often coupled with multiple factors such as the electronic structure, crystal phase, metal-support interaction, and heat and mass transfer characteristics of materials. Large-scale screening cannot be completed within a reasonable timeframe through manual design alone.
[0003] High-throughput experimental techniques enable the simultaneous or rapid sequential testing of a large number of candidate catalysts on automated reaction platforms, allowing for systematic perturbation of preparation parameters, component ratios, and processing techniques, and obtaining structure-performance data. Simultaneously, machine learning-based data-driven methods can pre-screen and predict properties across a vast material space, providing candidate sets and feature descriptions for experiments. However, in existing technologies, the experimental and computational ends are often separate, resulting in inconsistent data formats and untimely feedback. This makes it difficult for predictive models to provide closed-loop guidance for experimental optimization, and experimental data is insufficient to effectively enrich the model training set, thus limiting screening efficiency. Therefore, there is an urgent need for a catalyst screening method that integrates high-throughput experiments with data-driven modeling to achieve rapid generation of candidate catalysts, performance prediction, batch preparation and automated testing, and closed-loop data iteration.
[0004] Chinese patent application CN111177915A discloses a high-throughput calculation method and system for catalytic materials. This patent proposes a technical solution based on machine learning algorithms to construct a catalyst structure-activity relationship model. By combining crystal structure databases, first-principles calculations, and microkinetic analysis, a correlation model is established from the structural features of the catalyst surface model to adsorption energy and then to theoretical catalytic performance. This model is used to screen candidate materials that meet target catalytic performance. When the deviation between the predicted and calculated results exceeds a preset range, new data is fed back to the training set to update the model. This technical solution obtains parameters such as the catalyst's adsorption energy and reaction barrier through theoretical calculations as the model's input and output, achieving high-throughput screening of catalytic materials based on computational data. However, this method relies on first-principles calculations to evaluate the performance parameters of each candidate catalyst, resulting in a complex and time-consuming calculation process that limits the efficiency of screening large-scale candidate materials. Furthermore, there is a discrepancy between the theoretical calculation results and actual catalytic performance, making it difficult to accurately predict catalyst performance. Summary of the Invention
[0005] In view of this, the present invention proposes a method and system for predicting the performance of Fischer-Tropsch synthesis catalysts based on high-throughput data, which solves the problems of long screening cycle, low screening efficiency and insufficient prediction accuracy of machine learning models for high-performance catalysts in the prior art.
[0006] The technical solution of this invention is implemented as follows: This invention provides a method for predicting the performance of Fischer-Tropsch synthesis catalysts based on high-throughput data, comprising the following steps: S1. High-throughput evaluation tests were conducted on various catalysts under Fischer-Tropsch synthesis reaction conditions. Temperature deviations were obtained by acquiring temperature field images of the reaction bed of each catalyst through infrared thermal imaging. At the same time, structural characteristic data and thermodynamic property data of each catalyst were acquired to construct a high-throughput database of Fischer-Tropsch synthesis catalysts. S2. Based on the single-phase, composite, and supported structural characteristics of Fischer-Tropsch synthesis catalysts, the structural characteristic data and thermodynamic property data are normalized and encoded to construct a feature vector that can simultaneously characterize different structural characteristics of Fischer-Tropsch synthesis catalysts. S3. The feature vectors and corresponding temperature deviations are used to form a dataset and divided into a training set and a test set. The training set is input into the machine learning model, and cross-validation and multi-objective Bayesian optimization strategies are used for training and optimization to establish a catalyst performance prediction model. S4. Use the test set to verify the accuracy of the catalyst performance prediction model. If the accuracy does not meet the standard, return to step S3 to adjust the model parameters and retrain. S5. Input the feature vector of the catalyst to be screened into the catalyst performance prediction model with the required accuracy, output the predicted temperature deviation, and evaluate and select the candidate catalyst based on the predicted temperature deviation.
[0007] Based on the above technical solutions, preferably, in step S1, the structural feature data includes the active components of the Fischer-Tropsch synthesis catalyst, the type identification, mass fraction, volume fraction and crystal structure parameters of the promoter and support, and the thermodynamic property data includes thermal conductivity, specific heat capacity and density.
[0008] Based on the above technical solutions, preferably, the temperature deviation obtained in step S1 by acquiring temperature field images of each catalyst reaction bed through infrared thermal imaging specifically includes: High-throughput thermal imaging temperature sequences were continuously acquired during the Fischer-Tropsch synthesis reaction, and the highest temperature T of the reaction bed was extracted from the temperature field images. max Calculate the temperature deviation ΔT=T max -T set T set A temperature is set for the reaction, and the temperature deviation is used as a unified performance characterization metric to characterize the exothermic properties of the Fischer-Tropsch synthesis catalyst.
[0009] Based on the above technical solutions, preferably, step S2, which involves normalizing and encoding the feature vectors of Fischer-Tropsch synthesis catalysts with different structural characteristics, specifically includes: The active components, supports, and promoters of single-phase catalysts, composite catalysts, and supported catalysts are all digitally coded. The mass fraction, volume fraction, and crystal structure parameters are normalized. For composite catalysts and supported catalysts, the equivalent thermodynamic property parameters are calculated using a weighted average method based on the mass fraction and volume fraction of each component. The digital type code, the normalized mass fraction, volume fraction, crystal structure parameters, and thermodynamic property parameters or equivalent thermodynamic property parameters are combined to form a feature vector.
[0010] Based on the above technical solutions, preferably, step S3 specifically includes: S31. Combine the feature vectors with the corresponding temperature deviations to form a complete dataset, and divide it into a training set and a test set according to a preset ratio. S32. Input the training set into the machine learning model for training. The machine learning model consists of a temporal neural network module and a gradient boosting decision tree module. The temporal neural network module is used to process the temporal thermal signal features of the temperature field image, and the gradient boosting decision tree module is used to process the structural feature data and thermodynamic property data of the catalyst. After fusing the outputs of the two modules, the relationship between catalyst structure, formulation and exothermic performance is established. S33. The model is trained and optimized using cross-validation and multi-objective Bayesian optimization strategies to obtain the catalyst performance prediction model.
[0011] Based on the above technical solutions, preferably, in step S31, after forming a complete dataset, the method further includes: using a temperature-weighted feature importance evaluation method to filter features from the feature vector. The temperature-weighted feature importance evaluation method includes: introducing a sample weight based on its temperature deviation for each sample to enhance the contribution of high-temperature deviation samples to feature importance; calculating the temperature-weighted global SHAP value as a feature importance index; and using the temperature-weighted variance inflation coefficient method to remove collinear features. When the temperature-weighted VIF of a feature is greater than a preset threshold and the corresponding temperature-weighted SHAP value is low, the feature is preferentially removed. After obtaining a simplified feature vector, the training set and test set are then divided.
[0012] Based on the above technical solutions, the preferred multi-objective Bayesian optimization strategy in step S33 specifically includes: defining a multi-objective loss function combining the coefficient of determination, root mean square error, and mean absolute error of the high-temperature deviation sample subset as the optimization objective; approximating the objective function using a Gaussian process; selecting a Matérn5 / 2 kernel function with automatic correlation determination; setting different base length scales for each hyperparameter dimension according to the sensitivity of different hyperparameters to the model on the Fischer-Tropsch system; and introducing a temperature deviation adaptive factor to dynamically adjust the length scale of the kernel function, so that the Bayesian optimization process maintains a fine search in the high-temperature deviation region while quickly skipping over the low-temperature deviation region, and iteratively searching to obtain the optimal hyperparameter combination that minimizes the multi-objective loss function.
[0013] Based on the above technical solutions, preferably, step S4, which uses a test set to verify the accuracy of the catalyst performance prediction model, specifically includes: inputting the feature vector of the catalyst in the test set into the catalyst performance prediction model to obtain the predicted temperature deviation; comparing the predicted temperature deviation with the corresponding actual temperature deviation in the test set; calculating the coefficient of determination, root mean square error, and mean absolute error as accuracy evaluation indicators; determining whether the accuracy evaluation indicators meet the preset accuracy standard; if they do, the model accuracy is considered to be up to standard; if they do not, returning to step S3 to adjust the model parameters or feature vector and retraining.
[0014] Based on the above technical solutions, preferably, the prediction method further includes: The candidate catalysts predicted by the catalyst performance prediction model in step S5 are verified by high-throughput preparation and Fischer-Tropsch synthesis experiments. The actual measured temperature deviation is obtained and compared with the predicted temperature deviation. When the deviation is within the preset range, the catalyst is confirmed to have reached the target performance. When the deviation exceeds the preset range, the structural feature data, thermodynamic property data and actual measured temperature deviation of the catalyst are written into the high-throughput database and added to the training set. The catalyst performance prediction model is then retrained or incrementally trained to form a closed-loop self-evolutionary process of prediction-experiment-reflux-retraining.
[0015] This invention also provides a Fischer-Tropsch synthesis catalyst performance prediction system based on high-throughput data, comprising: The data acquisition unit is used to perform high-throughput evaluation tests on a variety of catalysts under Fischer-Tropsch synthesis reaction conditions. It acquires temperature field images of the reaction bed through infrared thermal imaging to obtain temperature deviations and obtains structural characteristic data and thermodynamic property data of the catalysts. The database module is used to store structural feature data, thermodynamic property data, and temperature deviation data to build a high-throughput database of Fischer-Tropsch synthesis catalysts; The feature encoding unit is used to normalize and encode structural feature data and thermodynamic property data to construct feature vectors that can simultaneously characterize single-phase, composite, and supported catalysts. The model training unit is used to construct a dataset by combining feature vectors and corresponding temperature deviations and divide it into training and test sets. The machine learning model is trained using cross-validation and multi-objective Bayesian optimization strategies to establish a catalyst performance prediction model. The model validation unit is used to validate the accuracy of the catalyst performance prediction model using a test set. The performance prediction unit is used to input the feature vector of the catalyst to be screened into a catalyst performance prediction model with sufficient accuracy, and output the predicted catalyst performance index. The model self-calibration unit is used to experimentally verify the candidate catalysts predicted in the output, and to feed the verification data back to the database module and the model training unit to trigger the retraining or incremental training of the catalyst performance prediction model.
[0016] The method and system for predicting the performance of Fischer-Tropsch synthesis catalysts based on high-throughput data of the present invention have the following advantages over the prior art: (1) This invention uses infrared thermal imaging technology to collect temperature field images of the reaction bed and uses temperature deviation as a unified performance characterization quantity, avoiding the complexity of measuring multiple indicators such as CO conversion rate, product selectivity, and carbon chain distribution in the traditional Fischer-Tropsch synthesis catalyst evaluation, and realizing rapid prediction of catalyst performance. By constructing a normalized coding system, single-phase, composite, and supported Fischer-Tropsch synthesis catalysts are incorporated into a unified feature characterization framework. Combined with a hybrid machine learning model, the static structure-property information and dynamic time-series thermal signal information of the catalyst are integrated to establish an accurate catalyst structure-formulation-exothermic performance relationship model. The temperature-weighted feature screening and multi-objective Bayesian optimization strategy are adopted to improve the prediction accuracy of the model for high-temperature deviation regions, i.e., high-activity catalysts. Through the closed-loop mechanism of "prediction-experiment-reflux-retraining", the model is continuously optimized and self-evolved, providing an efficient technical means for high-throughput screening of Fischer-Tropsch synthesis catalysts.
[0017] (2) This invention employs a hybrid machine learning model consisting of a temporal neural network module and a gradient boosting decision tree module. The temporal neural network module processes the temporal thermal signal features of the temperature field image sequence, capturing information such as the initial temperature rise rate, steady-state temperature level and fluctuation amplitude, and the dynamic evolution of local hot spots, reflecting the catalyst's activation process, activity stability, and heat and mass transfer characteristics. The gradient boosting decision tree module processes the catalyst's structural feature data and thermodynamic property data, possessing excellent processing capabilities for both categorical and numerical features, and can automatically learn nonlinear relationships and feature interactions. By fusing the outputs of the two modules, the model simultaneously utilizes the catalyst's static structure-property information and dynamic reaction process information, establishing a more comprehensive and accurate catalyst structure-formulation-exothermic performance relationship, thus improving prediction accuracy and model generalization ability.
[0018] (3) The multi-objective Bayesian optimization strategy proposed in this invention defines a multi-objective loss function that includes the coefficient of determination, root mean square error, and mean absolute error of the high-temperature deviation sample subset. This allows the Bayesian optimization process to maintain a fine search in the high-temperature deviation region while quickly skipping over the low-temperature deviation region, thus prioritizing the allocation of the limited optimization iteration budget to the hyperparameter search space corresponding to the high-performance catalyst. Compared with traditional grid search or random search methods, this strategy can find the optimal hyperparameter combination in fewer evaluation iterations, significantly improving the efficiency of hyperparameter tuning and enabling the final model to achieve higher prediction accuracy in the high-performance catalyst range. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a flowchart of the Fischer-Tropsch synthesis catalyst performance prediction method based on high-throughput data of the present invention; Figure 2 The SHAP value diagram shows the predicted performance of the iron-based catalyst synthesized by Fischer-Tropsch in this invention. Figure 3 This is a schematic diagram of the model evaluation for predicting the performance of the iron-based catalyst synthesized by Fischer-Tropsch in this invention. Detailed Implementation
[0021] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0022] like Figure 1 As shown, this invention provides a method for predicting the performance of Fischer-Tropsch synthesis catalysts based on high-throughput data, comprising the following steps: S1. High-throughput evaluation tests were conducted on various catalysts under Fischer-Tropsch synthesis reaction conditions. Temperature deviations were obtained by acquiring temperature field images of the reaction bed of each catalyst through infrared thermal imaging. At the same time, structural characteristic data and thermodynamic property data of each catalyst were acquired to construct a high-throughput database of Fischer-Tropsch synthesis catalysts. S2. Based on the single-phase, composite, and supported structural characteristics of Fischer-Tropsch synthesis catalysts, the structural characteristic data and thermodynamic property data are normalized and encoded to construct a feature vector that can simultaneously characterize different structural characteristics of Fischer-Tropsch synthesis catalysts. S3. The feature vectors and corresponding temperature deviations are used to form a dataset and divided into a training set and a test set. The training set is input into the machine learning model, and cross-validation and multi-objective Bayesian optimization strategies are used for training and optimization to establish a catalyst performance prediction model. S4. Use the test set to verify the accuracy of the catalyst performance prediction model. If the accuracy does not meet the standard, return to step S3 to adjust the model parameters and retrain. S5. Input the feature vector of the catalyst to be screened into the catalyst performance prediction model with the required accuracy, output the predicted temperature deviation, and evaluate and select the candidate catalyst based on the predicted temperature deviation.
[0023] In one embodiment of the present invention, step S1 includes: setting up multiple catalyst reaction beds on a high-throughput evaluation device, each reaction bed being filled with a Fischer-Tropsch synthesis catalyst of a different formulation, wherein the catalysts include three types: single-phase catalysts, composite catalysts, and supported catalysts. Single-phase catalysts are single metal components such as Fe, Co, and Ni; composite catalysts are mechanical mixtures or co-precipitated complexes of two or more active components; and supported catalysts are composed of active components supported on an oxide carrier. The reaction beds adopt a fixed-bed reactor structure, with the bed height controlled within the range of 5 to 20 mm to ensure that infrared thermal imaging can effectively acquire the temperature distribution of the entire bed. In a specific embodiment of the present invention, the datasets obtained from the Fischer-Tropsch synthesis reaction are all carried out under fixed reaction conditions of 1 MPa and 325 °C, with the synthesis gas containing 48.4% hydrogen, 48.1% carbon monoxide, and the remainder argon.
[0024] During the Fischer-Tropsch synthesis reaction, temperature field images of each catalyst bed were continuously acquired using an infrared thermal imager at a frequency of 1 to 10 frames per second, recording the temperature evolution for at least 30 minutes to 2 hours after the reaction started. The highest temperature of each catalyst bed at each moment was extracted from the acquired temperature field images. This highest temperature typically occurs in the center of the bed or in the highly reactive region near the reactant gas inlet. According to the formula... Calculate the temperature deviation, where The set temperature for the reaction is the temperature control setpoint of the reactor; temperature deviation. The temperature deviation reflects the actual exothermic intensity of the catalyst in the Fischer-Tropsch synthesis reaction. A larger temperature deviation indicates higher catalytic activity and stronger exothermic reaction. The calculated temperature deviation is used as a unified performance characterization metric to characterize the exothermic properties of Fischer-Tropsch synthesis catalysts. This metric avoids the complexity of traditional catalyst evaluation, which requires measuring multiple indicators such as conversion, selectivity, and product distribution. The catalytic performance of the catalyst can be quickly evaluated using a single temperature deviation index.
[0025] Simultaneously, structural characteristic data and thermodynamic property data for each catalyst were acquired. Structural characteristic data included the type identifiers of the active components, promoters, and supports, as well as the mass fraction, volume fraction, and crystal structure parameters of each component. The active components included metals such as Fe, Co, Ni, and Ru, and their oxides; promoters included alkali metals such as Cu and K, alkaline earth metals, and rare earth elements; and supports included porous materials such as activated carbon. The mass fraction represents the proportion of each component in the total mass of the catalyst, and the volume fraction represents the proportion of each component in the total volume of the catalyst, obtained through weighing and density conversion. Crystal structure parameters included lattice constant, cell volume, space group, and grain size, obtained through X-ray diffraction analysis. For amorphous components or supports, their specific surface area and pore structure parameters were recorded. Thermodynamic property data included thermal conductivity, specific heat capacity, and density. Thermal conductivity characterizes the catalyst material's ability to transfer heat and is measured using laser flash or hot wire methods. Specific heat capacity characterizes the catalyst material's ability to absorb heat and cause a temperature increase, determined using differential scanning calorimetry. Density, the mass-to-volume ratio of the catalyst material, is determined by water displacement or gas displacement methods. For thermodynamic properties that cannot be directly measured experimentally, data can be obtained from publicly available materials databases such as the Materials Project and Matminer, or from reported literature values. The aforementioned structural characteristic data, thermodynamic property data, and corresponding temperature deviation data are stored together to construct a high-throughput database for Fischer-Tropsch synthesis catalysts. This database is categorized and managed according to catalyst type, formulation number, and preparation batch, facilitating subsequent data retrieval and model training.
[0026] In one embodiment of the present invention, step S2 includes: for the three structural characteristics of Fischer-Tropsch synthesis catalysts—single-phase, composite, and supported—a unified coding rule is used to normalize and encode the structural feature data and thermodynamic property data, constructing a feature vector capable of simultaneously characterizing different structural features. For single-phase catalysts, the type identifiers of active components are digitally encoded, for example, Fe is encoded as 1, Co as 2, Ni as 3, and so on, while the positions of the support and promoters are filled with 0 or a predetermined missing value identifier. For composite catalysts, the type identifiers of each active component are digitally encoded separately, the support position is filled with 0, and any promoters are encoded. For supported catalysts, the active components, support, and promoters are all encoded according to a preset digital coding table. Digital coding transforms the discrete category information of the catalyst composition into numerical input acceptable to a machine learning model.
[0027] Mass fraction, volume fraction, and crystal structure parameters are normalized by using maximum-minimum normalization or standardization methods to scale each parameter to the range of 0 to 1 or the range of a standard normal distribution with a mean of 0 and a standard deviation of 1, thereby eliminating the influence of differences in the dimensions and numerical ranges of different parameters on model training.
[0028] For composite catalysts and supported catalysts, since they are composed of multiple components with different thermodynamic properties, it is necessary to calculate the equivalent thermodynamic property parameters using a weighted average method based on the mass and volume fractions of each component. The equivalent thermal conductivity is calculated using a volume fraction weighted average. ,in Components volume fraction, Components The thermal conductivity. The equivalent specific heat capacity is calculated using a weighted average based on mass fraction. ,in Components mass fraction, Components Specific heat capacity. Equivalent density is calculated as the reciprocal of the weighted average of volume fraction and the reciprocal of density. ,in Components The density. By using the weighted average method described above, the complex thermodynamic properties of multi-component catalysts are simplified into a single equivalent parameter, which not only preserves the contribution of each component to the overall thermal properties, but also facilitates unified modeling with single-phase catalysts within the same framework.
[0029] The digitized type encoding, normalized mass fraction, volume fraction, crystal structure parameters, and thermodynamic or equivalent thermodynamic property parameters are combined according to a preset feature vector structure to form a feature vector. The dimension of the feature vector depends on the catalyst type and the number of features, and the arrangement order of the feature vectors is fixed, for example, active component encoding, support encoding, auxiliary agent encoding, mass fraction, volume fraction, lattice constant, cell volume, thermal conductivity, specific heat capacity, density, etc., ensuring that the feature vectors of different catalysts have a consistent structure, facilitating batch input into the model. Through a unified encoding method, integrated feature characterization of single-phase, composite, and supported catalysts is achieved, avoiding the tediousness of building independent models for different types of catalysts, and improving the model's versatility and data utilization efficiency.
[0030] In one embodiment of the present invention, step S3 includes the following sub-steps: S31. Combine the feature vectors with the corresponding temperature deviations to form a complete dataset, and divide it into a training set and a test set according to a preset ratio. S32. Input the training set into the machine learning model for training. The machine learning model includes a hybrid machine learning model, which consists of a temporal neural network module and a gradient boosting decision tree module. The temporal neural network module is used to process the temporal features of the temperature field image sequence, and the gradient boosting decision tree module is used to process the structural feature data and thermodynamic property data of the catalyst. After fusing the outputs of the two modules, the relationship between catalyst structure, formulation and exothermic performance is established. S33. The model is trained and optimized using cross-validation and multi-objective Bayesian optimization strategies to obtain the catalyst performance prediction model.
[0031] In step S31, the feature vectors constructed in step S2 are paired one by one with the corresponding temperature deviations obtained in step S1 to form a complete training sample dataset. Each sample in the dataset contains a feature vector and a temperature deviation label. The complete dataset is divided into a training set and a test set according to a preset ratio, typically 70% to 80% of the samples are used as the training set and 20% to 30% as the test set. The partitioning method can be random partitioning, stratified partitioning by catalyst type, or using the K-fold cross-validation method to ensure that the proportion of single-phase, composite, and supported catalysts in the training set and test set is consistent with that of the complete dataset, avoiding data distribution bias that could lead to a decrease in model generalization ability. In one specific implementation, the KFold method of scikit-learn is used to perform five-fold cross-validation to partition the training set, further dividing the training set into a training subset and a validation subset for model performance evaluation under cross-validation.
[0032] After constructing the complete dataset, the following steps are also included: using a temperature-weighted feature importance assessment method to filter the feature vectors. This temperature-weighted feature importance assessment method includes: introducing a sample weight based on its temperature deviation for each sample to enhance the contribution of high-temperature deviation samples to feature importance; calculating the temperature-weighted global SHAP value as a feature importance index; and using the temperature-weighted variance inflation coefficient method to remove collinear features. When the temperature-weighted VIF of a feature is greater than a preset threshold and the corresponding temperature-weighted SHAP value is low, the feature is preferentially removed. After obtaining the simplified feature vector, the training set and test set are then divided.
[0033] Specifically, sample weights Set as follows: in For the sample Temperature deviation, The preset temperature deviation threshold is typically taken as the median or 75th percentile of the central temperature deviation in the dataset. The adjustment coefficient controls the magnitude of the weighting increase for high-temperature samples. This formula increases the weight of samples with larger temperature deviations, thus giving more attention to high-temperature deviation samples in subsequent feature importance calculations.
[0034] Based on the introduction of sample weights, a temperature-weighted global SHAP value is calculated as a feature importance indicator. The SHAP value, based on the Shapley value concept in game theory, is used to quantify the marginal contribution of each feature to the model's prediction results. For features... Its effect on samples The temperature-weighted SHAP value is defined as: in For the complete feature set, For features not included Feature subset, Indicates using only subsets The model of features in the sample The predicted value, Sample weights. Global importance index. The absolute SHAP values weighted by temperature are averaged over all samples: This means that the larger the indicator, the stronger the characteristic. It makes a greater contribution to the prediction of high-temperature deviation catalysts and should be given priority in feature screening.
[0035] To avoid collinearity among features affecting model stability and interpretability, the Temperature Weighted Variance Inflation Factor (VIF) method is used to remove collinear features. For each feature... We then performed a temperature-weighted regression using the remaining features to obtain the temperature-weighted determination coefficients: in For the sample Features The value, These are the fitted values from the temperature-weighted regression. Features In weight Weighted mean ; For the sample Temperature weighting.
[0036] The corresponding temperature-weighted VIF is defined as follows: in, The temperature-weighted variance expansion coefficient of feature j is... For tolerance, the two are inversely related. A larger VIF value indicates that the feature is more easily linearly represented by other features, i.e., stronger collinearity. When the temperature-weighted VIF of a feature is greater than a preset threshold, and the corresponding temperature-weighted SHAP value... When the VIF is low, the feature is preferentially removed. For features with clear physical meaning and strong correlation with catalytic performance, such as thermal conductivity and specific heat capacity, the threshold can be appropriately relaxed even if the VIF is high, to avoid accidentally deleting key information. Through the combined screening of temperature-weighted SHAP and temperature-weighted VIF, features that contribute significantly to the prediction of high-temperature deviation catalysts are retained, while redundant and collinear features are removed. After obtaining a simplified feature vector, the training and test sets are re-divided to ensure that the subsequent model training uses the optimized feature set.
[0037] Compared with the existing practice of directly applying the general SHAP / VIF, the improved formula has the following effects in the Fischer-Tropsch system of this invention: It focuses more on high-activity / high-risk samples: the feature selection results are more inclined to explain the behavior of high-temperature biased catalysts, rather than being diluted by a large number of "ordinary" samples; it reduces the false deletion of key physical information: by combining "temperature-weighted VIF + importance threshold", it controls collinearity and retains physical features that are strongly correlated with exothermic behavior; under small sample conditions, it improves the stability and generalization ability of the model, especially the prediction accuracy in the high-temperature range is significantly improved.
[0038] Figure 2The SHAP value analysis results of the performance prediction model for iron-based catalysts in Fischer-Tropsch synthesis are presented. The SHAP value plot visually illustrates the contribution and direction of influence of each feature on the model's prediction results. Figure 2 As can be observed, taking the supported material system as an example, the absolute values of the SHAP values for characteristics such as the type of support, space group category, crystal parameters, and density are relatively large, indicating that these characteristics have a significant impact on temperature deviation prediction. However, the degree of influence is relatively limited, and thermodynamic characteristics will be added for further testing in the future.
[0039] In step S32, the training set is input into the machine learning model for training. The machine learning model of this invention consists of a temporal neural network module and a gradient boosting decision tree module, giving full play to the advantages of the two model architectures to achieve comprehensive modeling of the relationship between catalyst structure, formulation, and exothermic performance.
[0040] The temporal neural network module is used to process the temporal features of the temperature field image sequence. Specifically, the temporal neural network module employs recurrent neural network structures such as Long Short-Term Memory (LSTM) or Gated Recurrent Unit (GRU) to process the temporal thermal signal features of the temperature field image sequence acquired in step S1. Using the temperature distribution or temperature deviation sequence at different times during the reaction process as temporal input, the temporal neural network module can capture the evolution of temperature over time, including the rate of temperature rise in the initial stage of the reaction, the temperature level and fluctuation amplitude in the steady-state stage, and the dynamic features such as the appearance and disappearance of local hot spots. These temporal features reflect the catalyst activation process, activity stability, and heat and mass transfer characteristics. The output of the temporal neural network module is the extracted temporal feature vector, typically with a dimension of 10 to 30.
[0041] The gradient boosting decision tree module is used to process the structural feature data and thermodynamic property data of catalysts. It has a natural ability to handle categorical features such as active component type and support type, allowing direct input without one-hot encoding. Simultaneously, for numerical features such as mass fraction, volume fraction, thermal conductivity, specific heat capacity, and density, it can automatically learn nonlinear relationships and feature interactions through tree splitting. The output of the gradient boosting decision tree module is a temperature deviation value predicted based on the structure and thermodynamic properties.
[0042] In the training process of the gradient boosting decision tree module, in order to better adapt to the importance of high temperature deviation samples in the Fischer-Tropsch system, this invention improves the feature importance calculation method built into CatBoost, including improved PredictionValueChange importance and LossFunctionChange importance.
[0043] Regarding the importance of the predicted value change in PredictionValueChange, while preserving the CatBoost tree structure, we consider the amount of predicted value change caused by each split (or leaf set). Multiplied by sample temperature weight Aggregate PVCs according to node weights. The improved PVC importance definition is: in For the set of all split nodes using feature j, Let s be the set of samples within node s. The temperature weight for sample i The contribution of node weights, such as the number of leaf samples or leaf weights, This is a normalization factor. This formula adds a temperature weight compared to the traditional PVC index. This makes the splitting process, which has a significant impact on the prediction of high-temperature deviation samples, more important.
[0044] Regarding the importance of changes in the LossFunctionChange loss function, this invention defines: in Let be the amount of loss reduction at node s caused by splitting using feature j. The weights are assigned to the proportion of high-temperature deviation samples in node s. Through the above formula, this invention ensures that the feature importance reflects its predictive contribution to highly exothermic, potentially highly active catalyst samples, rather than simply its average contribution to all samples.
[0045] The importance metrics for PredictionValueChange / LossFunctionChange provided by CatBoost were originally designed for general datasets, calculating the "average prediction change / loss change across all samples." For the Fischer-Tropsch catalysis problem of this invention, which uses temperature as the core characterization metric in high-throughput thermal imaging, further improvements have been made: In the PVC formula, sample weights are introduced. and node weights This gives a greater weight to the "predicted change caused by feature j" on high-temperature deviation samples. In the LFC formula, the loss decreases at node s. Apply weights related to the proportion of high-temperature samples. This makes "splitting that can significantly reduce errors in the high-temperature region" more important.
[0046] The resulting technical benefits include: the improved importance index is no longer a purely "global average contribution," but rather more accurately reflects the contribution of features to key physical phenomena such as high-temperature deviation and hotspot formation; it provides a more physically meaningful ranking for subsequent feature selection and mechanism analysis—the selected high-importance features often directly correspond to variables highly correlated with exothermic behavior, such as thermal conductivity, specific heat, density, and composite ratio. Compared to the unweighted version, using the improved importance index for feature selection on the same training set can reduce redundant features while maintaining or even improving the model. And high-temperature zone MAE.
[0047] The outputs of the temporal neural network module and the gradient boosting decision tree module are fused. Fusion methods can include weighted averaging, concatenating the outputs and then training a fully connected network layer, or stacking and ensembling. Through fusion, the model can simultaneously utilize the catalyst's static structure-property information and dynamic temporal thermal signal information to establish a more comprehensive and accurate relationship between catalyst structure, formulation, and exothermic performance. The output of the fused model is the final predicted value for the temperature deviation.
[0048] During training, mean squared error (MSE) or root mean square error (RMSE) is used as the loss function. The gradient of the loss function with respect to the model parameters is calculated using the backpropagation algorithm. Gradient descent optimization algorithms such as Adam optimizer and SGD optimizer are used to update the model parameters, so as to minimize the error between the predicted value and the actual temperature deviation.
[0049] In step S33, the model is trained and optimized using cross-validation and multi-objective Bayesian optimization strategies to obtain the catalyst performance prediction model. Cross-validation employs the K-fold cross-validation method, further dividing the training set into K subsets (typically K is 5 or 10). Each time, K-1 subsets are selected as the training subset, and the remaining subset is used as the validation subset. This process is repeated K times, ensuring each subset serves as a validation set. The average performance index, such as the coefficient of determination, is calculated based on the K validation results. Root mean square error (RMSE) and mean absolute error (MAE) are used to evaluate the model's generalization ability and stability. Cross-validation can make full use of limited training data, avoid the influence of randomness caused by a single split on model evaluation, and provide a more reliable model performance estimate.
[0050] The multi-objective Bayesian optimization strategy in step S33 specifically includes: defining a multi-objective loss function that combines the coefficient of determination, root mean square error, and mean absolute error of the high-temperature deviation sample subset as the optimization objective; approximating the objective function using a Gaussian process; selecting a Matérn5 / 2 kernel function with automatic correlation determination; setting different base length scales for each hyperparameter dimension based on the sensitivity of different hyperparameters to the model on the Fischer-Tropsch system; and introducing a temperature deviation adaptive factor to dynamically adjust the length scale of the kernel function, so that the Bayesian optimization process maintains a fine search in the high-temperature deviation region while quickly skipping over the low-temperature deviation region, and iteratively searching to obtain the optimal hyperparameter combination that minimizes the multi-objective loss function.
[0051] Specifically, hyperparameters include the number of iterations, tree depth, learning rate, L2 regularization coefficient (l2_leaf_reg), and subsampling ratio for the gradient boosting decision tree module, as well as the number of hidden layer units, learning rate, dropout ratio, and number of recurrent layers for the temporal neural network module. For hybrid models, hyperparameters such as fusion weights are also included. A multi-objective loss function is defined as the optimization objective. in For hyperparameter combination, The coefficient of determination under cross-validation. The root mean square error under cross-validation. The mean absolute error over the high-temperature deviation sample subset. , and All of these are weighting coefficients. By introducing a specific error term for the high-temperature deviation sample subset, the optimization process is more inclined to accurately predict highly active catalysts, which meets the key focus on high-performance samples in the screening of Fischer-Tropsch synthesis catalysts.
[0052] Gaussian process is used for multi-objective loss function Approximate modeling is performed. The Gaussian process is a non-parametric Bayesian method that provides a probabilistic approximation of the black-box function and estimates the uncertainty of the prediction. For the evaluated set of hyperparameters... and the corresponding target value , any new point The posterior mean and variance are: in The kernel matrix between the evaluated points. The kernel vector between the new point and the previously evaluated points. To observe the noise variance, It is the identity matrix. Posterior mean Give the best estimate of the target value for the new hyperparameter configuration, and the posterior variance. The uncertainty of the estimate was quantified.
[0053] Kernel function The Matérn5 / 2 kernel function with automatic correlation determination of ARD is selected, specifically in the following form: The adaptive distance metric is: Let be the number of hyperparameter dimensions. For the first The values of the hyperparameters, For the first The basic length scale of each hyperparameter dimension. This is the signal variance hyperparameter. This is the temperature deviation adaptive factor. In the above posterior mean and variance, the kernel matrix... The elements are This indicates that the points have been assessed. and Correlation between them; kernel vector The elements are , indicating a new point Compared with the assessed points Correlation between them; scalar = This represents the kernel function value for the new point itself. Compared to the commonly used radial basis function (RBF) kernel, the atérn5 / 2 kernel offers better smoothness and differentiability, making it suitable for modeling complex hyperparameter response surfaces. The ARD mechanism allows setting different length scales for different hyperparameter dimensions, enabling the model to automatically identify the importance of each hyperparameter.
[0054] Temperature deviation adaptive factor is defined as in, Configure the corresponding average temperature deviation for the two hyperparameters. To train the median of temperature deviation in the dataset, For adaptive strength coefficient, It is a hyperbolic tangent function. The introduction of the temperature deviation adaptive factor allows the effective length scale of the kernel function to be dynamically adjusted according to the corresponding prediction performance based on the hyperparameter configuration.
[0055] Based on the sensitivity of different hyperparameters to the model on the Fischer-Tropsch system, different base length scales are set for each hyperparameter dimension. For hyperparameters that significantly impact model performance, such as tree depth and learning rate, setting a smaller length scale (e.g., 0.1 to 0.5) makes the Gaussian process more sensitive to hyperparameter changes in that dimension, allowing for a more detailed characterization of local variations in the performance response surface. For hyperparameters with less impact, such as the regularization coefficient l2_leaf_reg, setting a larger length scale (e.g., 1 to 2) reduces unnecessary sampling and improves optimization efficiency. When configuring the two hyperparameters, the corresponding average temperature deviation... Much larger than the median of the dataset At this time, the hyperbolic tangent function approaches 1, the adaptive factor approaches 1, and the effective length scale approaches the basic length scale. The kernel function maintains a small length scale, allowing the Bayesian optimization process to perform a fine-grained search in the high-temperature deviation region, fully exploring the hyperparameter space of this region to find the configuration that best predicts the performance of the catalyst. Conversely, when much smaller When the hyperbolic tangent function approaches 0, the adaptive factor approaches 0. The effective length scale is enlarged to This enhances the correlation between sample points in the low-temperature deviation region within the kernel space, allowing the Bayesian optimization algorithm to quickly skip this region without lingering too long. This prioritizes allocating the limited optimization iteration budget to the hyperparameter search space corresponding to the high-performance catalyst.
[0056] In each iteration, the next hyperparameter combination to be evaluated is selected based on the posterior distribution of the Gaussian process. The selection strategy employs acquisition functions such as expected improvement (EI), confidence upper bound (UCB), or probabilistic improvement (PI), balancing the utilization of the current optimal solution with the exploration of unknown regions. The multi-objective loss function value under this hyperparameter combination is evaluated, and the new evaluation results are added to the evaluated set to update the Gaussian process model. The iterative search is performed a preset number of times, such as 50 to 200, to obtain the optimal hyperparameter combination that minimizes the multi-objective loss function. The model is then retrained on the complete training set using the optimal hyperparameter combination to obtain the final catalyst performance prediction model.
[0057] Compared to traditional grid search or random search methods, the aforementioned multi-objective Bayesian optimization strategy can find near-optimal hyperparameter combinations in fewer evaluation iterations, significantly improving hyperparameter tuning efficiency. By introducing a multi-objective loss function and a temperature deviation adaptive kernel function for the Fischer-Tropsch system, the optimization process focuses more on the prediction accuracy of high-temperature deviation catalysts, enabling the final model to achieve higher prediction accuracy in the high-performance catalyst range most relevant to industrial applications.
[0058] In one embodiment of the present invention, step S4 includes: inputting the feature vector of the catalyst in the test set into the catalyst performance prediction model to obtain the predicted temperature deviation; comparing the predicted temperature deviation with the corresponding actual temperature deviation in the test set; calculating the coefficient of determination, root mean square error and mean absolute error as accuracy evaluation indicators; determining whether the accuracy evaluation indicators meet the preset accuracy standard; if they meet the standard, the model accuracy is considered to be up to standard; if they do not meet the standard, returning to step S3 to adjust the model parameters or feature vector and retraining.
[0059] Set accuracy standards, for example , , Or, for specific accuracy requirements of high-temperature deviation sample subsets, such as When the performance metrics on the test set meet the above standards, the model accuracy is considered satisfactory, and the process proceeds to step S5 for practical application. If the accuracy does not meet the standards, the process returns to step S3 to adjust the model parameters and retrain. Adjustment strategies include: increasing the training sample size by supplementing with high-throughput experiments to obtain more catalyst data; adjusting feature engineering methods, such as modifying the normalization method, adding or deleting features, or constructing new feature combinations; modifying the model architecture, such as adjusting the number of neural network layers, changing the gradient boosting tree ensemble strategy, or optimizing the fusion method; expanding the hyperparameter search space and increasing the number of iterations in Bayesian optimization; and adjusting the weight coefficients of the multi-objective loss function to be more biased towards a specific performance metric. After adjustments, the model is retrained, and the accuracy is verified again on the test set. This process is iterated until the accuracy meets the standards.
[0060] In a specific embodiment of the present invention, Figure 3 Figure (a) shows the performance evaluation results of the CatBoost model on the training and test sets, and Figure (b) shows the performance evaluation results of the RandomForestRegressor model on the training and test sets. The coefficient of determination was used as the evaluation metric. ,from Figure 3 As can be seen from this, the CatBoost model performs well on both the training and test sets. Both results are superior to the RandomForestRegressor model, indicating that the CatBoost model has better fitting performance and generalization ability. This is mainly due to CatBoost's inherent advantage in handling categorical features, which can directly distinguish different catalyst types (such as active components like Fe, Co, and Ni, as well as different support types) without using one-hot encoding, avoiding the overfitting risk caused by high-dimensional sparse features. Furthermore, when new supporting materials or new active components are subsequently added, CatBoost can directly classify and model them as new categories without redesigning the feature encoding method, demonstrating good scalability and adaptability.
[0061] In one specific implementation, leave-one-out cross-validation or an external validation set can also be used for additional validation. Leave-one-out cross-validation involves setting aside one sample from the dataset each time as a validation sample, training the model with the remaining samples, calculating the prediction error of the set-off sample, and then iterating through all samples to obtain the overall prediction error. This method is suitable for small datasets. The external validation set consists of new catalyst samples completely independent of the training and test sets. Performance evaluation on the external validation set further confirms the model's generalization ability and practical application value.
[0062] In one embodiment of the present invention, step S5 includes: inputting the feature vector of the catalyst to be screened into a catalyst performance prediction model with sufficient accuracy, outputting the predicted temperature deviation, and evaluating and selecting the candidate catalyst based on the predicted temperature deviation.
[0063] Specifically, predicting temperature deviation The larger the value, the higher the catalytic activity and the stronger the exothermic reaction in the Fischer-Tropsch synthesis, indicating greater industrial application value. A temperature deviation threshold can be set, for example... or Candidate catalysts with predicted temperature deviations exceeding a threshold are selected as preferred formulations for subsequent small-scale or pilot-scale verification. For multiple candidate catalysts, they can be sorted from highest to lowest predicted temperature deviation, and the top-ranked formulations can be selected for experimental verification, significantly reducing the workload and cost of experimental screening.
[0064] In one specific implementation, uncertainty analysis can also be used to assess the confidence level of the prediction results. For models employing Bayesian methods or ensemble learning, the uncertainty or confidence interval of the prediction can be output. For example, Gaussian process regression or Bayesian neural networks can provide the prediction mean and prediction variance. For multiple base learners in ensemble models such as random forests or gradient boosting trees, uncertainty can be assessed through the variance or quantile range of the prediction results from the base learners. Candidate catalysts with lower prediction uncertainty indicate that the model is more confident in predicting their performance and can be prioritized for experimental validation; candidate catalysts with higher prediction uncertainty may be located in areas with less training data coverage, and the reliability of their predictions is lower. Whether to validate them or add them to the training set using an active learning strategy can be determined based on resource availability to improve the model. Through the above-described machine learning-based catalyst performance prediction and screening method, the performance of a large number of candidate catalyst formulations can be rapidly evaluated without conducting actual high-throughput experiments, significantly shortening the catalyst development cycle and reducing development costs.
[0065] In one embodiment of the present invention, the prediction method further includes: The candidate catalysts predicted by the catalyst performance prediction model in step S5 are verified by high-throughput preparation and Fischer-Tropsch synthesis experiments. The actual measured temperature deviation is obtained and compared with the predicted temperature deviation. When the deviation is within the preset range, the catalyst is confirmed to have reached the target performance. When the deviation exceeds the preset range, the structural feature data, thermodynamic property data and actual measured temperature deviation of the catalyst are written into the high-throughput database and added to the training set. The catalyst performance prediction model is then retrained or incrementally trained to form a closed-loop self-evolutionary process of prediction-experiment-reflux-retraining.
[0066] Specifically, following step S1, the selected candidate catalyst was tested for Fischer-Tropsch synthesis reaction on a high-throughput evaluation device, and the experimental temperature deviation was obtained by acquiring temperature field images using infrared thermal imaging. Calculate the prediction error. ,when When the temperature is below a preset threshold, such as 5°C, the prediction error is considered to be within an acceptable range, and the model prediction is accurate. When the error exceeds the preset threshold, the source of the analysis error may include batch differences in catalyst preparation, deviation of reaction conditions, inaccurate feature extraction, or insufficient model extrapolation ability. The feature engineering or model structure should be adjusted according to the analysis results.
[0067] The experimentally validated new catalyst sample data, including feature vectors and experimental temperature deviations, are added to the high-throughput database constructed in step S1. The expanded database is used to retrain the catalyst performance prediction model, or incremental learning methods are employed to update the parameters of the existing model, avoiding the high computational cost of training from scratch. Through a continuous "prediction-validation-feedback-update" closed loop, the model's prediction accuracy and generalization ability are continuously improved, gradually covering a wider range of catalyst formulations and achieving self-evolution of the model. This iterative optimization strategy aligns with the idea of active learning, maximizing the contribution of each new sample to model performance improvement by selectively validating samples with high model uncertainty or excellent expected performance, thus achieving maximum model improvement with minimal experimental cost.
[0068] This invention also provides a high-throughput data-based system for predicting the performance of Fischer-Tropsch synthesis catalysts. This system is used to perform the prediction method described above, including: The data acquisition unit is used to perform high-throughput evaluation tests on a variety of catalysts under Fischer-Tropsch synthesis reaction conditions. It acquires temperature field images of the reaction bed through infrared thermal imaging to obtain temperature deviations and obtains structural characteristic data and thermodynamic property data of the catalysts. The database module is used to store structural feature data, thermodynamic property data, and temperature deviation data to build a high-throughput database of Fischer-Tropsch synthesis catalysts; The feature encoding unit is used to normalize and encode structural feature data and thermodynamic property data to construct feature vectors that can simultaneously characterize single-phase, composite, and supported catalysts. The model training unit is used to construct a dataset by combining feature vectors and corresponding temperature deviations and divide it into training and test sets. The machine learning model is trained using cross-validation and multi-objective Bayesian optimization strategies to establish a catalyst performance prediction model. The model validation unit is used to validate the accuracy of the catalyst performance prediction model using a test set. The performance prediction unit is used to input the feature vector of the catalyst to be screened into a catalyst performance prediction model with sufficient accuracy, and output the predicted catalyst performance index. The model self-calibration unit is used to experimentally verify the candidate catalysts predicted in the output, and to feed the verification data back to the database module and the model training unit to trigger the retraining or incremental training of the catalyst performance prediction model.
[0069] In summary, this solution enables real-time updates to the database and performance prediction model, and establishes a mutually corroborating bridge between experimental results from high-throughput experiments and the prediction results of new catalyst modeling, thereby predicting the performance of new catalysts. This solution can be automated, shortening the experimental cycle and reducing economic costs.
[0070] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for predicting the performance of Fischer-Tropsch synthesis catalysts based on high-throughput data, characterized in that: Includes the following steps: S1. High-throughput evaluation tests were conducted on various catalysts under Fischer-Tropsch synthesis reaction conditions. Temperature deviations were obtained by acquiring temperature field images of the reaction bed of each catalyst through infrared thermal imaging. At the same time, structural characteristic data and thermodynamic property data of each catalyst were acquired to construct a high-throughput database of Fischer-Tropsch synthesis catalysts. S2. Based on the single-phase, composite, and supported structural characteristics of Fischer-Tropsch synthesis catalysts, the structural characteristic data and thermodynamic property data are normalized and encoded to construct a feature vector that can simultaneously characterize different structural characteristics of Fischer-Tropsch synthesis catalysts. S3. The feature vectors and corresponding temperature deviations are used to form a dataset and divided into a training set and a test set. The training set is input into the machine learning model, and cross-validation and multi-objective Bayesian optimization strategies are used for training and optimization to establish a catalyst performance prediction model. S4. Use the test set to verify the accuracy of the catalyst performance prediction model. If the accuracy does not meet the standard, return to step S3 to adjust the model parameters and retrain. S5. Input the feature vector of the catalyst to be screened into the catalyst performance prediction model with the required accuracy, output the predicted temperature deviation, and evaluate and select the candidate catalyst based on the predicted temperature deviation.
2. The method for predicting the performance of Fischer-Tropsch synthesis catalysts based on high-throughput data as described in claim 1, characterized in that: In step S1, the structural feature data includes the active components, type identification, mass fraction, volume fraction, and crystal structure parameters of the Fischer-Tropsch synthesis catalyst, and the thermodynamic property data includes thermal conductivity, specific heat capacity, and density.
3. The method for predicting the performance of Fischer-Tropsch synthesis catalysts based on high-throughput data as described in claim 1, characterized in that: In step S1, the temperature deviation is obtained by acquiring temperature field images of each catalyst reaction bed layer using infrared thermal imaging. Specifically, this includes: High-throughput thermal imaging temperature sequences were continuously acquired during the Fischer-Tropsch synthesis reaction, and the highest temperature T of the reaction bed was extracted from the temperature field images. max Calculate the temperature deviation ΔT=T max -T set T set A temperature is set for the reaction, and the temperature deviation is used as a unified performance characterization metric to characterize the exothermic properties of the Fischer-Tropsch synthesis catalyst.
4. The method for predicting the performance of Fischer-Tropsch synthesis catalysts based on high-throughput data as described in claim 2, characterized in that: Step S2, which involves normalizing and encoding Fischer-Tropsch synthesis catalysts with different structural features to construct feature vectors, specifically includes: The active components, supports, and promoters of single-phase catalysts, composite catalysts, and supported catalysts are all digitally coded. The mass fraction, volume fraction, and crystal structure parameters are normalized. For composite catalysts and supported catalysts, the equivalent thermodynamic property parameters are calculated using a weighted average method based on the mass fraction and volume fraction of each component. The digital type code, the normalized mass fraction, volume fraction, crystal structure parameters, and thermodynamic property parameters or equivalent thermodynamic property parameters are combined to form a feature vector.
5. The method for predicting the performance of Fischer-Tropsch synthesis catalysts based on high-throughput data as described in claim 2, characterized in that: Step S3 specifically includes: S31. Combine the feature vectors with the corresponding temperature deviations to form a complete dataset, and divide it into a training set and a test set according to a preset ratio. S32. Input the training set into the machine learning model for training. The machine learning model consists of a temporal neural network module and a gradient boosting decision tree module. The temporal neural network module is used to process the temporal thermal signal features of the temperature field image, and the gradient boosting decision tree module is used to process the structural feature data and thermodynamic property data of the catalyst. After fusing the outputs of the two modules, the relationship between catalyst structure, formulation and exothermic performance is established. S33. The model is trained and optimized using cross-validation and multi-objective Bayesian optimization strategies to obtain the catalyst performance prediction model.
6. The method for predicting the performance of Fischer-Tropsch synthesis catalysts based on high-throughput data as described in claim 5, characterized in that: In step S31, after the complete dataset is formed, the following steps are also included: using a temperature-weighted feature importance evaluation method to filter the feature vector. The temperature-weighted feature importance evaluation method includes: introducing a sample weight based on its temperature deviation for each sample to enhance the contribution of high-temperature deviation samples to feature importance; calculating the temperature-weighted global SHAP value as a feature importance index; and using the temperature-weighted variance inflation coefficient method to remove collinear features. When the temperature-weighted VIF of a feature is greater than a preset threshold and the corresponding temperature-weighted SHAP value is low, the feature is preferentially removed. After obtaining the simplified feature vector, the training set and test set are then divided.
7. The method for predicting the performance of Fischer-Tropsch synthesis catalysts based on high-throughput data as described in claim 5, characterized in that: The multi-objective Bayesian optimization strategy in step S33 specifically includes: defining a multi-objective loss function that combines the coefficient of determination, root mean square error, and mean absolute error of the high-temperature deviation sample subset as the optimization objective; approximating the objective function using a Gaussian process; selecting a Matérn5 / 2 kernel function with automatic correlation determination; setting different base length scales for each hyperparameter dimension based on the sensitivity of different hyperparameters to the model on the Fischer-Tropsch system; and introducing a temperature deviation adaptive factor to dynamically adjust the length scale of the kernel function, so that the Bayesian optimization process maintains a fine search in the high-temperature deviation region while quickly skipping over the low-temperature deviation region, and iteratively searching to obtain the optimal hyperparameter combination that minimizes the multi-objective loss function.
8. The method for predicting the performance of Fischer-Tropsch synthesis catalysts based on high-throughput data as described in claim 2, characterized in that: Step S4, which uses a test set to verify the accuracy of the catalyst performance prediction model, specifically includes: inputting the feature vector of the catalyst in the test set into the catalyst performance prediction model to obtain the predicted temperature deviation; comparing the predicted temperature deviation with the corresponding actual temperature deviation in the test set; calculating the coefficient of determination, root mean square error, and mean absolute error as accuracy evaluation indicators; determining whether the accuracy evaluation indicators meet the preset accuracy standard; if they do, the model accuracy is considered to be up to standard; if they do not, returning to step S3 to adjust the model parameters or feature vector and retraining.
9. The method for predicting the performance of Fischer-Tropsch synthesis catalysts based on high-throughput data as described in claim 1, characterized in that: The prediction method further includes: The candidate catalysts predicted by the catalyst performance prediction model in step S5 are verified by high-throughput preparation and Fischer-Tropsch synthesis experiments. The actual measured temperature deviation is obtained and compared with the predicted temperature deviation. When the deviation is within the preset range, the catalyst is confirmed to have reached the target performance. When the deviation exceeds the preset range, the structural feature data, thermodynamic property data and actual measured temperature deviation of the catalyst are written into the high-throughput database and added to the training set. The catalyst performance prediction model is then retrained or incrementally trained to form a closed-loop self-evolutionary process of prediction-experiment-reflux-retraining.
10. A performance prediction system for Fischer-Tropsch synthesis catalysts based on high-throughput data, characterized in that: The system is used to perform the prediction method as described in any one of claims 1-9, including: The data acquisition unit is used to perform high-throughput evaluation tests on a variety of catalysts under Fischer-Tropsch synthesis reaction conditions. It acquires temperature field images of the reaction bed through infrared thermal imaging to obtain temperature deviations and obtains structural characteristic data and thermodynamic property data of the catalysts. The database module is used to store structural feature data, thermodynamic property data, and temperature deviation data to build a high-throughput database of Fischer-Tropsch synthesis catalysts; The feature encoding unit is used to normalize and encode structural feature data and thermodynamic property data to construct feature vectors that can simultaneously characterize single-phase, composite, and supported catalysts. The model training unit is used to construct a dataset by combining feature vectors and corresponding temperature deviations and divide it into training and test sets. The machine learning model is trained using cross-validation and multi-objective Bayesian optimization strategies to establish a catalyst performance prediction model. The model validation unit is used to validate the accuracy of the catalyst performance prediction model using a test set. The performance prediction unit is used to input the feature vector of the catalyst to be screened into a catalyst performance prediction model with sufficient accuracy, and output the predicted catalyst performance index. The model self-calibration unit is used to experimentally verify the candidate catalysts predicted in the output, and to feed the verification data back to the database module and the model training unit to trigger the retraining or incremental training of the catalyst performance prediction model.
Citation Information
Patent Citations
Catalytic material high-flux calculation method and system
CN111177915A