Bio-oil yield prediction method based on interpretable machine learning
By applying the CTGAN-XGB-BO model and SHAP analysis in the field of biomass pyrolysis, the problems of difficulty in obtaining data and lack of interpretability in bio-oil yield prediction are solved, and high-precision prediction and process optimization are achieved.
Patent Information
- Application Number
- CN202510622389.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-06-13
AI Technical Summary
In the prediction of bio-oil yields during biomass pyrolysis, the existing technology has problems such as high data acquisition cost, small sample size, and lack of interpretability of the model, making it difficult to accurately capture complex nonlinear relationships and multivariable interactions.
The machine learning algorithm based on CTGAN-XGB-BO is used to enhance the data through a conditional table generation adversarial network (CTGAN), combined with the extreme gradient enhancement (XGB) model and Bayesian optimization (BO) tuning model, a bio-oil yield prediction model was constructed, and the model was interpreted using SHAP analysis.
Improves the accuracy of bio-oil yield prediction, reduces the time and cost of experimental research, provides model interpretability, and allows engineers and researchers to better understand and optimize the production process.
Smart Images

Figure CN120148703A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of biomass pyrolysis technology and machine learning, and particularly relates to a method for predicting biomass pyrolysis yield based on interpretable machine learning. Background Art
[0002] Biomass pyrolysis is a thermochemical process that converts biomass raw materials into bio-oil, bio-pyrolysis gas, and biochar. Among them, bio-oil, as the main liquid product in the biomass pyrolysis process, has a high energy density and is considered a promising renewable energy source that can be used to replace traditional fossil fuels. However, due to the diversity of biomass raw materials and the complexity of the pyrolysis process, the bio-oil yield is affected by various factors, such as biomass type, pyrolysis temperature, heating rate, particle size, etc. Therefore, accurately predicting the bio-oil yield is of great significance for optimizing the biomass pyrolysis process, increasing the bio-oil yield and quality.
[0003] Traditional methods for predicting bio-oil yield mainly rely on experimental research and empirical formulas. Although these methods can provide prediction results to a certain extent, they have the following limitations. First, experimental research is usually time-consuming and laborious, and it is difficult to cover all possible experimental conditions. Second, empirical formulas are often based on data under specific experimental conditions, with limited applicability and difficulty in generalizing to different biomass raw materials and pyrolysis conditions. In addition, traditional methods have certain limitations in dealing with complex non-linear relationships and multi-variable interactions, and it is difficult to accurately capture the influence of various factors on the bio-oil yield during the biomass pyrolysis process.
[0004] With the development of machine learning technology, data-driven prediction methods have gradually been applied to the field of biomass pyrolysis. Machine learning models can automatically capture the complex non-linear relationship between input parameters and bio-oil yield by learning a large amount of experimental data, and have high prediction accuracy. Although data-driven methods have made progress, there are still challenges in data acquisition. First, the acquisition cost of experimental data is high and the cycle is long, resulting in a generally small sample size, making its generalization ability insufficient in new scenarios outside the training set. Second, biomass pyrolysis involves multi-variable non-linear coupling, and it is difficult for the original data to cover all possible parameter combinations, especially the data for extreme conditions or rare raw material types is severely lacking. In addition, traditional machine learning models such as neural networks and support vector machines are usually regarded as "black box" models, and it is difficult to explain their internal working principles and prediction results. This lack of interpretability limits the promotion and trust of machine learning models in practical applications, especially in scenarios where it is necessary to understand and optimize the biomass pyrolysis process. Summary of the Invention
[0005] To overcome the deficiencies in the existing technologies and address the dual bottlenecks of biomass pyrolysis data and model transparency, the present invention proposes a method for predicting bio-oil yield based on interpretable machine learning, designs a machine learning algorithm based on the CTGAN-XGB-BO model, improves the accuracy of bio-oil yield prediction, reduces the time and cost of experimental research, and uses SHAP analysis to interpret the model, enabling engineers and researchers to better understand and optimize the production process, enhance the process control level, and contribute to the technological progress and sustainable development of the biomass pyrolysis industry.
[0006] To achieve the above object, the present invention provides the following technical solution: A method for predicting bio-oil yield based on interpretable machine learning, comprising:
[0007] Step 1: Obtain a bio-oil yield data set of biomass raw material samples and perform preprocessing.
[0008] Step 1.1: Process the missing values and outliers in the data set constructed for predicting bio-oil yield. Fill the missing values of the corresponding attributes with the average values of the attributes. Use box plot analysis to detect the outliers in the data, define the normal fluctuation range of the data, so as to accurately identify the outliers beyond this range.
[0009] Step 1.2: Standardize the data to make different features have the same scale, thereby improving the training effect and prediction performance of the model.
[0010] Step 2: Enhance the preprocessed data by applying the conditional tabular generative adversarial network CTGAN.
[0011] Step 2.1: In CTGAN, the generator receives the conditional variable and the random noise vector , generates a new data sample , and the conditional variable is used to guide the generator to generate data samples under specific conditions. Select the pyrolysis temperature PT and the gas purge flow rate FR as the conditional variables, and define the conditional variable as c = [ PT , FR ] . The random noise vector is specifically a standard normal distribution with a mean of 0 and a variance of 1 , which makes each component of the random noise vector independent and identically distributed, thus providing the necessary random input for to explore the latent space of the data during the learning process and generate diverse data samples. receives the input e = [ c , z ] After that, new data samples are generated through the internal neural network structure .
[0012] Step 2.2: Discriminator Receives real data samples and generated data samples , and at the same time receives the conditional variable , that is the input of is d = [ c , x ] , where is the overall input sample. Upon receiving the input vector , outputs a probability value , indicating the authenticity of the data sample.
[0013] Step 2.3: Train the discriminator. First, fix the parameters of the generator and update the parameters of the discriminator. The goal of the discriminator is to maximize its loss function :[[]]
[0014] m a x L D = E c , x r [ D ( c , x r ) ] − E c , z [ D ( G ( c , z ) , c ) ] + λ E x ⌢ ~ P x ⌢ [ ( ∥ ∇ x ⌢ D ( c , x ⌢ ) ∥ 2 − 1 ) 2 ] ,
[0015] The loss function consists of two parts. One part is the discrimination result of real data and generated data, and the other part is the gradient penalty term, where is the expected value of the conditional variable and real data, is the expected value of the conditional variable and random noise, is the expected value of , is a sample sampled from real data and generated data through linear interpolation, is 's probability distribution, is the gradient penalty coefficient, used to control the weight of the gradient penalty term, is the discriminator's gradient of . Specifically, randomly sample and the corresponding conditional variable from the real dataset and randomly sample a batch from the standard normal distribution:[[]]
[0016] ,
[0017] where is the batch size, is the th real data sample, is the conditional variable corresponding to the th real data sample, is the A random noise vector. Generate a batch of data samples using the generator :
[0018] ,
[0019] where is the th data sample generated by the generator. Sample a batch of samples from the real data and the generated data through linear interpolation :
[0020] ,
[0021] where is a random number uniformly distributed between 0 and 1, used to control the interpolation ratio, is the th sample obtained through linear interpolation, is the random number in the th real data sample. Finally, calculate the losses of the discriminator for the real data, the generated data, and the interpolated samples, and update the parameters of the discriminator.
[0022] Step 2.4: Train the generator. First, fix the parameters of the discriminator and update the parameters of the generator. The goal of the generator is to minimize its loss function :
[0023] min L G =− E c , z [ D ( G ( c , z ) , c ) ] ,
[0024] Specifically, randomly sample a batch of and the corresponding :
[0025] ,
[0026] Use the generator to generate a batch of data samples :
[0027] ,
[0028] Finally, calculate the loss of the generator and update the parameters of the generator. By alternately training the generator and the discriminator, the generator can finally generate data samples that meet the conditions.
[0029] Step 3: Divide the enhanced dataset into a training set and a test set, train using the Extreme Gradient Boosting XGB model, optimize the hyperparameters of the Bayesian Optimization BO model, and construct a CTGAN-XGB-BO prediction model.
[0030] Step 3.1: The enhanced dataset Randomly divided into a training set and a test set in a ratio of 8:2, the training set is used to train the XGB model, and the test set is used to evaluate the performance of the model. Among them is the number of samples, represents the th sample's feature vector, represents the th sample's label.
[0031] Step 3.2: Use the training set to train the XGB model. The XGB model consists of additive trees and is used to predict the output of the sample, which can be expressed as:
[0032] ,
[0033] where is the predicted output of the th sample, is the th sample's feature vector, is the th decision tree, represents the set of regression trees, is the function that maps the data to the regression tree leaf nodes, is the weight of the leaf node. The objective function of XGB is:
[0034] ,
[0035] where is the loss function, which is used to calculate the sample's predicted value and the residual sum between the actual value . is the regularization term, the smaller the value, the better the tree structure, and its definition is:
[0036] ,
[0037] where and are hyperparameters that control the complexity of the model structure, is the number of leaf nodes, is the weight parameter. XGB makes the model more accurate by adding a tree in each iteration. Therefore, in the th iteration, since the result of the th iteration has been obtained, the objective function can be transformed into:
[0038] ,
[0039] where is the value increased by the regression tree in the th iteration. Solving is equivalent to solving in the current state.
[0040] For , performing a second-order Taylor expansion can determine the gain generated by tree splitting after iteration. The larger the value, the more the objective function can be simplified and the better the model. Then the change in the loss function for this iteration
[0041] can be expressed as:
[0042] where and are the first and second derivatives of the loss function respectively, is the sample set of the left child node, is the sample set of the right child node, represents the sample set of the current node.
[0043] Step 3.3: Apply Bayesian optimization BO to optimize the model. Define the hyperparameter optimization problem of the XGB model as the objective function, and set the coefficient of determination and the root mean square error as the objective function of the model. Its specific calculation expression is:
[0044] ,
[0045] where is the mean of the true values. Set the learning rate, maximum depth, number of trees, regularization parameters alpha and lambda as hyperparameters. The process of optimizing the hyperparameters of the BO model is as follows:
[0046] ,
[0047] where is the objective function, is the prior distribution, represents the posterior distribution of the objective function, is the sampling point information marginal probability distribution, is the hyperparameter at the th step, is the observed value at the th step, is the observation error, is the The likelihood distribution. First, construct a probabilistic surrogate model to approximately represent the behavior of the objective function. Then, based on the current surrogate model, use the acquisition function to search for candidate points in the hyperparameter space that are most likely to improve the performance of the objective function under the current model prediction. Apply the selected hyperparameter configuration to the actual objective function, evaluate it, and obtain the true performance metrics. Update the surrogate model using the newly obtained evaluation results. Repeat the above steps until the preset termination condition is reached.
[0048] Step 3.4: Extract the optimal hyperparameter combination from the Bayesian optimizer, retrain the XGB model using these hyperparameters, and evaluate the performance of the model using the test set.
[0049] Step 4: Interpret the CTGAN-XGB-BO bio-oil yield prediction model after training and optimization.
[0050] Step 4.1: Apply SHAP analysis to interpret the model. It helps understand the model's decision-making process by calculating the contribution of each feature to the prediction result. The calculation of SHAP values is based on the concept of Shapley values, which are used to allocate the benefits of cooperation. For each feature , its Shapley value represents the average marginal contribution of feature to the prediction result. The specific process is as follows:
[0051] Step 4.2: For a specific feature combination , calculate the difference in the model output between the case that includes feature and the case that does not include feature :
[0052] ,
[0053] where is the number of all features, is the model output using only the feature combination , is the model output using the feature combination and feature .
[0054] Step 4.3: Perform a weighted average over all possible feature combinations to calculate the Shapley value of feature :
[0055] ,
[0056] where represents the size of the feature combination , represents feature The marginal contribution is calculated to obtain the SHAP values for each feature, which represent the contribution of each feature to the prediction result.
[0057] Step 4.4: Display the SHAP values through a visualization tool. Through these visualization results, the features that have the greatest impact on the prediction of bio-oil yield can be identified, and the positive and negative impacts of each feature on the prediction result can be shown. Description of the Drawings
[0058] Figure 1 is the overall flowchart of the embodiments of the present invention.
[0059] Figure 2 is the structural framework diagram of the conditional tabular generative adversarial network. Detailed Embodiments
[0060] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0061] Appendix Figure 1 is the overall flowchart of the embodiments of the present invention. This embodiment provides a method for predicting bio-oil yield based on interpretable machine learning. The specific process includes: obtaining a data set and performing preprocessing, applying a conditional tabular generative adversarial network to augment the original data, constructing a CTGAN-XGB-BO model to predict the bio-oil yield, and using SHAP analysis to interpret the model.
[0062] Appendix Figure 2 is the structural framework diagram of the conditional tabular generative adversarial network of the present invention, which consists of a generator and a discriminator module. The generator receives a random variable and a conditional vector, processes the original data through pattern-specific normalization, and generates data samples. The discriminator samples the mixed data multiple times to judge the authenticity of the data. The two are in mutual confrontation to generate high-quality synthetic data and improve the prediction performance of the model.
[0063] Step 1: Obtain a data set of bio-oil yields of biomass raw material samples and perform preprocessing.
[0064] Step 1.1: Process the missing values and outliers in the constructed data set for predicting bio-oil yield. Fill the missing values of the corresponding attributes with the average values of the attributes. Use box plot analysis to detect the outliers in the data and define the normal fluctuation range of the data, so as to accurately identify the outliers beyond this range.
[0065] Step 1.2: Standardize the data to make different features have the same scale, thereby improving the training effect and prediction performance of the model.
[0066] Step 2: Augment the preprocessed data through the conditional tabular generative adversarial network CTGAN.
[0067] Step 2.1: In CTGAN, the generator receives the conditional variable and the random noise vector , and generates new data samples , where the conditional variable is used to guide the generator to generate data samples under the condition. The pyrolysis temperature PT and the gas purge flow rate FR are selected as the conditional variables, and the conditional variable is defined as c = [ PT , FR ] . The random noise vector is specifically a standard normal distribution with a mean of 0 and a variance of 1 , which makes each component of the random noise vector independent and identically distributed, thus providing the necessary random input for to explore the latent space of the data during the learning process and generate diverse data samples. After receiving the input e = [ c , z ] , it generates new data samples through its internal neural network structure.
[0068] Step 2.2: The discriminator receives the real data samples and the generated data samples , and at the same time receives the conditional variable , that is, the input of d = [ c , x ] , where is the overall input sample. After receiving the input vector , it outputs a probability value , indicating the authenticity of the data sample.
[0069] Step 2.3: Train the discriminator. First, fix the parameters of the generator and update the parameters of the discriminator. The goal of the discriminator is to maximize its loss function :
[0070] m a x L D = E c , x r [ D ( c , x r ) ] − E c , z [ D ( G ( c , z ) , c ) ] + λ E x ⌢ ~ P x ⌢ [ ( ∥ ∇ x ⌢ D ( c , x ⌢ ) ∥ 2 − 1 ) 2 ] ,
[0071] The loss function consists of two parts. One part is the discrimination result of the real data and the generated data, and the other part is the gradient penalty term, where is the expected value of the conditional variable and the real data, is the expected value of the conditional variable and the random noise, is The expected value of is a sample sampled from real data and generated data through linear interpolation. is the probability distribution of is the gradient penalty coefficient, which is used to control the weight of the gradient penalty term. is the gradient of the discriminator with respect to Specifically, randomly sample from the real dataset and the corresponding conditional variable and randomly sample a batch of from the standard normal distribution:
[0072] ,
[0073] where is the batch size, is the th real data sample, is the conditional variable corresponding to the th real data sample, is the th random noise vector. Use the generator to generate a batch of data samples :
[0074] ,
[0075] where is the th data sample generated by the generator. Sample a batch of samples from real data and generated data through linear interpolation :
[0076] ,
[0077] where is a random number uniformly distributed between 0 and 1, which is used to control the interpolation ratio, is the th sample obtained through linear interpolation, is the random number in the th real data sample. Finally, calculate the losses of the discriminator for real data, generated data, and interpolated samples, and update the parameters of the discriminator.
[0078] Step 2.4: Train the generator. First, fix the parameters of the discriminator and update the parameters of the generator. The goal of the generator is to minimize its loss function :
[0079] min L G =− E c , z [ D ( G ( c , z ) , c ) ] ,
[0080] Specifically, a batch of samples is randomly sampled from the standard normal distribution and the corresponding :
[0081] ,
[0082] Use a generator to generate a batch of data samples :
[0083] ,
[0084] Finally, calculate the loss of the generator and update the parameters of the generator. By alternately training the generator and the discriminator, the generator can finally generate data samples that meet the conditions.
[0085] Step 3: Divide the enhanced dataset into a training set and a test set, and use the Extreme Gradient Boosting XGB model for training. Optimize the hyperparameters of the BO model through Bayesian optimization to construct a CTGAN-XGB-BO prediction model.
[0086] Step 3.1: Randomly divide the enhanced dataset into a training set and a test set at a ratio of 8:2. The training set is used to train the XGB model, and the test set is used to evaluate the performance of the model. Among them is the number of samples, represents the feature vector of the th sample, represents the label of the th sample.
[0087] Step 3.2: Use the training set to train the XGB model. The XGB model consists of additive trees and is used to predict the output of the sample, which can be expressed as:
[0088] ,
[0089] where is the predicted output of the th sample, is the feature vector of the th sample, is the th decision tree, represents the set of regression trees, is the function that maps the data to the leaf nodes of the regression tree, is the weight of the leaf node. The objective function of XGB is:
[0090] ,
[0091] where is the loss function used to calculate The predicted value of a sample and the actual value The sum of the residuals between them is the regularization term The smaller the value, the better the tree structure. It is defined as
[0092] ,
[0093] where and are hyperparameters that control the complexity of the model structure is the number of leaf nodes is the weight parameter. XGB makes the model more accurate by adding a tree in each iteration. Therefore, in the th iteration, since the result of the th iteration has been obtained, the objective function can be transformed into
[0094] ,
[0095] where is the value added by the regression tree in the th iteration. Solving is equivalent to solving in the current state
[0096] Performing a second-order Taylor expansion on can determine the gain generated by the tree split after iteration. The larger the value, the more the objective function can be simplified and the better the model. Then the change in the loss function for this iteration can be expressed as
[0097] ,
[0098] where and are the first and second derivatives of the loss function respectively is the sample set of the left child node is the sample set of the right child node represents the sample set of the current node
[0099] Step 3.3: Apply Bayesian optimization BO to optimize the model. Define the hyperparameter optimization problem of the XGB model as the objective function, and set the coefficient of determination and the root mean square error as the objective function of the model. Its specific calculation expression is
[0100] ,
[0101] Among them is the mean of the true values. Set the learning rate, maximum depth, number of trees, regularization parameters alpha and lambda as hyperparameters. The process of BO optimizing the model hyperparameters is as follows:
[0102] ,
[0103] Among them is the objective function, is the prior distribution of, represents the posterior distribution of the objective function, is the sampling point information the marginal probability distribution of, is the hyperparameter at the th step, is the observation value at the th step, is the observation error, is the likelihood distribution of. First, construct a probabilistic surrogate model to approximately represent the behavior of the objective function. Then, based on the current surrogate model, use the acquisition function to search for candidate points in the hyperparameter space that are most likely to improve the performance of the objective function under the current model prediction, and apply the selected hyperparameter configuration to the actual objective function for evaluation and obtain the true performance metrics. Update the surrogate model using the newly obtained evaluation results. Repeat the above steps until the preset termination condition is reached.
[0104] Step 3.4: Extract the optimal hyperparameter combination from the Bayesian optimizer, retrain the XGB model using these hyperparameters, and evaluate the performance of the model using the test set.
[0105] Step 4: Interpret the trained and optimized CTGAN-XGB-BO bio-oil yield prediction model.
[0106] Step 4.1: Apply SHAP analysis to interpret the model. It helps to understand the decision-making process of the model by calculating the contribution of each feature to the prediction result. The calculation of SHAP values is based on the concept of Shapley values, which are used to allocate cooperative benefits. For each feature , its Shapley value represents the average marginal contribution of feature to the prediction result. The specific process is as follows:
[0107] Step 4.2: For a specific feature combination , calculate the difference in the model output with and without feature :
[0108] ,
[0109] where is the number of all features, is the model output using only the feature combination . is the model output using the feature combination and the feature .
[0110] Step 4.3: Perform weighted averaging on all possible feature combinations to calculate the Shapley value of the feature :
[0111] ,
[0112] where represents the size of the feature combination , represents the marginal contribution of the feature , thus obtaining the SHAP values of each feature, and these values represent the contribution of each feature to the prediction result.
[0113] Step 4.4: Display the SHAP values through a visualization tool. Through these visualization results, the features that have the greatest impact on the prediction of bio - oil yield can be identified, and the positive and negative impacts of each feature on the prediction result can be shown.
Claims
1. A bio-oil yield prediction method based on interpretable machine learning, characterized in that: The following steps are involved: Step 1: Obtain the bio-oil yield dataset of biomass raw material samples and perform preprocessing; Step 2: Enhance the preprocessed data by applying the Conditional Table Generative Adversarial Network (CTGAN). Step 3: Divide the enhanced dataset into training set and test set, use extreme gradient boosting XGB model for training, and build CTGAN-XGB-BO prediction model by optimizing the hyperparameters of BO model through Bayesian optimization; Step 4: Model interpretation of the trained and optimized CTGAN-XGB-BO bio-oil yield prediction model.
2. The method for predicting bio-oil yield based on interpretable machine learning according to claim 1, characterized in that: Step 2, as described in step 2, enhances the preprocessed data by applying the conditional table generation adversarial network CTGAN, which is specifically implemented as follows: Step 2.1: In CTGAN, the generator Receiving condition variables and a random noise vector , generate new data samples , select the characteristic variables pyrolysis temperature PT and gas purge flow rate FR in the data set as conditional variables, and define the conditional variables as , a random noise vector is a standard normal distribution with a mean of 0 and a variance of 1 , Receiving Input After that, new data samples are generated through its internal neural network structure. ; Step 2.2: Discriminator Receive real data samples and generate data samples , while receiving the condition variable , The input is ,in is the input sample of the population, When receiving the input vector , output a probability value , indicating the authenticity of the data sample; Step 2.3: Train the discriminator. The goal of the discriminator is to maximize its loss function : , The loss function consists of two parts: one is the discrimination result of real data and generated data, and the other is the gradient penalty term. is the expected value operator, is the gradient penalty coefficient, which is used to control the weight of the gradient penalty term. are samples sampled from the real data and generated data by linear interpolation, is the discriminator pair The gradients of are randomly sampled from the real data set. and the corresponding conditional variables and randomly sample a batch from a standard normal distribution : , in is the batch size, For the A real data sample, For the The conditional variables corresponding to the real data samples are For the A random noise vector is generated using the generator to generate a batch of data samples : , in The first data samples, and a batch of samples are sampled from the real data and the generated data by linear interpolation : , in is a random number uniformly distributed between 0 and 1, used to control the ratio of interpolation. is the value obtained by linear interpolation. samples, For the Random numbers in real data samples are finally calculated, and the discriminator's loss for real data, generated data, and interpolated samples is updated. Step 2.4: Train the generator. The goal of the generator is to minimize its loss function : , Specifically, a batch of random samples are sampled from a standard normal distribution and the corresponding : , Generate a batch of data samples using the generator : , Finally, the loss of the generator is calculated and the parameters of the generator are updated. By alternately training the generator and the discriminator, the generator can eventually generate qualified data samples.
3. The method for predicting bio-oil yield based on interpretable machine learning according to claim 1, characterized in that: In step 3, the enhanced data set is divided into a training set and a test set, and the extreme gradient enhancement XGB model is used for training. The CTGAN-XGB-BO prediction model is constructed by optimizing the hyperparameters of the BO model through Bayesian optimization. The specific implementation steps are as follows: Step 3.1: The enhanced dataset The data is divided into training set and test set in a ratio of 8:
2. is the number of samples, Indicates The feature vector of the samples, Indicates The labels of the samples; Step 3.2: Use the training set to train the XGB model. The XGB model consists of It is composed of an addition tree and is used to predict the output of the sample, which can be expressed as: , in For the The predicted output of samples is For the The feature vector of the samples, For the A decision tree, represents a collection of regression trees, is a function that maps data to regression tree leaf nodes, is the weight of the leaf node, and the objective function of XGB is: , in is the loss function, used to calculate The predicted value of samples With actual value The residual between is the regularization term, which is defined as: , in and is a hyperparameter that controls the complexity of the model structure. Points for leaves, As the weight parameter, XGB adds a tree in each iteration to make the model more accurate. In the first iteration, since we have obtained The result of the iteration, the objective function can be converted to: , in is the regression tree in The value increased in the iterations, solving This is equivalent to solving the current state ; right Performing a second-order Taylor expansion, we can determine the gain produced by the iterative tree splitting. The larger the value, the simpler the objective function is and the better the model performance is. The change in the loss function for this iteration is It can be expressed as: , in and are the first and second order derivatives of the loss function, respectively. is the sample set of the left child node, is the sample set of the right child node, Represents the sample set of the current node, and finally extracts the optimal hyperparameter combination from the Bayesian optimizer, retrains the XGB model with these hyperparameters, and uses the test set to evaluate the performance of the model.
4. The method for predicting bio-oil yield based on interpretable machine learning according to claim 1, characterized in that: The model interpretation of the trained and optimized CTGAN-XGB-BO bio-oil yield prediction model described in step 4 is specifically implemented in the following steps: Step 4.1: Apply SHAP analysis to explain the model. The calculation of SHAP value is based on the concept of Shapley value, which is used to distribute the benefits of cooperation. For each feature , its Shapley value Features The average marginal contribution to the predicted outcome; Step 4.2: For a specific feature combination , calculate the inclusion feature and does not contain features The model output difference is: , in is the number of all features, To use only feature combinations The model output is To use feature combinations And features The model output of Step 4.3: For all possible feature combinations Perform weighted averaging and calculate features The Shapley value of: , in Represents a combination of features The size of Representation characteristics The marginal contribution of is used to obtain the SHAP value of each feature.
Citation Information
Patent Citations
Generation method of wind and light typical operation scene based on improved generative adversarial network
CN118673327A
Photovoltaic abnormal data reconstruction method and system based on feature weighted local abnormal factors
CN119337069A
Biomass pyrolysis product prediction method based on machine learning model
CN119626362A
Bearing heat treatment quality prediction method combined with model interpretation technology
CN119808523A
Cited By
Foamed aluminum process parameter optimization method based on physical information deep learning
CN120656620A
Train operation situation prediction method based on intelligent system
CN120911765A
A train operation situation prediction method based on embodied intelligence
CN120911765B
Grouting data enhancement method and system for improving table generative adversarial network
CN121919457A
Improved grouting data augmentation method and system for table generating adversarial network
CN121919457B