Machine learning based prediction of carbon sequestration by iron salt promoted biochar and process optimization method

CN122598845APending Publication Date: 2026-08-18KUNMING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610690009.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-19
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

传统经验公式或统计回归往往难以捕捉原料组成、结构特征、催化剂种类与热解参数之间的高维非线性与交互效应,因此在异质原料与多催化体系下表现受限

Benefits of technology

本发明通过构建基于支持向量回归(SVR)等机器学习算法的预测模型,实现了对铁盐改性生物质热解固碳过程的精准量化与工艺寻优。首先,针对传统经验公式难以刻画原料种类(松木/水稻秸秆)、铁基添加剂类型(FeCl3·6H2O、FeSO4·7H2O、Fe(NO3)3·9H2O、Fe3O4、Fe2O3)与热解参数(温度、升温速率、保温时间)之间高维非线性交互作用的局限,本发明通过使用Optuna贝叶斯优化框架进行超参数自动寻优,SVR模型在测试集上拟合优度R²高达0.94、均方根误差RMSE仅3.00,显著优于传统回归及集成树模型,实现了对生物炭碳保留率的高精度预测。其次,引入SHAP可解释性分析揭示了影响固碳效率的关键特征排序:碳含量、产率、氢含量及H/C比为主导因子,灰分与挥发分次之,从而突破了过去“试错式”筛选催化剂的盲目性。最后,基于SHAP依赖图和部分依赖图明确了工艺参数的协同调控方向,验证了FeCl3·6H2O和Fe(NO3)3·9H2O在促进纤维素/半纤维素催化降解、提前裂解温区并提高残炭量方面的显著优势。本发明所提方法将机器学习与催化热解深度融合,为铁基添加剂的选择及热解工艺优化提供了理论依据与数据支撑,有效提升了生物炭的固碳潜力与制备效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122598845A_ABST
    Figure CN122598845A_ABST
Patent Text Reader

Abstract

The application discloses a kind of iron salt promotes carbon sequestration prediction and process optimization method based on machine learning, belongs to the technical field of biochar preparation process optimization.The method takes pine and rice straw as raw material, introduces FeCl3·6H2O, FeSO4·7H2O, Fe (NO3) 3·9H2O, Fe3O4, Fe2O3 five kinds of iron-based additives, carries out pyrolysis experiment under different pyrolysis temperature, heating rate and holding time condition, and establishes a variety of machine learning models based on experimental data to predict the carbon retention rate of biochar.The results show that the SVR model performs best, the comprehensive prediction performance of SVR model is best, and the R2 value reaches 0.9399, and RMSE=3.0022.Further analysis shows that the feature importance analysis shows that: carbon content> yield> hydrogen content> H / C> ash content> volatile matter.The application provides a new method for systematic prediction of the influence of iron-based additives on biochar retention rate, and has high application ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biochar preparation process optimization technology, specifically a machine learning-based method for predicting and optimizing the carbon fixation process of iron salt-promoted biochar. Background Technology

[0002] With the increasing severity of global climate change, reducing greenhouse gas emissions and enhancing carbon sequestration capacity have become important directions in environmental science research. Soil organic carbon is a crucial component of the global carbon cycle, and its stability directly affects atmospheric CO2 concentration and ecosystem carbon balance. Biomass, through oxygen-limited pyrolysis, can form carbon-rich solid products, retaining some of the raw material carbon in a relatively stable solid phase, thus endowing biochar with potential carbon sequestration functions. Biochar, as a carbon-rich material obtained from biomass pyrolysis, has received widespread attention due to its potential in soil improvement, pollutant adsorption, and carbon sequestration management. Its stable carbon structure not only helps mitigate atmospheric CO2 accumulation but also improves soil physicochemical properties.

[0003] However, during pyrolysis, a significant portion of the carbon in biomass is lost as gas and tar, limiting the carbon retention level and carbon fixation capacity of biochar. Therefore, improving the carbon retention rate during pyrolysis has become a key issue in biochar research. The carbon retention rate and structural properties of biochar are not fixed but are influenced by a combination of factors, including raw material composition, pyrolysis conditions, and exogenous additives. Therefore, accurately predicting and effectively controlling the carbon retention rate of biochar has become a core issue in this field of research and application. Factors affecting biochar carbon fixation include the biochar's own properties, such as biomass raw materials, pyrolysis temperature, and carbon content. Currently, research on the "systematic comparison of the effects of different iron-based additives on biochar carbon retention rate" is still limited, especially lacking systematic studies that simultaneously consider the interaction between raw material differences and catalyst types. Traditional empirical formulas or statistical regressions often fail to capture the high-dimensional nonlinearities and interaction effects between raw material composition, structural characteristics, catalyst type, and pyrolysis parameters, thus limiting their performance in heterogeneous raw materials and multi-catalytic systems.

[0004] Regarding the specific issue of "raw materials – iron-based additives – biochar carbon retention rate", related research has two problems: first, there is a lack of systematic datasets covering multiple raw materials and multiple catalysts; second, it is impossible to improve the interpretability of the model and clarify the relative contributions and interactions of various factors while ensuring the accuracy of prediction. Summary of the Invention

[0005] To address the aforementioned issues, this invention provides a machine learning-based method for predicting and optimizing carbon sequestration in iron-based biochar. This invention selects two representative biomass feedstocks and introduces five iron-based additives: FeCl3·6H2O, FeSO4·7H2O, Fe(NO3)3·9H2O, Fe3O4, and Fe2O3. Biochar is prepared under controlled pyrolysis conditions, and the carbon retention rate is measured. Based on this, various machine learning models are constructed and compared to evaluate their performance in predicting biochar carbon retention rate. Through feature importance analysis, the contribution of feedstock type and iron-based additive type to the difference in carbon retention rate is further explored.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a method for predicting and optimizing the carbon sequestration process of iron salt-promoted biochar based on machine learning, specifically including the following steps: S1. Based on pine sawdust, rice straw and five iron-based additives, raw data were obtained by combining orthogonal experiments and supplementary experiments to carry out oxygen-limited pyrolysis and construct a dataset. The raw data includes: biomass type, type of iron-based additive, iron addition ratio, pyrolysis temperature, heating rate, holding time, N2 flow rate, and biochar yield, moisture, ash, volatile matter, C content, H content, N content, S content, H / C atomic ratio, and carbon retention rate.

[0007] The input features of the dataset are biomass type, type of iron-based additive, iron addition ratio, pyrolysis temperature, heating rate, holding time, N2 flow rate, and biochar yield, moisture, ash, volatile matter, C content, H content, N content, S content, and H / C atomic ratio. The output variable is carbon retention rate.

[0008] S2. Perform preprocessing on the dataset, output the standardized dataset, perform feature analysis, and divide it into training set and test set according to a preset ratio. The preset ratio is specifically 9:1; The preprocessing operations include one-hot encoding of categorical variables and standardization of numerical features; The feature analysis specifically involves calculating the linear correlation coefficients between each input feature and between each feature and the carbon retention rate using the Pearson correlation test.

[0009] S3. Based on the training and test sets, the Optuna Bayesian optimization framework is adopted to define the hyperparameter search space of the machine learning model with the goal of minimizing the cross-validation RMSE, and automatically search and determine the optimal hyperparameter combination for each model.

[0010] S4. Based on the optimal combination of hyperparameters, 11 machine learning models were constructed and trained, including Random Forest, Extremely Random Tree, Gradient Boosting, AdaBoost, Decision Tree, XGBoost, SVR, Ridge Regression, Lasso, Elastic Network, and MLP. The predicted values ​​were output respectively. Based on the performance comparison table of R² and RMSE evaluation metrics and the box plot, SVR was determined to be the optimal training model.

[0011] S5. Based on the SVR training model and the standardized dataset, the SHAP method is used to calculate the average absolute SHAP value of each input feature. After drawing the feature importance ranking diagram and SHAP summary diagram, the nonlinear interaction effects between variables such as volatile matter and ash content, hydrogen content and heat preservation time, and metal valence state and heat preservation time are analyzed through SHAP dependency diagram and partial dependency diagram. Combined with experimental results, the effects of different iron salt types and concentrations on the yield and carbon retention rate of rice straw / pine biochar are compared. Feature importance ranking and optimization scheme of iron salt modified biochar preparation process are obtained, and the prediction and process optimization method of iron salt promoted biochar carbon fixation based on machine learning is completed.

[0012] Compared with existing technologies, this invention provides a machine learning-based method for predicting and optimizing the carbon sequestration process of iron salt-promoted biochar, which has the following beneficial effects: This invention achieves precise quantification and process optimization of the carbon sequestration process of iron-modified biomass pyrolysis by constructing a predictive model based on machine learning algorithms such as Support Vector Regression (SVR). Firstly, addressing the limitations of traditional empirical formulas in characterizing the high-dimensional nonlinear interactions between raw material types (pine / rice straw), iron-based additive types (FeCl3·6H2O, FeSO4·7H2O, Fe(NO3)3·9H2O, Fe3O4, Fe2O3), and pyrolysis parameters (temperature, heating rate, holding time), this invention uses the Optuna Bayesian optimization framework for automatic hyperparameter optimization. The SVR model achieves a goodness-of-fit R² of 0.94 and a root mean square error (RMSE) of only 3.00 on the test set, significantly outperforming traditional regression and ensemble tree models, thus realizing high-precision prediction of biochar carbon retention rate. Secondly, the introduction of SHAP interpretability analysis revealed the ranking of key features affecting carbon fixation efficiency: carbon content, yield, hydrogen content, and H / C ratio were the dominant factors, followed by ash and volatile matter, thus overcoming the blindness of past "trial and error" catalyst screening. Finally, based on SHAP dependency and partial dependency plots, the synergistic regulation direction of process parameters was clarified, verifying the significant advantages of FeCl3·6H2O and Fe(NO3)3·9H2O in promoting the catalytic degradation of cellulose / hemicellulose, advancing the pyrolysis temperature range, and increasing residual char. The method proposed in this invention deeply integrates machine learning with catalytic pyrolysis, providing a theoretical basis and data support for the selection of iron-based additives and the optimization of pyrolysis processes, effectively improving the carbon fixation potential and preparation efficiency of biochar. Attached Figure Description

[0013] Figure 1 This is a flowchart of a machine learning-based method for predicting and optimizing the carbon fixation process of iron salt-promoted biochar. Figure 2 The correlation network and correlation coefficient heatmap are shown below. The lower triangular heatmap shows the Pearson correlation coefficients between various process parameters, industrial analysis and elemental composition indicators, and the significance level is marked with an asterisk (*** p<0.001, ** p<0.01, * p<0.05). Figure 3 To illustrate the different machine learning models of this invention in terms of RMSE and R 2 Performance comparison charts under various metrics, where... Figure 3 (a) shows a performance comparison of different machine learning models under the RMSE metric. Figure 3 (b) shows different machine learning models in R 2 Performance comparison chart under the specified indicators; Figure 4 This is a comparison chart of the prediction performance of the SVR model of this invention on the training set and the test set, wherein, Figure 4(a) is a scatter plot of predicted and true values ​​in the training set. Figure 4 (b) is a scatter plot of predicted and actual values ​​in the test set; Figure 5 This is a graph showing the feature importance analysis results based on the SHAP method, where... Figure 5 In the middle (a), the overall ranking of the influence of features on the model output is shown. Figure 5 (b) SHAP summary diagram (beeswarm); Figure 6 The graph shows the dependencies of the three most significant sets of variables in the variable interactions on SHAP. Figure 6 In the middle (a), volatile matter and ash are represented. Figure 6 (b) represents the hydrogen content and holding time. Figure 6 In the middle (c), the heat preservation time and the metal valence state are represented. Figure 6 In the middle (d), the partial dependence diagram of volatile matter and ash content in three dimensions is shown. Figure 6 (e) is a partial dependence plot of hydrogen content and holding time. Figure 6 In the middle (f), the partial dependence of metal valence state and holding time is shown. Figure 7 A graph showing biochar yield and carbon retention, where, Figure 7 The middle (a) graph shows the rice yield. Figure 7 (b) is a graph showing the yield of pine wood. Figure 7 The middle (c) diagram shows the carbon retention of rice. Figure 7 (d) shows the charcoal retention diagram of pine wood; Figure 8 The graph shows the biomass pyrolysis behavior under different iron-based catalytic methods. Figure 8 In the middle (a), the thermogravimetric curve is shown. Figure 8 In the middle (b), the differential curve is shown. Detailed Implementation

[0014] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0015] Please see Figures 1-8 A machine learning-based method for predicting and optimizing the carbon sequestration process of iron salt-promoted biochar includes the following steps: S1. Based on pine sawdust, rice straw and five iron-based additives, raw data were obtained by combining orthogonal experiments and supplementary experiments to carry out oxygen-limited pyrolysis and construct a dataset. The raw data includes: biomass type, type of iron-based additive, iron addition ratio, pyrolysis temperature, heating rate, holding time, N2 flow rate, and biochar yield, moisture, ash, volatile matter, C content, H content, N content, S content, H / C atomic ratio, and carbon retention rate.

[0016] The input features of the dataset are biomass type, type of iron-based additive, iron addition ratio, pyrolysis temperature, heating rate, holding time, N2 flow rate, and biochar yield, moisture, ash, volatile matter, C content, H content, N content, S content, and H / C atomic ratio. The output variable is carbon retention rate.

[0017] The formula for calculating the yield of biochar is as follows: The formula for calculating the carbon retention rate is: Where, m a For the mass of dried biochar after pyrolysis, m b For the quality of the raw material before pyrolysis (dried basis), Y char For biochar yield, C solid and C biomass The carbon content on a dry basis is for both solid products and biomass. Raw materials: rice straw and pine sawdust, both pre-treated by drying and crushing (<2mm). Five iron-based additives: (FeCl3·6H2O, FeSO4·7H2O, Fe(NO3)3·9H2O, Fe3O4, Fe2O3). The loading amounts based on iron mass were uniformly 0.5wt% and 1wt%. The iron-based additives were prepared as solutions or uniformly suspended, then thoroughly mixed with the raw materials, stirred at room temperature for 2 hours, and then dried for later use. The mixture was dried in a forced-air oven for 24 hours, and then pyrolyzed using a box-type resistance furnace. A three-factor, three-level L9 (3) model was used. 3 An orthogonal experimental design was used to investigate the effects of heating rate (10, 15, 20°C / min), holding time (30, 60, 90 min), and pyrolysis temperature (400, 600, 800°C) on carbon retention.

[0018] Table 1: Factor Level Table for Orthogonal Experiment Building upon orthogonal experiments, several non-orthogonal pyrolysis experiments were conducted to further expand the experimental parameter space and construct the dataset required for machine learning models. These included experiments with different combinations of pyrolysis temperatures and other parameters under fixed heating rates or holding times. All experimental conditions were controlled within the parameter range set by the orthogonal design. Therefore, this study employs a combination of orthogonal experiments and supplementary experiments to construct an experimental dataset containing various pyrolysis parameters, raw material types, and iron-based additive modification conditions for subsequent training and validation of machine learning models.

[0019] The method for determining the moisture, ash, and volatile matter of biomass is as follows: Place approximately 2.0000 g of sample in a 60 mm × 30 mm weighing bottle, place it in an oven at 105 ± 2°C for 4 h, remove it and place it in a silica gel desiccator, cool it to room temperature, weigh it and record the mass, and calculate the moisture content of the sample based on the difference in mass before and after.

[0020] The method for determining volatile matter is as follows: Weigh approximately 2.0000 g of the sample and place it in a 30 mL crucible. Place the crucible in a muffle furnace at 950±25°C and measure for 7 min under a nitrogen atmosphere. Cool to room temperature, weigh and record the mass. Calculate the volatile matter content of the sample based on the difference in mass and the moisture content.

[0021] The method for determining ash content is as follows: Weigh approximately 2.0000 g of the sample and place it in a 30 mL crucible. Place the crucible in a muffle furnace at 575 ± 25 ℃ for 4 h, cool it to room temperature in a silica gel desiccator, weigh and record the mass. Calculate the ash content of the sample based on the mass difference and the moisture content. Three parallel samples are set up for each batch of experiments.

[0022] Elemental analysis of biochar: The percentage content of C, H, N, and S in the sample particles was determined using an elemental analyzer (vario Mico Cube, Elementar, Germany), and the atomic ratio of H / C was calculated based on the percentage content of each element. The specific procedure was as follows: approximately 2 mg of sample was weighed into a special tin boat, pressed into shape, and placed sequentially into a sample plate, with three replicates for each sample. During the determination, the combustion furnace temperature was set to 1150°C, and the reduction furnace temperature was set to 850°C.

[0023] Thermogravimetric analyzer was used to heat the sample to 800 K at a heating rate of 10 K / min. High-purity nitrogen (99.999%) was used as the carrier gas at a flow rate of 20 mL / min. The mass of the thermogravimetric sample was 10 ± 0.5 mg, and each batch of experiments was repeated three times. Before the pyrolysis reaction, the reaction cell was purged with high-purity nitrogen for 30 min to remove air from the apparatus.

[0024] S2. Perform preprocessing on the dataset, output the standardized dataset, perform feature analysis, and divide it into training set and test set according to a preset ratio. The preset ratio is specifically 9:1; The preprocessing operations include one-hot encoding of categorical variables and standardization of numerical features; The feature analysis specifically involves calculating the linear correlation coefficients between each input feature and between the feature and the carbon retention rate using the Pearson correlation test. The experimental parameters were as follows: heating rate: 10, 15, 20 °C / min; pyrolysis temperature: 400, 600, 800 °C; holding time: 30, 60, 90 min. Experiments were conducted under nitrogen protection (flow rates of 250 mL / min and 100 mL / min) according to the above combinations. Data are shown in Supplementary Material Table S1. Feature coding: The types of raw materials and iron-based additives were encoded using unique thermal coding; numerical features (heating rate, temperature, time) were standardized. Data splitting: Data were randomly divided into a training set (90%) and a test set (10%).

[0025] To ensure the effectiveness of model development, all raw experimental data were first randomly rearranged and then randomly divided into a training set (90%) and a test set (10%) in a 9:1 ratio. The training set was used for model training and selection of optimal hyperparameters, while the test set was used to independently evaluate model performance.

[0026] like Figure 2 As shown, there are different degrees of correlation between each variable and carbon retention rate. Among them, the yield is significantly positively correlated with the carbon retention rate (r = 0.52, p<0.001), that is, a higher yield usually corresponds to a higher carbon retention rate. This result is basically consistent with existing studies. Nan et al.

[13] pointed out that the addition of exogenous Ca can simultaneously improve the biochar yield and carbon retention rate at different pyrolysis temperatures, and increase the carbon retention rate from 49.2%-68.3% of the original biochar to 66.1%-79.7%, which shows that improving the production and retention capacity of solid carbon products may be an important way to enhance the carbon retention rate. In addition, H content (r = 0.39, p<0.001) and H / C ratio (r = 0.28, p<0.001) are also positively correlated with carbon retention rate, while volatile matter content (r = 0.38, p<0.001) and temperature (r = A significant negative correlation was observed (r = 0.31, p < 0.001), indicating that higher temperatures may decrease carbon retention. A very strong positive correlation was found between volatile matter and ash (r = 0.91, p < 0.001), while both showed a significant negative correlation with yield. Some correlations also existed among elemental composition variables: H and H / C (r = 0.88) and C and volatile matter (r = 0.81). Overall, this correlation analysis suggests that yield, elemental composition (H, H / C), and volatile matter content may be important factors affecting carbon retention. Due to strong collinearity among some variables, the Pearson correlation coefficient was only used for preliminary screening of variable associations and does not represent independent contributions.

[0027] S3. Based on the training and test sets, the Optuna Bayesian optimization framework is adopted to define the hyperparameter search space of the machine learning model with the goal of minimizing the cross-validation RMSE, and automatically search and determine the optimal hyperparameter combination for each model. The optimal combination of hyperparameters for the model is as follows: Random Forest (n_estimators = 340, max_depth = 14, min_samples_split =2, min_samples_leaf = 1); Extra Trees (n_estimators = 227, max_depth = 17, min_samples_split = 3); Gradient Boosting (n_estimators = 309, max_depth = 2, learning_rate =0.0696, subsample = 0.6201); AdaBoost(n_estimators = 200, learning_rate = 1.9575); SVR (kernel = poly, C = 19.0380, gamma = scale, epsilon = 0.4865, degree =2); XGBoost (n_estimators: 462, max_depth: 2, learning_rate:0.09944758051555969, 'subsample': 0.6091368096230909, colsample_bytree:0.9172139290144911, reg_alpha: 0.09747003149291002, 'reg_lambda':1.697747255296283e-06); Ridge (alpha = 1.3237); Lasso(alpha = 0.0286); ElasticNet (alpha = 0.0305, l1_ratio = 0.7871); MLP (2 hidden layers, 144 neurons per layer, activation = relu, alpha = 0.00037, learning_rate = constant).

[0028] The optimal hyperparameters for the Random Forest model are: 340 decision trees, a maximum depth of 14 layers, a minimum number of samples required for internal node subdivision of 2, and a minimum number of samples required for leaf nodes of 1. The optimal hyperparameters for the Extra Trees model are: 227 decision trees, a maximum depth of 17 layers, and a minimum number of samples required for internal node subdivisions of 3. The optimal hyperparameters for the Gradient Boosting model are: 309 weak learners, a maximum depth of 2 layers, a learning rate of 0.0696, and a subsampling ratio of 0.6201. The optimal hyperparameters for the AdaBoost model are: 200 weak learners and a learning rate of 1.9575. The optimal hyperparameters for the Support Vector Regression (SVR) model are: a polynomial kernel, a penalty coefficient C of 19.0380, a kernel coefficient gamma calculated automatically using the "scale" method, an insensitive interval epsilon of 0.4865, and a polynomial degree of 2. The optimal hyperparameters for the K-Nearest Neighbors (KNN) model are: 5 nearest neighbors, weights based on distance, and Manhattan distance as the distance metric. The optimal hyperparameters for the Ridge regression model are: regularization intensity alpha = 1.3237; The optimal hyperparameter for the Lasso regression model is: regularization strength alpha = 0.0286; The optimal hyperparameters for the ElasticNet model are: regularization strength alpha = 0.0305, L1 regularization ratio l1_ratio = 0.7871; The optimal hyperparameters for the multilayer perceptron (MLP) model are: two hidden layers, each with 144 neurons, ReLU activation function, regularization intensity alpha = 0.00037, and constant learning rate. S4. Based on the optimal combination of hyperparameters, 11 machine learning models were constructed and trained, including Random Forest, Extreme Random Tree, Gradient Boosting, AdaBoost, Decision Tree, XGBoost, SVR, Ridge Regression, Lasso, Elastic Network, and MLP. The predicted values ​​were output respectively. Based on the performance comparison table of R² and RMSE evaluation metrics and the box plot, SVR was determined to be the optimal training model. This invention compares the predicted values ​​with the corresponding experimental true values ​​and uses the goodness of fit (R2) and root mean square error (RMSE) to evaluate the model's performance.

[0029] (1) Root Mean Square Error (RMSE) The root mean square error (RMSE) represents the square root of the ratio of the sum of squares of the differences between predicted and true values ​​to the number of observations, n. RMSE is highly sensitive to the maximum or minimum errors in a set of measurements, thus providing a good reflection of the precision of the prediction results. The calculation formula is as follows: (2) Fit coefficient (R) 2 ) R is the ratio of the regression sum of squares to the total sum of squares, representing the proportion of the total sum of squares that can be explained by the regression sum of squares. The larger its value, the better the model's predictive accuracy and regression performance. 2 The value of is between 0 and 1. The closer it is to 1, the better the regression fit. When the value is greater than 0.8, we consider the model to have a good fit. The calculation formula is as follows: like Figure 3 As shown in Table 2, different regression models exhibit significant performance differences on the training and test sets. Overall, SVR performs best on the test set, with a test set RMSE of 3.0021 and R0. 2 = 0.9376, which is better than other models, indicating that it has a strong predictive ability for carbon retention rate. Ridge, Lasso and ElasticNet are next, with their test set R... 2All three exceeded 0.91, specifically 0.9162, 0.9128, and 0.9128, and the RMSE remained between 3.48 and 3.55, demonstrating good stability and generalization ability.

[0030] In contrast, while some tree models perform exceptionally well on the training set, their performance drops significantly on the test set. For example, Extra Trees, Random Forest, and Decision Tree show significantly lower R-values ​​on the training set. 2 The scores reached 0.9970, 0.9701, and 0.9191 respectively, but the test set R... 2 The RMSE values ​​were only 0.7844, 0.8020, and 0.7090 respectively, and the test errors were relatively large, indicating that these models have varying degrees of overfitting tendency. Among them, Decision Tree had the highest RMSE on the test set (6.4835), but the weakest generalization ability.

[0031] Furthermore, XGBoost and Gradient Boosting perform relatively well in ensemble learning models, with good performance on the test set R. 2 The RMS values ​​were 0.8828 and 0.8761 respectively, and the RMSE values ​​on the test set were 4.1146 and 4.2310 respectively. This indicates that Boosting-type models can balance fitting ability and generalization performance to some extent, but overall they are still inferior to SVR and regularized linear models. The MLP's test set performance was generally poor (RMSE = 5.6154, R...). 2 = 0.7817), indicating that under the current sample size and feature conditions, the neural network model did not show a significant advantage.

[0032] Table 2: Performance comparison of different regression models on the training and test sets. like Figure 4 As shown, the SVR model's predicted values ​​almost perfectly match the true values ​​on the training set, with the scatter points basically distributed along y=x, exhibiting R0. 2 =1.0000, RMSE=0.0011. This result shows that the SVR model has a very strong fitting ability to the training samples and can learn the mapping relationship between input features and target variables with high accuracy.

[0033] On the test set, although the scatter distribution is slightly more discrete than that on the training set, it still closely follows the y=x distribution, achieving high prediction performance (R²). 2=0.9376, RMSE=3.0022). This indicates that the SVR model can still maintain good predictive stability and accuracy on unknown samples, and has strong generalization ability. Overall, SVR can accurately characterize the changing trend of the target variable and maintain high prediction accuracy on the test set.

[0034] S5. Based on the SVR training model and the standardized dataset, the SHAP method is used to calculate the average absolute SHAP value of each input feature. After drawing the feature importance ranking chart and SHAP summary chart, the nonlinear interaction effects between variables such as volatile matter and ash content, hydrogen content and heat preservation time, and metal valence state and heat preservation time are analyzed through SHAP dependency chart and partial dependency chart. Combined with the experimental results, the effects of different iron salt types and concentrations on the yield and carbon retention rate of rice straw / pine biochar are compared. Feature importance ranking and iron salt modified biochar preparation process optimization scheme are obtained, and the machine learning-based iron salt-promoted biochar carbon fixation prediction and process optimization method is completed. like Figure 5 As shown, SHAP analysis reveals the differences in the contribution of each feature to the model output from two aspects: global importance ranking and local effect distribution. Figure 5 (a) shows that the average absolute |SHAP| value of C content is the highest, reaching 7.61, significantly higher than the other variables, indicating that it has the strongest explanatory power in carbon retention rate prediction and is the dominant factor driving changes in the model output. C content is not only an important indicator characterizing the properties of biochar, but may also be a core factor affecting the carbon retention rate prediction results. The average absolute |SHAP| value of yield is the second highest, at 5.83. This indicates that in the current model, feedstock carbon content and yield play a dominant role in carbon retention rate prediction.

[0035] Looking at other characteristics, the average absolute |SHAP| values ​​for H content, H / C ratio, and ash content were 2.06, 1.51, and 1.18, respectively, indicating that these elemental composition characteristics have some influence on carbon retention rate prediction, but their contribution is significantly lower than that of C content and yield. Volatile matter and moisture content are relatively less important, but still contribute to the model output. In contrast, process parameters such as addition ratio, holding time, N2 flow rate, and heating rate are all less important, with the heating rate being the least important (0.07), indicating that within the data range of this study, these variables have limited marginal contributions to the model output. The importance of raw material category variables is also low, indicating that compared to raw material type, the physicochemical properties of the raw material, especially its elemental composition characteristics, are more critical for carbon retention rate prediction.

[0036] Combination Figure 5(b) Further analysis shows that high values ​​for C content and yield mainly correspond to positive SHAP values, indicating that increases in these two parameters typically promote an increase in carbon retention. Conversely, high values ​​for volatile matter and ash content are more prevalent in the negative SHAP value region, suggesting a generally negative correlation with carbon retention. The H / C ratio and H content show a certain positive contribution trend overall, but the impact is relatively small. In general, the SHAP analysis results are largely consistent with the correlation analysis conclusions, further validating the model's reliability in identifying key driving factors and providing a basis for subsequent mechanistic explanations and variable selection.

[0037] like Figure 6 As shown, SHAP dependency analysis reveals the nonlinear influence of key variables on the prediction results. With increasing volatile matter content, the SHAP value generally shows a decreasing trend (…). Figure 6 (a) indicates that the increase in volatile matter content has an inhibitory effect on the target variable and there is a certain interaction relationship with ash content. H content is significantly positively correlated with the prediction results ( Figure 6 In (b), the SHAP value increased significantly with increasing H content, indicating that it promotes carbon retention, and this effect is synergistically enhanced with holding time. Holding time ( Figure 6 (c) shows a phased impact characteristic, contributing positively to the predicted value in a short time interval, while the SHAP value tends to be negative over a longer period, indicating that excessively long heat preservation time may be detrimental to improving the target variable. The three-dimensional response surface further verifies the interaction effects between variables. Figure 6 As shown in (d), the volatile matter content and ash content together affect the carbon retention rate, with higher predicted values ​​under conditions of lower volatile matter and moderate ash content. Figure 6 (e) indicates that there is a significant nonlinear coupling relationship between H content and heat preservation time, and the carbon retention rate reaches its peak under the conditions of medium heat preservation time and high H content. Figure 6 (f) shows that there is also an interaction between the metal valence state and the holding time; the carbon retention rate is higher under conditions of shorter holding time and higher metal valence state. Overall, this figure systematically reveals the nonlinear influence and interaction mechanism of key physicochemical indicators and process parameters on the target variable, indicating that the model can not only capture single-factor effects but also multi-variable synergistic effects, providing a visual basis for process optimization.

[0038] like Figure 7 As shown, after adding different iron-based additives, the yield and carbon retention rate of rice and pine biochar showed similar fluctuation trends with changes in experimental conditions, but there were significant differences in yield levels between different treatments. Figure 7(a) This demonstrates the effect of different iron salts (FeCl3·6H2O, FeSO4·7H2O, Fe(NO3)3·9H2O, Fe3O4, Fe2O3) and their concentrations (0.5 and 1 mass ratio) on yield (%) when using rice as a raw material. Pine wood generally has a higher yield than rice; overall, woody plants may have a higher solids yield than herbaceous plants. The yield of rice decreases with increasing temperature, which is consistent with the previous... Figure 2 The Pearson correlation yielded consistent conclusions. Figure 7 (c) Effects of different iron salts (FeCl3·6H2O, FeSO4·7H2O, Fe(NO3)3·9H2O, Fe3O4, Fe2O3) and their concentrations (0.5 and 1 mass ratio) on carbon retention when using rice as a raw material. The carbon retention rates of rice and pine wood decreased with decreasing temperature. Figure 7 (d) Regarding carbon retention, pine wood also showed a related trend after adding different iron-based additives.

[0039] Figure 8 These are the thermogravimetric curves and corresponding differential (DTG) curves of rice and pine wood catalyzed by different concentrations of ferric chloride. Figure 8 (a) For pyrolysis without iron salts, less than 2 wt.% mass loss was observed at 260 °C, which is a deep drying process. Significant mass loss occurred in the temperature range of 220–375 °C, leaving only 35 wt.% residue. This was due to the decomposition of hemicellulose and cellulose. Slower weight loss was then observed above 375 °C, due to the slow pyrolysis and carbonization of lignin. Significant weight loss was observed around 250 °C when FeCl3 was introduced, as FeCl3 promotes the depolymerization of cellulose and hemicellulose, accompanied by the release of some water and CO2. However, more residue was retained at 800 °C; of which, approximately 34–36 wt.% residue was retained for ferric chloride, compared to 30 wt.% for the original biomass. The increase in residue could be due to two possible reasons: the introduction of additional, non-volatile iron salts, and the iron salts promoting carbonization and increasing carbon yield.

[0040] Figure 8 As shown in (b), in the differential curve of the original biomass pyrolysis, there is a peak corresponding to cellulose pyrolysis at 348℃ and a shoulder peak related to hemicellulose pyrolysis at 290℃. When FeCl3 is introduced, the weight loss peaks related to cellulose and hemicellulose pyrolysis are advanced to 332℃, which may be due to the decrease in the crystallinity of biomass during impregnation and the fact that iron promotes the breaking of glycosidic bonds. Chloride salts may have a greater ability to disrupt the crystal structure of biomass and promote the removal of internal moisture.

[0041] It should be noted that the above are merely preferred embodiments of this application and do not limit the scope of patent protection of this application. Any equivalent structural or procedural changes made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of this application.

Claims

1. A machine learning-based method for predicting and optimizing the carbon sequestration process of iron salt-promoted biochar, characterized in that, Includes the following steps: S1. Based on pine sawdust, rice straw and five iron-based additives, raw data were obtained by combining orthogonal experiments and supplementary experiments to carry out oxygen-limited pyrolysis and construct a dataset. The raw data includes: biomass type, type of iron-based additive, iron addition ratio, pyrolysis temperature, heating rate, holding time, N2 flow rate, and biochar yield, moisture, ash, volatile matter, C content, H content, N content, S content, H / C atomic ratio, and carbon retention rate. The input features of the dataset are biomass type, type of iron-based additive, iron addition ratio, pyrolysis temperature, heating rate, holding time, N2 flow rate, and biochar yield, moisture, ash, volatile matter, C content, H content, N content, S content, H / C atomic ratio. The output variable is carbon retention rate. S2. Perform preprocessing on the dataset, output the standardized dataset, perform feature analysis, and divide it into training set and test set according to a preset ratio. S3. Based on the training and test sets, the Optuna Bayesian optimization framework is adopted to define the hyperparameter search space of the machine learning model with the goal of minimizing the cross-validation RMSE, and automatically search and determine the optimal hyperparameter combination for each model. S4. Based on the optimal combination of hyperparameters, 11 machine learning models were constructed and trained, including Random Forest, Extreme Random Tree, Gradient Boosting, AdaBoost, Decision Tree, XGBoost, SVR, Ridge Regression, Lasso, Elastic Network, and MLP. The predicted values ​​were output respectively. Based on the performance comparison table of R² and RMSE evaluation metrics and the box plot, SVR was determined to be the optimal training model. S5. Based on the SVR training model and the standardized dataset, the SHAP method is used to calculate the average absolute SHAP value of each input feature. After drawing the feature importance ranking diagram and SHAP summary diagram, the nonlinear interaction effects between variables such as volatile matter and ash content, hydrogen content and heat preservation time, and metal valence state and heat preservation time are analyzed through SHAP dependency diagram and partial dependency diagram. Combined with experimental results, the effects of different iron salt types and concentrations on the yield and carbon retention rate of rice straw / pine biochar are compared. Feature importance ranking and optimization scheme of iron salt modified biochar preparation process are obtained, and the prediction and process optimization method of iron salt promoted biochar carbon fixation based on machine learning is completed.

2. The method for predicting and optimizing the carbon sequestration process of iron salt-promoted biochar based on machine learning according to claim 1, characterized in that, In S1, the iron-based additives are loaded at 0.5 wt% and 1 wt% by weight of iron, respectively. The orthogonal experiment adopts a three-factor, three-level L9 (3 3 The pyrolysis temperatures are 400°C, 600°C, and 800°C, the heating rates are 10°C / min, 15°C / min, and 20°C / min, and the holding times are 30 min, 60 min, and 90 min, respectively.

3. The method for predicting and optimizing the carbon sequestration process of iron salt-promoted biochar based on machine learning according to claim 1, characterized in that, In S1, the formula for calculating the carbon retention rate is: in, For dry basis solids yield, C solid and C biomass The figures represent the dry basis carbon content of solid products and biomass, respectively.

4. The method for predicting and optimizing the carbon sequestration process of iron salt-promoted biochar based on machine learning according to claim 1, characterized in that, In S2, the preset ratio is specifically 9:1; the preprocessing operation includes one-hot encoding of categorical variables and standardization of numerical features. The feature analysis specifically involves calculating the linear correlation coefficients between each input feature and between each feature and the carbon retention rate using the Pearson correlation test.

5. The method for predicting and optimizing the carbon sequestration process of iron salt-promoted biochar based on machine learning according to claim 1, characterized in that, In S3, the optimal hyperparameters for the Support Vector Regression (SVR) model in the optimal hyperparameter combination are: a polynomial kernel is selected as the kernel function, the penalty coefficient C is 19.0380, the kernel coefficient gamma is calculated automatically using the "scale" method, the insensitive interval epsilon is 0.4865, and the polynomial degree is 2.

6. The method for predicting and optimizing the carbon sequestration process of iron salt-promoted biochar based on machine learning according to claim 1, characterized in that, In step S5, the nonlinear interaction effect analysis between variables includes: the interaction effect between volatile matter and ash content, the interaction effect between hydrogen content and holding time, and the interaction effect between metal valence state and holding time.

7. The method for predicting and optimizing the carbon sequestration process of iron salt-promoted biochar based on machine learning according to claim 1, characterized in that, In S5, the optimization results of the iron salt modified biochar preparation process are as follows: FeCl3·6H2O and Fe(NO3)3·9H2O have the most significant effects on improving the yield and carbon retention rate of rice straw and pine wood. The carbon retention rate is higher under shorter holding time and higher metal valence state conditions, and reaches its peak under medium holding time and higher hydrogen content conditions.