Process optimization method for separating levoglucosenone from bio-oil

By constructing a lightweight gradient boosting machine learning model, the process optimization and effect prediction problems of LGO separation and purification from complex bio-oil by selective pyrolysis of biomass were solved, efficient and accurate LGO separation was achieved, the purity and yield of bio-based chemicals were improved, an online quality monitoring tool was provided, and the intelligent preparation process of bio-based chemicals was promoted.

CN120673878APending Publication Date: 2025-09-19NORTH CHINA ELECTRIC POWER UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510764581.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

In the existing separation and purification technology of complex bio-oil obtained by selective pyrolysis of biomass, there is a lack of means to optimize the process of LGO and predict the separation effect, especially the lack of quantitative characterization of the dynamic influence mechanism of acidic components. As a result, the traditional separation process has high energy consumption and limited purity improvement threshold, making it difficult to meet the purity requirements of pharmaceutical-grade products.

Method used

A lightweight gradient boosting machine learning model was constructed, and data was processed by combining multiple interpolation, Mahalanobis distance outlier detection, principal component analysis, and one-hot encoding. A multi-dimensional experimental parameter database was constructed, and the model hyperparameters were optimized through an adaptive weight distribution mechanism to achieve process optimization and effect prediction for the distillation separation and purification of LGO from complex bio-oils by selective pyrolysis of biomass.

Benefits of technology

It achieves the rapid and accurate separation and purification of complex bio-oil obtained from the selective pyrolysis of biomass, improves the purity and yield of LGO, provides an online quality monitoring tool, guides the dynamic optimization of separation process parameters, and improves the intelligence level of the bio-based chemical preparation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673878A_ABST
    Figure CN120673878A_ABST
Patent Text Reader

Abstract

The invention relates to a process optimization method for separating levoglucosenone from bio-oil. The method comprises the following steps: firstly, collecting experimental data in a process of rectifying, separating and purifying LGO (Levoglucosenone) from complex bio-oil obtained by selective pyrolysis of biomass; performing standardization processing on the original data to form a processed data set; and finally, constructing a lightweight gradient lifting machine learning framework, balancing the training set and the verification set through a self-adaptive weight distribution mechanism, and realizing process optimization and effect accurate prediction of complex bio-oil rectification separation and purification of the LGO. By constructing a multi-dimensional experimental parameter database and an intelligent prediction model, a novel digital tool is provided for online quality monitoring of the LGO separated and purified from the complex bio-oil obtained by biomass selective pyrolysis, and dynamic optimization of separation process parameters can be guided through real-time prediction feedback; the method has important engineering application value for improving the intelligent level of the bio-based chemical preparation process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of separation and purification of bio-based chemicals, and particularly relates to a process optimization method for separating levoglucosone from bio-oil. Background Art

[0002] Thermochemical conversion systems of biomass have become a research focus due to their closed carbon cycle characteristics. Among them, the rapid pyrolysis technology of lignocellulosic biomass can achieve the directional conversion of biomass components into high-value-added chemicals due to its short reaction time and the synergistic effect of medium and high temperatures. As a characteristic component of pyrolysis products, levulinic ketone (LGO) has irreplaceable application value in the fields of pharmaceutical intermediate synthesis and asymmetric catalysis due to its unique pyran ring chiral structure and multi-functional group characteristics. However, existing industrial production still relies on a multi-step chemical synthesis route using galactose as raw material, facing problems such as high raw material costs, high process complexity, and limited yield, which seriously restricts its large-scale application.

[0003] Although biomass selective catalytic pyrolysis technology has significant advantages in terms of raw material economy and process simplicity, its industrialization process is still limited by the problem of product separation and purification. The complex bio-oil obtained by selective pyrolysis of biomass is a multi-component mixture. In addition to the target product LGO, it also contains a large number of organic acids, phenols, furan derivatives and other compounds. The physical and chemical properties of these components are highly similar, resulting in low efficiency of traditional distillation separation. More importantly, the coexistence of acidic components and LGO will trigger side reactions such as hydroxymethyl condensation, causing degradation of the target product. Although existing separation processes (such as multi-stage distillation, solvent extraction, etc.) can achieve initial enrichment, they generally have technical defects such as high energy consumption intensity and limited purity improvement threshold, making it difficult to meet the purity requirements of pharmaceutical-grade products.

[0004] In recent years, machine learning methods based on multimodal data fusion have provided a new paradigm for the optimization of complex chemical separation processes. Compared with traditional mechanistic models based on thermodynamic equilibrium, ensemble learning algorithms represented by light gradient boosting machine (lightGBM) have shown unique advantages in the optimization of industrial process parameters through feature importance analysis and nonlinear relationship modeling. This technology can construct a predictive model for process optimization and target product separation effect by integrating multi-dimensional variables such as the physical and chemical properties of complex bio-oil obtained by selective pyrolysis of biomass, pretreatment methods, and distillation tower operating conditions. The portability of its model parameters makes it possible to quickly deploy different production units. However, existing research has not yet established a dedicated model for the process optimization and prediction of separation effect of LGO in the complex bio-oil separation and purification system obtained by selective pyrolysis of biomass. In particular, there is a lack of quantitative characterization of the dynamic influence mechanism of acidic components, which limits the applicability of the predictive model in real industrial scenarios. Summary of the Invention

[0005] The present invention addresses the lack of process optimization and separation effect prediction methods for LGO in existing technologies for separating and purifying complex bio-oils obtained by selective pyrolysis of biomass. It provides a process optimization method for separating L-glucose ketone from bio-oil, which can more quickly and accurately achieve process optimization and effect prediction for separating and purifying LGO from complex bio-oils obtained by selective pyrolysis of biomass.

[0006] To solve the above technical problems, an embodiment of the present invention provides a process optimization method for separating levoglucosone from bio-oil, comprising the following steps:

[0007] S1: Collect experimental data from the distillation separation and purification of LGO from complex bio-oil obtained by selective pyrolysis of biomass, including the physicochemical properties of the bio-oil, pretreatment methods, distillation tower operating conditions, and separation effects, to form the original experimental data set. The physicochemical properties of the bio-oil include LGO content, moisture content, acidity, viscosity, sediment, and main component composition. The pretreatment methods include deacidification and extraction. The deacidification process parameters include the type and dosage of the deacidification reagent, and the extraction parameters include the type and dosage of the extractant. The distillation tower operating conditions include tower height, tower diameter, temperature, vacuum degree, packing type, packing density, reflux ratio, theoretical plate equivalent height, and feed position. The separation effect includes LGO yield and purity.

[0008] S2: Data processing was performed on the original experimental data set. Missing values ​​were filled using multiple interpolation, outliers were detected and eliminated based on the Mahalanobis distance, and principal component analysis was used to eliminate multicollinearity interference. The string data of the deacidification reagent type, extractant type, and distillation tower packing type were converted using one-hot encoding, ultimately forming a processed data set consisting of numerical data.

[0009] S3: Based on the processed data set, 80% of it is randomly selected as the training set and the remaining 20% ​​is used as the test set. A lightweight gradient boosting lightGBM machine learning model is constructed. An adaptive weight distribution mechanism is introduced to balance the training set and the training set. Model hyperparameters are screened based on the training set data. Bayesian optimization and progressive cross-validation methods are used for hyperparameter screening. The root mean square error is used as the evaluation indicator to screen the optimal hyperparameters. The model is then trained based on the optimal hyperparameters, which include the number of trees, the maximum depth of the tree, the learning rate, the random sampling ratio, and the regularization parameters L1 and L2.

[0010] S4: Perform model performance evaluation. Based on the trained machine learning model, test its performance on the test set data. Use the root mean square error and correlation coefficient as evaluation indicators of the model prediction performance to evaluate the model prediction performance. If the root mean square error is greater than or equal to 5 or the correlation coefficient is less than or equal to 0.9, it indicates that the prediction effect of the machine learning model is not ideal and steps 1 to 3 need to be repeated. If the root mean square error is less than 5 and the correlation coefficient is greater than 0.9, it indicates that the machine learning model has a better prediction effect and can more quickly and accurately achieve the process optimization and effect prediction of the separation and purification of complex bio-oil obtained by selective pyrolysis of biomass.

[0011] Preferably, the types of the acid removal reagent and the extractant in step S1 are distinguished by chemical formula.

[0012] Preferably, the similar data in step S2 are mainly judged according to the distillation tower packing type. If the distillation tower packing type data is missing, the judgment is made in the order of tower height, tower diameter, temperature, vacuum degree, packing density, reflux ratio, theoretical plate equivalent height, and feed position. If all of the above data are missing, the data is considered to be seriously missing and will not be used.

[0013] Preferably, the screening range configuration dimensions of the hyperparameters in step S3 are as follows: the number of trees is 300, 500, 800, 1000 and 2000, the maximum depth of the tree is 3, 5, 7, 9 and 10, the learning rate is 0.01, 0.05, 0.1, 0.15 and 0.2, the random sampling ratio is 0.6, 0.7, 0.8 and 1.0, the regularization parameter L1 is 0, 0.1, 0.2, 0.3, 0.4 and 0.5, and the regularization parameter L2 is 0, 0.1, 0.2, 0.3, 0.4 and 0.5.

[0014] Preferably, the root mean square error and the correlation coefficient in step S4 are calculated according to the following formulas (1) and (2), respectively:

[0015]

[0016]

[0017] In the above formula, RMSE is the root mean square error, R 2 is the correlation coefficient, n represents the nth group of data, N represents the total number of data groups, is the predicted value for the nth group of data, y is the actual value of the nth group of data, is the average value of the predicted data.

[0018] The technical solutions of the present invention address the lack of process optimization and separation effect prediction methods for LGO in existing technologies for the separation and purification of complex bio-oils obtained through selective pyrolysis of biomass. A process optimization method for the separation of LGO from bio-oil is provided. This method utilizes a multi-dimensional experimental parameter database and an intelligent prediction model to achieve process optimization and accurate prediction of the effects of LGO during the distillation separation of complex bio-oils obtained through selective pyrolysis of biomass. First, experimental data from the distillation separation and purification of LGO from complex bio-oils obtained through selective pyrolysis of biomass are collected. The original experimental dataset is then processed using multiple interpolation to fill missing values, Mahalanobis distance-based outlier detection and removal, and principal component analysis to eliminate multicollinearity. String data for the types of deacidifying agents, extractants, and distillation tower packings are converted to one-hot encoding to further enhance model generalization. This ultimately results in a processed dataset composed of numerical data. A lightweight gradient boosting machine learning framework is then constructed, balancing the training and validation sets through an adaptive weight allocation mechanism to achieve accurate prediction of process optimization and effects for the distillation separation and purification of LGO from complex bio-oils. The present invention provides a new digital tool for online quality monitoring of LGO (Long-Term Oxygen Glutamate) separation and purification of complex bio-oil obtained by selective pyrolysis of biomass. It can guide the dynamic optimization of separation process parameters through real-time prediction and feedback, and has important engineering application value for improving the intelligent level of the bio-based chemical preparation process. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 The present invention provides a flowchart of a process optimization method for separating levoglucosone from bio-oil. DETAILED DESCRIPTION

[0020] In order to make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments. The following embodiments are used to illustrate the present invention but are not intended to limit the scope of the present invention.

[0021] The present invention addresses existing issues and provides a process optimization method for separating levoglucosone from bio-oil. This example uses experimental data from the distillation separation and purification of LGO from complex bio-oil obtained by selective pyrolysis of biomass as an example to optimize the process and predict the results of the separation and purification of levoglucosone from complex bio-oil using the method of the present invention.

[0022] like Figure 1 As shown, the specific steps of this embodiment are as follows:

[0023] Step S1:

[0024] Experimental data were collected during the distillation separation and purification of LGO (LGO) from complex bio-oil obtained by selective pyrolysis of biomass. The data included four aspects: the physicochemical properties of the bio-oil, pretreatment methods, distillation tower operating conditions, and separation effects. The original experimental data set consisted of 618 groups. The physicochemical properties of the bio-oil included LGO content, moisture content, acidity, viscosity, sediment, and main component composition. The pretreatment method included deacidification and extraction. The deacidification process parameters included the type and dosage of the deacidification reagent, and the extraction parameters included the type and dosage of the extractant. The distillation tower operating conditions included tower height, tower diameter, temperature, vacuum degree, packing type, packing density, reflux ratio, theoretical plate equivalent height, and feed position. The separation effect included LGO yield and purity. The types of the deacidification reagent and the extractant in step S1 were distinguished by chemical formulas.

[0025] Step S2:

[0026] The original experimental data set was processed, and missing values ​​were filled using the multiple interpolation method. Outliers were detected and eliminated based on the Mahalanobis distance, and principal component analysis was used to eliminate multicollinearity interference. The string type data of the deacidification reagent type, extractant type, and distillation tower packing type were converted using one-hot encoding, and finally a processed data set consisting of numerical type data was formed. The same type of data was mainly judged according to the distillation tower packing type. If the distillation tower packing type data was missing, it was judged in the order of tower height, tower diameter, temperature, vacuum degree, packing density, reflux ratio, theoretical plate equivalent height, and feed position. If all of the above data were missing, the data was considered to be seriously missing and was not used. The processed data set contained 618 groups of data.

[0027] Step S3:

[0028] Based on the processed data set, 80% of it is randomly selected as the training set and the remaining 20% ​​is used as the test set. A lightweight gradient boosting 1ightGBM machine learning model is constructed. An adaptive weight distribution mechanism is introduced to balance the training set and the training set. Model hyperparameters are screened based on the training set data. Bayesian optimization and progressive cross-validation methods are used for hyperparameter screening. The root mean square error is used as the evaluation indicator to screen and obtain the optimal hyperparameters. The model is then trained based on the optimal hyperparameters, which include the number of trees, the maximum depth of the tree, the learning rate, the random sampling ratio, and the regularization parameters L1 and L2.

[0029] The results of Bayesian optimization automatic screening show that the model effect is best when the number of trees is 800, the maximum depth of the tree is 10, the learning rate is 0.05, the random sampling ratio is 0.8, the regularization parameter L1 is 0.1, and the regularization parameter L1 is 0.2. At this time, the root mean square error of the validation set is 1.45.

[0030] Step S4:

[0031] Perform model performance evaluation. Based on the trained machine learning model, test its performance on the test set data. Use the root mean square error and correlation coefficient as the evaluation indicators of the model prediction performance to evaluate the model prediction performance. Calculate the root mean square error RMSE and correlation coefficient R according to the following formula 2 :

[0032]

[0033] Among them, n represents the nth group of data, N represents the total number of data groups, is the predicted value for the nth group of data, y is the actual value of the nth group of data, is the average value of the predicted data.

[0034] The calculation results show (see Table 1 below) that the root mean square error is less than 5 and the correlation coefficient is greater than 0.9, indicating that the machine learning model has a good prediction effect and can accurately predict the process optimization and effect of LGO purification by distillation separation of complex bio-oil.

[0035] Table 1 Prediction results of the lightGBM model

[0036] Prediction object Root mean square error Correlation coefficient training set 0.435 0.995 Test set 0.923 0.981

[0037] Based on machine learning technology, the present invention constructs a numerical model for process optimization and effect prediction for the separation and purification of LGO from complex bio-oil obtained by selective pyrolysis of biomass. The model is suitable for online quality monitoring of LGO in industrial production and can guide the dynamic optimization of separation process parameters through real-time prediction feedback. It has important engineering application value for improving the intelligent level of the bio-based chemical preparation process.

[0038] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be considered to fall within the scope of protection of the present invention.

Claims

1. A process optimization method for separating levoglucosone from bio-oil, characterized in that: The following steps are involved: Step 1: Experimental data from the distillation separation and purification of LGO (LGO) from complex bio-oil obtained by selective pyrolysis of biomass were collected. The data included the bio-oil's physicochemical properties, pretreatment methods, distillation tower operating conditions, and separation performance, forming the original experimental data set. The physicochemical properties of the bio-oil included LGO content, moisture content, acidity, viscosity, sediment, and major component composition. The pretreatment methods included deacidification and extraction. The deacidification process parameters included the type and dosage of the deacidification reagent, and the extraction parameters included the type and dosage of the extractant. The distillation tower operating conditions included tower height, tower diameter, temperature, vacuum level, packing type, packing density, reflux ratio, theoretical plate equivalent height, and feed position. The separation performance included LGO yield and purity. Step 2: The original experimental data set was processed, and missing values ​​were filled using multiple interpolation. Outliers were detected and eliminated based on the Mahalanobis distance. Principal component analysis was used to eliminate multicollinearity interference. The string data of the deacidification reagent type, extractant type, and distillation tower packing type were converted using one-hot encoding, ultimately forming a processed data set consisting of numerical data. Step 3: Based on the processed data set, 80% of it is randomly selected as the training set and the remaining 20% ​​is used as the test set. A lightweight gradient boosting lightGBM machine learning model is constructed. An adaptive weight distribution mechanism is introduced to balance the training set and the training set. Model hyperparameters are screened based on the training set data. Bayesian optimization and progressive cross-validation methods are used for hyperparameter screening. The root mean square error is used as the evaluation indicator to screen and obtain the optimal hyperparameters. The model is then trained based on the optimal hyperparameters, which include the number of trees, the maximum depth of the tree, the learning rate, the random sampling ratio, and the regularization parameters L1 and L2. Step 4: The model performance was evaluated. Based on the trained machine learning model, its performance on the test set data was tested. The root mean square error and correlation coefficient were used as evaluation indicators of the model prediction performance. If the root mean square error was greater than or equal to 5 or the correlation coefficient was less than or equal to 0.9, it indicated that the prediction effect of the machine learning model was not ideal and steps 1 to 3 needed to be repeated. If the root mean square error was less than 5 and the correlation coefficient was greater than 0.9, it indicated that the machine learning model had a good prediction effect and could achieve accurate prediction of the process optimization and effect of LGO purification by distillation of complex bio-oil.

2. The method according to claim 1, characterized in that The types of the acid removal reagent and the extractant in step 1 are distinguished by chemical formula.

3. The method according to claim 1, characterized in that The similar data in step 2 are mainly judged according to the type of distillation tower packing. If the data of the type of distillation tower packing is missing, the judgment is made in the order of tower height, tower diameter, temperature, vacuum degree, packing density, reflux ratio, theoretical plate equivalent height, and feed position. If all of the above data are missing, it is considered that the data is seriously missing and will not be used.

4. The method according to claim 1, wherein The screening range configuration dimensions of the hyperparameters in step 3 are as follows: the number of trees is 300, 500, 800, 1000 and 2000, the maximum depth of the tree is 3, 5, 7, 9 and 12, the learning rate is 0.01, 0.05, 0.1, 0.15 and 0.2, the random sampling ratio is 0.6, 0.7, 0.8 and 1.0, the regularization parameter L1 is 0, 0.1, 0.2, 0.3, 0.4 and 0.5, and the regularization parameter L2 is 0, 0.1, 0.2, 0.3, 0.4 and 0.

5.

5. The method according to claim 1, wherein The root mean square error and the correlation coefficient in step 4 are calculated according to the following formulas (1) and (2), respectively: In the above formula, RMSE is the root mean square error, R 2 is the correlation coefficient, n represents the nth group of data, N represents the total number of data groups, is the predicted value for the nth group of data, y is the actual value of the nth group of data, is the average value of the predicted data.

Citation Information

Patent Citations

  • Method for co-producing coke and levoglucosenone by pyrolyzing biomass

    CN111849525A

  • Intelligent oil distillation optimization method and system based on machine learning

    CN117744831A

  • Prediction method and device of biomass pyrolysis product, computer equipment and storage medium

    CN119862780A

  • Storage medium, microwave flash pyrolysis process optimization method, device and equipment

    CN119864095A