Process optimization method for preparing levoglucosenone by separating and purifying bio-oil
By constructing the lightGBM model, the process optimization and prediction problems of LGO separation and purification of complex bio-oil by selective pyrolysis of biomass were solved, efficient LGO separation effect monitoring and process optimization were achieved, and the industrial application of complex bio-oil obtained by selective pyrolysis of biomass was supported.
Patent Information
- Application Number
- CN202510760629.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-09-19
AI Technical Summary
In the existing technology, there is a lack of process optimization and separation effect prediction methods for the separation and purification of L-glucose ketone (LGO) from complex bio-oil obtained by selective pyrolysis of biomass. In particular, traditional separation methods are inefficient and acidic components easily trigger LGO degradation side reactions, making industrialization difficult to achieve.
A lightweight gradient boosting machine learning model (lightGBM) was used to construct a process optimization method. By collecting and processing bio-oil separation experimental data, data cleaning, feature engineering conversion and model training were performed, and hyperparameters were screened to achieve process optimization and effect prediction for the distillation separation and purification of LGO from complex bio-oil molecules.
It achieves rapid and accurate prediction of LGO purification by molecular distillation of complex bio-oil, supports LGO separation effect monitoring and process optimization in industrial production, and reduces repeated modeling time and economic costs.
Smart Images

Figure CN120673876A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of separation and purification of bio-based chemicals, and particularly relates to a process optimization method for preparing levoglucosone by separation and purification of bio-oil. Background Art
[0002] As a key component of a sustainable carbon cycle, biomass, with its renewable nature and carbon-enriching advantages, continues to unleash its potential in the development of clean energy and fine chemicals. Rapid pyrolysis technology, in particular, has attracted considerable attention due to its efficient conversion properties. This process, through instantaneous high-temperature cracking, converts biomass into pyrolysis oil rich in high-value-added compounds. L-glucose ketone (LGO), a signature product of this process, demonstrates exceptional value in high-end pharmaceutical fields, such as asymmetric synthesis catalysts and anti-tumor drug precursors, due to its unique chiral pyran ring structure and multifunctional properties.
[0003] At present, industrial production mainly relies on the synthesis of LGO through multi-step chemical reactions using high-priced sugar sources such as galactose. Although high-purity products can be obtained, the cost of raw materials and the complexity of the process seriously restrict its large-scale development. In contrast, the pyrolysis route using lignocellulosic biomass as raw materials has become a research hotspot due to its easy availability of raw materials and simple process. For example, the catalytic pyrolysis technology disclosed in Chinese patents CN116274248, CN202110604000.2, and CN201910278528.8 has significantly improved the selectivity of LGO by optimizing the reaction conditions and catalyst system, providing a new path for large-scale production.
[0004] However, the complex composition of pyrolysis products remains a bottleneck for the industrialization of bio-based LGO. Experimental analysis shows that, in addition to the target product LGO, bio-oil also contains dozens of compounds, including organic acids, phenols, aldehydes, and ketones. The similar physicochemical properties of these components make traditional separation methods inefficient. In particular, acidic components are prone to triggering LGO degradation side reactions, further increasing the separation difficulty. Therefore, the development of new and efficient separation technologies has become a core research direction for the industrialization of bio-based LGO.
[0005] In recent years, data-driven machine learning modeling has provided innovative solutions for predicting target product separation performance and optimizing processes in industrial distillation. This approach establishes a nonlinear mapping relationship between multidimensional input variables, such as the physicochemical properties of bio-oil, pretreatment methods, and molecular distillation operating conditions, and separation performance, enabling process optimization and intelligent prediction of target product separation performance. Compared to traditional mechanistic modeling methods, machine learning models possess the ability to rapidly transfer parameters after data training. Model parameters can be replicated to enable instant deployment of prediction capabilities across different production equipment, significantly reducing the time and cost of repetitive modeling. This technical feature has demonstrated significant application value in areas such as industrial process digital twins, industrial process optimization, and separation performance prediction. In the field of industrial process modeling, the light gradient boosting machine (lightGBM) demonstrates superior performance due to its unique decision tree ensemble architecture, significantly reducing memory usage while maintaining prediction accuracy. Compared to single decision tree models or shallow neural networks, lightGBM maintains high model generalization capabilities even with small sample sizes. Furthermore, its low-code deployment, supported by an open-source framework, makes it technically feasible to rapidly build prediction systems on-site. However, machine learning methods are currently lacking in process optimization and performance prediction for the separation and purification of LGO from complex bio-oils derived from selective pyrolysis of biomass. Furthermore, the high complexity of bio-oil components and the secondary condensation of LGO initiated by acidic components lead to a strong nonlinear correlation between process variables and separation performance indicators, which also poses challenges for machine learning modeling. Currently, there are no publicly reported lightGBM-based methods for process optimization and performance prediction for the molecular distillation separation and purification of LGO from complex bio-oils derived from selective pyrolysis of biomass. Summary of the Invention
[0006] The present invention addresses the lack of process optimization and separation effect prediction methods for LGO in existing technologies for separating and purifying complex bio-oils obtained by selective pyrolysis of biomass. The invention provides a process optimization method for separating and purifying bio-oil to prepare L-glucose ketone, which can more quickly and accurately achieve process optimization and effect prediction for the molecular distillation separation and purification of LGO from complex bio-oils obtained by selective pyrolysis of biomass.
[0007] To solve the above technical problems, an embodiment of the present invention provides a process optimization method for separating and purifying bio-oil to prepare levoglucosone, comprising the following steps:
[0008] S1: Collect experimental data from the molecular distillation separation and purification of LGO (LGO) from complex bio-oil obtained by selective pyrolysis of biomass, including the physicochemical properties of the bio-oil, pretreatment methods, molecular distillation operating conditions, and separation effects, to form the original experimental data set. The physicochemical properties of the bio-oil include LGO content, moisture content, acidity, viscosity, sediment, and main component composition. The pretreatment methods include deacidification and extraction. The deacidification process parameters include the type and amount of deacidification reagents, the extraction parameters include the type and amount of extractants, the molecular distillation operating conditions include temperature, pressure, feed rate, and scraper speed, and the separation effect includes LGO yield and purity.
[0009] S2: Data processing is performed on the original experimental data set. The data cleaning module removes redundant and repeated data and data with too many missing values. The string data of the types of acid removal reagents and extractants are converted into feature engineering using one-hot encoding. Data with fewer missing values are interpolated using the average value of similar data. Finally, a fully numerical standardized data set is constructed through multi-level collaborative processing.
[0010] S3: Based on the processed dataset, 80% of it is randomly selected as the training set and the remaining 20% as the test set. A lightweight gradient boosting lightGBM machine learning model is constructed. Model hyperparameters are screened based on the training set data. Hyperparameter screening uses grid search and five-fold cross-validation methods, and the root mean square error is used as the evaluation metric to screen out the optimal hyperparameters. The model is then trained based on the optimal hyperparameters, which include the number of trees, the maximum depth of the tree, the learning rate, the random sampling ratio, and the regularization parameters L1 and L2.
[0011] S4: Perform model performance evaluation. Based on the trained machine learning model, test its performance on the test set data. Use the root mean square error and correlation coefficient as evaluation indicators of the model prediction performance to evaluate the model prediction performance. If the root mean square error is greater than or equal to 5 or the correlation coefficient is less than or equal to 0.9, it indicates that the prediction effect of the machine learning model is not ideal and steps 1 to 3 need to be repeated. If the root mean square error is less than 5 and the correlation coefficient is greater than 0.9, it indicates that the machine learning model has a better prediction effect and can achieve accurate prediction of the process optimization and effect of molecular distillation separation and purification of LGO from complex bio-oil.
[0012] Preferably, the types of the acid removal reagent and the extractant in step S1 are distinguished by chemical formula.
[0013] Preferably, the similar data in step S2 are mainly judged according to the molecular distillation conditions. If the molecular distillation condition data is missing, the judgment is made in the order of temperature, pressure, feed rate and scraper speed. If all the above data are missing, it is considered that the data is seriously missing and will not be used.
[0014] Preferably, the screening range of the hyperparameters in step S3 is set as follows: the number of trees is 100, 200, 300, 400 and 500, the maximum depth of the tree is 0, 5, 10, 15 and 20, the learning rate is 0.01, 0.05, 0.1, 0.15 and 0.2, the random sampling ratio is 0.5, 0.6, 0.7, 0.8 and 1.0, the regularization parameter L1 is 0, 0.1, 0.2, 0.3, 0.4 and 0.5, and the regularization parameter L2 is 0, 0.1, 0.2, 0.3, 0.4 and 0.5.
[0015] Preferably, the root mean square error and the correlation coefficient in step S4 are calculated according to the following formulas (1) and (2), respectively:
[0016]
[0017] In the above formula, RMSE is the root mean square error, R 2 is the correlation coefficient, n represents the nth group of data, N represents the total number of data groups, is the predicted value for the nth group of data, y is the actual value of the nth group of data, is the average value of the predicted data.
[0018] The present invention addresses the lack of process optimization and separation effect prediction methods for LGO in existing technologies for separating and purifying complex bio-oils obtained by selective pyrolysis of biomass. The present invention provides a process optimization method for preparing LGO by separating and purifying bio-oils, which is used for monitoring LGO separation effects and process optimization in industrial production, and analyzes the influence of process parameters on LGO separation effects. First, experimental data from the molecular distillation separation and purification of LGO from complex bio-oils obtained by selective pyrolysis of biomass are collected. Subsequently, data processing is performed on the original experimental dataset, including data cleaning, feature engineering conversion, and missing value data interpolation. A fully digitized, standardized dataset is constructed through multi-level collaborative processing. A lightweight gradient boosting machine learning model is constructed. This model is an efficient ensemble learning algorithm that uses training data to optimize model parameter configuration and achieve accurate prediction of process optimization and effects for the molecular distillation separation and purification of LGO from complex bio-oils. Based on machine learning technology, the present invention constructs a numerical model for process optimization and separation effect prediction in the molecular distillation separation and purification of LGO from complex bio-oils, providing strong support for the development of technologies for separating and purifying LGO from complex bio-oils. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 The present invention provides a flowchart of a process optimization method for separating and purifying bio-oil to prepare levoglucosone. DETAILED DESCRIPTION
[0020] In order to make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments. The following embodiments are used to illustrate the present invention but are not intended to limit the scope of the present invention.
[0021] The present invention addresses existing challenges and provides a process optimization method for separating and purifying bio-oil to produce L-glucose ketone. This example uses experimental data from the molecular distillation separation and purification of LGO (L-glucose ketone) from complex bio-oil obtained through selective pyrolysis of biomass as an example to demonstrate the process optimization and accurate prediction of the results of the molecular distillation separation and purification of LGO from complex bio-oil.
[0022] like Figure 1 As shown, the specific steps of this embodiment are as follows:
[0023] Step S1:
[0024] Experimental data were collected during the molecular distillation separation and purification of LGO (LGO) from complex bio-oil obtained by selective pyrolysis of biomass. The data included four aspects: the physicochemical properties of the bio-oil, pretreatment methods, molecular distillation operating conditions, and separation effects. This data set consisted of 435 original experimental data sets. The physicochemical properties of the bio-oil included LGO content, moisture content, acidity, viscosity, sediment, and main component composition. The pretreatment methods included deacidification and extraction. The deacidification process parameters included the type and dosage of the deacidification reagent, and the extraction parameters included the type and dosage of the extractant. The molecular distillation operating conditions included temperature, pressure, feed rate, and scraper speed. The separation effects included LGO yield and purity. The types of the deacidification reagent and the extractant in step S1 were distinguished using chemical formulas.
[0025] Step S2:
[0026] The original experimental data set was processed, and the data cleaning module eliminated redundant and repeated data and data with too many missing values. The string type data of the acid removal reagent type and the extractant type were converted into feature engineering using one-hot encoding. The data with fewer missing values were interpolated using the average value of similar data. Finally, a fully numerical standardized data set was constructed through multi-level collaborative processing. The similar data were mainly judged based on the molecular distillation conditions. If the molecular distillation condition data were missing, they were judged in the order of temperature, pressure, feed rate and scraper speed. If all of the above data were missing, it was considered that the data was seriously missing and would not be used. The processed data set contained 435 groups of data.
[0027] Step S3:
[0028] Based on the processed data set, 80% of it was randomly selected as the training set and the remaining 20% as the test set. A lightweight gradient boosting 1ightGBM machine learning model was constructed. Model hyperparameters were screened based on the training set data. Grid search and five-fold cross-validation methods were used for hyperparameter screening. The root mean square error was used as the evaluation metric to screen out the optimal hyperparameters. The model was then trained based on the optimal hyperparameters, which included the number of trees, the maximum depth of the tree, the learning rate, the random sampling ratio, and the regularization parameters L1 and L2.
[0029] The automatic screening results show that the model works best when the number of trees is 300, the maximum depth of the tree is 5, the learning rate is 0.1, the random sampling ratio is 0.6, the regularization parameter L1 is 0.1, and the regularization parameter L1 is 0.2. At this time, the root mean square error of the validation set is 3.89.
[0030] Step S4:
[0031] Perform model performance evaluation. Based on the trained machine learning model, test its performance on the test set data. Use the root mean square error and correlation coefficient as the evaluation indicators of the model prediction performance to evaluate the model prediction performance. Calculate the root mean square error RMSE and correlation coefficient R according to the following formula 2 :
[0032]
[0033] Among them, n represents the nth group of data, N represents the total number of data groups, is the predicted value for the nth group of data, y is the actual value of the nth group of data, is the average value of the predicted data.
[0034] The calculation results show (see Table 1 below) that the root mean square error is less than 5 and the correlation coefficient is greater than 0.9, indicating that the machine learning model has a good prediction effect and can accurately predict the process optimization and effect of molecular distillation separation and purification of LGO from complex bio-oil.
[0035] Table 1 Prediction results of the lightGBM model
[0036] Prediction object Root mean square error Correlation coefficient training set 0.561 0.993 Test set 1.849 0.952
[0037] Based on machine learning technology, the present invention constructs a numerical model for process optimization and effect prediction for the separation and purification of LGO from complex bio-oil obtained by selective pyrolysis of biomass. The model is suitable for monitoring the LGO separation effect and process optimization in industrial production, and analyzes the influence of process parameters on the LGO separation effect. This model provides strong support for the development of LGO separation technology for complex bio-oil obtained by selective pyrolysis of biomass.
[0038] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be considered to fall within the scope of protection of the present invention.
Claims
1. A process optimization method for preparing levoglucosone by separation and purification of bio-oil, characterized in that: The following steps are involved: Step 1: Experimental data from the molecular distillation separation and purification of LGO (LGO) from complex bio-oil obtained by selective pyrolysis of biomass were collected. The data included the bio-oil's physicochemical properties, pretreatment methods, molecular distillation operating conditions, and separation performance, forming an original experimental dataset. The physicochemical properties of the bio-oil included LGO content, moisture content, acidity, viscosity, sediment, and major component composition. The pretreatment methods included deacidification and extraction. The deacidification process parameters included the type and dosage of the deacidification reagent, and the extraction parameters included the type and dosage of the extractant. The molecular distillation operating conditions included temperature, pressure, feed rate, and scraper speed. The separation performance included LGO yield and purity. Step 2: The original experimental data set was processed. The data cleaning module removed redundant and repeated data and data with excessive missing values. The string data of the types of acid removal reagents and extractants were converted into feature engineering data using one-hot encoding. The data with fewer missing values was interpolated using the average value of similar data. Finally, a fully numerical and standardized data set was constructed through multi-level collaborative processing. Step 3: Based on the processed data set, 80% of it was randomly selected as the training set and the remaining 20% as the test set. A lightweight gradient boosting lightGBM machine learning model was constructed. Model hyperparameters were screened based on the training set data. Grid search and five-fold cross-validation methods were used for hyperparameter screening. The root mean square error was used as the evaluation metric to screen out the optimal hyperparameters. The model was then trained based on the optimal hyperparameters, which included the number of trees, the maximum depth of the tree, the learning rate, the random sampling ratio, and the regularization parameters L1 and L2. Step 4: The model performance was evaluated. Based on the trained machine learning model, its performance on the test set data was tested. The root mean square error and correlation coefficient were used as evaluation indicators of the model prediction performance. If the root mean square error was greater than or equal to 5 or the correlation coefficient was less than or equal to 0.9, it indicated that the prediction effect of the machine learning model was not ideal and steps 1 to 3 needed to be repeated. If the root mean square error was less than 5 and the correlation coefficient was greater than 0.9, it indicated that the machine learning model had a good prediction effect and could achieve accurate prediction of the process optimization and effect of molecular distillation separation and purification of LGO from complex bio-oil.
2. The method according to claim 1, characterized in that The types of the acid removal reagent and the extractant in step 1 are distinguished by chemical formula.
3. The method according to claim 1, characterized in that The similar data in step 2 are mainly judged according to the molecular distillation conditions. If the molecular distillation condition data is missing, the judgment is made in the order of temperature, pressure, feed rate and scraper speed. If all of the above data are missing, it is considered that the data is seriously missing and will not be used.
4. The method according to claim 1, wherein The screening range of the hyperparameters in step 3 is set as follows: the number of trees is 100, 200, 300, 400 and 500, the maximum depth of the tree is 0, 5, 10, 15 and 20, the learning rate is 0.01, 0.05, 0.1, 0.15 and 0.2, the random sampling ratio is 0.5, 0.6, 0.7, 0.8 and 1.0, the regularization parameter L1 is 0, 0.1, 0.2, 0.3, 0.4 and 0.5, and the regularization parameter L2 is 0, 0.1, 0.2, 0.3, 0.4 and 0.
5.
5. The method according to claim 1, characterized in that The root mean square error and the correlation coefficient in step 4 are calculated according to the following formulas (1) and (2), respectively: In the above formula, RMSE is the root mean square error, R 2 is the correlation coefficient, n represents the nth group of data, N represents the total number of data groups, is the predicted value for the nth group of data, y is the actual value of the nth group of data, is the average value of the predicted data.
Citation Information
Patent Citations
Green preparation method of levoglucosone having high added value
CN103626808A
Method for selectively preparing high-value product through rapid pyrolysis of cassava residues
CN112010824A
Characterization of phase separation of coating compositions
CN114199753A
Method and system for predicting thermoplastic polyimide film preparation process through machine learning
CN118876475A
Storage medium, microwave flash pyrolysis process optimization method, device and equipment
CN119864095A