Method for predicting lignin separation and extraction effect of eutectic solvent based on machine learning
By constructing multi-dimensional data models and machine learning technology, the problems of resource waste and inefficiency in the screening process of eutectic solvents are solved, accurate prediction of lignin removal rate and optimization of solvent system are achieved, and the efficiency and resource utilization of lignin separation are improved.
Patent Information
- Application Number
- CN202510359226.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-08
AI Technical Summary
In the prior art, the screening process of eutectic solvents is time-consuming and resource waste is severe, making it difficult to comprehensively examine the multi-dimensional characteristics of the solvent and its impact on lignin removal rate, and there is a lack of efficient and accurate solvent screening and separation optimization scheme.
A multi-dimensional data model containing the characteristics of low eutectic solvents and experimental conditions was constructed. Combined with machine learning technology, a predictive model was established through Stacking integrated algorithm, key parameters affecting lignin removal rate were analyzed, and the model was explained using the SHAP method to optimize solvent composition and experimental conditions.
It realizes efficient prediction and optimization of the eutectic solvent system, reduces resource waste, and improves the screening efficiency and separation process efficiency of lignin separation.
Smart Images

Figure CN120280036A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for predicting the effect of separating and extracting lignin by deep eutectic solvents based on machine learning, and belongs to the technical field of high-value utilization of biomass. Background Art
[0002] With the continuous growth of global energy demand and the increasing depletion of fossil energy resources, the efficient utilization of renewable energy and biomass resources has become an important part of the national energy strategy. As a rich and renewable biomass resource, the high-value utilization of lignocellulose is of great significance in promoting green energy and sustainable development.
[0003] Lignocellulose is the main component of plant cell walls and is mainly composed of three macromolecules: cellulose, lignin, and hemicellulose. Cellulose is a linear polysaccharide mainly composed of glucose units, with high crystallinity and mechanical strength; lignin is composed of aromatic macromolecular compounds, with strong chemical stability and anti-degradability, and usually acts as an "adhesive" in plant cell walls; while hemicellulose is composed of polysaccharides and plays a role in connecting cellulose and lignin. Due to the complex and stable structure of lignin, its presence significantly increases the degradation difficulty of lignocellulose. Therefore, the removal of lignin is a key step in improving the utilization efficiency of lignocellulose. Deep eutectic solvents show good application prospects in the field of lignin separation due to their environmental friendliness, low cost, simple preparation, etc. However, in the prior art, the screening of deep eutectic solvents usually relies on experimental test methods, which results in a very time-consuming screening process and serious waste of resources. The main technical problems of the existing methods include: low screening efficiency, especially when facing a large number of different solvent systems, the time cost and resource consumption of manual experiments are unbearable; due to the complex interactions between solvents, it is often difficult to comprehensively investigate the multi-dimensional characteristics of solvents and their specific effects on the lignin removal rate in experiments. The prior art mainly relies on single experimental data, lacks in-depth data-based analysis and prediction, and is difficult to provide efficient and accurate solvent screening and separation optimization schemes. Summary of the Invention To solve these problems, the present invention proposes a method for predicting the effect of separating and extracting lignin by deep eutectic solvents based on machine learning. By constructing a multi-dimensional data model including the characteristics of deep eutectic solvents and experimental conditions, combined with machine learning techniques, it can efficiently predict the influence of different solvent systems on the lignin removal rate, thereby greatly improving the screening efficiency, reducing the waste of experimental resources, and providing a scientific basis and theoretical support for the optimization of lignin separation processes. Through this method, the problems of resource waste and low screening efficiency in traditional experimental methods can be effectively avoided, and the development of lignin separation technology can be promoted.
[0004] The technical solution of the present invention is as follows: A method for predicting the effect of separating and extracting lignin with deep eutectic solvents based on machine learning, comprising the following steps: S1. Data construction: Collect and construct a sample database for separating and extracting lignin with deep eutectic solvents. The sample database includes the hydrogen bond acidity ( ), hydrogen bond basicity ( ) of the deep eutectic solvents, the molar ratio of hydrogen bond donors to acceptors (Molar Ratio), the molecular polarity index (MPI), the polar surface area (PSA), the polar surface area ratio (PSARatio), and the removal rate of lignin (DR); S2. Feature variable extraction and quantitative analysis: Based on the molecular structures of hydrogen bond donors and acceptors in deep eutectic solvents, calculate characteristic parameters such as hydrogen bond acidity ( ), hydrogen bond basicity ( ), and molecular polarity index (MPI) through quantum chemistry; integrate parameters such as the molar ratio (Molar Ratio) of the solvent system and the experimentally measured removal rate of lignin (DR) into the sample database, and screen out variables highly correlated with the removal rate of lignin through correlation analysis; S3. Machine learning model construction and training: Input the sample database into a machine learning model, adopt the Stacking integration algorithm to construct a prediction model with the removal rate of lignin as the target variable, and perform training. To enhance the generalization ability of the model, introduce a regularization technique to prevent overfitting; S4. Model performance evaluation and optimization: After the model training is completed, use five-fold cross-validation to evaluate the model performance. Adopt the coefficient of determination (R 2 ), mean square error (MSE), root mean square error (RMSE), and mean absolute error (MAE) as model performance evaluation indicators, and optimize the model accuracy by adjusting hyperparameters; S5. Model interpretation: Use the SHAP method to interpret the trained machine learning model. By calculating the contribution value of each feature to the model prediction result, generate a feature importance ranking graph to intuitively display the importance and action mode of key variables; at the same time, analyze the positive and negative effects and action ranges of features on the removal rate through the SHAP value distribution graph; S6. Prediction and result optimization: Use the optimized model to predict the lignin extraction effect of different deep eutectic solvent systems. Combine the analysis results of the feature importance calculated by the SHAP method to identify key variables that have a greater impact on the removal rate of lignin, such as hydrogen bond acidity, hydrogen bond basicity, and molecular polarity index. According to the analysis results, adjust the composition ratio of hydrogen bond donors to acceptors in the deep eutectic solvents and optimize the experimental conditions to further improve the removal rate of lignin and the separation process efficiency. The optimized results can reduce resource consumption and improve the separation efficiency.
[0005] Further, the hydrogen bond acceptor includes choline chloride, tetramethylammonium chloride, tetraethylammonium chloride, tetrabutylammonium chloride, tetramethylammonium bromide, tetrabenzylammonium chloride, hydroxyethyltrimethylammonium chloride, chlormequat chloride, and ethylene glycol; The hydrogen bond donors include formic acid, lactic acid, urea, glycerol, oxalic acid, citric acid, succinic acid or amino acids, levulinic acid, malonic acid, phenylpropionic acid, xylitol, D-isosorbide and D-sorbitol, acetamide, urea, ethylene glycol.
[0006] Furthermore, in step S1, Nile red, 4-nitroaniline, N,N-diethyl-4-nitroaniline and an ultraviolet spectrophotometer are used to measure the hydrogen bond acidity, hydrogen bond basicity and solvent polarization parameters respectively. Specifically, Nile red, 4-nitroaniline and N,N-diethyl-4-nitroaniline solvent indicators are dissolved in methanol to prepare three concentrations of 1.0×10 - 3 mol / L solution, take 3 centrifuge tubes and add 50 μL of the above indicator solution respectively, put them in a vacuum oven at 40°C for 30 min to remove methanol, then add 2 mL of the low eutectic solvent into the 3 centrifuge tubes respectively, mix them thoroughly and transfer them into a quartz cuvette, and measure the absorption spectrum with a UV-visible spectrophotometer at room temperature; Furthermore, the sample database of step S1 also includes the sum of hydrogen bond acidity and basicity of the low eutectic solvent, the molecular polarizability of the low eutectic solvent, the solvent polarity of the low eutectic solvent, the thermal stability of the low eutectic solvent, the ash content of the biomass raw material, the initial density of the biomass raw material, the particle size distribution of the biomass raw material, the pretreatment temperature, the pretreatment time, the pretreatment pressure, the pretreatment heating rate, and the structural change rate of the biomass after pretreatment.
[0007] Furthermore, in step S3, the constructed sample database is divided into a training set and a test set in a ratio of 7:3, and a machine learning model is constructed using the Stacking ensemble algorithm; the model is composed of Xgboost, RF and SVM as base learners and CNN as a meta-learner; the removal rate of extracted lignin is set as the target variable, and other solvent parameters and experimental conditions are set as input variables.
[0008] Furthermore, in step S4, the model performance test uses five-fold cross validation to perform hyperparameter tuning, determine the optimal hyperparameter combination, and calculate the determination coefficient (R 2 ), mean square error (MSE), root mean square error (RMSE) and mean absolute error (MAE) to measure model performance.
[0009] Further, in step S5, the results of the trained Stacking model are interpreted using the SHAP method to calculate the contribution value of each feature to the model prediction result. A feature importance ranking graph is generated through SHAP value analysis to intuitively display the positive and negative impacts and contribution degrees of each feature on the lignin removal rate. At the same time, the range of action and significance of different feature variables are analyzed by combining the SHAP value distribution graph, and key variables that have a greater impact on the lignin removal rate are screened out, including hydrogen bond acidity, hydrogen bond basicity, and molecular polarity index, etc., and the influence mechanism of these variables on the prediction result is further evaluated.
[0010] Advantages of the present invention: By constructing a sample database containing the characteristic parameters of deep eutectic solvents and extraction process conditions, and using the Stacking integration algorithm to establish a prediction model, with the lignin removal rate as the target variable, the present invention realizes the accurate prediction of the separation effect of deep eutectic solvent systems. By using quantum chemical calculations and feature variable extraction methods, the key parameters affecting the separation effect are analyzed, and the model is explained by the SHAP method to quantify the contribution of each feature to the prediction result and intuitively display the importance of key variables. The present invention guides the adjustment of the composition ratio of hydrogen bond donors and acceptors of deep eutectic solvents and the optimization of extraction conditions through the prediction results, providing a scientific basis for the efficient separation of lignin, and having important application value and promotion potential. Description of the drawings
[0011] Figure 1 is the flowchart of the effect prediction and optimization of the present invention; Figure 2 is the framework diagram of the Stacking model of the present invention; Figure 3 is the SHAP framework diagram of the present invention. Detailed implementation manners
[0012] The present invention will be further described in detail below in conjunction with the drawings and specific implementation manners: As Figure 1 shown, an embodiment of the present invention discloses a method for predicting the effect of separating and extracting lignin from deep eutectic solvents based on machine learning, including the following steps: S1. Data construction: Collect and construct a sample database for separating and extracting lignin from deep eutectic solvents. The sample database includes the hydrogen bond acidity of the deep eutectic solvent of the solvent ( ), the hydrogen bond basicity of the deep eutectic solvent ( ), the molar ratio of hydrogen bond donors to acceptors (Molar Ratio), the molecular polarity index (MPI), the polar surface area (PSA), the polar surface area ratio (PSARatio), and the removal rate of lignin (DR); S2. Feature Variable Extraction and Quantitative Analysis: Based on the molecular structures of hydrogen bond donors and acceptors in the deep eutectic solvent, quantum chemistry is used to calculate characteristic parameters such as hydrogen bond acidity ( ), hydrogen bond basicity ( ), and molecular polarity index (MPI); parameters such as the molar ratio (Molar Ratio) of the solvent system and the experimentally measured lignin removal rate (DR) are integrated into the sample database, and variables highly correlated with the lignin removal rate are screened out through correlation analysis; S3. As shown in Figure 2 , Machine Learning Model Construction and Training: The sample database is input into the machine learning model, and the Stacking ensemble algorithm is used to construct a prediction model with the lignin removal rate as the target variable and perform training. To enhance the generalization ability of the model, regularization techniques are introduced to prevent overfitting; S4. Model Performance Evaluation and Optimization: After the model training is completed, five-fold cross-validation is used to evaluate the model performance. The coefficient of determination (R 2 ), mean square error (MSE), root mean square error (RMSE), and mean absolute error (MAE) are used as model performance evaluation indicators, and the model accuracy is optimized by adjusting hyperparameters; S5. As shown in Figure 3 , Model Explanation: The SHAP method is used to explain the trained machine learning model. By calculating the contribution value of each feature to the model prediction result, a feature importance ranking graph is generated to intuitively display the importance and action mode of key variables; at the same time, the positive and negative effects and action ranges of features on the lignin removal rate are analyzed through the SHAP value distribution graph; S6. Prediction and Result Optimization: The optimized model is used to predict the lignin removal effect of different deep eutectic solvent systems. Combining the analysis results of the feature importance calculated by the SHAP method, key variables that have a greater impact on the lignin removal rate are identified, such as hydrogen bond acidity, hydrogen bond basicity, and molecular polarity index. According to the analysis results, the composition ratio of hydrogen bond donors and acceptors in the deep eutectic solvent is adjusted, and the experimental conditions are optimized to further improve the lignin removal rate and the separation process efficiency. The optimized results can reduce resource consumption and improve the removal efficiency.
[0013] The hydrogen bond acceptors include choline chloride, tetramethylammonium chloride, tetraethylammonium chloride, tetrabutylammonium chloride, tetramethylammonium bromide, tetrabenzylammonium chloride, 2-hydroxyethyltrimethylammonium chloride, chlormequat chloride, and ethylene glycol; The hydrogen bond donors include formic acid, lactic acid, urea, glycerol, oxalic acid, citric acid, succinic acid or amino acids, levulinic acid, malonic acid, phenylpropionic acid, xylitol, D-isosorbide, and D-sorbitol, acetamide, urea, and ethylene glycol.
[0014] Nile red, 4-nitroaniline, and N,N-diethyl-4-nitroaniline solvent indicators were dissolved in methanol to prepare three concentrations of 1.0×10 -3 mol / L solution, take 3 centrifuge tubes and add 50μL of the above indicator solution respectively, put them in a vacuum oven at 40℃ for 30 min to remove methanol, then add 2mL of low eutectic solvent into the 3 centrifuge tubes respectively, mix them thoroughly and transfer them into quartz cuvettes, and measure the absorption spectrum with UV-visible spectrophotometer at room temperature; The sample database in step S1 also includes the sum of the hydrogen bond acidity and basicity of the low eutectic solvent, the molecular polarizability of the low eutectic solvent, the solvent polarity of the low eutectic solvent, the thermal stability of the low eutectic solvent, the ash content of the biomass raw material, the initial density of the biomass raw material, the particle size distribution of the biomass raw material, the pretreatment temperature, the pretreatment time, the pretreatment pressure, the pretreatment heating rate, and the structural change rate of the biomass after pretreatment.
[0015] The constructed sample database was divided into training set and test set in a ratio of 7:3, and a machine learning model was constructed using the Stacking ensemble algorithm; the model consisted of Xgboost, RF and SVM as base learners and CNN as meta-learner; the removal rate of extracted lignin was set as the target variable, and other solvent parameters and experimental conditions were set as input variables.
[0016] The parameters of each base model are optimized through cross-validation to achieve the best fitting effect on the training set and ensure the generalization ability of the model on the test set.
[0017] The trained Stacking model results were used to perform dimensionality reduction analysis on the feature variables using the t-SNE method, and the influence of the variables on the lignin removal effect was intuitively displayed through the characteristic distribution law. At the same time, the independent variables in the sample database were evaluated according to the model results, and the key variables with a greater impact on the lignin removal rate were screened out, including hydrogen bond acidity, hydrogen bond alkalinity and molecular polarity index, and the contribution of these variables to the lignin removal rate was further analyzed.
[0018] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents, and any modifications or equivalent replacements that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A prediction method for the effect of separating and extracting lignin by deep eutectic solvents based on machine learning, characterized in that: It includes the following steps: S1. Data construction: Collect and construct a sample database for the separation and extraction of lignin by deep eutectic solvents. The sample database includes the hydrogen bond acidity of the deep eutectic solvent, the hydrogen bond basicity of the deep eutectic solvent, the molar ratio of the hydrogen bond donor to the acceptor, the molecular polarity index, the polar surface area, the proportion of the polar surface area, and the lignin removal rate; S2. Feature variable extraction and quantitative analysis: Based on the solvent molecular structure, calculate the hydrogen bond acidity, hydrogen bond basicity, and molecular polarity index, and extract the key variables affecting the lignin removal rate; S3. Model construction and training: Input the sample database into a machine learning model, and use the Stacking integration algorithm to construct a prediction model with the lignin removal rate as the target variable and perform training; S4. Model Performance Evaluation and Optimization: Five-fold cross-validation is used for hyperparameter tuning to determine the optimal hyperparameter combination, and the coefficient of determination (R 2 )), mean squared error (MSE), root mean squared error (RMSE), and mean absolute error (MAE) are calculated respectively to measure the model performance; S5. Model interpretation: Use the SHAP method to interpret the model, calculate the contribution value of each feature to the prediction result, and generate a feature importance ranking graph, a SHAP value distribution graph, and a force graph to visually display the importance of the key variables and their influence on the prediction result; S6. Prediction and result optimization: Use the optimized model to predict the lignin extraction effect of different deep eutectic solvent systems, and combine the feature importance analysis results calculated by the SHAP method to adjust the composition ratio of the hydrogen bond donor to the acceptor of the deep eutectic solvent and the experimental conditions to optimize the separation process.
2. A method for predicting the effect of separating and extracting lignin with a deep eutectic solvent based on machine learning according to claim 1, characterized in that: The hydrogen bond acceptor includes choline chloride, tetramethylammonium chloride, tetraethylammonium chloride, tetrabutylammonium chloride, tetramethylammonium bromide, tetrabenzylammonium chloride, 2-hydroxyethyltrimethylammonium chloride, chlormequat chloride, and ethylene glycol; The hydrogen bond donor includes formic acid, lactic acid, urea, glycerol, oxalic acid, citric acid, succinic acid or amino acids, levulinic acid, malonic acid, phenylpropionic acid, xylitol, D-isosorbide and D-sorbitol, acetamide, urea, and ethylene glycol.
3. The method according to claim 1, characterized in that: In step S1, the hydrogen bond acidity, hydrogen bond basicity, and solvent polarization parameters are measured respectively using Nile red, 4-nitroaniline, N,N-diethyl-4-nitroaniline, and an ultraviolet spectrophotometer.
4. The method according to claim 1, characterized in that: The sample database in step S1 further includes the sum of the hydrogen bond acidity and basicity of the deep eutectic solvent, the molecular polarizability of the deep eutectic solvent, the solvent polarity of the deep eutectic solvent, the thermal stability of the deep eutectic solvent, the ash content of the biomass raw material, the initial density of the biomass raw material, the particle size distribution of the biomass raw material, the pretreatment temperature, the pretreatment time, the pretreatment pressure, the pretreatment heating rate, and the structural change rate of the biomass after pretreatment.
5. The method according to claim 1, wherein: In step S3, the constructed sample database is divided into a training set and a test set in a ratio of 7:3, and a machine learning model is constructed using the Stacking integration algorithm; the lignin removal rate is set as the target variable, and other solvent parameters and experimental conditions are set as input variables.
6. The method according to claim 1, wherein: In step S4, five-fold cross-validation is used for hyperparameter tuning to determine the optimal hyperparameter combination, and a machine learning prediction model is constructed according to the optimal hyperparameter results.
7. The method according to claim 1, characterized in that: In step S5, the results of the trained Stacking model are interpreted using the SHAP method, and the contribution value of each feature to the model prediction result is calculated; Generate a feature importance ranking graph through SHAP value analysis to visually display the positive and negative impacts and contribution degrees of each feature on the lignin removal rate; at the same time, analyze the action range and significance of different feature variables in combination with the SHAP value distribution graph, and screen out the key variables that have a greater impact on the lignin removal rate, including hydrogen bond acidity, hydrogen bond basicity, and molecular polarity index, and further evaluate the influence mechanism of these variables on the prediction results.
Citation Information
Cited By
Prediction method for lignin separation and extraction based on deep neural network model
CN122346667A