Lignocellulose biomass component separation method based on machine learning

By combining machine learning with various detection instruments and models, the problem of insufficient monitoring in traditional lignocellulose biomass component separation methods has been solved, achieving efficient and accurate separation results and improving separation efficiency.

CN120954569APending Publication Date: 2025-11-14GUANGXI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511103588.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-07
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Traditional methods for separating components from lignocellulose biomass are time-consuming, involve complex data, lack dynamic monitoring tools, resulting in large calculation errors, difficulty in accurately analyzing the nonlinear relationship between solvent functional groups and process parameters, and low separation efficiency.

Method used

Machine learning methods are employed to acquire data through Fourier transform infrared spectrometer, infrared thermometer, chromatograph, and thermogravimetric analyzer, constructing a closed-loop data system for the entire process. Solvent and reaction data are analyzed in real time, and a predictive model is built by combining random forest, XGBoost, and TabPFN models to dynamically monitor and accurately predict the separation effect.

Benefits of technology

It achieves high accuracy in multi-dimensional monitoring, excellent intelligent prediction and separation effect, eliminates moisture interference, ensures the reliability of separation rate calculation, and significantly improves separation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120954569A_ABST
    Figure CN120954569A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of acid pretreatment, and discloses a lignocellulose biomass component separation method based on machine learning, which comprises the following steps: 1, obtaining detection data of an acid solvent, an acid pretreatment environment and lignocellulose biomass components, and classifying to form a data set; 2, analyzing the change trend of each acidic solvent in the lignocellulose biomass component separation process in real time, generating a reaction data set, and dynamically monitoring the solvent degradation process; 3, the lignin separation rate, the hemicellulose separation rate and the cellulose separation rate in the lignocellulose biomass component separation process are analyzed in real time, the reliability of experimental data is ensured, and the multi-dimensional monitoring precision is high; and 4, analyzing the influence of different acidic pretreatment environments on the lignocellulose biomass component separation effect, constructing a prediction model, evaluating the accuracy, accurately analyzing the nonlinear relationship among solvent functional groups, process parameters and separation efficiency, and achieving a good intelligent prediction separation effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of acidic pretreatment technology, specifically a method for separating lignocellulose biomass components based on machine learning. Background Technology

[0002] Against the backdrop of the global fossil fuel crisis, developing renewable lignocellulosic biomass resources has become a key path for energy transition. Lignocellulosic biomass resources, due to their abundant content and excellent sustainable conversion characteristics, have gradually attracted the attention of researchers and the energy industry. Biomass generally refers specifically to cellulose, hemicellulose, and lignin contained in fibers. These are naturally occurring biomolecules widely distributed in nature and, as a potential green energy source, can, to some extent, replace traditional fossil fuels. However, the macromolecules in lignocellulosic biomass are usually tightly bound, which poses a challenge to achieving their high-value utilization. Research has found that suitable acidic pretreatment technology can effectively break down the tight bonds between lignocellulosic biomass molecules. The number of oxygen-containing functional groups in the solvent molecules, the appropriate reaction temperature, and the concentration of the acidic solvent play a crucial role in the separation efficiency of lignin. Through separation and purification techniques, natural biomass macromolecules can be effectively separated. To promote the development of biomass refining technology, there is an urgent need to utilize machine learning tools to output control strategies that can improve the separation rate.

[0003] Currently, research experiments on traditional methods for separating components from lignocellulose biomass are time-consuming, and the index data are complex and high-dimensional. Traditional experimental trial-and-error methods lack dynamic monitoring means, making it difficult to accurately analyze high-dimensional data. In the process of separating lignocellulose biomass components, the calculation deviation is easily caused by the moisture absorption of residues, making it difficult to accurately analyze the complex nonlinear relationship between solvent functional groups, process parameters and separation efficiency, resulting in low separation efficiency. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a machine learning-based method for separating lignocellulose biomass components. This method has advantages such as high accuracy of multidimensional monitoring and excellent intelligent prediction separation effect, solving the problems of lack of dynamic monitoring means and large index deviation in traditional lignocellulose biomass component separation methods.

[0005] To achieve the above objectives, the present invention provides the following technical solution: a machine learning-based method for separating lignocellulosic biomass components, comprising the following steps: Step 1: Connect the Fourier transform infrared spectrometer, infrared thermometer, chromatograph and thermogravimetric analyzer via network to acquire detection data of all acidic solvents, detection data of acidic pretreatment environment and detection data of lignocellulose biomass components, and classify them into solvent dataset, reaction dataset and component dataset. Step 2: Based on the solvent dataset and reaction dataset, analyze the changing trends of each acidic solvent in the separation process of lignocellulose biomass components in real time, and generate the corresponding reaction dataset. ; Step 3: Based on the component dataset, analyze the lignin separation rate in real time during the separation process of lignocellulose biomass components. hemicellulose separation rate and cellulose separation rate ; Step 4: Based on the solvent dataset, reaction dataset, component dataset, and reaction data set Lignin separation rate hemicellulose separation rate and cellulose separation rate The effects of different acidic pretreatment environments on the separation efficiency of lignocellulose biomass components were analyzed, and a predictive model was constructed. And assess the accuracy.

[0006] Preferably, in step one, the solvent dataset includes the number and concentration of oxygen-containing functional groups for each acidic solvent.

[0007] Preferably, in step one, the reaction dataset includes the reaction time and reaction temperature in the acidic pretreatment environment.

[0008] Preferably, in step one, the component dataset includes the oven-dry raw material mass, oven-dry residue mass, raw material lignin content, residue lignin content, raw material hemicellulose content, residue hemicellulose content, raw material cellulose content, and residue cellulose content of the lignocellulose biomass components.

[0009] Preferably, in step two, the reaction data set The calculation process is as follows: Based on the reaction dataset, the reaction time in the acidic pretreatment environment was labeled as... The reaction temperature of the acidic pretreatment environment is marked as... ; Based on the solvent dataset, extract the first... The detection data of the acidic solvent, and the data before separation and processing, the first The number of oxygen-containing functional groups in an acidic solvent is labeled as follows: Before separation processing, the first The concentration of the acidic solvent is labeled as After separation processing, the first The number of oxygen-containing functional groups in an acidic solvent is labeled as follows: After separation processing, the first The concentration of the acidic solvent is labeled as ; In the formula, Indicates the use of the first The change in reaction temperature during separation treatment with an acidic solvent. Indicates the first The rate of change of the number of oxygen-containing functional groups in an acidic solvent. Indicates the first The rate of change of the concentration of the acidic solvent, Indicates the use of the first A set of reaction data during the separation process using acidic solvents.

[0010] Preferably, in step three, the lignin separation rate is... The calculation process is as follows: Based on the component dataset, the mass of the oven-dried raw material before separation treatment is marked as follows: The lignin content of the raw material before separation treatment is marked as The mass of the oven-dried residue after separation is marked as follows: The lignin content of the residue after separation treatment is labeled as ; In the formula, This indicates the quality of the raw material lignin before separation treatment. This indicates the quality of the lignin in the oven-dried residue after separation treatment.

[0011] Preferably, in step three, the hemicellulose separation rate is... The calculation process is as follows: Based on the component dataset, the hemicellulose content of the raw material before separation treatment was marked as... The hemicellulose content of the residue after separation is labeled as ; In the formula, This indicates the mass of the raw material hemicellulose before separation processing. This indicates the mass of the oven-dried hemicellulose residue after separation and processing.

[0012] Preferably, in step three, the cellulose separation rate is... The calculation process is as follows: Based on the component dataset, the cellulose content of the raw material before separation treatment was marked as... The cellulose content of the residue after separation is labeled as ; In the formula, This indicates the mass of the raw material cellulose before separation and processing. This indicates the mass of the oven-dried cellulose residue after separation and treatment.

[0013] Preferably, in step four, the prediction model The build process is as follows: S1. Based on the solvent dataset, reaction dataset, component dataset, and reaction data set. Lignin separation rate hemicellulose separation rate and cellulose separation rate Experimental cases were statistically analyzed and compiled into an experimental dataset. In this dataset, the number of oxygen-containing functional groups in the acidic solvent, the concentration of the acidic solvent, the reaction time, and the reaction temperature varied for each experimental case. Each experimental case included the following parameters of the lignocellulose biomass components: oven-dry raw material mass, oven-dry residue mass, lignin content of the raw material, lignin content of the residue, hemicellulose content of the raw material, hemicellulose content of the residue, cellulose content of the raw material and cellulose content of the residue, change in reaction temperature, rate of change in the number of oxygen-containing functional groups in the acidic solvent, rate of change in the concentration of the acidic solvent, and lignin separation rate. hemicellulose separation rate and cellulose separation rate ; S2. The experimental cases are shuffled using a data shuffling method. The random seed number is set to 42. 80% of the experimental cases in the experimental dataset are divided into the training set, and 20% of the experimental cases in the experimental dataset are divided into the test set. S3. Select three models—Random Forest (RF), XGBoost, and TabPFN—to construct a prediction model. The training process involves focusing on the following parameters for each experimental case: number of oxygen-containing functional groups in the acidic solvent, concentration of the acidic solvent, reaction time, reaction temperature, oven-dry weight of the lignocellulose biomass components, oven-dry weight of the residue, lignin content of the raw material, lignin content of the residue, hemicellulose content of the raw material, hemicellulose content of the residue, cellulose content of the raw material, and cellulose content of the residue. These parameters are then used as the basis for the prediction model. The input features will be the reaction temperature change, the rate of change in the number of oxygen-containing functional groups in the acidic solvent, the rate of change in the concentration of the acidic solvent, and the lignin separation rate for each experimental case in the training set. hemicellulose separation rate and cellulose separation rate As a prediction model The prediction target.

[0014] Preferably, in step four, the accuracy judgment process is as follows: The number of oxygen-containing functional groups in the acidic solvent, the concentration of the acidic solvent, the reaction time, the reaction temperature, the oven-dry weight of the lignocellulose biomass components, the oven-dry weight of the residue, the lignin content of the raw material, the lignin content of the residue, the hemicellulose content of the raw material, the hemicellulose content of the residue, the cellulose content of the raw material, and the cellulose content of the residue are input into the prediction model for each experimental case in the test set. ; If the prediction model The output prediction targets and test sets include the reaction temperature change, the rate of change in the number of oxygen-containing functional groups in the acidic solvent, the rate of change in the acidic solvent concentration, and the lignin separation rate for each experimental case. hemicellulose separation rate and cellulose separation rate If they are different, it indicates that the prediction model The accuracy is low, and the prediction model should be rebuilt. ; If the prediction model The output prediction targets and test sets include the reaction temperature change, the rate of change in the number of oxygen-containing functional groups in the acidic solvent, the rate of change in the acidic solvent concentration, and the lignin separation rate for each experimental case. hemicellulose separation rate and cellulose separation rate If they are the same, it means the prediction model is the same. High accuracy is the key to the final prediction model. .

[0015] Compared with existing technologies, this invention provides a machine learning-based method for separating lignocellulosic biomass components, which has the following advantages: 1. This invention connects a Fourier transform infrared spectrometer, an infrared thermometer, a chromatograph, and a thermogravimetric analyzer via a network to acquire detection data for all acidic solvents, the acidic pretreatment environment, and the lignocellulosic biomass components. These data are then categorized into solvent datasets, reaction datasets, and component datasets. This breaks through the limitations of traditional single-device detection, constructing a closed-loop data process. Based on the solvent and reaction datasets, the invention analyzes in real-time the changing trends of each acidic solvent during the separation of lignocellulosic biomass components, generating corresponding reaction datasets. The solvent degradation process is dynamically monitored, and the lignin separation rate during the separation of lignocellulose biomass components is analyzed in real time based on the component dataset. hemicellulose separation rate and cellulose separation rate It corrects the absolute dry weight before and after separation treatment, eliminates moisture interference, ensures the reliability of separation rate calculation results, solves the calculation deviation problem caused by residue moisture absorption in traditional methods, and has high accuracy in multidimensional monitoring.

[0016] 2. This invention utilizes solvent datasets, reaction datasets, component datasets, and reaction data sets. Lignin separation rate hemicellulose separation rate and cellulose separation rate We statistically analyzed experimental cases and compiled them into an experimental dataset. We then shuffled the cases using a data shuffling method, dividing the dataset into training and test sets. Finally, we selected three models—Random Forest (RF), XGBoost, and TabPFN—to construct a prediction model. It is trained and learned to accurately analyze the complex nonlinear relationship between solvent functional groups, process parameters and separation efficiency, and intelligently predicts excellent separation results. Attached Figure Description

[0017] Figure 1 This is a diagram illustrating the steps of the method of the present invention. Detailed Implementation

[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] Because traditional methods for separating components from lignocellulose biomass are time-consuming, involve complex and high-dimensional data, and lack dynamic monitoring mechanisms, traditional trial-and-error methods struggle to accurately analyze high-dimensional data. Furthermore, the separation process is prone to significant calculation errors due to residue hygroscopicity, making it difficult to precisely resolve the complex nonlinear relationships between solvent functional groups, process parameters, and separation efficiency, resulting in low separation efficiency. Therefore, please refer to [the relevant documentation / reference needed]. Figure 1 This invention provides a machine learning-based method for separating lignocellulose biomass components, comprising the following steps: Step 1: Connect Fourier transform infrared spectrometer, infrared thermometer, chromatograph and thermogravimetric analyzer via network to acquire detection data of all acidic solvents, detection data of acidic pretreatment environment and detection data of lignocellulose biomass components, and classify them into solvent dataset, reaction dataset and component dataset, breaking the limitations of traditional single-device detection and building a closed loop of data for the whole process; Fourier transform infrared spectroscopy analyzer is used to acquire solvent datasets, which include the number and concentration of oxygen-containing functional groups for each acidic solvent; Infrared thermometers are used to collect reaction datasets, which include the reaction time and temperature in the acidic pretreatment environment. Chromatography and thermogravimetric analysis were used to collect component datasets, which included the oven-dry raw material mass, oven-dry residue mass, raw material lignin content, residue lignin content, raw material hemicellulose content, residue hemicellulose content, raw material cellulose content, and residue cellulose content of lignocellulose biomass components. Step 2: Based on the solvent dataset and reaction dataset, analyze the changing trends of each acidic solvent in the separation process of lignocellulose biomass components in real time, and generate the corresponding reaction dataset. ; Reaction Data Set The calculation process is as follows: Based on the reaction dataset, the reaction time in the acidic pretreatment environment was labeled as... The reaction temperature of the acidic pretreatment environment is marked as... ; Based on the solvent dataset, extract the first... The detection data of the acidic solvent, and the data before separation and processing, the first The number of oxygen-containing functional groups in an acidic solvent is labeled as follows: Before separation processing, the first The concentration of the acidic solvent is labeled as After separation processing, the first The number of oxygen-containing functional groups in an acidic solvent is labeled as follows: After separation processing, the first The concentration of the acidic solvent is labeled as ; In the formula, Indicates the use of the first The change in reaction temperature during separation using an acidic solvent reflects the intensity of exothermic or endothermic heat generation during the separation process. Indicates the first The rate of change in the number of oxygen-containing functional groups in an acidic solvent reflects the rate of decline in the chemical activity of the acidic solvent. Indicates the first The rate of change in the concentration of the acidic solvent reflects the rate of consumption of the acidic solvent. Indicates the use of the first The reaction data set during the separation and treatment of acidic solvents was used to dynamically monitor the solvent degradation process, providing data for subsequent training and prediction models. Data support was provided; Step 3: Based on the component dataset, analyze the lignin separation rate in real time during the separation process of lignocellulose biomass components. hemicellulose separation rate and cellulose separation rate ; Lignin separation rate The calculation process is as follows: Based on the component dataset, the mass of the oven-dried raw material before separation treatment is marked as follows: The lignin content of the raw material before separation treatment is marked as The mass of the oven-dried residue after separation is marked as follows: The lignin content of the residue after separation treatment is labeled as ; In the formula, This indicates the quality of the raw material lignin before separation treatment. This indicates the mass of lignin in the oven-dried residue after separation treatment; hemicellulose separation rate The calculation process is as follows: Based on the component dataset, the hemicellulose content of the raw material before separation treatment was marked as... The hemicellulose content of the residue after separation is labeled as ; In the formula, This indicates the mass of the raw material hemicellulose before separation processing. This indicates the mass of the oven-dried hemicellulose residue after separation and processing; Cellulose separation rate The calculation process is as follows: Based on the component dataset, the cellulose content of the raw material before separation treatment was marked as... The cellulose content of the residue after separation is labeled as ; In the formula, This indicates the mass of the raw material cellulose before separation and processing. This indicates the mass of the oven-dried cellulose residue after separation and processing; Specifically, by correcting the oven-dry mass before and after separation, moisture interference is eliminated, ensuring the reliability of the separation rate calculation results. This solves the calculation deviation problem caused by residue moisture absorption in traditional methods, and the multi-dimensional monitoring has high accuracy. Step 4: Based on the solvent dataset, reaction dataset, component dataset, and reaction data set Lignin separation rate hemicellulose separation rate and cellulose separation rate The effects of different acidic pretreatment environments on the separation efficiency of lignocellulose biomass components were analyzed, and a predictive model was constructed. And assess the accuracy; Predictive Model The build process is as follows: S1. Based on the solvent dataset, reaction dataset, component dataset, and reaction data set. Lignin separation rate hemicellulose separation rate and cellulose separation rate Experimental cases were statistically analyzed and compiled into an experimental dataset. In this dataset, the number of oxygen-containing functional groups in the acidic solvent, the concentration of the acidic solvent, the reaction time, and the reaction temperature varied for each experimental case. Each experimental case included the following parameters of the lignocellulose biomass components: oven-dry raw material mass, oven-dry residue mass, lignin content of the raw material, lignin content of the residue, hemicellulose content of the raw material, hemicellulose content of the residue, cellulose content of the raw material and cellulose content of the residue, change in reaction temperature, rate of change in the number of oxygen-containing functional groups in the acidic solvent, rate of change in the concentration of the acidic solvent, and lignin separation rate. hemicellulose separation rate and cellulose separation rate The experimental dataset comprehensively covers all key metrics in the acidic pretreatment process. S2. The experimental cases are shuffled using a data shuffling method. The random seed number is set to 42. 80% of the experimental cases in the experimental dataset are divided into the training set, and 20% of the experimental cases in the experimental dataset are divided into the test set. S3. Select three models—Random Forest (RF), XGBoost, and TabPFN—to construct a prediction model. The model combines the high robustness of Random Forest (RF), the feature ranking accuracy of XGBoost, and the efficient learning of TabPFN. The training set incorporates the following parameters for each experimental case: number of oxygen-containing functional groups in the acidic solvent, acidic solvent concentration, reaction time, reaction temperature, oven-dry raw material mass of lignocellulose biomass, oven-dry residue mass, lignin content of raw material, lignin content of residue, hemicellulose content of raw material, hemicellulose content of residue, cellulose content of raw material, and cellulose content of residue, all used as prediction models. The input features will be the reaction temperature change, the rate of change in the number of oxygen-containing functional groups in the acidic solvent, the rate of change in the concentration of the acidic solvent, and the lignin separation rate for each experimental case in the training set. hemicellulose separation rate and cellulose separation rate As a prediction model The predicted target; Predictive models with high generalization It can accurately analyze the complex nonlinear relationship between solvent functional groups, process parameters and separation efficiency, and can precisely control the amount of acidic solvent, significantly improve the separation yield, and thus provide strong data support for reactor design and continuous production. The accuracy assessment process is as follows: The number of oxygen-containing functional groups in the acidic solvent, the concentration of the acidic solvent, the reaction time, the reaction temperature, the oven-dry weight of the lignocellulose biomass components, the oven-dry weight of the residue, the lignin content of the raw material, the lignin content of the residue, the hemicellulose content of the raw material, the hemicellulose content of the residue, the cellulose content of the raw material, and the cellulose content of the residue are input into the prediction model for each experimental case in the test set. ; If the prediction model The output prediction targets and test sets include the reaction temperature change, the rate of change in the number of oxygen-containing functional groups in the acidic solvent, the rate of change in the acidic solvent concentration, and the lignin separation rate for each experimental case. hemicellulose separation rate and cellulose separation rate If they are different, it indicates that the prediction model The accuracy is low, and the prediction model should be rebuilt. ; If the prediction model The output prediction targets and test sets include the reaction temperature change, the rate of change in the number of oxygen-containing functional groups in the acidic solvent, the rate of change in the acidic solvent concentration, and the lignin separation rate for each experimental case. hemicellulose separation rate and cellulose separation rate If they are the same, it means the prediction model is the same. High accuracy is the key to the final prediction model. The intelligent prediction and separation effect is excellent.

[0020] Example 1: In this experiment, oxalic acid solution was selected as the experimental subject. The reaction time for separating lignocellulosic biomass components using oxalic acid solution was determined to be 5 hours, with a reaction temperature of 110℃-120℃. Before separation, the oxalic acid solution contained 100 oxygen-containing functional groups / mol and had a concentration of 2 mol / L. After separation, the oxalic acid solution contained 80 oxygen-containing functional groups / mol and had a concentration of 1.5 mol / L. The reaction data for separating lignocellulosic biomass components using oxalic acid solution were analyzed. The calculation process is as follows: In the formula, This indicates the change in reaction temperature when using oxalic acid solution for separation. This indicates the rate of change in the number of oxygen-containing functional groups in oxalic acid solution. This indicates the rate of change in the concentration of oxalic acid solution.

[0021] Example 2: In this experiment, a lignocellulosic biomass component with an oven-dried raw material mass of 100g was selected as the experimental subject. The lignin content of the raw material before separation treatment was 20%, and the oven-dried residue after separation treatment had a lignin content of 5% (10g of residue). The lignin separation rate of this lignocellulosic biomass component was determined. The calculation process is as follows: In the formula, This indicates the quality of the raw material lignin before separation treatment. This indicates the mass of lignin in the oven-dried residue after separation treatment, representing the lignin separation rate of this lignocellulosic biomass component. for .

[0022] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A machine learning-based method for separating lignocellulose biomass components, characterized in that, Includes the following steps: Step 1: Connect the Fourier transform infrared spectrometer, infrared thermometer, chromatograph and thermogravimetric analyzer via network to acquire detection data of all acidic solvents, detection data of acidic pretreatment environment and detection data of lignocellulose biomass components, and classify them into solvent dataset, reaction dataset and component dataset. Step 2: Based on the solvent dataset and reaction dataset, analyze the changing trends of each acidic solvent in the separation process of lignocellulose biomass components in real time, and generate the corresponding reaction dataset. ; Step 3: Based on the component dataset, analyze the lignin separation rate in real time during the separation process of lignocellulose biomass components. hemicellulose separation rate and cellulose separation rate ; Step 4: Based on the solvent dataset, reaction dataset, component dataset, and reaction data set Lignin separation rate hemicellulose separation rate and cellulose separation rate The effects of different acidic pretreatment environments on the separation efficiency of lignocellulose biomass components were analyzed, and a predictive model was constructed. And assess the accuracy.

2. The method for separating lignocellulosic biomass components based on machine learning according to claim 1, characterized in that: In step one, the solvent dataset includes the number and concentration of oxygen-containing functional groups for each acidic solvent.

3. The method for separating lignocellulosic biomass components based on machine learning according to claim 2, characterized in that: In step one, the reaction dataset includes the reaction time and reaction temperature in the acidic pretreatment environment.

4. The method for separating lignocellulosic biomass components based on machine learning according to claim 3, characterized in that: In step one, the component dataset includes the oven-dry raw material mass, oven-dry residue mass, raw material lignin content, residue lignin content, raw material hemicellulose content, residue hemicellulose content, raw material cellulose content, and residue cellulose content of the lignocellulose biomass components.

5. The method for separating lignocellulosic biomass components based on machine learning according to claim 4, characterized in that: In step two, the reaction data set The calculation process is as follows: Based on the reaction dataset, the reaction time in the acidic pretreatment environment was labeled as... The reaction temperature of the acidic pretreatment environment is marked as... ; Based on the solvent dataset, extract the first... The detection data of the acidic solvent, and the data before separation and processing, the first The number of oxygen-containing functional groups in an acidic solvent is labeled as follows: Before separation processing, the first The concentration of the acidic solvent is labeled as After separation processing, the first The number of oxygen-containing functional groups in an acidic solvent is labeled as follows: After separation processing, the first The concentration of the acidic solvent is labeled as ; In the formula, Indicates the use of the first The change in reaction temperature during separation treatment with an acidic solvent. Indicates the first The rate of change of the number of oxygen-containing functional groups in an acidic solvent. Indicates the first The rate of change of the concentration of the acidic solvent, Indicates the use of the first A set of reaction data during the separation process using acidic solvents.

6. The method for separating lignocellulosic biomass components based on machine learning according to claim 5, characterized in that: In step three, the lignin separation rate The calculation process is as follows: Based on the component dataset, the mass of the oven-dried raw material before separation treatment is marked as follows: The lignin content of the raw material before separation treatment is marked as The mass of the oven-dried residue after separation is marked as follows: The lignin content of the residue after separation treatment is labeled as ; In the formula, This indicates the quality of the raw material lignin before separation treatment. This indicates the quality of the lignin in the oven-dried residue after separation treatment.

7. The method for separating lignocellulosic biomass components based on machine learning according to claim 6, characterized in that: In step three, the hemicellulose separation rate The calculation process is as follows: Based on the component dataset, the hemicellulose content of the raw material before separation treatment was marked as... The hemicellulose content of the residue after separation is labeled as ; In the formula, This indicates the mass of the raw material hemicellulose before separation processing. This indicates the mass of the oven-dried hemicellulose residue after separation and processing.

8. The method for separating lignocellulosic biomass components based on machine learning according to claim 7, characterized in that: In step three, the cellulose separation rate The calculation process is as follows: Based on the component dataset, the cellulose content of the raw material before separation treatment was marked as... The cellulose content of the residue after separation is labeled as ; In the formula, This indicates the mass of the raw material cellulose before separation and processing. This indicates the mass of the oven-dried cellulose residue after separation and treatment.

9. The method for separating lignocellulosic biomass components based on machine learning according to claim 8, characterized in that: In step four, the prediction model The build process is as follows: S1. Based on the solvent dataset, reaction dataset, component dataset, and reaction data set. Lignin separation rate hemicellulose separation rate and cellulose separation rate Experimental cases were statistically analyzed and compiled into an experimental dataset. In this dataset, the number of oxygen-containing functional groups in the acidic solvent, the concentration of the acidic solvent, the reaction time, and the reaction temperature varied for each experimental case. Each experimental case included the following parameters of the lignocellulose biomass components: oven-dry raw material mass, oven-dry residue mass, lignin content of the raw material, lignin content of the residue, hemicellulose content of the raw material, hemicellulose content of the residue, cellulose content of the raw material and cellulose content of the residue, change in reaction temperature, rate of change in the number of oxygen-containing functional groups in the acidic solvent, rate of change in the concentration of the acidic solvent, and lignin separation rate. hemicellulose separation rate and cellulose separation rate ; S2. The experimental cases are shuffled using a data shuffling method. The random seed number is set to 42. 80% of the experimental cases in the experimental dataset are divided into the training set, and 20% of the experimental cases in the experimental dataset are divided into the test set. S3. Select three models—Random Forest (RF), XGBoost, and TabPFN—to construct a prediction model. The training process involves focusing on the following parameters for each experimental case: number of oxygen-containing functional groups in the acidic solvent, concentration of the acidic solvent, reaction time, reaction temperature, oven-dry weight of the lignocellulose biomass components, oven-dry weight of the residue, lignin content of the raw material, lignin content of the residue, hemicellulose content of the raw material, hemicellulose content of the residue, cellulose content of the raw material, and cellulose content of the residue. These parameters are then used as the basis for the prediction model. The input features will be the reaction temperature change, the rate of change in the number of oxygen-containing functional groups in the acidic solvent, the rate of change in the concentration of the acidic solvent, and the lignin separation rate for each experimental case in the training set. hemicellulose separation rate and cellulose separation rate As a prediction model The prediction target.

10. The method for separating lignocellulosic biomass components based on machine learning according to claim 9, characterized in that: In step four, the accuracy assessment process is as follows: The number of oxygen-containing functional groups in the acidic solvent, the concentration of the acidic solvent, the reaction time, the reaction temperature, the oven-dry weight of the lignocellulose biomass components, the oven-dry weight of the residue, the lignin content of the raw material, the lignin content of the residue, the hemicellulose content of the raw material, the hemicellulose content of the residue, the cellulose content of the raw material, and the cellulose content of the residue are input into the prediction model for each experimental case in the test set. ; If the prediction model The output prediction targets and test sets include the reaction temperature change, the rate of change in the number of oxygen-containing functional groups in the acidic solvent, the rate of change in the acidic solvent concentration, and the lignin separation rate for each experimental case. hemicellulose separation rate and cellulose separation rate If they are different, it indicates that the prediction model The accuracy is low, and the prediction model should be rebuilt. ; If the prediction model The output prediction targets and test sets include the reaction temperature change, the rate of change in the number of oxygen-containing functional groups in the acidic solvent, the rate of change in the acidic solvent concentration, and the lignin separation rate for each experimental case. hemicellulose separation rate and cellulose separation rate If they are the same, it means the prediction model is the same. High accuracy is the key to the final prediction model. .