A method, device and storage medium for predicting biomass pyrolysis efficiency based on machine learning

By constructing a biomass pyrolysis efficiency prediction model through machine learning, the problems of traditional methods being time-consuming, labor-intensive and costly are solved, and efficient and accurate prediction of biomass pyrolysis efficiency is achieved, which reduces experimental costs and expands the scope of application.

CN116386748BActive Publication Date: 2025-09-30NANJING UNIV OF INFORMATION SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310366541.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-07
Publication Date
2025-09-30
Estimated Expiration
2043-04-07

AI Technical Summary

Technical Problem

Existing technologies are time-consuming, labor-intensive, costly, and have strong limitations in predicting biomass pyrolysis efficiency. Traditional experimental methods cannot fully cover all biochar preparation situations.

Method used

A machine learning method was used to construct a biomass pyrolysis efficiency prediction model. The data set was processed by data cleaning, completion and normalization, and the random forest algorithm was used to construct a biomass pyrolysis efficiency model to predict biochar yield and pyrolysis energy yield.

Benefits of technology

It effectively saves manpower and time costs, improves prediction accuracy, can cover a wider range of biomass pyrolysis conditions, and reduces experimental costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116386748B_ABST
    Figure CN116386748B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, device, and storage medium for predicting biomass pyrolysis efficiency based on machine learning. The method comprises: obtaining a data set, the data set comprising a biomass material and experimental condition descriptor data set and a multi-type biomass characterization descriptor data set; cleaning the data set and supplementing missing values ​​to obtain a normalized biomass material and experimental condition descriptor data set; performing data descriptor screening on the normalized biomass material and experimental condition descriptor data set to determine a biomass pyrolysis efficiency data set; constructing and evaluating a biomass pyrolysis efficiency model using the biomass pyrolysis efficiency data set to obtain a trained biomass pyrolysis efficiency model; and predicting a target biomass pyrolysis efficiency using the trained biomass pyrolysis efficiency model to obtain a target biomass pyrolysis efficiency prediction result. The present invention solves the problem of missing biomass data on modeling, and the established model has a high degree of accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of machine learning and biomass material preparation, and specifically relates to a method for predicting biomass pyrolysis efficiency based on machine learning. Background Art

[0002] Biochar is a solid carbonaceous material obtained by high-temperature pyrolysis, dehydration, and carbonization of biomass. Biochar has good stability, high porosity, and a high specific surface area, and therefore has broad application prospects. Biochar yield and pyrolysis energy yield, as important indicators for evaluating biomass pyrolysis efficiency, are affected by many factors, including the type of raw biomass, pyrolysis temperature, and time. Higher pyrolysis temperatures and longer pyrolysis times generally result in higher energy yields, but may also lead to decreased biochar quality and yield and increased energy consumption. Therefore, when determining the optimal production conditions for biomass, many factors need to be considered to achieve a high biochar material yield and biomass pyrolysis energy yield. High pyrolysis efficiency means that relatively less raw material can produce more biochar material and energy, thereby reducing production costs and resource consumption.

[0003] Compared with obtaining the biomass pyrolysis efficiency value through traditional trial and error experiments, the machine learning-assisted modeling method is more efficient. At the same time, predicting biomass pyrolysis efficiency through machine learning models can better save labor costs. Therefore, how to establish a model to predict biomass pyrolysis efficiency is an urgent problem that needs to be solved.

[0004] The existing technology has the following problems: First, it is time-consuming and labor-intensive. In the traditional biomass pyrolysis efficiency, the determination of biochar yield requires burning the biomass material into biochar and then calculating the biochar yield by the mass ratio before and after burning. The determination of biomass pyrolysis energy yield requires obtaining the calorific value data before and after pyrolysis of the biomass and multiplying the ratio by the biochar yield. The specific calculation formula is as follows:

[0005]

[0006]

[0007] Therefore, in the actual experimental process, certain equipment, instruments, and laboratory environment are required, which involves a large amount of work and requires repeated experiments. This makes experimental measurements time-consuming and laborious. Second, the cost is high. Experimental measurements require the use of professional equipment and reagents, which are usually expensive. In addition, multiple measurements and repeated experiments are required during the experiment, which also increases costs. Third, the scope is narrow. Experimental measurements can usually only be performed on specific biomass and preparation conditions, and cannot fully cover all biochar preparation conditions. Therefore, there are certain limitations. Summary of the Invention

[0008] In order to solve the problems raised in the above background technology, the present invention provides a method, device and storage medium for predicting biomass pyrolysis efficiency based on machine learning, which has the characteristics of low testing cost and high accuracy.

[0009] To achieve the above object, the present invention provides the following technical solutions:

[0010] Technical solution: To solve the above technical problems, the technical solution adopted by the present invention is:

[0011] In a first aspect, the present invention provides a method for predicting biomass pyrolysis efficiency based on machine learning, comprising:

[0012] Step (1) obtaining a data set, wherein the data set includes a biomass material and experimental condition descriptor data set and a multi-type biomass characterization descriptor data set;

[0013] Step (2) cleaning the data set and filling in missing values ​​to obtain a normalized biomass material and experimental condition descriptor data set;

[0014] Step (3) performing data descriptor screening on the normalized biomass material and experimental condition descriptor data set to determine a biomass pyrolysis efficiency data set;

[0015] Step (4) constructing and evaluating a biomass pyrolysis efficiency model using the biomass pyrolysis efficiency dataset to obtain a trained biomass pyrolysis efficiency model;

[0016] Step (5) predicts the target biomass pyrolysis efficiency using the trained biomass pyrolysis efficiency model to obtain a target biomass pyrolysis efficiency prediction result, wherein the target biomass pyrolysis efficiency prediction result includes a biochar yield prediction result and a pyrolysis energy yield prediction result.

[0017] In step (1), the biomass material and experimental condition descriptor dataset is a dataset collected from existing public Chinese and English literature, and the descriptors are complete. The complete descriptors mean that they include biomass characterization descriptors, biomass pyrolysis condition descriptors, and biomass pyrolysis efficiency descriptors; the data in the biomass characterization descriptors may contain missing values, while the data in the biomass pyrolysis condition descriptors do not contain missing values.

[0018] In some embodiments, in step (1), the multi-class biomass characterization descriptor dataset is from a publicly available dataset in the GitHub database; and the data in the multi-class biomass characterization descriptor dataset must include biomass characterization descriptors.

[0019] In some embodiments, step (2) cleans the data set and fills in missing values ​​to obtain a normalized biomass material and experimental condition descriptor data set, including:

[0020] The biomass material and experimental condition descriptor dataset was preliminarily screened, and data with descriptor missing values ​​greater than 40% were deleted to obtain a preliminarily screened biomass material and experimental condition descriptor dataset;

[0021] The biomass material and experimental condition descriptor datasets that were initially screened were subjected to secondary data collation, and the missing item datasets were supplemented to obtain a supplementary version of the multi-type biomass characterization descriptor dataset;

[0022] Add the missing item datasets that have been completed in the supplementary version of the multi-type biomass characterization descriptor dataset to the preliminarily screened biomass material and experimental condition descriptor dataset to obtain the completed biomass material and experimental condition descriptor dataset;

[0023] The completed biomass material and experimental condition descriptor dataset is subjected to one-hot encoding and normalization processing to obtain a normalized biomass material and experimental condition descriptor dataset.

[0024] In some embodiments, the biomass material and experimental condition descriptor datasets obtained from the preliminary screening are subjected to secondary data collation to supplement the missing item datasets to obtain supplementary versions of multiple biomass characterization descriptor datasets, including:

[0025] S1. selecting all data with missing items in the biomass characterization descriptor data from the preliminarily screened biomass material and experimental condition descriptor dataset as a missing item dataset;

[0026] S2. Integrate the data of the missing item dataset and the multi-class biomass characterization descriptor dataset, and use the KNN algorithm to supplement the missing values ​​to obtain a supplemented version of the multi-class biomass characterization descriptor dataset.

[0027] In some embodiments, using the KNN algorithm to fill in missing values ​​includes:

[0028]

[0029] Where d is the Euclidean distance, x i is the value of the new sample with missing values, y i is the value of the sample with no missing values, i is the number of samples of the i-th feature, and n represents the total number of samples.

[0030] In some embodiments, step (3) performs data descriptor screening on the normalized biomass material and experimental condition descriptor dataset to determine a biomass pyrolysis efficiency dataset, including:

[0031] Calculate the linear correlation coefficients of pairwise descriptor data in the normalized biomass material and experimental condition descriptor datasets;

[0032] Descriptors with linear correlation coefficients greater than 0.9 were deleted, and features with strong correlation with biomass pyrolysis efficiency descriptors were retained as the biomass pyrolysis efficiency dataset;

[0033] The calculation of the linear correlation coefficients of the pairwise descriptor data in the normalized biomass material and experimental condition descriptor dataset includes:

[0034]

[0035] Where r represents the linear correlation coefficient between two variables x and y. and Represent the mean of x and y respectively, x i 、y i They represent the eigenvalues ​​corresponding to any random descriptor, and n represents the total number of samples.

[0036] In some embodiments, step (4) constructs and evaluates a biomass pyrolysis efficiency model using the biomass pyrolysis efficiency dataset to obtain a trained biomass pyrolysis efficiency model, including:

[0037] The biomass pyrolysis efficiency model is constructed using a random forest algorithm using a biomass pyrolysis efficiency dataset;

[0038] The biomass pyrolysis efficiency model is characterized by characteristic contribution, linear correlation coefficient R 2 The best prediction model was evaluated using the root mean square error (RMSE).

[0039] In a second aspect, the present invention provides a device for predicting biomass pyrolysis efficiency based on machine learning, comprising a processor and a storage medium;

[0040] The storage medium is used to store instructions;

[0041] The processor is configured to operate according to the instructions to perform the method according to the first aspect.

[0042] In a third aspect, the present invention provides a device comprising:

[0043] Memory;

[0044] processor;

[0045] as well as

[0046] computer programs;

[0047] The computer program is stored in the memory and is configured to be executed by the processor to implement the method described in the first aspect above.

[0048] In a fourth aspect, the present invention provides a storage medium having a computer program stored thereon, wherein the computer program implements the method described in the first aspect when executed by a processor.

[0049] Compared with the prior art, the beneficial effects (innovations) of the present invention are: the present invention proposes a method, device, equipment and storage medium for predicting biomass pyrolysis efficiency based on machine learning:

[0050] (1) By collecting existing data, a prediction model for biomass pyrolysis efficiency was constructed, including the prediction of biochar yield and pyrolysis energy yield, which can effectively save the manpower and time costs required to determine the pyrolysis efficiency parameters of biomass materials.

[0051] (2) A data cleaning and completion method was proposed. By rationally splitting the missing parts of the biomass pyrolysis data in a targeted manner, a new solution was provided for the missing problem of the biomass pyrolysis efficiency dataset constructed by machine learning. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 Schematic diagram of a method flow in an embodiment of the present invention;

[0053] Figure 2 Figure 2 shows the linear correlation between the characteristic values ​​of the biochar pyrolysis efficiency prediction model in the embodiment of the present invention (the left figure shows the linear correlation between the characteristics of biochar yield prediction, and the right figure shows the linear correlation between the characteristics of energy yield prediction);

[0054] Figure 3 Graph showing the contribution of each descriptor to the biochar pyrolysis efficiency prediction model in an embodiment of the present invention (the left graph shows biochar yield prediction, and the right graph shows energy yield prediction);

[0055] Figure 4 This is a linear correlation diagram between the predicted value and the true value in an embodiment of the present invention. DETAILED DESCRIPTION

[0056] The present invention will be further described below in conjunction with the accompanying drawings and examples. The following examples are only used to more clearly illustrate the technical solutions of the present invention and are not intended to limit the scope of protection of the present invention.

[0057] In the description of the present invention, "several" means more than one, "plurality" means more than two, "greater than," "less than," and "exceed" are understood to exclude the number itself, while "above," "below," and "within" are understood to include the number itself. The use of "first" and "second" in the description is solely for the purpose of distinguishing technical features and should not be construed as indicating or implying relative importance, implicitly specifying the number of the indicated technical features, or implicitly specifying the order of the indicated technical features.

[0058] In the description of the present invention, reference to terms such as "one embodiment," "some embodiments," "illustrative embodiments," "examples," "specific examples," or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the exemplary expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0059] Example 1

[0060] like Figure 1 As shown, a method for predicting biomass pyrolysis efficiency based on machine learning includes:

[0061] Step (1) obtaining a data set, wherein the data set includes a biomass material and experimental condition descriptor data set and a multi-type biomass characterization descriptor data set;

[0062] Step (2) cleaning the data set and filling in missing values ​​to obtain a normalized biomass material and experimental condition descriptor data set;

[0063] Step (3) performing data descriptor screening on the normalized biomass material and experimental condition descriptor data set to determine a biomass pyrolysis efficiency data set;

[0064] Step (4) constructing and evaluating a biomass pyrolysis efficiency model using the biomass pyrolysis efficiency dataset to obtain a trained biomass pyrolysis efficiency model;

[0065] Step (5) predicts the target biomass pyrolysis efficiency using the trained biomass pyrolysis efficiency model to obtain a target biomass pyrolysis efficiency prediction result, wherein the target biomass pyrolysis efficiency prediction result includes a biochar yield prediction result and a pyrolysis energy yield prediction result.

[0066] In some embodiments, in step (1), the biomass material and experimental condition descriptor dataset is a dataset collected from existing public Chinese and English literature, and the descriptors are complete. The complete descriptors refer to the dataset containing biomass characterization descriptors, biomass pyrolysis condition descriptors, and biomass pyrolysis efficiency descriptors; the data in the biomass characterization descriptors may contain missing values, while the data in the biomass pyrolysis condition descriptors do not contain missing values.

[0067] In some embodiments, the biomass characterization descriptors specifically include: biomass material name - Biomass material (BM), fixed carbon content of biochar raw material - Fixed carbon (FC), volatile matter content - Volatile matter (VM), ash content of biomaterial - ash content (Ash), carbon content (C), hydrogen content (H), oxygen content (O), nitrogen content (N), sulfur content (S), cellulose - Cellulose (Cel), hemicellulose - Hemicellulose (Hem), lignin - Lignin (Lig); biomass pyrolysis condition descriptors specifically include residence time - Residence time (RT), pyrolysis temperature - Temperature (PT), heating rate - Heating rate (HR); biomass pyrolysis efficiency descriptors specifically include: biochar yield - Biochar yield (BY), energy yield - Energy yield (EY).

[0068] In some embodiments, in step (1), the multi-class biomass characterization descriptor dataset is from a publicly available dataset in the GitHub database; and the data in the multi-class biomass characterization descriptor dataset must include biomass characterization descriptors.

[0069] By adopting the above technical solution, a variety of biomass characterization descriptor datasets are obtained by retrieving keywords from the GitHub database and searching Biochar. The descriptors collected in the variety of biomass characterization descriptor datasets must be complete and include biomass characterization descriptors. Since the data in the variety of biomass characterization descriptor datasets are biomass characterization data, the acquisition scope is wider than that of the biomass material and experimental condition descriptor dataset during the collection process. It is not necessary to cover biomass characterization descriptors, biomass pyrolysis condition descriptors, and biomass pyrolysis efficiency descriptors like the biomass material and experimental condition descriptor dataset. It can be understood that the descriptors of the biomass characterization descriptor dataset are included in the descriptors of the biomass material and experimental condition descriptor dataset, so it provides the possibility for data supplementation.

[0070] In some embodiments, step (2) cleans the data set and fills in missing values ​​to obtain a normalized biomass material and experimental condition descriptor data set, including:

[0071] The biomass material and experimental condition descriptor dataset was preliminarily screened, and data with descriptor missing values ​​greater than 40% were deleted to obtain a preliminarily screened biomass material and experimental condition descriptor dataset;

[0072] The biomass material and experimental condition descriptor datasets that were initially screened were subjected to secondary data collation, and the missing item datasets were supplemented to obtain a supplementary version of the multi-type biomass characterization descriptor dataset;

[0073] Add the missing item datasets that have been completed in the supplementary version of the multi-type biomass characterization descriptor dataset to the preliminarily screened biomass material and experimental condition descriptor dataset to obtain the completed biomass material and experimental condition descriptor dataset;

[0074] The completed biomass material and experimental condition descriptor dataset is subjected to one-hot encoding and normalization processing to obtain a normalized biomass material and experimental condition descriptor dataset.

[0075] In some embodiments, the biomass material and experimental condition descriptor datasets obtained from the preliminary screening are subjected to secondary data collation to supplement the missing item datasets to obtain supplementary versions of multiple biomass characterization descriptor datasets, including:

[0076] S1. selecting all data with missing items in the biomass characterization descriptor data from the preliminarily screened biomass material and experimental condition descriptor dataset as a missing item dataset;

[0077] S2. Integrate the data of the missing item dataset and the multi-class biomass characterization descriptor dataset, and use the KNN algorithm to supplement the missing values ​​to obtain a supplemented version of the multi-class biomass characterization descriptor dataset.

[0078] By adopting the above technical solution, the biomass material and experimental condition descriptor dataset can be cleaned. The cleaned dataset still contains missing values, but the data integrity has been greatly improved compared to the dataset before cleaning.

[0079] By adopting the above technical solution, the secondary data sorting helps to supplement the biomass material and experimental condition descriptor dataset that still has missing values. Since the biomass material and experimental condition descriptor dataset has been cleaned, and only the biomass characterization descriptor is missing in the biomass material and experimental condition descriptor dataset, the descriptor corresponding to the selected missing value is only the biomass characterization descriptor. The missing item dataset and multiple types of biomass characterization descriptor datasets are integrated and put into the same CSV format file to facilitate software calling. Since the missing values ​​of the biomass material and experimental condition descriptor dataset only come from the data in the biomass characterization descriptor, and the biomass characterization descriptor is an independent dataset relative to the biomass pyrolysis condition descriptor, the two are parallel factors. They form a mapping relationship with the biomass pyrolysis efficiency descriptor data, thereby constructing a biomass pyrolysis efficiency model, so they can be extracted separately;

[0080] The use of KNN algorithm to supplement and integrate the data set is beneficial to improve the accuracy of biomass pyrolysis efficiency modeling.

[0081] In some embodiments, using the KNN algorithm to fill in missing values ​​includes:

[0082]

[0083] Where d is the Euclidean distance, x i is the value of the new sample with missing values, y i is the value of the sample with no missing values, i is the number of samples of the i-th feature, and n represents the total number of samples.

[0084] By adopting the above technical solution, in the preliminary screened biomass material and experimental condition descriptor dataset, the name of the biochar material is in text type and cannot be recognized by the algorithm, so it is processed with one-hot encoding: as shown in Table 1:

[0085]

[0086] Table 1

[0087] Using one-hot encoding, the text feature of the biomass material name in the biomass characterization descriptor is converted into a numerical feature. In the preliminarily screened biomass material and experimental condition descriptor dataset, the value ranges of different descriptors vary greatly, which may lead to the overestimation of the importance of some features. In order to improve the stability, accuracy and calculation speed of the model, normalization is used to reduce the overestimation of feature importance, as shown in the following formula:

[0088]

[0089] In the above formula, y is the normalized data, x is the data value in the descriptor, and X minFor the minimum value in this descriptor data, X max The maximum value in this descriptor's data.

[0090] In some embodiments, step (3) performs data descriptor screening on the normalized biomass material and experimental condition descriptor dataset to determine a biomass pyrolysis efficiency dataset, including:

[0091] Calculate the linear correlation coefficients of pairwise descriptor data in the normalized biomass material and experimental condition descriptor datasets;

[0092] Descriptors with linear correlation coefficients greater than 0.9 were deleted, and features with strong correlation with biomass pyrolysis efficiency descriptors were retained as the biomass pyrolysis efficiency dataset.

[0093] The calculation formula for the linear correlation coefficient of the pairwise descriptor data in the normalized biomass material and experimental condition descriptor dataset is:

[0094]

[0095] Where r represents the linear correlation coefficient between two variables x and y. and Represent the mean of x and y respectively, x i 、y i They represent the eigenvalues ​​corresponding to any random descriptor, and n represents the total number of samples.

[0096] By adopting the above technical solutions, it helps to improve analysis efficiency: reduce multicollinearity: and improve model performance.

[0097] like Figure 2 As shown, the left figure is a linear correlation diagram between the characteristic values ​​of biochar yield prediction, and the right figure is a linear correlation diagram between the characteristic values ​​of pyrolysis energy yield prediction.

[0098] In some embodiments, step (4) constructs and evaluates a biomass pyrolysis efficiency model using the biomass pyrolysis efficiency dataset to obtain a trained biomass pyrolysis efficiency model, including:

[0099] The biomass pyrolysis efficiency model is constructed using a random forest algorithm using a biomass pyrolysis efficiency dataset; the construction formula is as follows:

[0100]

[0101] in represents the prediction of the final tree, B is the number of decision trees, b represents the current tree, and x is the training sample.

[0102] The biomass pyrolysis efficiency model is characterized by characteristic contribution, linear correlation coefficient R 2 The best prediction model is evaluated by using the root mean square error (RMSE);

[0103]

[0104] in y、 and are the predicted, actual, and average values ​​of the object descriptor, respectively, n is the number of data points for any given instance, and N is the total number of data points.

[0105] By adopting the above technical solution, the data is substituted into the random forest model, and Gridsearch is used to tune the hyperparameters. Finally, a trained biomass pyrolysis efficiency model is established. The contribution of each descriptor feature is as follows: Figure 3 (The left figure shows the contribution of each descriptor to the prediction of biochar yield, and the right figure shows the contribution of each descriptor to the prediction of energy yield), linear correlation coefficient R 2 The root mean square error RMSE is shown in Table 2: Figure 4 is the linear correlation graph between the predicted value and the true value.

[0106] Table 2

[0107]

[0108] Example 2

[0109] In a second aspect, this embodiment provides a device for predicting biomass pyrolysis efficiency based on machine learning, comprising a processor and a storage medium;

[0110] The storage medium is used to store instructions;

[0111] The processor is configured to operate according to the instructions to perform the method according to embodiment 1.

[0112] Example 3

[0113] In a third aspect, this embodiment provides a device, including:

[0114] Memory;

[0115] processor;

[0116] as well as

[0117] computer programs;

[0118] The computer program is stored in the memory and is configured to be executed by the processor to implement the method described in embodiment 1.

[0119] Example 4

[0120] In a fourth aspect, this embodiment provides a storage medium on which a computer program is stored. When the computer program is executed by a processor, the method described in Example 1 is implemented.

[0121] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0122] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0123] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0124] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0125] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

[0126] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A method for predicting biomass pyrolysis efficiency based on machine learning, characterized in that: include: Step (1) obtaining a data set, wherein the data set includes a biomass material and experimental condition descriptor data set and a multi-type biomass characterization descriptor data set; the biomass material and experimental condition descriptor data set is a data set obtained by collecting existing public Chinese and English literature, and the descriptors are complete, and the complete descriptors refer to the data set including biomass characterization descriptors, biomass pyrolysis condition descriptors, and biomass pyrolysis efficiency descriptors; the data in the biomass characterization descriptors may contain missing values, while the data in the biomass pyrolysis condition descriptors do not contain missing values; Step (2) cleaning the data set and filling in missing values ​​to obtain a normalized biomass material and experimental condition descriptor data set, including: The biomass material and experimental condition descriptor dataset was preliminarily screened, and data with descriptor missing values ​​greater than 40% were deleted to obtain a preliminarily screened biomass material and experimental condition descriptor dataset; The preliminary screened biomass material and experimental condition descriptor dataset is subjected to secondary data collation to supplement the missing item dataset to obtain a supplementary version of the multi-class biomass characterization descriptor dataset, specifically comprising: S1, selecting the data of the biomass characterization descriptor data in the preliminary screened biomass material and experimental condition descriptor dataset that still has missing items as a whole, as the missing item dataset; S2, integrating the data of the missing item dataset with the multi-class biomass characterization descriptor dataset, and using the KNN algorithm to supplement the missing values ​​to obtain a supplementary version of the multi-class biomass characterization descriptor dataset; Add the missing item datasets that have been completed in the supplementary version of the multi-type biomass characterization descriptor dataset to the preliminarily screened biomass material and experimental condition descriptor dataset to obtain the completed biomass material and experimental condition descriptor dataset; Performing one-hot encoding and normalization on the completed biomass material and experimental condition descriptor dataset to obtain a normalized biomass material and experimental condition descriptor dataset; Step (3) performing data descriptor screening on the normalized biomass material and experimental condition descriptor data set to determine a biomass pyrolysis efficiency data set; Step (4) constructing and evaluating a biomass pyrolysis efficiency model using the biomass pyrolysis efficiency dataset to obtain a trained biomass pyrolysis efficiency model; Step (5) predicts the target biomass pyrolysis efficiency using the trained biomass pyrolysis efficiency model to obtain a target biomass pyrolysis efficiency prediction result, wherein the target biomass pyrolysis efficiency prediction result includes a biochar yield prediction result and a pyrolysis energy yield prediction result.

2. The method for predicting biomass pyrolysis efficiency based on machine learning according to claim 1, characterized in that: In step (1), the multi-class biomass characterization descriptor dataset is from a publicly available dataset in the GitHub database; and the data in the multi-class biomass characterization descriptor dataset must contain biomass characterization descriptors.

3. The method for predicting biomass pyrolysis efficiency based on machine learning according to claim 1, characterized in that: Using the KNN algorithm to fill missing values ​​includes: Where d is the Euclidean distance, x i is the value of the new sample with missing values, y i is the value of the sample without missing values, i is the number of samples of the i-th feature, and n represents the total number of samples.

4. The method for predicting biomass pyrolysis efficiency based on machine learning according to claim 1, characterized in that: Step (3) performs data descriptor screening on the normalized biomass material and experimental condition descriptor dataset to determine the biomass pyrolysis efficiency dataset, including: Calculate the linear correlation coefficients of pairwise descriptor data in the normalized biomass material and experimental condition descriptor datasets; Descriptors with linear correlation coefficients greater than 0.9 were deleted, and features with strong correlation with biomass pyrolysis efficiency descriptors were retained as the biomass pyrolysis efficiency dataset; The calculation of the linear correlation coefficients of the pairwise descriptor data in the normalized biomass material and experimental condition descriptor dataset includes: Where r represents the linear correlation coefficient between two variables x and y. and Represent the mean of x and y respectively, x i 、y i They represent the eigenvalues ​​corresponding to any random descriptor, and n represents the total number of samples.

5. The method for predicting biomass pyrolysis efficiency based on machine learning according to claim 1, characterized in that: Step (4) constructs and evaluates a biomass pyrolysis efficiency model using the biomass pyrolysis efficiency dataset to obtain a trained biomass pyrolysis efficiency model, including: The biomass pyrolysis efficiency model is constructed using a random forest algorithm using a biomass pyrolysis efficiency dataset; The biomass pyrolysis efficiency model is characterized by characteristic contribution, linear correlation coefficient R 2 The model is evaluated using the root mean square error (RMSE).

6. An electronic device, characterized in that: include: Memory; processor; as well as computer programs; The computer program is stored in the memory and configured to be executed by the processor to implement the method according to any one of claims 1 to 5.

7. A storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Method for interpolating and supplementing missing values based on neighbor algorithm in power load prediction

    CN111768034A

  • Method for efficiently predicting band gap of semiconductor material based on machine learning

    CN114724652A