Machine learning based biomass composition prediction method, system, device, and medium

CN118197461BActive Publication Date: 2026-09-15CENT SOUTH UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410256159.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-06
Publication Date
2026-09-15
Estimated Expiration
2044-03-06

AI Technical Summary

Technical Problem

对生物质化学组成的了解是对其高效利用的必要前提,但在传统实验方法下测定生物质化学组成数据存在工作量大、时间及经济成本高、数据误差较大等问题

Benefits of technology

[0011] Firstly, by inputting the first elemental composition data into the first prediction model, the predicted biochemical composition of the biomass to be tested can be obtained quickly and easily, reducing time and economic costs. Secondly, by inputting the first elemental composition data and the predicted biochemical composition results into the second prediction model, further predictions can be made using the predicted biochemical composition results and the first elemental composition data to obtain the second elemental composition data for each biochemical component. This refines the elemental composition data for each biochemical component, providing a good foundation for subsequent predictions of monomer content. Finally, by inputting each second elemental composition data into the third prediction model, the monomers of each biochemical component can be predicted. Since the biochemical components are distinguished, the prediction accuracy is further improved, avoiding a large number of experimental and computational steps, reducing the complexity of biomass composition prediction, and improving the efficiency of biomass composition prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118197461B_ABST
    Figure CN118197461B_ABST
Patent Text Reader

Abstract

The application discloses a biomass composition prediction method, system, device and medium based on machine learning, and relates to the technical field of biomass composition prediction.The method comprises the following steps: acquiring first element composition data of a to-be-tested biomass; inputting the first element composition data into a preset first prediction model to obtain a biochemical composition prediction result corresponding to the to-be-tested biomass; inputting the first element composition data and the biochemical composition prediction result into a preset second prediction model to obtain second element composition data of each biochemical composition part corresponding to the biochemical composition prediction result; and inputting each second element composition data into a preset third prediction model to obtain biomass composition monomer content data of the to-be-tested biomass.The application can reduce the time and economic cost required by traditional experiments, and has the advantages of fast output, accurate results, no limitation on experimental conditions and simple operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biomass composition prediction technology, and in particular to a biomass composition prediction method, system, device and medium based on machine learning. Background Technology

[0002] Biomass refers to various biological organisms formed directly or indirectly through photosynthesis in nature, including animals, plants, and microorganisms. Due to its advantages such as high yield, low pollution, wide availability, and low price, biomass energy is considered an excellent renewable and clean energy source. The complexity of the chemical composition of biomass is considered a key issue in the field of biomass research. Biomass generated under different environments varies significantly in elemental and biochemical composition, and the types and amounts of each constituent monomer significantly affect its characteristic properties and potential value for production. Understanding the chemical composition of biomass is a necessary prerequisite for its efficient utilization; however, traditional experimental methods for determining the chemical composition of biomass suffer from problems such as high workload, high time and economic costs, and significant data errors.

[0003] To address this challenge, researchers have been trying to develop simpler and cheaper methods for determining the chemical composition of biomass. Examples include using Fourier transform infrared spectroscopy to determine the relative content of proteins, lipids, and polysaccharides in microalgal biomass; using high-performance liquid chromatography-electrospray ionization (HPLC-ECIS) to determine monosaccharides in cinnamon polysaccharides; and using hyperspectral imaging combined with chemometrics to rapidly determine the content of lignocellulose plant chemical components. However, these techniques are all derivatives of traditional experimental methods and still generally suffer from limitations imposed by raw materials and equipment, the need for extensive experiments and data comparisons, complex procedures, and high learning curves. They have failed to truly propose a method for detecting biomass chemical composition data that is applicable to a wide range of biomass, easy to operate, and not limited by experimental conditions. Summary of the Invention

[0004] This invention aims to address at least one of the technical problems existing in the prior art. To this end, this invention proposes a method, system, device, and medium for predicting biomass composition based on machine learning, which can reduce the time and economic costs required by traditional experiments, and provides fast results, is not limited by experimental conditions, and is simple to operate.

[0005] In a first aspect, embodiments of the present invention provide a biomass composition prediction method based on machine learning, comprising:

[0006] Obtain the first elemental composition data of the biomass to be tested;

[0007] Input the first element composition data into a preset first prediction model to obtain the biochemical composition prediction result corresponding to the biomass to be tested.

[0008] The first elemental composition data and the biochemical composition prediction result are input into a preset second prediction model to obtain the second elemental composition data of each biochemical component corresponding to the biochemical composition prediction result.

[0009] Each of the second element composition data is input into a preset third prediction model to obtain the biomass composition monomer content data of the biomass to be tested.

[0010] The method according to embodiments of the present invention has at least the following beneficial effects:

[0011] Firstly, by inputting the first elemental composition data into the first prediction model, the predicted biochemical composition of the biomass to be tested can be obtained quickly and easily, reducing time and economic costs. Secondly, by inputting the first elemental composition data and the predicted biochemical composition results into the second prediction model, further predictions can be made using the predicted biochemical composition results and the first elemental composition data to obtain the second elemental composition data for each biochemical component. This refines the elemental composition data for each biochemical component, providing a good foundation for subsequent predictions of monomer content. Finally, by inputting each second elemental composition data into the third prediction model, the monomers of each biochemical component can be predicted. Since the biochemical components are distinguished, the prediction accuracy is further improved, avoiding a large number of experimental and computational steps, reducing the complexity of biomass composition prediction, and improving the efficiency of biomass composition prediction.

[0012] According to some embodiments of the present invention, the first prediction model is obtained through the following steps:

[0013] Acquire the first theoretical element composition data, the theoretical biochemical composition data corresponding to the first theoretical element composition data, and the theoretical composition unit data corresponding to the theoretical biochemical composition data;

[0014] A biomass theoretical database is constructed using the first theoretical element composition data, the theoretical biochemical composition data, and the theoretical composition unit data.

[0015] A first machine learning model is constructed by randomly calling the first theoretical element composition data from the biomass theoretical database as the input of the first machine learning model, and using the theoretical biochemical composition data corresponding to the first theoretical element composition data as the output of the first machine learning model, thereby training the first machine learning model to obtain the first prediction model.

[0016] According to some embodiments of the present invention, the second prediction model is obtained through the following steps:

[0017] Construct a second machine learning model;

[0018] The second theoretical element composition data of each biochemical component of the theoretical biochemical composition data is calculated using the theoretical component data;

[0019] The theoretical elemental composition data and the theoretical biochemical composition data are used as inputs to the second machine learning model, and the second theoretical elemental composition data are used as outputs to train the second machine learning model to obtain the second prediction model.

[0020] According to some embodiments of the present invention, the third prediction model is obtained through the following steps:

[0021] Construct a third machine learning model;

[0022] The third machine learning model is trained by using the data composed of the second theoretical elements as input and the data composed of the theoretical elements as output.

[0023] According to some embodiments of the present invention, constructing a biomass theoretical database using the first theoretical elemental composition data, the theoretical biochemical composition data, and the theoretical component data includes the following steps:

[0024] Randomly obtain the constituent units from the theoretical constituent unit data;

[0025] The expected value and variance of the constituent units are randomly assigned within a preset data range, and the expected value and variance of other constituent units are randomly generated using a normal distribution.

[0026] Normalize each component by using the expected value and variance of all the component components to obtain the relative content data corresponding to each component component;

[0027] The biomass theoretical database is constructed using the first theoretical elemental composition data, the theoretical biochemical composition data, and the relative content data.

[0028] According to some embodiments of the present invention, after inputting the composition data of each second element into a preset third prediction model to obtain the biomass composition monomer content data of the biomass to be tested, the biomass composition prediction method based on machine learning further includes the following steps:

[0029] The first prediction performance evaluation result of the first prediction model, the second prediction performance evaluation result of the second prediction model, and the third prediction performance evaluation result of the third prediction model are calculated using the coefficient of determination and the root mean square error.

[0030] The training regression algorithms corresponding to the first prediction model, the second prediction model, and the third prediction model are adjusted based on the first prediction effect evaluation result, the second prediction effect evaluation result, and the third prediction effect evaluation result.

[0031] According to some embodiments of the present invention, the optimal hyperparameters of the first prediction model, the second prediction model, and the first prediction model are calculated by cross-validation error and optimization algorithms.

[0032] Secondly, embodiments of the present invention provide a biomass composition prediction system based on machine learning, comprising:

[0033] The first element composition data acquisition unit is used to acquire the first element composition data of the biomass to be tested.

[0034] A biochemical composition prediction unit is used to input the first elemental composition data into a preset first prediction model to obtain the biochemical composition prediction result corresponding to the biomass to be tested.

[0035] The second element composition data prediction unit is used to input the first element composition data and the biochemical composition prediction result into a preset second prediction model to obtain the second element composition data of each biochemical component corresponding to the biochemical composition prediction result.

[0036] The biomass composition monomer content data prediction unit is used to input the composition data of each second element into a preset third prediction model to obtain the biomass composition monomer content data of the biomass to be tested.

[0037] Thirdly, embodiments of the present invention provide an electronic device including at least one control processor and a memory for communicatively connecting to the at least one control processor; the memory stores instructions executable by the at least one control processor, the instructions being executed by the at least one control processor to enable the at least one control processor to perform the machine learning-based biomass composition prediction method as described in the first aspect.

[0038] Fourthly, embodiments of the present invention provide a computer storage medium storing computer-executable instructions for causing a computer to perform the machine learning-based biomass composition prediction method as described in the first aspect.

[0039] It should be noted that the beneficial effects of the second to fourth aspects of the present invention compared with the prior art are the same as the beneficial effects of the machine learning-based biomass composition prediction method of the first aspect, and will not be described in detail here.

[0040] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. Attached Figure Description

[0041] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which:

[0042] Figure 1 This is a flowchart of a biomass composition prediction method based on machine learning provided in an embodiment of the present invention;

[0043] Figure 2 This is a flowchart of obtaining the first prediction model according to an embodiment of the present invention;

[0044] Figure 3 This is a flowchart of obtaining the second prediction model according to an embodiment of the present invention;

[0045] Figure 4 This is a flowchart of obtaining the third prediction model according to an embodiment of the present invention;

[0046] Figure 5 This is a flowchart of an embodiment of the present invention for constructing a biomass theoretical database using first theoretical element composition data, theoretical biochemical composition data, and theoretical composition unit data;

[0047] Figure 6 This is a flowchart of a biomass composition prediction method based on machine learning, provided in an embodiment of the present invention, after inputting the composition data of each second element into a preset third prediction model to obtain the content data of biomass composition monomers of the biomass to be tested.

[0048] Figure 7 This is a schematic diagram of a specific embodiment of a biomass composition prediction method based on machine learning provided by an embodiment of the present invention;

[0049] Figure 8 This is a schematic diagram provided by an embodiment of the present invention, showing the proportion of the elemental composition corresponding to each biochemical component to the total elemental composition calculated using a biomass theoretical database;

[0050] Figure 9 This is a structural diagram of a biomass composition prediction system based on machine learning, provided in an embodiment of the present invention.

[0051] Figure 10 This is a structural diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0052] Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0053] In the description of this invention, the use of terms such as "first," "second," etc., is for the purpose of distinguishing technical features only and should not be construed as indicating or implying relative importance, or implicitly indicating the number of technical features indicated, or implicitly indicating the order of the technical features indicated.

[0054] In the description of this invention, it should be understood that the orientation descriptions, such as up, down, etc., are based on the orientation or positional relationship shown in the drawings and are only for the convenience of describing this invention and simplifying the description, and are not intended to indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this invention.

[0055] In the description of this invention, it should be noted that, unless otherwise explicitly defined, terms such as "setting," "installation," and "connection" should be interpreted broadly, and those skilled in the art can reasonably determine the specific meaning of the above terms in this invention in conjunction with the specific content of the technical solution.

[0056] The technical solution of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described below are some embodiments of the present invention, not all embodiments.

[0057] Reference Figure 1 In some embodiments of the present invention, a method for predicting biomass composition based on machine learning is provided, comprising:

[0058] Step S100: Obtain the first elemental composition data of the biomass to be tested.

[0059] It should be noted that the first elemental composition data refers to the composition ratio of all elements in the biomass to be tested. Preferably, the molecular composition and structural characteristics of the biomass only consider the organic components, i.e., dry and ash-free. The elemental composition includes five elements: carbon, hydrogen, oxygen, nitrogen, and sulfur.

[0060] Step S200: Input the first element composition data into the preset first prediction model to obtain the biochemical composition prediction result corresponding to the biomass to be tested.

[0061] It should be noted that the predicted biochemical composition results include proteins, lipids, cellulose, hemicellulose, carbohydrates, polysaccharides, lignin, and chitin.

[0062] Step S300: Input the first element composition data and the biochemical composition prediction results into the preset second prediction model to obtain the second element composition data of each biochemical component corresponding to the biochemical composition prediction results.

[0063] It should be noted that each biochemical component refers to a single biochemical component in the biochemical composition prediction result, such as a protein. Correspondingly, the second elemental composition data refers to the elemental composition ratio of the single biochemical component in the biochemical composition prediction result, such as the elemental composition ratio of a protein.

[0064] Step S400: Input the data of each second element composition into the preset third prediction model to obtain the biomass composition monomer content data of the biomass to be tested.

[0065] It should be noted that the biomass composition monomer content data refers to the proportion of each of the constituent monomers in a biochemical component. For example, proteins are composed of amino acids, including cysteine, tryptophan, methionine, histidine, tyrosine, isoleucine, threonine, arginine, glycine, serine, phenylalanine, proline, lysine, valine, alanine, leucine, aspartic acid, glutamic acid, asparagine, and glutamine. Therefore, the biomass composition monomer content data of proteins refers to the proportion of each of these monomers.

[0066] It should be noted that in this embodiment, when calculating data related to the number of atoms and molecular mass, the bonds and linkages between molecules are all included in the calculation. For example, these include: the dehydration condensation of the carboxyl group (C) between amino acids and the nitrogen group (N) to form a peptide bond; covalent bonding between complex sulfur atoms; hydrogen bonding between the side chain H and other amino acid O or N; esterification between glycerol and fatty acids to form ester bonds; dehydration condensation of hydroxyl groups between monosaccharides to form glycosidic bonds; dehydrogenation polymerization of phenylpropane units to form ether bonds or C / C bonds; and deacetylglucosamine to form glycosidic bonds.

[0067] This method first inputs the first elemental composition data into a first prediction model, which quickly and easily yields the predicted biochemical composition of the biomass to be tested, reducing time and economic costs. Secondly, the first elemental composition data and the predicted biochemical composition are input into a second prediction model, which further predicts the second elemental composition data for each biochemical component, refining the elemental composition data for each component and providing a solid foundation for subsequent predictions of monomer content. Finally, the second elemental composition data is input into a third prediction model, which predicts the monomer content of each biochemical component. Because the predictions are differentiated among the biochemical components, the accuracy is further improved. This method avoids numerous experimental and computational steps, reducing the complexity of biomass composition prediction and increasing its efficiency.

[0068] Reference Figure 2 In some embodiments of the present invention, the first prediction model is obtained through the following steps:

[0069] Step S210: Obtain the first theoretical element composition data, the theoretical biochemical composition data corresponding to the first theoretical element composition data, and the theoretical composition unit data corresponding to the theoretical biochemical composition data.

[0070] It should be noted that the data on the first theoretical elemental composition, the corresponding theoretical biochemical composition data, and the corresponding theoretical constituent unit data can be obtained through existing traditional methods, or by searching literature and other means to acquire existing data. Furthermore, the data on the first theoretical elemental composition, the corresponding theoretical biochemical composition data, and the corresponding theoretical constituent unit data are of the same data type as the aforementioned data on the first elemental composition, biochemical composition prediction results, and biomass constituent unit content. The difference lies in that the data on the first theoretical elemental composition, the corresponding theoretical biochemical composition data, and the corresponding theoretical constituent unit data are used for model training, while the data on the first elemental composition, biochemical composition prediction results, and biomass constituent unit content are used for actual prediction.

[0071] Step S220: Construct a biomass theoretical database using the first theoretical element composition data, theoretical biochemical composition data, and theoretical component data.

[0072] Preferably, the biomass theoretical database is composed of biomass theoretical sub-databases, each of which is independent of the others. The biomass theoretical sub-databases are obtained by dividing the biomass theoretical database according to each biochemical component in the theoretical biochemical composition data. That is, the biomass theoretical sub-database corresponding to each biochemical component includes the data of the biochemical component itself and the data of the theoretical component unit corresponding to the biochemical component.

[0073] Step S230: Construct a first machine learning model by randomly calling the first theoretical element composition data from the biomass theoretical database as the input of the first machine learning model, and using the theoretical biochemical composition data corresponding to the first theoretical element composition data as the output of the first machine learning model, and training the first machine learning model to obtain the first prediction model.

[0074] Reference Figure 3 In some embodiments of the present invention, the second prediction model is obtained through the following steps:

[0075] Step S310: Construct the second machine learning model.

[0076] Step S320: Calculate the theoretical biochemical composition data and the second theoretical element composition data of each biochemical component using the theoretical component data.

[0077] It should be noted that, correspondingly, the data type of the second theoretical elemental composition data is the same as that of the first theoretical elemental composition data, but the second theoretical elemental composition data refers to the elemental composition ratio of each biochemical component.

[0078] Step S330: Use the theoretical elemental composition data and theoretical biochemical composition data as input to the second machine learning model, and use the second theoretical elemental composition data as output to train the second machine learning model to obtain the second prediction model.

[0079] Reference Figure 4 In some embodiments of the present invention, the third prediction model is obtained through the following steps:

[0080] Step S410: Construct the third machine learning model.

[0081] Step S420: Use the data composed of the second theoretical elements as the input of the third machine learning model, and use the data of the theoretical components as the output of the third machine learning model to train the third machine learning model and obtain the third prediction model.

[0082] Reference Figure 5 In some embodiments of the present invention, a biomass theoretical database is constructed using first theoretical elemental composition data, theoretical biochemical composition data, and theoretical component data, including the following steps:

[0083] Step S221: Randomly obtain the constituent units from the theoretical constituent unit data.

[0084] Step S222: Randomly assign values ​​to the expected value and variance of the constituent units through a preset data range, and randomly generate the expected value and variance of other constituent units through a normal distribution.

[0085] It should be noted that the preset data range is derived from the data ranges of monomers corresponding to various biochemical components in nature. For example, taking proteins as an example, one amino acid is first assigned a mathematical expectation between 0-100% and a gradient between 0.1-10%. Then, the values ​​of other amino acids are randomly generated according to a normal distribution, keeping the sum of amino acids in each data set at 100%. For example, amino acids (mathematical expectation * 100 (%), variance * 100) include: cysteine ​​(1.26, 0.67); tryptophan (0.67, 0.39); methionine (1.08, 0.45); histidine (1.59, 0.26); tyrosine (1.92, 0.43); Isoleucine (3.86, 0.34); Threonine (5.13, 0.52); Arginine (3.44, 0.68); Glycine (10.03, 0.94); Serine (5.83, 0.71); Phenylalanine (3.81, 0.38); Proline (5.35, 0.88); Lysine (4.51, 0.7); Valine (6.07, 0.47); Alanine (10.46, 1.27); Leucine (7.66, 0.6); Aspartic acid (9.49, 0.97); Glutamic acid (10.14, 1.44); Asparagine (9.57, 0.98); Glutamine (10.22, 1.45).

[0086] Step S223: Normalize each component unit by using the mathematical expectation and variance of all component units to obtain the relative content data corresponding to each component unit.

[0087] It should be noted that the normalization calculation formula is as follows:

[0088]

[0089] in, x represents the result of normalizing the data in the i-th row and j-th column. ij S represents the original value of the data in the i-th row and j-th column. i This represents the sum of the data in the i-th row within the range of normalized data.

[0090] Step S224: Construct a biomass theoretical database using the first theoretical elemental composition data, theoretical biochemical composition data, and relative content data.

[0091] By limiting the data range, the proportion of constituent units in the biomass theoretical database is made reasonable. At the same time, the normal distribution reduces the error of the proportion of constituent units in the biomass theoretical database. Finally, by normalizing each constituent unit to obtain relative content data, the reasonableness of the proportion of constituent units in the biomass theoretical database is further improved. This provides good training data for the training of the first, second and third prediction models, and further improves the accuracy of prediction.

[0092] Reference Figure 6 In some embodiments of the present invention, after inputting the composition data of each second element into a preset third prediction model to obtain the content data of biomass constituent monomers of the biomass to be tested, the biomass composition prediction method based on machine learning further includes the following steps:

[0093] Step S500: Calculate the first prediction performance evaluation result of the first prediction model, the second prediction performance evaluation result of the second prediction model, and the third prediction performance evaluation result of the third prediction model using the coefficient of determination and root mean square error.

[0094] It should be noted that the formulas for calculating the coefficient of determination and root mean square error are as follows:

[0095]

[0096]

[0097] Among them, R 2 The coefficient of determination is represented by y, RMSE represents the root mean square error, and y represents the root mean square error. i This represents the theoretical value. Indicates the predicted value. The coefficient of determination represents the theoretical average value. The larger the coefficient of determination and the smaller the root mean square error, the better the prediction model performs.

[0098] Step S600: Adjust the training regression algorithms corresponding to the first prediction model, the second prediction model, and the third prediction model based on the first prediction effect evaluation results, the second prediction effect evaluation results, and the third prediction effect evaluation results.

[0099] It should be noted that the training regression algorithm includes random forest, artificial neural network and gradient boosting regression, etc. The training regression algorithm can be changed based on the first prediction effect evaluation result, the second prediction effect evaluation result and the third prediction effect evaluation result. The training regression algorithm of the first prediction model, the second prediction model and the third prediction model can be changed uniformly, or the training regression algorithm of the first prediction model, the second prediction model and the third prediction model can be changed separately. There are no specific restrictions here.

[0100] The prediction performance of the first, second, and third prediction models is calculated using the coefficient of determination and root mean square error, which avoids major prediction errors and allows for real-time adjustment of the training regression algorithm to improve the accuracy of each prediction.

[0101] In some embodiments of the present invention, the optimal hyperparameters of the first prediction model, the second prediction model, and the third prediction model are calculated using cross-validation error and optimization algorithms.

[0102] It should be noted that the optimization algorithms include particle swarm optimization, simulated annealing optimization, and Bayesian optimization. Hyperparameters include the number of regression trees, regression tree depth, loss function selection, learning rate, number of hidden layers, activation function, optimization weight metric, and maximum number of iterations.

[0103] By using cross-validation error and optimization algorithms, the optimal hyperparameters of the first, second, and third prediction models are calculated, further improving the accuracy of biomass composition prediction.

[0104] Reference Figure 7 To enable those skilled in the art to better understand this method, a specific embodiment of a machine learning-based biomass composition prediction method is provided, including:

[0105] Example 1:

[0106] A biomass theory sub-database was constructed using 20 amino acids, 26 fatty acids, and 2 monosaccharides as constituent monomers. The biomass theory sub-database was then merged to obtain the biomass theory database.

[0107] The specific number of atoms for the 20 amino acids included in the protein database is shown in Table 1 below:

[0108] Table 1. Examples of the number of atoms in amino acid monomers.

[0109]

[0110]

[0111] Based on the statistical range of amino acid content in proteins, the 20 amino acids mentioned above are considered to follow a normal distribution under specific parameters. First, each amino acid is assigned a mathematical expectation and variance, resulting in 101 data points. Then, the values ​​of other amino acids are proportionally adjusted according to random numbers under the assumed normal distribution, maintaining a 100% sum for each data set and ensuring each amino acid has an equal probability of being selected and assigned a value. Next, the elemental composition corresponding to each data set is calculated. This results in a data matrix where the first 5 columns represent elemental composition and the last 20 columns represent amino acid content, forming the biomass theory sub-database for proteins.

[0112] The fat database contains 26 fatty acids obtained by mixing glycerol, and their specific atomic numbers are shown in Table 2 below:

[0113] Table 2. Examples of Fatty Acid Monomer Atom Count

[0114]

[0115]

[0116] All oils and fats are considered as triglycerides formed by the dehydration condensation of 1 unit of glycerol and 3 units of fatty acids. The data presented in Table 2 already includes 1 / 3 of the atoms of glycerol in each fatty acid. Based on existing research, the proportion of each fatty acid in oils and fats is summarized. For each data set, the 26 fatty acids are randomly selected within their corresponding intervals, ensuring the sum of their values ​​is 100%. The elemental composition of each data set is then calculated. The resulting data matrix consists of the first 5 columns representing elemental composition and the last 26 columns representing fatty acid content, forming the sub-database of the oil and fat biomass theory.

[0117] The databases for cellulose, hemicellulose, carbohydrates, and polysaccharides are combined in this embodiment, and the specific atomic numbers of the two monosaccharides included are shown in Table 3 below:

[0118] Table 3. Examples of Monosaccharide Atom Counts

[0119]

[0120] Since pentose and hexose are the most common types of monosaccharides in nature, this embodiment only considers these two types when establishing the relevant theoretical database, without considering the influence of isomerism. The content of the two monosaccharides is treated as following a random distribution from 0-100%, with the sum of the values ​​remaining at 100%, and the elemental composition of each data set is calculated. This results in a data matrix where the first five columns represent the elemental composition and the last two columns represent the monosaccharide content, thus forming the biomass theoretical sub-database of carbohydrates.

[0121] Reference Figure 8 The above-mentioned biomass theory sub-databases are combined to obtain a complete biomass theory database. When combining databases, such as... Figure 8The left figure shows the content proportions of each biochemical component based on 36 possibilities. By adding random numbers, 2000 allocations are derived for each possibility, with each allocation maintaining a sum of 100%. This means that each biochemical component has the same probability of taking a value within the 0-100% range, ultimately resulting in 72,000 data points, as shown in the right figure (Figure 8). Each data point represents the chemical composition of a biomass, including the overall elemental composition, biochemical composition, and the content of individual monomers such as amino acids, fatty acids, and monosaccharides. Based on the overall biomass theoretical database, the proportion of each biochemical component's elemental composition to the total elemental composition is calculated, with each data set maintaining a sum of 100%.

[0122] Predictive model training:

[0123] Using the elemental composition of biomass as input variables and the biochemical composition of biomass as predictor variables, a first predictive model was trained. Using the elemental composition and biochemical composition of biomass as input variables and the proportion of each biochemical composition in the elemental composition as predictor variables, a second predictive model was trained. Using the elemental composition of protein as input variables and the content of 20 amino acids as predictor variables, a third predictive model for protein was trained. Using the elemental composition of cellulose, hemicellulose, carbohydrates, and polysaccharides as input variables and the content of two monosaccharides as predictor variables, a third training model for carbohydrates was trained. Using the elemental composition of lipids as input variables and the content of 26 fatty acids as predictor variables, a third training model for lipids was trained.

[0124] The random forest algorithm is used to construct the prediction model. Each model is cross-validated to adjust the two hyperparameters of the number of regression trees and the depth of regression trees to achieve the best prediction effect. The coefficient of determination and root mean square error are used as evaluation indicators for the prediction effect of each prediction model.

[0125] Example 2:

[0126] A biomass theory sub-database was constructed using pentose sugars, hexose sugars, hydrogen-rich lignin monomers, carbon-rich lignin monomers, and oxygen-rich lignin monomers as constituent monomers. All biomass theory sub-databases were combined to obtain the biomass theory database.

[0127] The other procedures are the same as in Example 1, the difference being in the training of the prediction model. Since the number of atoms and molecular mass of the monomers have already taken into account the intermolecular bonds and connection methods in the calculation, the prediction model training in Example 2 is as follows:

[0128] Using the C, H, and O content of biomass as input variables and the total cellulose content and total lignin content of biomass as prediction variables, the first prediction model was trained (the first prediction model here is labeled as Model 1-1).

[0129] The C, H, and O content of biomass elements were used as input variables, and the C1, H1, and O1 elements of biomass cellulose monomers and the C2, H2, and O2 elements of lignin monomers were used as prediction variables. The first prediction model was trained and obtained (the first prediction model here is labeled as Model 1-2).

[0130] Using the C, H, and O content of biomass elements as input variables, and the pentose sugars, hexose sugars, hydrogen-rich lignin monomers, carbon-rich lignin monomers, and oxygen-rich lignin monomers of biomass as prediction variables, the first prediction model was trained (the first prediction model here is labeled as Model 1-3).

[0131] The biomass elemental composition C, H, O content, total cellulose content, and total lignin content were used as input variables, and the biomass cellulose monomer elemental composition C1, H1, O1 and lignin monomer elemental composition C2, H2, O2 were used as prediction variables. The second prediction model (labeled as Model 2) was trained.

[0132] The C1, H1, and O1 content of biomass cellulose and hemicellulose elements were used as input variables, and the pentose and hexose components of biomass monomers were used as prediction variables. The third prediction model (labeled as Model 3-1 here) was trained.

[0133] Using the C2, H2, and O2 content of biomass lignin elemental composition as input variables, and the hydrogen-rich lignin monomer, carbon-rich lignin monomer, and oxygen-rich lignin monomer composition of biomass as prediction variables, a third prediction model (labeled as Model 3-2 here) was trained.

[0134] Reference Figure 9 An embodiment of the present invention also provides a biomass composition prediction system based on machine learning, including a first elemental composition data acquisition unit 1001, a biochemical composition prediction unit 1002, a second elemental composition data prediction unit 1003, and a biomass monomer content data prediction unit 1004, wherein:

[0135] The first element composition data acquisition unit 1001 is used to acquire the first element composition data of the biomass to be tested.

[0136] The biochemical composition prediction unit 1002 is used to input the first element composition data into a preset first prediction model to obtain the biochemical composition prediction result corresponding to the biomass to be tested.

[0137] The second element composition data prediction unit 1003 is used to input the first element composition data and the biochemical composition prediction results into a preset second prediction model to obtain the second element composition data of each biochemical component corresponding to the biochemical composition prediction results.

[0138] The biomass composition monomer content prediction unit 1004 is used to input the composition data of each second element into a preset third prediction model to obtain the biomass composition monomer content data of the biomass to be tested.

[0139] It should be noted that since the biomass composition prediction system based on machine learning in this embodiment is based on the same inventive concept as the biomass composition prediction method based on machine learning described above, the corresponding content in the method embodiment is also applicable to this device embodiment, and will not be described in detail here.

[0140] refer to Figure 10 In another embodiment of the present invention, an electronic device 6000 is also provided, which can be any type of smart terminal, such as a personal computer.

[0141] Specifically, the electronic device 6000 includes: one or more control processors 6001 and memory 6002. Figure 10 Taking a control processor 6001 and a memory 6002 as an example, the control processor 6001 and the memory 6002 can be connected via a bus or other means. Figure 10 Taking the bus connection between China and Israel as an example.

[0142] The memory 6002, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as program instructions / modules corresponding to an electronic device in an embodiment of the present invention.

[0143] The control processor 6001 executes various functional applications and data processing of a machine learning-based biomass composition prediction method by running non-transient software programs, instructions, and modules stored in the memory 6002, thereby implementing a machine learning-based biomass composition prediction method according to the above method embodiment.

[0144] The memory 6002 may include a program storage area and a data storage area. The program storage area may store an operating system and applications required for at least one function; the data storage area may store data created using a machine learning-based biomass composition prediction method, etc. Furthermore, the memory 6002 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory 6002 may optionally include memory remotely located relative to the control processor 6001, and these remote memories can be connected to the electronic device 6000 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0145] One or more modules are stored in memory 6002. When executed by one or more control processors 6001, a machine learning-based biomass composition prediction method from the above method embodiments is executed, such as the method described above. Figures 1 to 6 The method and steps.

[0146] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0147] It should be noted that since the electronic device in this embodiment is based on the same inventive concept as the above-described biomass composition prediction method based on machine learning, the corresponding content in the method embodiment is also applicable to this device embodiment, and will not be described in detail here.

[0148] One embodiment of the present invention also provides a computer-readable storage medium storing computer-executable instructions for performing: the machine learning-based biomass composition prediction method as described in the above embodiments.

[0149] It should be noted that since the computer-readable storage medium in this embodiment is based on the same inventive concept as the above-described biomass composition prediction method based on machine learning, the corresponding content in the method embodiment is also applicable to this device embodiment, and will not be described in detail here.

[0150] Those skilled in the art will understand that all or some of the steps and systems in the methods disclosed above can be implemented as software, firmware, hardware, and suitable combinations thereof. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing data (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired data and is accessible to a computer. Furthermore, as is known to those skilled in the art, communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any data delivery medium.

[0151] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0152] Although embodiments of the invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.

Claims

1. A method for predicting biomass composition based on machine learning, characterized in that, The machine learning-based biomass composition prediction method includes: Obtain the first elemental composition data of the biomass to be tested; The first element composition data is input into a preset first prediction model to obtain the biochemical composition prediction result corresponding to the biomass to be tested. The first prediction model is obtained through the following steps: Acquire the first theoretical element composition data, the theoretical biochemical composition data corresponding to the first theoretical element composition data, and the theoretical composition unit data corresponding to the theoretical biochemical composition data; A biomass theoretical database is constructed using the first theoretical element composition data, the theoretical biochemical composition data, and the theoretical component data, specifically as follows: Randomly obtain the constituent units from the theoretical constituent unit data; The expected value and variance of the constituent units are randomly assigned within a preset data range, and the expected value and variance of other constituent units are randomly generated using a normal distribution. Normalize each component by using the expected value and variance of all the component components to obtain the relative content data corresponding to each component component; The biomass theoretical database is constructed using the first theoretical elemental composition data, the theoretical biochemical composition data, and the relative content data. A first machine learning model is constructed by randomly calling the first theoretical element composition data in the biomass theoretical database as the input of the first machine learning model and using the theoretical biochemical composition data corresponding to the first theoretical element composition data as the output of the first machine learning model, thereby training the first machine learning model to obtain the first prediction model. The first elemental composition data and the biochemical composition prediction result are input into a preset second prediction model to obtain the second elemental composition data of each biochemical component corresponding to the biochemical composition prediction result. Each of the second element composition data is input into a preset third prediction model to obtain the biomass composition monomer content data of the biomass to be tested.

2. The biomass composition prediction method based on machine learning according to claim 1, characterized in that, The second prediction model is obtained through the following steps: Construct a second machine learning model; The second theoretical element composition data of each biochemical component of the theoretical biochemical composition data is calculated using the theoretical component data; The theoretical elemental composition data and the theoretical biochemical composition data are used as inputs to the second machine learning model, and the second theoretical elemental composition data are used as outputs to train the second machine learning model to obtain the second prediction model.

3. The biomass composition prediction method based on machine learning according to claim 2, characterized in that, The third prediction model is obtained through the following steps: Construct a third machine learning model; The third machine learning model is trained by using the data composed of the second theoretical elements as input and the data composed of the theoretical elements as output.

4. The biomass composition prediction method based on machine learning according to claim 1, characterized in that, After inputting the data of each second element into a preset third prediction model to obtain the biomass composition monomer content data of the biomass to be tested, the machine learning-based biomass composition prediction method further includes the following steps: The first prediction performance evaluation result of the first prediction model, the second prediction performance evaluation result of the second prediction model, and the third prediction performance evaluation result of the third prediction model are calculated using the coefficient of determination and the root mean square error. The training regression algorithms corresponding to the first prediction model, the second prediction model, and the third prediction model are adjusted based on the first prediction effect evaluation result, the second prediction effect evaluation result, and the third prediction effect evaluation result.

5. The biomass composition prediction method based on machine learning according to claim 1, characterized in that, The optimal hyperparameters of the first prediction model, the second prediction model, and the first prediction model are calculated using cross-validation error and optimization algorithms.

6. A biomass composition prediction system based on machine learning, characterized in that, The machine learning-based biomass composition prediction system includes: The first element composition data acquisition unit is used to acquire the first element composition data of the biomass to be tested. A biochemical composition prediction unit is used to input the first elemental composition data into a preset first prediction model to obtain the biochemical composition prediction result corresponding to the biomass to be tested. The first prediction model is obtained through the following steps: Acquire the first theoretical element composition data, the theoretical biochemical composition data corresponding to the first theoretical element composition data, and the theoretical composition unit data corresponding to the theoretical biochemical composition data; A biomass theoretical database is constructed using the first theoretical element composition data, the theoretical biochemical composition data, and the theoretical component data, specifically as follows: Randomly obtain the constituent units from the theoretical constituent unit data; The expected value and variance of the constituent units are randomly assigned within a preset data range, and the expected value and variance of other constituent units are randomly generated using a normal distribution. Normalize each component by using the expected value and variance of all the component components to obtain the relative content data corresponding to each component component; The biomass theoretical database is constructed using the first theoretical elemental composition data, the theoretical biochemical composition data, and the relative content data. A first machine learning model is constructed by randomly calling the first theoretical element composition data in the biomass theoretical database as the input of the first machine learning model and using the theoretical biochemical composition data corresponding to the first theoretical element composition data as the output of the first machine learning model, thereby training the first machine learning model to obtain the first prediction model. The second element composition data prediction unit is used to input the first element composition data and the biochemical composition prediction result into a preset second prediction model to obtain the second element composition data of each biochemical component corresponding to the biochemical composition prediction result. The biomass composition monomer content data prediction unit is used to input the composition data of each second element into a preset third prediction model to obtain the biomass composition monomer content data of the biomass to be tested.

7. An electronic device, characterized in that: It includes at least one control processor and a memory for communicatively connecting to the at least one control processor; the memory stores instructions executable by the at least one control processor, which, when executed by the at least one control processor, enable the at least one control processor to perform the machine learning-based biomass composition prediction method as described in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions for causing a computer to perform the machine learning-based biomass composition prediction method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Hydrothermal carbon property regulation and control method, system and equipment based on machine learning and medium

    CN117393073A