A method, system and electronic device for identifying biochar

By constructing high-dimensional multi-class sample data of biochar and utilizing the random subspace nearest neighbor clustering ensemble learning algorithm, the problem of biochar identification was solved, achieving efficient and accurate biochar identification and promoting the diversified and standardized application of biochar products.

CN116721707BActive Publication Date: 2025-10-31ZHEJIANG UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310692030.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-12
Publication Date
2025-10-31
Estimated Expiration
2043-06-12

AI Technical Summary

Technical Problem

The lack of efficient and accurate methods for identifying biochar in existing technologies makes it difficult to standardize the quality of biochar products, affecting their applicability in different application fields and their industrial development.

Method used

By acquiring the physicochemical property index data of target individuals, a high-dimensional multi-class sample data of biochar is constructed. Outlier sample testing and standardization are performed. A biochar identity discrimination model is established using a random subspace nearest neighbor clustering ensemble learning algorithm to achieve the identification of biochar identity information.

Benefits of technology

It achieves efficient and accurate identification of biochar identity with high accuracy, good result stability, and strong generalization ability, supporting the diversified and standardized production of biochar products.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116721707B_ABST
    Figure CN116721707B_ABST
Patent Text Reader

Abstract

This invention discloses a method, system, and electronic device for biochar identification, relating to the field of biochar detection and pattern recognition technology. The method includes: inputting the physicochemical properties of a target individual into a biochar identification model to obtain the target biochar's identification information; wherein the process of determining the biochar identification model involves: performing outlier detection and standardization on the sample input variable data matrix, and constructing a feature data matrix based on the processed sample input variable data matrix; obtaining a random subspace nearest neighbor clustering ensemble learning classifier based on the feature data matrix, the sample identity multi-class label column vector, and a random subspace nearest neighbor clustering ensemble learning algorithm; the random subspace nearest neighbor clustering ensemble learning classifier is the biochar identification model. This invention can efficiently and accurately identify biochar identification information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biochar detection and pattern recognition technology, and in particular to a biochar identification method, system and electronic device. Background Technology

[0002] Biochar is a carbon-rich, porous solid material obtained by placing biomass raw materials in a high-temperature (usually between 300°C and 700°C) oxygen-free or low-oxygen environment and then pyrolyzing or oxidizing them. It has good antioxidant properties, heat resistance, and adsorption capacity.

[0003] Dethermally carbonizing waste biomass (such as crop straw, livestock manure, forestry waste, perishable waste, and industrial sludge) into biochar not only effectively reduces pollution and greenhouse gas emissions from waste biomass, but also allows biochar to be used as a soil conditioner, increasing soil organic matter content, improving soil structure, promoting plant growth, and achieving green and sustainable development through environmental protection and resource recycling. Simultaneously, biochar production will provide employment opportunities in rural areas, promoting rural economic development and common prosperity.

[0004] Due to differences in raw material types, technical methods, and pyrolysis processes, biochar exhibits significant variations in its physicochemical properties, such as structure, composition, pore volume, and specific surface area, resulting in different environmental effects. As the applications of biochar continue to expand, classifying biochar based on properties such as carbon storage value, fertilizer value, lime equivalent value, and particle size is crucial for standardizing product quality, facilitating the selection of suitable biochar for application, promoting diversified, standardized, and serialized production of biochar products, and contributing to the sustainable development of the biochar industry. However, currently, there is no efficient and accurate method for identifying the type of biochar. Summary of the Invention

[0005] The purpose of this invention is to provide a method, system, and electronic device for identifying biochar, which can efficiently and accurately identify biochar identity information.

[0006] To achieve the above objectives, the present invention provides the following solution:

[0007] In a first aspect, the present invention provides a method for identifying biochar, comprising:

[0008] Acquire physicochemical property data of the target individual; the target individual includes target waste biomass and corresponding target biochar; the target biochar is a solid material obtained by carbonizing the target waste biomass; the physicochemical property data includes hydrogen content, organic carbon concentration, nitrogen content, phosphorus content, potassium content, pH value, specific surface area, and pore volume.

[0009] The physicochemical properties of the target individual are input into the biochar identification model to obtain the identity information of the target biochar.

[0010] The process for determining the biochar identification model is as follows:

[0011] Construct high-dimensional multi-class biochar sample data; the high-dimensional multi-class biochar sample data includes a sample input variable data matrix and a sample identity multi-class label column vector;

[0012] Outlier test and standardization were performed on the sample input variable data matrix in the high-dimensional multi-class sample data of biochar to obtain the processed sample input variable data matrix.

[0013] Construct a feature data matrix based on the processed sample input variable data matrix;

[0014] Based on the feature data matrix, the multi-category label column vector of sample identity, and the random subspace nearest neighbor clustering ensemble learning algorithm, a random subspace nearest neighbor clustering ensemble learning classifier is obtained; the random subspace nearest neighbor clustering ensemble learning classifier is a biochar identity discrimination model.

[0015] Secondly, the present invention provides a biochar identification system, comprising:

[0016] The target individual data acquisition module is used to acquire the physicochemical property index data of the target individual; the target individual includes the target waste biomass and the corresponding target biochar; the target biochar is a solid material obtained by carbonizing the target waste biomass; the physicochemical property index data includes hydrogen content, organic carbon concentration, nitrogen content, phosphorus content, potassium content, pH value, specific surface area, and pore volume.

[0017] The target biochar identity information determination module is used to input the physicochemical property index data of the target individual into the biochar identity discrimination model to obtain the identity information of the target biochar.

[0018] The process for determining the biochar identification model is as follows:

[0019] Construct high-dimensional multi-class biochar sample data; the high-dimensional multi-class biochar sample data includes a sample input variable data matrix and a sample identity multi-class label column vector;

[0020] Outlier test and standardization are performed on the sample input variable data matrix in the high-dimensional multi-class sample data of biochar to obtain the processed sample input variable data matrix.

[0021] Construct a feature data matrix based on the processed sample input variable data matrix;

[0022] Based on the feature data matrix, the multi-class label column vector of sample identity, and the random subspace nearest neighbor clustering ensemble learning algorithm, a random subspace nearest neighbor clustering ensemble learning classifier is obtained; the random subspace nearest neighbor clustering ensemble learning classifier is a biochar identity discrimination model.

[0023] Thirdly, the present invention provides an electronic device, including a memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to cause the electronic device to perform the biochar identification method according to the first aspect.

[0024] According to specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0025] This invention first experimentally samples and analyzes the main physicochemical properties of waste biomass and its carbonized biochar from different regions, with different raw materials, and at different temperatures. Then, it uses a random subspace nearest neighbor clustering ensemble learning method to establish a biochar identity discrimination model, which has the advantages of high accuracy, good result stability, strong generalization ability, and good scalability. Attached Figure Description

[0026] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0027] Figure 1 A flowchart of the biochar identification method provided in this embodiment of the invention;

[0028] Figure 2 This is a schematic diagram illustrating the structure and function of a random subspace nearest neighbor clustering ensemble learning classifier provided in an embodiment of the present invention;

[0029] Figure 3 This is a flowchart illustrating the implementation of a random subspace nearest neighbor clustering ensemble learning classifier for biochar identification provided in this embodiment of the invention.

[0030] Figure 4 Importance ranking chart of input variables provided for embodiments of the present invention;

[0031] Figure 5 The image shows the results of the random subspace nearest neighbor clustering ensemble learning classifier used to identify biochar in sample data, as provided in an embodiment of the present invention. Detailed Implementation

[0032] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0033] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0034] Example 1

[0035] like Figure 1 As shown in this embodiment, a biochar identification method includes:

[0036] Step 100: Obtain the physicochemical property index data of the target individual; the target individual includes the target waste biomass and the corresponding target biochar; the target biochar is a solid material obtained by carbonizing the target waste biomass; the physicochemical property index data includes hydrogen content, organic carbon concentration, nitrogen content, phosphorus content, potassium content, pH value, specific surface area and pore volume.

[0037] Step 200: Input the physicochemical property data of the target individual into the biochar identification model to obtain the identity information of the target biochar.

[0038] The process for determining the biochar identification model is as follows:

[0039] Step S1: Construct high-dimensional multi-class sample data for biochar;

[0040] Step S2: Perform outlier detection and standardization on the input variable data matrix in the high-dimensional multi-class sample data of biochar;

[0041] Step S3: Construct a feature data matrix based on the processed sample input variable data matrix.

[0042] Step S4: Based on the feature data matrix, the multi-class label column vector of sample identity, and the random subspace nearest neighbor clustering ensemble learning algorithm, obtain the random subspace nearest neighbor clustering ensemble learning classifier; the random subspace nearest neighbor clustering ensemble learning classifier is a biochar identity discrimination model.

[0043] In this embodiment, step S1: Several important physicochemical property data collected from the experimental waste biomass and its carbonized biomass samples are organized, and a high-dimensional input vector of the sample data and multi-category attribute labels for the waste biomass and its carbonized biomass samples are established accordingly. The specific process is as follows:

[0044] Step S11: Screening the physicochemical properties of individuals; the individuals include waste biomass and biochar obtained after carbonizing waste biomass; in this step, based on the application value of waste biomass raw materials and its carbonized biochar, different types of physicochemical properties such as carbon storage value, fertility value, pH and particle size distribution are selected.

[0045] Step S12: Define specific physicochemical property parameters to obtain multiple physicochemical property indicators; among them, the physicochemical property indicators for carbon storage value include hydrogen (H) content (%), organic carbon concentration (Corg) (%), etc.; the physicochemical property indicators for fertility value include the content (%) of nitrogen (N), phosphorus (P), potassium (K), etc.; the physicochemical property indicator for pH value includes pH value; and the physicochemical property indicator for particle size distribution includes specific surface area BET (m²). 2 / g), pore size Vpore (cm) 3 / g), etc.

[0046] Step S13: Compile a high-dimensional input vector, that is, sequentially encode the multiple physicochemical property indices defined in step S12 as x1, x2, ..., x p (p is the number of physicochemical property indicators), and the input vector of an individual is composed of rows and columns x = [x1, x2, ..., x...]. p ].

[0047] Step S14 involves collecting the input variable data matrix. This involves selecting waste biomass samples from different regions and using different raw materials, along with biochar samples prepared at different carbonization temperatures. The p physicochemical properties defined in step S12 are measured experimentally, and the input vector x for each individual sample is compiled according to step S13. The p physicochemical properties of each sample are then arranged in row-order to obtain the sample input variable data matrix X. The number of sample individuals is called the sample size, denoted as n, meaning the sample input variable data matrix is ​​an n-row, p-column data matrix. Each sample individual includes waste biomass and its corresponding biochar sample.

[0048] Step S15: Encode the multi-category identity tags of waste biomass and its carbonized biochar. The sources of waste biomass include different categories such as straw, perishable waste, pecan shells, and livestock manure, with the number of categories denoted as ω. The identity category information of waste biomass is classified according to its source, while the identity category information of biochar is attributed to the identity category information of its raw materials. Thus, the identity category y of waste biomass and biochar prepared at different carbonization temperatures is assigned. ν (ν=1,2,…,ω) are labeled with numerical serial numbers 1, 2, 3, … respectively. Based on the row numbers of the sample input variable data matrix in step S14, the identity category labels of each individual sample are labeled one by one, forming a multi-category label column vector y for the waste biomass and biochar samples.

[0049] Step S16: Construct high-dimensional multi-class biochar sample data, that is, combine the sample input variable data matrix X and the sample identity multi-class label column vector y to obtain high-dimensional multi-class biochar sample data {X, y}, where the dimensions of the sample input variable data matrix X and the sample identity class label column vector y are n×p and n×1, respectively.

[0050] In this embodiment, step S2, outlier detection and input variable standardization, involves outlier detection of the sample data to ensure the accuracy and reliability of the ensemble learning classifier model. Simultaneously, to reduce the impact of the dimensions and orders of magnitude of various physicochemical property parameters on the discriminant model, standardization of each input variable is performed, specifically including:

[0051] Step S21, outlier detection, involves implementing an outlier identification algorithm based on statistical methods for possible random outliers in the sample input variable data matrix X. Individual samples with outliers are identified as outliers and removed from the sample input variable data matrix X, resulting in a sample input variable data matrix after outlier removal.

[0052] Step S22: Standardize the sample input variable data matrix after removing outliers to obtain the processed sample input variable data matrix. That is, calculate the mean and standard deviation of each physicochemical property index based on the sample input variable data matrix after removing outliers. The detailed process is as follows: Take the input variable data matrix X of the sample data from step S14, remove the outliers from step S21, and denot it as... Then, calculate the mean values ​​of each physicochemical property index data column by column. and standard deviation s l (l=1,2,…,p).

[0053] Input variable standardization: Based on the sample input variable data matrix after outlier removal Implement the column vector x of each physicochemical property parameter according to equation (1). l The standardized processing is performed, and the standardized input variable data matrix is ​​denoted as... These are standardized physicochemical property index data.

[0054]

[0055] In this embodiment, step S3, importance ranking and feature selection of input variables: To reduce model complexity and improve model interpretability and predictive performance, input variables that have a significant impact on category attribute labels are identified and selected, thus completing the importance ranking and feature selection of input variables. The detailed process is as follows:

[0056] Step S31, Weight Initialization and Parameter Setting: Initialize the weights and set the parameters for each input variable x1, x2, ..., x in the processed sample input variable data matrix. p The weights are initialized to 0, i.e., w (0) (x l ) = 0 (l = 1, 2, ..., p), and set the number of random samplings m, the number of nearest neighbor samples k1, and the variable importance threshold θ.

[0057] Step S32: Randomly select a sample individual from the processed sample input variable data matrix, calculate the distance between the selected sample individual and the remaining sample individuals, and use the distance as a similarity metric to search for k1 nearest neighbor sample individuals with the same category attribute as the selected sample individual and form a sample subset of the same category. Search for k1 nearest neighbor sample individuals with different category attributes than the selected sample individual and form a sample subset of different categories. Specifically, this involves taking the processed sample input variable data matrix from step S22. Randomly select one sample individual from them. Calculate its relationship with the remaining sample individuals using equation (2). Distance on the l-th dimension (l=1,2,…,p) of the input variable Using distance as a similarity metric, search for similarities with... The k1 nearest neighbor samples with the same class attribute (within-class distance) are selected and a subset H is formed. i,j A subset M is formed from (j = 1, 2, ..., k1) and ω-1 samples with different class attributes but all having k1 nearest neighbor (inter-class distance) values. i,j (y v )(j=1,2,…,k1, v=1,2,…,ω,y v ≠y i );

[0058]

[0059] Step S33: Based on the subsets of samples with the same category and the subsets of samples with different categories, iteratively calculate the weight of each input variable in the processed sample input variable data matrix, specifically: input variable x l Iterative weight w (h+1) (x l ) is calculated by equation (3), where l = 1, 2, ..., p and h = 0, 1, 2, ..., m.

[0060]

[0061] Where p(y v ) represents the category attribute y in the training sample set. v The prior probability of a subset of samples.

[0062] Step S34, determine the final weights of each input variable: Repeat steps S32 and S33 until the number of random samplings h+1 reaches the m set in step S31, and the final weights w of each input variable will be obtained. (m) (x l (l=1,2,…,p).

[0063] Step S35: Based on the final weights of each input variable, sort the input variables from highest to lowest. Specifically, retrieve the final weights w of each input variable from step S34. (m) (x l (l = 1, 2, ..., p). Sort the input variables from highest to lowest according to their final weights, and denote the sorted input variables as x′. l (l=1,2,…,p).

[0064] Step S36, Feature selection of input variables: Compare the final weights of the sorted input variables with the variable importance threshold θ one by one, and eliminate invalid and redundant input variables. Then, from the remaining s input variables, increase the number of input variables one by one, and calculate the corresponding cumulative contribution rate η of each variable by equation (4). r When η r When the percentage increases to 95% or higher, determine the number of features r selected for the input variables.

[0065]

[0066] Step S37: Construct a feature data matrix based on the number of features selected for the input variables: retrieve the processed sample input variable data matrix from step S23. Based on the sorted input variable x′ in step S35 l (l=1,2,…,p) and the number of features r determined in step S36, construct the feature data matrix of the input variables.

[0067] In this embodiment, step S4, the Random Subspace Nearest Neighbors Clustering Ensemble Learning Classifier, is a technique that combines random subspaces, nearest neighbor (KNN) clustering, and ensemble learning methods and steps to improve the predictive performance of the classifier. Random subspaces not only reduce the dimensionality of the feature space and decrease the complexity of the base classifiers, but also increase the diversity of each base classifier by allowing different random permutations of input variables while maintaining the subspace dimensionality. Meanwhile, KNN clustering can effectively handle the similarity and relationships between individual samples. Finally, the different base classifiers are merged into a more generalized ensemble learning classifier through a voting process. However, parameters such as the noise of the sample dataset, the dimensionality of the subspace, the number of random sampling attempts, and the number of nearest neighbors in the KNN cluster significantly affect the performance of the ensemble learning classifier. Therefore, reasonable preprocessing and parameter tuning are necessary for the Random Subspace Nearest Neighbors Clustering Ensemble Learning Classifier. Figure 2 The functional architecture diagram of the random subspace nearest neighbor clustering ensemble learning classifier is shown. The detailed process is as follows:

[0068] To enhance the adaptability of the ensemble learning classifier to diverse samples, several random subspaces with the same or different dimensions are first randomly selected from the set of input variables after feature selection. Then, the nearest neighbor (KNN) clustering algorithm is used to construct base classifiers in each random subspace. Finally, the voting method is used to merge these base classifiers into an ensemble learning classifier, thereby improving the overall discriminative performance and robustness of the classifier.

[0069] Step S41, Create a random subspace: from the r-dimensional feature data matrix of the input variables Randomly select q features (1≤q<r) and repeat the process u times to obtain u q-dimensional subspace sample data matrices T. c (c = 1, 2, ..., u).

[0070] Step S42, KNN clustering in random subspaces: Take each q-dimensional random subspace sample matrix T from step S41. c (c = 1, 2, ..., u), with a nearest neighbor sample size of k2, the KNN clustering algorithm is implemented to generate u base classifiers.

[0071] Step S43, Ensemble learning of base classifiers: The u base classifiers in step S42 are fused using the relative majority voting method. The ensemble learning classifier constructed is shown in Equation 5. That is, the result of ensemble learning is the class with the most votes. If multiple classes receive the most votes at the same time, one of them is randomly selected.

[0072]

[0073] Step S44, Optimization of parameters for the random subspace nearest neighbor clustering ensemble learning classifier: The dimension q of the random subspace in step S41 and k2 in the KNN clustering algorithm in step S42 have a significant impact on the ensemble learning performance of the u base classifiers in step S43, and the values ​​of the above two parameters are positive integers. Now, using the accuracy of the ensemble learning classifier in judging the sample data as the optimization index, gridding is used to find the optimal q. opt and

[0074] Step S45, Random Subspace Nearest Neighbor Clustering Ensemble Learning Classifier Discriminant Model: Extract the best q found in step S44. opt and Thus, the u base classifiers in learning step S43 are integrated, and a biochar identity discrimination model as shown in equation (6) is established.

[0075]

[0076] In this embodiment, step 200 specifically includes: performing outlier testing and standardization on the physicochemical property index data of the target individual to obtain processed physicochemical property index data of the target individual; constructing a feature data matrix of the target individual based on the processed physicochemical property index data of the target individual; and inputting the feature data matrix of the target individual into the biochar identity discrimination model to obtain the identity information of the target biochar. The detailed process is as follows:

[0077] Step S201, Preprocessing of the target individual: When it is necessary to preprocess a target individual x new When classifying biochar, outlier detection is first performed in step S21, and then the mean values ​​of each physicochemical property parameter are obtained from step S22. and standard deviation s l (l=1,2,…,p), and substitute into equation (1) for standardization. Finally, based on the r feature variables selected in step S36, the dimensionality is reduced to...

[0078] Step S202, the base classifier predicts the biochar identity category of the target individual: taken from step S201. Substituting this into the u base classifiers established in step S42 (the random subspace dimension q of each base classifier and k2 in the KNN clustering method come from q in step S45) opt and This allows us to obtain the predicted biochar identity category of the target individual in each base classifier.

[0079] Step S203, Determining the biochar identity of the sample individual: Statistically analyze the u results of the predicted biochar identity category of the target individual in step S202, use the relative majority voting method, and substitute them into the discrimination model in formula (6) in step S45 to determine the category with the most votes as the biochar identity category of the sample individual.

[0080] The random subspace nearest neighbor clustering ensemble learning classifier for biochar identification provided by this invention uses experimental data on the physicochemical properties of biochar to identify biochar. This classifier can perform the task of biochar identification well, with a prediction accuracy of 100%, and has the advantages of strong specificity, good repeatability, and reliable results.

[0081] Example 2

[0082] like Figure 3 As shown, the random subspace nearest neighbor clustering ensemble learning classifier for biochar identity discrimination provided by the present invention includes the following steps performed in sequence:

[0083] Step S1, Construction of high-dimensional multi-class sample data of biochar: Organize the data such as several important physicochemical properties obtained from the experiments of waste biomass and its carbonized biochar, and establish a high-dimensional input vector of sample data and multi-class attribute labels for the identity of waste biomass and its carbonized products.

[0084] In step S1, the method for constructing high-dimensional multi-class biochar sample data is as follows:

[0085] Step S11: Screening the physicochemical properties of waste biomass and its carbonized biochar: Based on the raw materials of biochar and the application value of carbonized biochar, select four types of physicochemical properties: carbon storage value, fertilizer value, pH, and particle size distribution.

[0086] Step S12, define specific physicochemical property parameters: In this implementation case, the carbon storage value index is selected from hydrogen (H) content (%) and organic carbon concentration (Corg) (%); the fertility value index is selected from the content (%) of nitrogen (N), phosphorus (P), potassium (K), sulfur (S), calcium (Ca), and magnesium (Mg); the acidity / alkalinity index is selected from pH value; and the particle size distribution index is selected from specific surface area (BET) (m²). 2 / g) and pore size Vpore (cm) 3 / g).

[0087] Step S13, Encode high-dimensional input vectors: Sequentially encode the 11 physicochemical property parameters defined in step S12 as x1, x2, ..., x 11 The input vector of each sample individual is composed of rows and columns, namely x = [x1, x2, ..., x]. 11 ].

[0088] Step S14: Experimentally collect the input variable data matrix: Select waste biomass from different regions and using different raw materials, and biochar prepared at three carbonization temperatures: low temperature (350℃), medium temperature (500℃), and high temperature (650℃). Experimentally measure the p=11 physicochemical property parameters defined in step S12, and compile the input vector x for each sample according to step S13. Then, compile the input vectors of each sample in row order to obtain the input variable data matrix X for the waste biomass and biochar samples. In this implementation case, the sample size n=40.

[0089] Step S15: Encode the multi-category identity tags of waste biomass and its carbonized biochar. The sources of waste biomass include six categories: corn stalks, rice stalks, perishable waste, pecan shells, cow dung, and pig dung, i.e., ω = 6. The identity category information of waste biomass is classified according to its source, while the identity category information of biochar is attributed to the identity category information of its raw materials. Thus, the identity category y of waste biomass and biochar produced at three different carbonization temperatures is assigned. ν (ν=1,2,…,6) are labeled with numerical serial numbers 1, 2, 3, 4, 5, and 6 respectively. Based on the row numbers of the input variable data matrix X of the biochar samples obtained in step S14, the identity category labels of each individual sample are labeled one by one, forming a multi-category label column vector y of the biochar sample data.

[0090] Step S16: Construct high-dimensional multi-class biochar sample data: Combine the input variable data matrix X and the identity category label column vector y to obtain waste biomass and biochar sample data {X,y}. In this implementation case, the dimensions of X and y are 40×11 and 40×1, respectively.

[0091] Step S2, Outlier Detection and Input Variable Standardization: To ensure the accuracy and reliability of the data-driven ensemble learning classifier model, outlier detection of the sample data is performed. Simultaneously, to reduce the impact of the dimensions and orders of magnitude of various physicochemical properties on the discriminant model, standardization of each input variable is executed.

[0092] In step S2, the outlier test and input variable standardization methods are as follows:

[0093] Step S21, Outlier Detection: For any random outliers that may exist in the input variable data matrix X, the clustering-based DBSCAN algorithm is used to identify outliers. Individual samples with outliers are identified as outliers and removed from the sample data. In this implementation case, the DBSCAN algorithm did not find any outliers.

[0094] Step S22: Calculate the mean and standard deviation of each physicochemical property parameter based on the input variable data matrix after outlier removal: Obtain the input variable matrix X of the sample data from step S14. Since no outliers were detected in step S21, the outlier-removed data is used to calculate the mean and standard deviation of each physicochemical property parameter. Then calculate the mean value of each physicochemical property parameter for each column. and standard deviation s l (l=1,2,…,11).

[0095] Step S23, Input variable standardization: Based on the input variable data matrix X, implement the standardization of the column vectors x of each physicochemical property parameter according to equation (1). l The standardized processing is performed, and the standardized input variable data matrix is ​​denoted as...

[0096] Step S3, Importance Ranking and Feature Selection of Input Variables: In order to reduce model complexity and improve model interpretability and prediction performance, we identify and select input variables that have a significant impact on category attribute labels, and complete the importance ranking and feature selection of input variables.

[0097] In step S3, the method for ranking the importance of input variables and selecting features is as follows:

[0098] Step S31, Weight initialization and parameter setting: Initialize each input variable x1,x2,...,x 11 The weight is 0, that is, w (0) (x l ) = 0 (l = 1, 2, ..., 11), and set the number of random samplings m = n = 40, the number of nearest neighbor samples k1 = 3, and the variable importance threshold 0.

[0099] Step S32, construct a subset H of samples with the same category. i,j (y i ) and a subset M of samples with different categories i,j (y v ): Retrieve the standardized input variable data matrix from step S22. Randomly select one sample individual from them. The difference between it and the remaining sample individuals is calculated using equation (2). Distance on the l-th dimension (l=1,2,…,11) of the input variable Using distance as a similarity metric, search for similarities with... The k1 nearest neighbor samples with the same class attribute (within-class distance) are selected and a subset H is formed. i,j (j=1,2,3) and ω-1=5 samples with different class attributes but all of which are k1=3 nearest neighbors (inter-class distances) form a subset M. i,j (y v )(j=1,2,3,v=1,2,…,6,y v ≠y i );

[0100] Step S33, calculate the iterative weights of each input variable: input variable x l Iterative weight w (h+1) (x l The prior probabilities p(y) of the six categories are calculated by equation (3), where l = 1, 2, ..., 11 and h = 0, 1, 2, ..., 39. v The values ​​are 0.2, 0.2, 0.1, 0.1, 0.2 and 0.2 respectively.

[0101] Step S34, determine the final weights of each input variable: Repeat steps S32 and S33 until the number of random samplings h+1 reaches m=40 set in step S31, and the final weights w of each input variable will be obtained. (40) (x l (l = 1, 2, ..., 11). In this implementation case, the final weights of each input variable are as follows: Figure 4 As shown.

[0102] Step S35, Ranking the Importance of Input Variables: Obtain the final weight w of each input variable from step S34. (40) (x l (l = 1, 2, ..., 11), the input variables with larger weights are more important. Sort the input variables from highest to lowest weight, and denote the sorted input variables as x′1 = x 10 , x′2=x7, x′3=x4, x′4=x 11 , x′5=x8, x′6=x6, x′7=x3, x′8=x2, x′9=x9, x′ 10 =x1, x′ 11 = x5.

[0103] Step S36, Feature selection of input variables: Compare the final weights of each input variable with the variable importance threshold θ = 0, and eliminate the 5 invalid and redundant input variables: x′ 10 =x1, x′8=x2, x′7=x3, x′ 11=x5, x′9 =x9. Then, from the remaining s = 11 - 5 = 6 input variables, the number of input variables is increased one by one, and the corresponding cumulative contribution rate η is calculated by equation (4). r In this implementation case, the value η is taken as... r =96% determines the number of features of the input variables r=6.

[0104] Step S37, construct the feature data matrix of the input variables: obtain the standardized input variable data matrix from step S23. Based on the sorted input variable x′ in step S35 l (l=1,2,…,11) and the number of features r=6 determined in step S36, construct the feature data matrix of the input variables.

[0105] Step S4, Construction of the random subspace nearest neighbor clustering ensemble learning classifier: To increase the adaptability of the ensemble learning classifier to diverse samples, several input variables with the same or different dimensions are randomly selected from the input variable set after feature selection to form their own random subspaces. Then, the nearest neighbor (KNN) clustering algorithm is used to construct base classifiers in each random subspace. Finally, the voting method is used to merge these base classifiers into an ensemble learning classifier, thereby improving the overall discriminative performance and robustness of the classifier.

[0106] In step S4, the method for constructing the random subspace nearest neighbor clustering ensemble learning classifier is as follows:

[0107] Step S41, random subspace creation: from the r = 6-dimensional feature data matrix of the input variables Randomly select q = 3 features and repeat the process u = 30 times to obtain 30 3D subspace sample data matrices T. c (c = 1, 2, ..., 30).

[0108] Step S42, KNN clustering in random subspaces: Take each q = 3-dimensional random subspace sample matrix T from step S41. c (c = 1, 2, ..., 30), with the number of nearest neighbor samples k2 = 3, the KNN clustering algorithm is implemented to generate 30 base classifiers.

[0109] Step S43, Ensemble learning of base classifiers: The ensemble classifiers (u=30) from step S42 are fused using a relative majority voting method. The output of the ensemble classifier is then constructed. The category with the most votes is shown in equation (5). If multiple categories receive the most votes, one of them is randomly selected.

[0110] Step S44, Optimization of parameters for the random subspace nearest neighbor clustering ensemble learning classifier: The dimension q of the random subspace in step S41 and k2 in the KNN clustering algorithm in step S42 have a significant impact on the ensemble learning performance of the u base classifiers in step S43, and the values ​​of the above two parameters are positive integers. The accuracy of the ensemble learning classifier in judging the sample data is now considered. (n v The optimization metric is the number of correctly classified samples of class v. A gridded approach is used to find the optimal q. opt and In this implementation, the optimization interval for q is set to [1, 5] with a step size of 1, and the optimization interval for k2 is set to [1, 6] with a step size of 1. The different levels of q and k2 form 30 permutations and combinations, and the desired result is found. There are multiple optimal results; we will now select one of them, q. opt =3 and

[0111] Step S45, Random Subspace Nearest Neighbor Clustering Ensemble Learning Classifier Discriminant Model: Extract the best q found in step S44. opt =3 and Thus, the u=30 base classifiers in learning step S43 are integrated, and a biochar identification model as shown in equation (6) is established:

[0112] Step S5: Determine biochar identity based on random subspace nearest neighbor clustering ensemble learning classifier: Use the random subspace nearest neighbor clustering ensemble learning classifier constructed in step S4 to determine the biochar identity of any sample individual.

[0113] In step S5, the method for identifying biochar based on a random subspace nearest neighbor clustering ensemble learning classifier is as follows:

[0114] Step S51, Preprocessing of individual samples: For a new individual sample x new When classifying biochar, outlier detection is first performed in step S21, and then the mean values ​​of each physicochemical property parameter are obtained from step S22. and standard deviation s l (l=1,2,…,11), and substituted into equation (1) for standardization. Finally, based on the 6 feature variables r selected in step S36, the dimensionality was reduced to...

[0115] Step S52, the base classifier predicts the biochar identity category of the sample individual: taken from step S51. Substituting this into the u=30 base classifiers established in step S42 (the random subspace dimension q of each base classifier and k2 in the KNN clustering method come from q in step S45) opt =3 and This allows us to obtain the predicted biochar identity category of the sample individual in each base classifier.

[0116] Step S53, Determining the biochar identity of the sample individual: Statistically analyze the 30 predicted biochar identity categories for the sample individual from step S52, and use the relative majority voting method. Substitute these results into the discrimination model in step S45 (6) to determine the category with the most votes as the biochar identity category of the sample individual.

[0117]

[0118] The results of biochar identification in this implementation case are as follows: Figure 5 As shown, the accuracy rate for each category is... Overall sample data discrimination accuracy

[0119] Example 3

[0120] In order to implement the method corresponding to Embodiment 1 above and achieve the corresponding functions and technical effects, a biochar identification system is provided below.

[0121] This embodiment provides a biochar identification system, including:

[0122] The target individual data acquisition module is used to acquire the physicochemical property index data of the target individual; the target individual includes the target waste biomass and the corresponding target biochar; the target biochar is a solid material obtained by carbonizing the target waste biomass; the physicochemical property index data includes hydrogen content, organic carbon concentration, nitrogen content, phosphorus content, potassium content, pH value, specific surface area and pore volume.

[0123] The target biochar identity information determination module is used to input the physicochemical property index data of the target individual into the biochar identity discrimination model to obtain the identity information of the target biochar.

[0124] The process for determining the biochar identification model is as follows:

[0125] Construct high-dimensional multi-class biochar sample data; the high-dimensional multi-class biochar sample data includes a sample input variable data matrix and a sample identity multi-class label column vector.

[0126] Outlier test and standardization are performed on the sample input variable data matrix in the high-dimensional multi-class sample data of biochar to obtain the processed sample input variable data matrix.

[0127] Based on the processed sample input variable data matrix, construct the feature data matrix.

[0128] Based on the feature data matrix, the multi-class label column vector of sample identity, and the random subspace nearest neighbor clustering ensemble learning algorithm, a random subspace nearest neighbor clustering ensemble learning classifier is obtained; the random subspace nearest neighbor clustering ensemble learning classifier is a biochar identity discrimination model.

[0129] Example 4

[0130] This invention provides an electronic device including a memory and a processor. The memory stores a computer program, and the processor runs the computer program to cause the electronic device to execute a biochar identification method according to Embodiment 1. Optionally, the electronic device may be a server.

[0131] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple; relevant parts can be referred to the method section.

[0132] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A method for identifying biochar, characterized in that, include: Acquire physicochemical property data of the target individual; the target individual includes the target waste biomass and the corresponding target biochar; The target biochar is a solid material obtained by carbonizing the target waste biomass; the physicochemical property data include hydrogen content, organic carbon concentration, nitrogen content, phosphorus content, potassium content, pH value, specific surface area, and pore volume. The physicochemical properties of the target individual are input into the biochar identification model to obtain the identity information of the target biochar. The process for determining the biochar identification model is as follows: Construct high-dimensional multi-class biochar sample data; the high-dimensional multi-class biochar sample data includes a sample input variable data matrix and a sample identity multi-class label column vector; Outlier test and standardization were performed on the sample input variable data matrix in the high-dimensional multi-class sample data of biochar to obtain the processed sample input variable data matrix. Construct a feature data matrix based on the processed sample input variable data matrix; Based on the feature data matrix, the multi-category label column vector of sample identity, and the random subspace nearest neighbor clustering ensemble learning algorithm, a random subspace nearest neighbor clustering ensemble learning classifier is obtained; the random subspace nearest neighbor clustering ensemble learning classifier is a biochar identity discrimination model.

2. The method for identifying biochar according to claim 1, characterized in that, Constructing high-dimensional, multi-class sample data for biochar, specifically including: The physicochemical properties of individuals are screened; the individuals include waste biomass and biochar obtained by carbonizing waste biomass; the physicochemical properties include carbon storage value, fertility value, pH and particle size distribution. Specific physicochemical property parameters were defined, resulting in multiple physicochemical property indicators; these indicators include: hydrogen content, organic carbon concentration, nitrogen content, phosphorus content, potassium content, pH value, specific surface area, and pore volume. The physicochemical properties are numbered sequentially and grouped into individual input vectors in the form of row vectors; Waste biomass samples from different regions and using different raw materials, along with biochar samples prepared at different carbonization temperatures, were selected. The values ​​corresponding to each physicochemical property index were experimentally determined. Following the input vector format for each individual sample, the physicochemical property index data for each sample were arranged in row-order to obtain a sample input variable data matrix. This matrix is ​​an n-row, p-column data matrix; p represents the number of physicochemical property indices, and n represents the number of sample individuals. Each sample individual includes waste biomass samples and their corresponding biochar samples. Based on the row numbers of the sample input variable data matrix, the identity category labels of each individual sample are labeled one by one, forming a multi-category label column vector of sample identities; Based on the sample input variable data matrix and the sample identity multi-class label column vector, a high-dimensional multi-class sample data of biochar is constructed.

3. The method for identifying biochar according to claim 2, characterized in that, Outlier detection and standardization were performed on the sample input variable data matrix in the high-dimensional multi-class sample data of biochar to obtain the processed sample input variable data matrix, which specifically includes: An outlier identification algorithm based on statistical methods identifies and removes outliers from the sample input variable data matrix containing outliers, thus obtaining a sample input variable data matrix after outlier removal. The sample input variable data matrix after removing outliers is standardized to obtain the processed sample input variable data matrix.

4. The method for identifying biochar according to claim 3, characterized in that, Based on the processed sample input variable data matrix, a feature data matrix is ​​constructed, specifically including: Step 1: Initialize the weights of each input variable in the processed sample input variable data matrix to 0, and set the number of random sampling times m, the number of nearest neighbor samples k1, and the variable importance threshold θ; Step 2: Randomly select a sample individual from the processed sample input variable data matrix, calculate the distance between the selected sample individual and the remaining sample individuals, and use the distance as a similarity metric to search for the k1 nearest neighbor sample individuals with the same category attribute as the selected sample individual and form a sample subset with the same category. Search for the k1 nearest neighbor sample individuals with different category attributes as the selected sample individual and form a sample subset with different categories. Step 3: Based on the subsets of samples with the same category and the subsets of samples with different categories, iteratively calculate the weights of each input variable in the processed sample input variable data matrix; Step 4: Repeat steps 2 and 3 until the number of random samplings reaches the set number of random samplings m, and obtain the final weights of each input variable in the processed sample input variable data matrix; Step 5: Based on the final weights of each input variable, sort the input variables from highest to lowest. Step 6: Compare the final weights of the sorted input variables with the variable importance threshold θ one by one, eliminate invalid and redundant input variables, obtain the remaining sorted input variables, and calculate the cumulative contribution rate by increasing the number of input variables one by one from the remaining sorted input variables. When the cumulative contribution rate is greater than the set value, determine the number of features selected for the input variables. Step 7: Construct a feature data matrix based on the number of features selected from the input variables.

5. The method for identifying biochar according to claim 4, characterized in that, Based on the feature data matrix, the multi-class label column vector of sample identity, and the random subspace nearest neighbor clustering ensemble learning algorithm, a random subspace nearest neighbor clustering ensemble learning classifier is obtained. The random subspace nearest neighbor clustering ensemble learning classifier is a biochar identity discrimination model, specifically including: Randomly select q features from the feature data matrix and repeat u times to obtain u q-dimensional subspace sample data matrices; Given the number of nearest neighbor samples k2, perform KNN clustering algorithm on each q-dimensional subspace sample data matrix to generate u base classifiers; An ensemble learning classifier is obtained by fusing u base classifiers using a relative majority voting method. Using the accuracy of the ensemble learning classifier in judging the sample data as the optimization metric, a gridded approach is used to find the optimal q and k2, thereby obtaining a random subspace nearest neighbor clustering ensemble learning classifier; the sample data includes the feature data matrix and the multi-class label column vector of the sample identity.

6. The method for identifying biochar according to claim 1, characterized in that, The physicochemical properties of the target biochar are input into the biochar identification model to obtain the identification information of the target biochar, specifically including: Outlier test and standardization were performed on the physicochemical property index data of the target individual to obtain the processed physicochemical property index data of the target individual. Based on the processed physicochemical property index data of the target individuals, a feature data matrix of the target individuals is constructed; The feature data matrix of the target individual is input into the biochar identity discrimination model to obtain the identity information of the target biochar.

7. A biochar identification system, characterized in that, include: The target individual data acquisition module is used to acquire the physicochemical property index data of the target individual; the target individual includes the target waste biomass and the corresponding target biochar; the target biochar is a solid material obtained by carbonizing the target waste biomass; the physicochemical property index data includes hydrogen content, organic carbon concentration, nitrogen content, phosphorus content, potassium content, pH value, specific surface area, and pore volume. The target biochar identity information determination module is used to input the physicochemical property index data of the target individual into the biochar identity discrimination model to obtain the identity information of the target biochar. The process for determining the biochar identification model is as follows: Construct high-dimensional multi-class biochar sample data; the high-dimensional multi-class biochar sample data includes a sample input variable data matrix and a sample identity multi-class label column vector; Outlier test and standardization are performed on the sample input variable data matrix in the high-dimensional multi-class sample data of biochar to obtain the processed sample input variable data matrix. Construct a feature data matrix based on the processed sample input variable data matrix; Based on the feature data matrix, the multi-class label column vector of sample identity, and the random subspace nearest neighbor clustering ensemble learning algorithm, a random subspace nearest neighbor clustering ensemble learning classifier is obtained; the random subspace nearest neighbor clustering ensemble learning classifier is a biochar identity discrimination model.

8. An electronic device, characterized in that, The device includes a memory and a processor, the memory being used to store a computer program, and the processor running the computer program to cause the electronic device to perform a biochar identification method according to any one of claims 1 to 6.