Broadleaf arbor biomass parameter model construction method based on dumb variable optimization
By introducing a dumb variable optimization method into the biomass parameter model of forest broadleaf trees, the problem of practicality reduction caused by the increase in model complexity in the prior art is solved, and higher biomass estimation accuracy and model applicability are achieved.
Patent Information
- Application Number
- CN202510563387.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-06-20
AI Technical Summary
When constructing a biomass parameter model of forest broadleaf trees, although the increase in independent variables can improve the estimation accuracy, it will increase the complexity of the model and limit its practicality.
Using a method based on dummy variable optimization, principal component factors and sample categories are extracted through principal component analysis and systematic clustering, and the clustered category is introduced as dummy variables into the specified parameter location of the optimal biomass base model to construct a dummy variable optimization model.
By introducing dummy variables, the problem of reducing the accuracy of the additive model is compensated, so that the model can satisfy the intrinsic relationship between broadleaf tree biomass and environmental factors and forest variables while maintaining accuracy, and improve the applicability of the model.
Smart Images

Figure CN120180756A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of forest resource investigation and monitoring, and particularly to a method for constructing a biomass parameter model of broad-leaved arbors optimized based on dummy variables. Background Art
[0002] As the main body of the global carbon cycle, the carbon sequestration potential of forest ecosystems has attracted much attention. Forest biomass data is the prerequisite for understanding forest carbon sequestration. Coupled with the vast territory of our country, it is crucial to construct accurate biomass models under the action of different environmental factors.
[0004] Traditional parameter models mainly include two structures: linear and non-linear. Among them, the non-linear model is most widely used in the estimation of both the total biomass and each component. According to the number of independent variables, it can be further divided into univariate or multivariate model equations. The selection of independent variable factors and the number can be determined according to the characteristics of the total amount and each component. Among them, the univariate model with only diameter at breast height as the variable is most widely used. Zeng Weisheng and Tang Shouzheng studied the model M = a·D b It was found that when the estimated value of parameter b is 2.33, the estimation accuracy of above-ground biomass is better. Although the increase of independent variables usually improves the estimation accuracy of biomass statistically, it also makes the model more complex and weakens its practicability.
[0005] Therefore, the present invention proposes a method for constructing a biomass parameter model of broad-leaved arbors optimized based on dummy variables. Summary of the Invention
[0006] To solve at least one of the above technical problems, the present invention proposes a method for constructing a biomass parameter model of broad-leaved arbors optimized based on dummy variables.
[0007] The present invention is realized through the following technical solutions: A method for constructing a biomass parameter model of broad-leaved arbors optimized based on dummy variables, comprising the following steps:
[0008] Step 1: Collect biomass data of target broad-leaved arbors and associated environmental factors and tree variable data;
[0009] Step 2: Conduct principal component analysis, extract principal component factors, and divide sample categories through hierarchical clustering;
[0010] Step 3: Use the clustered categories as dummy variables and introduce them to the specified parameter positions of the optimal biomass basic model to construct a dummy variable optimization model;
[0011] Step 4: Verify the fitting effect of the dummy variable optimization model through accuracy indicators and select the optimal parameter model.
[0012] Furthermore, the broad-leaved tree is a poplar tree, the environmental factors include latitude, climate variables and terrain variables, and the forest variables include diameter at breast height, tree height and stand density.
[0013] Furthermore, the principal component analysis in Step 2 includes:
[0014] Perform the KMO sampling adequacy measure test and the Bartlett spherical test to confirm the correlation between variables;
[0015] Extract the principal component factors with eigenvalues greater than 1 and a cumulative variance contribution rate exceeding 80%.
[0016] Furthermore, the hierarchical clustering in Step 2 is specifically as follows:
[0017] Based on the principal component factor score matrix, calculate the Euclidean distance between samples;
[0018] Set the scale distance threshold through the dendrogram to determine the number of classification groups.
[0019] Furthermore, the optimal biomass basic model is the binary allometric equation:
[0020] W = a·D b ·H c ;
[0021] where W is the biomass, D is the diameter at breast height, H is the tree height, and a, b, c are model parameters.
[0022] Furthermore, the dummy variable is introduced into the parameter position of the binary allometric equation, and the model expression is:
[0023]
[0024] In the formula, W is the biomass, x i is the dummy variable, a0 and b i are model parameters (i = 1, 2,..., n), D is the diameter at breast height, H is the tree height, and ε is the error term.
[0025] Furthermore, the dummy variable categories are divided into 2 categories or 3 categories according to the hierarchical clustering results, corresponding to different combinations of environmental or forest growth conditions.
[0026] Furthermore, the accuracy indicators include the coefficient of determination, adjusted coefficient of determination, relative standard error, root mean square error, and Akaike information criterion, and the model after dummy variable optimization needs to satisfy that the increase amplitude of the coefficient of determination ≥ 2% and the decrease amplitude of the relative standard error ≥ 1.5%.
[0027] Compared with the prior art, the present invention has the following beneficial effects:
[0028] The dummy variable factor introduced in this application can compensate for the reduction in accuracy caused by the additive model, enabling the additive model with dummy variables to satisfy the internal relationship between the total biomass of poplar and its components while maintaining accuracy. Therefore, this paper selects the binary additive model with dummy variables as the optimal biomass parameter model for poplar. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 It is a scree plot of the principal component analysis of poplar. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0030] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0031] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this invention belongs. The terms used in the description of the present invention in this specification are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The term "or / and" used herein includes any and all combinations of one or more of the related listed items.
[0032] Please refer to Figure 1 , this embodiment provides a method for constructing a biomass parameter model of broad-leaved arbors optimized based on dummy variables, which specifically includes:
[0033] Perform principal component analysis on 21 characteristic variables in the poplar dataset. First, test the correlation between the characteristic factors of poplar. As shown in Table 1, the KMO sampling adequacy measure = 0.737 > 0.5, the approximate chi-square of the Bartlett spherical test is 4399.356, the degree of freedom is 171, and the P value = 0.000 < 0.001. The correlation matrix of the characteristic variables of poplar does not conform to the spherical hypothesis, that is, there is a certain degree of correlation among the 21 factors, which is suitable for principal component analysis.
[0034] Table 1 KMO and Bartlett spherical tests of poplar influencing factors:
[0035]
[0036]
[0037] As Figure 1As the scree plot of the main component factors of poplar, it can be seen that there is a very large eigenvalue (10.603) for the main component F1, while the eigenvalue of F2 drops steeply, and the eigenvalues of the remaining factors decrease at a slower rate until they approach 0. Therefore, based on the principal component analysis, the score matrix and eigenvalues are calculated, and two main component factors with eigenvalues greater than 1 and a cumulative variance contribution rate of 83.148% are constructed as shown in Table 2. Among them, F1 has 14 significant loadings, while F2 has relatively fewer, only 8.
[0038] Table 2 Poplar factor data:
[0039]
[0040]
[0041] Table 3 Score matrix of poplar main component factors:
[0042] Factor F1 F2 Lon 0.016 -0.275 Lat -0.939* 0.151 TAVE_Y 0.916* 0.247 PAVE_Y 0.828* -0.195 TMAX_M 0.790* 0.541* TMIN_M 0.978* 0.179 TAVE_M 0.955* 0.257 PMIN_M 0.855* -0.384* PMAX_M 0.815* -0.497* PAVE_M 0.859* -0.492* EAVE_M 0.905* -0.333* TD -0.901* 0.038 AHM -0.545* 0.776* EAVE_Y 0.884* 0.351* NFFD 0.736* -0.014 Ele -0.051 -0.033 Asp -0.019 0.095 Slo -0.003 -0.120 S_thickness 0.264 -0.080 Age -0.121 -0.008 Den -0.296 0.378* Eigenvalue 10.603 2.160 Variance contribution rate 67.675 15.473
[0043] Note: * indicates that the variable load of this main component is significant, that is, the load coefficient > 0.30.
[0044] Since the significance of the two main component factors for the total poplar biomass and its components is within 0.1, systematic clustering is performed using the two main component factors as clustering factors. Based on the drawn classification dendrogram, a scale distance of 15 is set as the reference line, and the number of clustering groups is determined to be 2 through the intersection point. The final classification result is that 152 poplar data are assigned to the first category (79.58%), and 39 are assigned to the second category (20.42%).
[0045] Optimal basic model of poplar:
[0046] As shown in Table 3, for the comparison of the verification indicators of the three basic biomass models of poplar, generally speaking, the estimation accuracies of the three models for the total poplar and its components are relatively low compared to Chinese fir and Masson pine. Among them, the highest fitting accuracy of the single-variable allometry model does not exceed 0.9. Similarly, the accuracy of the whole plant, aboveground part, and trunk is the highest, and the accuracy of the leaves is the worst. Among them, the accuracy index range of the single-variable allometry model for the total poplar biomass and its components is 0.590 - 0.860 (R 2 ), 2.017 - 64.760 (RSE), 406.894 - 1066.008 (AIC). Compared with the single-variable allometry model, both of the two two-variable allometry models have a certain improvement in the model accuracy. And compared with the W = a·(D 2 ·H) b model, through W = a·D b ·H cThe method of introducing tree height in the form can significantly improve the R2, RSE, and AIC of all biomass models of poplar. The percentage range of improvement for each index is 2.31% - 6.71% (R 2 ) and 1.89% - 20.21% (RSE), 1.28% - 2.90% (AIC). Therefore, in this study, W = a·D b ·H c is selected as the optimal basic model for the total and component biomasses of poplar.
[0047] Table 4 Comparison of the accuracy of the basic models for the biomass of the whole poplar plant and its various organs:
[0048]
[0049] Introduction of dummy variable factors:
[0050] Previous studies have shown that different environmental factors, as well as two forest tree variables, forest age and stand density, are also important factors affecting the accuracy of biomass models. Obvious differences between factors can make the growth environment of forest trees unstable, further affecting the growth of forest trees. From the perspective of the theory of model construction, if all influencing factors are introduced into the model, it can ensure that the constructed model has a wider scope of application, especially for organs such as branches and leaves where the prediction accuracy is usually low. However, too many factors will lead to an increase in the complexity of the model and easy overfitting. In order to simply but fully quantify complex environmental variables and forest tree variables, this application uses qualitative factors after principal component plus systematic clustering as dummy variables to distinguish different environmental variables and forest tree variables (forest age and stand density), and introduces them into the optimal basic model of each tree species to establish a dummy variable model, and explores the improvement effect of the introduction of dummy variables on the model.
[0051] Taking the introduction of n dummy variables of the optimal binary biomass basic model as an example, the formula can be expressed as:
[0052]
[0053] In the formula, W is the biomass, x i is the dummy variable, a0 and b i are model parameters (i = 1, 2,..., n), D is the diameter at breast height, H is the tree height, and ε is the error term.
[0054] Table 4 Types and values of dummy variable factors:
[0055] Value Category x1 = 1, x2 = 0, x3 = 0 Category 1 x1 = 0, x2 = 1, x3 = 0 Category 2 x1 = 0, x2 = 0, x3 = 1 Category 3
[0056] Poplar dummy variable model:
[0057] The parameter estimation values and standard deviations of the optimal basic model for poplar are shown in Table 5. It can be seen that compared with Chinese fir and Masson pine, the parameter significance of the optimal biomass basic model for poplar is relatively poor, especially for parameter a and parameter c. Although the standard deviations of the coefficients between the total amount and different components vary greatly, parameter a is the most stable in all the optimal biomass basic models, with its estimated value ranging from 0.025 to 0.270 and the standard deviation being below 0.10. The standard deviations of parameter b and parameter c are both relatively large, but the standard deviation of parameter c is greater than that of parameter b in both the total amount and component models. Therefore, in the biomass models for the whole plant and each organ of poplar, the position of the parameter c with the largest standard deviation, that is, the most unstable parameter, is selected to set dummy variables.
[0058] Table 5 Parameter estimation values and standard deviations of the optimal biomass basic model for the whole poplar plant and each organ:
[0059]
[0060]
[0061] Two types of qualitative factors for poplar constructed based on the dimensionality reduction introduction method are introduced as dummy variables into the position of parameter c in the optimal basic model, and the accuracy is tested and compared with the optimal basic model without introducing dummy variables. The results are shown in Table 6. It can be seen that the prediction fitting accuracy rankings of the two models for the total biomass and each component are: aboveground > whole plant > trunk > crown > root > branch > leaf. The evaluation indexes of all biomass models have been improved to a certain extent after introducing dummy variables, among which the improvement effect of the leaf biomass model is the best, with R 2 increased by 7.51%, RSE decreased by 14.08%, and AIC decreased by 2.02%. This shows that introducing dummy variable factors has a certain improvement effect on the optimal basic model of poplar, enabling the model to more comprehensively describe the differences in biomass caused by non-timber measurement factors.
[0062] Table 6 Accuracy comparison between the optimal basic model for the whole poplar plant and each organ and the optimal basic model with introduced dummy variables:
[0063]
[0064] Additive optimal biomass parameter model for poplar:
[0065] The comparison of the test indexes of the additive biomass models for poplar with and without introducing dummy variables is shown in Table 7. It can be seen that after introducing dummy variables, the accuracy of the additive models for the total biomass and some components (roots, crowns, branches, and leaves) has been improved, and the fitting accuracy improvement of the leaf biomass model is relatively obvious, with its R 2 and Adj - R 2They were increased by 0.045 and 0.037 respectively, significantly improving the error accuracy of the whole-plant biomass model. The RSE and RMSE were reduced by 1.750 and 4.723 respectively. For the additive model with dummy variables introduced, from the overall accuracy of the total amount and each component, the R 2 and Adj-R 2 ranged from 0.646 to 0.916 and from 0.638 to 0.912 respectively, and the RSE and RMSE ranged from 1.624 to 54.607 and from 1.486 to 45.818 respectively. Among them, the biomass models of the whole plant, aboveground part and trunk had good fitting effects, and the R 2 was greater than 0.90. The R 2 of the leaf biomass model was slightly lower, only 0.646. In general, the fitting effect is better when dummy variables are introduced into the additive model of the total poplar biomass and its components.
[0066] Table 7 Comparison of the accuracies of the basic additive model and the additive model with dummy variables for the whole poplar plant and its organs:
[0067]
[0068] Table 8 Estimated values of the optimal biomass parameter model parameters for the whole poplar plant and its organs:
[0069] Parameter estimate value <![CDATA[a n > <![CDATA[b n > <![CDATA[c n1 > <![CDATA[c n2 > Q 0.568104 1.904882 0.028661 0.002216 ZW 0.193404 -0.028292 -0.050363 -0.021803 LR 3.189445 -0.775315 0.090420 -0.006042 OP 0.498154 -0.308624 -0.126934 0.008991
[0070] Comparing with Table 7, it can be seen that whether dummy variables are introduced or not, synchronous fitting will slightly reduce the model accuracy compared with independent fitting. However, compared with the model independently fitted directly based on the optimal basic model, the accuracy of the model synchronously fitted based on the optimal basic model with dummy variables introduced is still improved. Among them, the R 2 of the leaf biomass model has the largest increase, rising by 0.043, and the RSE of the trunk biomass model has the largest improvement, decreasing by 5.377. It shows that introducing the dummy variable factor can make up for the accuracy reduction caused by the additive model, enabling the additive model with dummy variables introduced to satisfy the internal relationship between the total poplar biomass and its components while maintaining the accuracy. Therefore, in this paper, the binary additive model with dummy variables introduced is selected as the optimal biomass parameter model for poplar, and Table 8 shows the estimated values of the parameters of the optimal biomass parameter model for poplar.
[0071] The calculation equations of the poplar biomass parameter model are shown in Table 9.
[0072] Table 9 Biomass parameter models for the whole poplar plant and its organs:
[0073]
[0074]
[0075] Note: Among them, X1, X2, X3, X4, and X5 are respectively the five component factors F1, F2, F3, F4, and F5 in the principal component analysis, D is the diameter at breast height, H is the tree height, Bio whole plant, Bio aboveground, Bio underground, Bio trunk, Bio crown, Bio branches, and Bio leaves are the whole plant biomass, aboveground biomass, underground biomass, trunk biomass, crown biomass, branch biomass, and leaf biomass.
[0076] The above-described embodiments merely represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the appended claims.
Claims
1. A method for constructing a broad-leaved tree biomass parameter model based on dummy variable optimization, characterized in that: The following steps are involved: Step 1: Collect biomass data of target broad-leaved trees and related environmental factors and forest variables; Step 2: Conduct principal component analysis, extract principal component factors, and divide sample categories through systematic clustering; Step 3: The clustered categories are used as dummy variables and introduced into the specified parameter positions of the optimal biomass basic model to construct a dummy variable optimization model; Step 4: Verify the fitting effect of the dummy variable optimization model through accuracy indicators and select the optimal parameter model.
2. The method for constructing a broad-leaved tree biomass parameter model based on dummy variable optimization according to claim 1, characterized in that: The broad-leaved tree is poplar, the environmental factors include latitude, climate variables and topographic variables, and the forest variables include breast diameter, tree height and forest stand density.
3. The method for constructing a broad-leaved tree biomass parameter model based on dummy variable optimization according to claim 2, characterized in that: The principal component analysis in step 2 includes: KMO sampling suitability test and Barlett's sphericity test were performed to confirm the correlation between variables; The principal component factors with eigenvalue greater than 1 and cumulative variance contribution rate exceeding 80% were extracted.
4. The method for constructing a broad-leaved tree biomass parameter model based on dummy variable optimization according to claim 3, characterized in that: The system clustering in step 2 is specifically as follows: Based on the principal component factor score matrix, the Euclidean distance between samples is calculated; The scale distance threshold is set through the pedigree diagram to determine the number of classification groups.
5. The method for constructing a broad-leaved tree biomass parameter model based on dummy variable optimization according to claim 3, characterized in that: The optimal biomass basic model is a binary allometric growth equation: W=a·D b ·H c ; Among them, W is biomass, D is diameter at breast height, H is tree height, and a, b, and c are model parameters.
6. The method for constructing a broad-leaved tree biomass parameter model based on dummy variable optimization according to claim 5, characterized in that: In step 3, the dummy variables are introduced into the parameter positions of the binary allometric growth equation, and the model expression is: Where W is biomass, x i are dummy variables, a0 and b i are model parameters (i=1, 2,…, n), D is the DBH, H is the tree height, and ε is the error term.
7. The method for constructing a broad-leaved tree biomass parameter model based on dummy variable optimization according to claim 1, characterized in that: The dummy variable categories are divided into 2 or 3 categories according to the system clustering results, corresponding to different combinations of environments or forest growth conditions.
8. The method for constructing a broad-leaved tree biomass parameter model based on dummy variable optimization according to claim 1, characterized in that: The accuracy indicators include the coefficient of determination, the corrected coefficient of determination, the relative standard error, the root mean square error and the Akaike information criterion, and the model after the dummy variable optimization must meet the requirements that the increase in the coefficient of determination is ≥2% and the decrease in the relative standard error is ≥1.5%.
Citation Information
Cited By
Method for estimating biomass of single trees of catalpa bungei with different genotypes based on remote sensing of unmanned aerial vehicle
CN120877157A