Natural grassland biomass remote sensing estimation method and system based on spatial layering heterogeneity

By adopting the combination of spatial hierarchical heterogeneity analysis and machine learning models in grassland biomass inversion technology, the problems of insufficient spatial resolution and neglect of heterogeneity in the existing technology are solved, and the accuracy and scientificity of biomass prediction are significantly improved.

CN120014469APending Publication Date: 2025-05-16INNER MONGOLIA UNIV OF TECH

Patent Information

Application Number
CN202510177577.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing grassland biomass inversion technology has problems such as insufficient spatial resolution, neglecting the spatial heterogeneity of the grassland internally and relying on experience to select inversion scales, resulting in low prediction accuracy and strong subjectivity of the results.

Method used

Using a natural grassland biomass remote sensing estimation method based on spatial hierarchical heterogeneity, the driving indicators of grassland biomass are analyzed through geographic detectors, the optimal remote sensing indicator combination of each sub-region is screened out, and the machine learning model is used for training, and the optimal prediction model is selected for biomass inversion.

Benefits of technology

It significantly improves the prediction accuracy of grassland biomass, provides high-correlation indicators and parameters, and provides more accurate data for grassland ecological monitoring and animal husbandry production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014469A_ABST
    Figure CN120014469A_ABST
Patent Text Reader

Abstract

The invention discloses a natural grassland biomass remote sensing estimation method and system based on spatial layering heterogeneity, and relates to the technical field of biomass remote sensing prediction. A geographic detector method is used for carrying out driving force sensitivity screening on multiple remote sensing indexes according to the spatial layering heterogeneity of biomass; the defect that a traditional index screening method which only depends on experience judgment cannot accurately provide high-correlation indexes of the biomass continuation mechanism in a regional self-adaptive reaction region is overcome, and a regional self-adaptive high-driving-force index data set which accurately reflects the biomass continuation mechanism can be screened out. Meanwhile, according to the method, a biomass machine learning remote sensing prediction model is optimized, the model is constructed by using index combinations with clear physical significance, and optimization is carried out in various machine learning model methods, so that the estimation precision of the biomass is remarkably improved, and a biomass raster data set which is high in precision and conforms to the actual situation can be produced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of biomass remote sensing prediction, and more particularly to a natural grassland biomass remote sensing estimation method and system based on spatial stratified heterogeneity. Background Art

[0002] Grassland aboveground biomass (gAGB) is an important monitoring indicator of grassland ecosystems and a key parameter for measuring the growth status and material cycle of grassland communities. It directly reflects the net primary productivity of grassland ecosystems. Dynamic monitoring of gAGB and analysis of its spatiotemporal changes under the combined effects of natural and human factors are important foundations for analyzing changes in grassland ecosystem structure and function, studying grassland ecosystem carbon storage, and assessing grassland productivity and health. Therefore, building a high-precision gAGB inversion model is of great significance to China's ecological environment restoration and sustainable utilization of grassland resources.

[0003] At present, there are many patent technologies related to the aboveground biomass of grazing grasslands, such as the patent "A method and system for estimating biomass of alpine grasslands based on satellite-driven models (CN113297904B)", which obtains data information, preprocesses, extracts indicators, builds an indicator system based on the XGboost algorithm and correlation analysis, optimizes the selection of satellite-driven models, and then conducts spatiotemporal dynamic analysis of alpine grassland biomass. Another example is "A method and system for measuring aboveground biomass of grassland plant populations (CN108663483A)", which only needs to measure the height and coverage of plant species, establish a regression equation between the total aboveground biomass of the measured sample plot and the estimated total aboveground biomass of the sample plot, and then the aboveground biomass of the plant population to be measured on the sample plot to be measured can be obtained. There is also "A method and system for remote sensing monitoring of aboveground biomass of grassland (CN115372284B)", which recursively counts biological samples at multiple sampling points on the grassland within a preset time period, converts and processes the remote sensing data and divides the elements, thereby calculating the remote sensing prediction data of the aboveground biomass of the grassland.

[0004] However, there are many problems with the existing inversion technology for grazing grassland biomass. The first is the limitation of spatial resolution. Existing technologies mostly rely on high-resolution remote sensing data to estimate grazing grassland biomass, but the spatial resolution of these data is often insufficient to capture subtle changes in the grass layer. Traditional remote sensing images have low resolution and cannot effectively reflect the slight terrain changes, vegetation distribution and local differences in biomass within the grassland. Important environmental features such as differences in grass height, local growth status and areas affected by grazing are easily overlooked during evaluation. Secondly, existing technologies often ignore the spatial heterogeneity within the grassland in estimating grazing grassland biomass. The biomass and ecological characteristics of different regions vary significantly. Traditional methods rely on the overall NDVI value or a single biomass model, ignoring micro-environmental differences, resulting in underestimation of biomass in some areas or overestimation of carrying capacity in other areas. In addition, existing prediction models lack targeted adjustments for different management scenarios, and the reliability of predictions is reduced. Finally, when selecting the inversion scale, the existing technology has the problem of relying on experience or trial and error methods to select the scale or ignoring the scale research. Researchers often choose a fixed scale based on past experience, or screen the best scale by trying multiple spatial scales, or directly select a fixed scale. This makes the inversion results highly subjective and lacks scientific basis. The screening process is blind and inefficient. At the same time, it ignores the spatial heterogeneity characteristics of variables and cannot reveal the specific impact of scale changes on variable characteristics and model performance.

[0005] Therefore, how to provide a natural grassland biomass remote sensing estimation method and system based on spatial stratified heterogeneity, provide highly relevant indicators and parameters for grassland ecological monitoring and animal husbandry production, and at the same time select the best machine learning model to significantly improve the prediction accuracy of regionalized biomass is an urgent problem that technical personnel in this field need to solve. Summary of the invention

[0006] In view of this, the present invention provides a natural grassland biomass remote sensing estimation method and system based on spatial stratified heterogeneity, using geographic detectors to perform spatial stratified heterogeneity analysis on driving indicators of grassland biomass in various climate zones within the study area, obtaining the ranking of driving indicators of biomass quantified by Q value within each sub-region within the study area, and finally obtaining key indicators that adaptively reflect the survival of biomass in each region. This method can provide highly relevant indicators and parameters for grassland ecological monitoring and animal husbandry production. Use remote sensing data and a large amount of ground-measured biomass data to optimize the prediction accuracy of machine learning prediction models, and use the highest accuracy as the best biomass prediction model to significantly improve the prediction accuracy of regionalized biomass.

[0007] In order to achieve the above object, the present invention adopts the following technical scheme: A natural grassland biomass remote sensing estimation method based on spatial stratified heterogeneity, comprising:

[0008] Based on the climate zoning grid data, the study area was divided into several sub-regions according to the Köppen-Geiger climate zoning method;

[0009] Collect and process multi-source remote sensing datasets in the study area to generate coordinate-matched datasets of biomass and its driving indicators;

[0010] The geographic detector model is used to perform regional adaptive screening of biomass driving indicators, and the best remote sensing indicator combination of each sub-region is obtained by combining the measured biomass data to obtain the best remote sensing indicator combination data set;

[0011] The best remote sensing indicator dataset is divided into training set and test set using stratified random sampling;

[0012] Constructing multiple machine learning models, using a random search hyperparameter set method combined with a ten-fold cross validation method to train the machine learning models respectively, and screening and encapsulating the optimal biomass prediction model;

[0013] Biomass inversion is performed based on the biomass optimal prediction model to obtain a biomass inversion data set.

[0014] Preferably, the driving indicators of biomass are screened regionally and adaptively using the geographic detector model, including: using the geographic detector to perform spatial stratified heterogeneity analysis on the driving indicators of biomass in each climate zone in the study area, thereby performing quantitative driving force analysis on the driving indicators of biomass in each sub-region in the study area based on the Q value, and screening out adaptive ecological key indicators in each sub-region. This method can provide highly relevant indicators and parameters for grassland ecological monitoring and animal husbandry production.

[0015] Preferably, spatially stratified heterogeneity analysis of biomass drivers is performed using geographic probes for each climate zone within the study area, including:

[0016] Optimize the discretization of each continuous driving index;

[0017] The factor detector in the geographic detector was used to analyze the biomass driving force of individual driving indicators of biomass, and the Q value of each driving indicator was obtained;

[0018] The interaction detector in the geographic detector was used to conduct sensitivity analysis of biomass driving force under the interaction of two different driving indicators, and the Q value of each two driving indicators under the interaction was obtained;

[0019] Compare the Q value of each two driving indicators under interaction with the Q value of each corresponding driving indicator to obtain a driving indicator combination sorted by Q value;

[0020] Based on the interaction between each two driving indicators, the ecological detector in the geographic detector is used to determine the significance of the difference in spatial distribution of the driving effects of the two driving indicators on biomass, and to obtain a driving indicator combination that is sensitive to spatial changes in biomass.

[0021] Preferably, the adaptive ecological key indicators in each sub-area are screened and obtained, including:

[0022] The m driving indicators with the highest single Q value, all indicators in the k groups of driving indicator combinations with the highest Q value under interaction, and all indicators with insignificant differences in spatial distribution of driving effects on biomass are taken as key ecological indicators for adaptation in each sub-region.

[0023] Preferably, the Q value of each driving indicator is calculated as follows:

[0024]

[0025] In the formula, h represents the number of categories formed after the driving indicator data is discretized. The number of categories h is spatially reflected as the number of partitions; N h and N represent the number of units of biomass in sub-region h and the whole region, respectively; and σ 2 are the variances of biomass in sub-region h and the entire region, respectively.

[0026] If Q = 0, there is no correlation between biomass and driving indicators; Q = 1 means that biomass is completely determined by the spatial distribution of driving indicators. Therefore, the larger the Q value, the more significant the impact of driving indicators on biomass change.

[0027] Preferably, the significance of the difference between the driving effects of the two driving indicators on biomass in spatial distribution is determined by the F statistic:

[0028]

[0029] Among them, N u and N v is the number of samples of the two indicators, M u and M v is the number of sub-regions of the two indicators, and are the sum of the sub-region variances within the two variables, respectively;

[0030] At a given significance level, set H0: The F distribution table was used to test the significance of the differences in the spatial distribution of biomass driving forces.

[0031] Preferably, the multiple machine learning models include: random forest, support vector machine, artificial neural network and linear regression.

[0032] Preferably, performing biomass inversion based on the biomass optimal prediction model comprises:

[0033] The best remote sensing indicator combination dataset is sampled item by item and input into the optimal biomass prediction model to output the biomass inversion dataset.

[0034] Preferably, a natural grassland biomass remote sensing estimation system based on spatial hierarchical heterogeneity comprises: a partitioning module for dividing a study area into a number of sub-areas based on climate partitioning grid data;

[0035] The dataset generation module is used to collect and process multi-source remote sensing datasets in the study area and generate coordinate-matched datasets of biomass and its driving indicators;

[0036] The adaptive screening module is used to use the geographic detector model to perform regional adaptive screening of biomass driving indicators, and combine the measured biomass data to obtain the best remote sensing indicator combination for each sub-region, thereby obtaining the best remote sensing indicator combination data set;

[0037] The stratified random sampling module is used to divide the best remote sensing indicator data set into training set and test set using stratified random sampling;

[0038] A model packaging module is used to construct a variety of machine learning models, and train the machine learning models respectively by using a random search hyperparameter set method combined with a ten-fold cross validation method, and screen and package to obtain the optimal biomass prediction model;

[0039] The biomass inversion module is used to perform biomass inversion based on the biomass optimal prediction model to obtain a biomass inversion data set.

[0040] It can be seen from the above technical scheme that compared with the prior art, the present invention discloses a natural grassland biomass remote sensing estimation method and system based on spatial stratified heterogeneity, including: dividing the study area into several sub-areas according to climate zoning raster data; collecting and processing multi-source remote sensing data sets in the study area to generate coordinate-matched biomass and its driving indicator data sets; using a geographic detector model to perform regional adaptive screening of biomass driving indicators, combining measured biomass data to obtain the best remote sensing indicator combination of each sub-area, and obtaining the best remote sensing indicator combination data set; dividing the best remote sensing indicator data set into a training set and a test set using a stratified random sampling method; constructing a variety of machine learning models, using a random search hyperparameter set method combined with a ten-fold cross-validation method to train the machine learning models respectively, screening and encapsulating the optimal biomass prediction model; performing biomass inversion based on the optimal biomass prediction model to obtain a biomass inversion data set. The present invention uses a geographic detector method to screen the driving force sensitivity of multiple remote sensing indicators based on their spatial stratified heterogeneity of biomass, overcoming the deficiency that the traditional indicator screening method that relies solely on empirical judgment cannot accurately provide high-correlation indicators that are regionally adaptive and accurately reflect the survival mechanism of biomass in the region. The present invention can screen out a high driving force indicator data set that is regionally adaptive and accurately reflects the survival mechanism of biomass. At the same time, the present invention optimizes the biomass machine learning remote sensing prediction model, uses a combination of indicators with clear physical meanings to build a model and simultaneously seeks the best among multiple machine learning model methods, significantly improving the estimation accuracy of biomass, and can produce a high-precision biomass raster data set that conforms to the actual situation. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.

[0042] Figure 1 A schematic flow chart of a natural grassland biomass remote sensing estimation method based on spatial stratified heterogeneity provided by the present invention.

[0043] Figure 2 A schematic diagram of the inversion process of aboveground biomass of natural grassland provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0044] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0045] The embodiment of the present invention discloses a natural grassland biomass remote sensing estimation method based on spatial hierarchical heterogeneity, such as Figure 1 As shown, including:

[0046] Based on the climate zoning grid data, the study area was divided into several sub-regions according to the Köppen-Geiger climate zoning method;

[0047] Collect and process multi-source remote sensing datasets in the study area to generate coordinate-matched datasets of biomass and its driving indicators;

[0048] The geographic detector model is used to perform regional adaptive screening of biomass driving indicators, and the best remote sensing indicator combination of each sub-region is obtained by combining the measured biomass data to obtain the best remote sensing indicator combination data set;

[0049] The best remote sensing indicator dataset is divided into training set and test set using stratified random sampling;

[0050] Constructing multiple machine learning models, using a random search hyperparameter set method combined with a ten-fold cross validation method to train the machine learning models respectively, and screening and encapsulating the optimal biomass prediction model;

[0051] Biomass inversion is performed based on the biomass optimal prediction model to obtain a biomass inversion data set.

[0052] In a specific embodiment of the present invention, the present invention inverts the biomass of grazing grassland:

[0053] (1) Screening of remote sensing indicators for regional adaptive natural grassland aboveground biomass prediction model:

[0054] The spatial stratified heterogeneity analysis of driving indicators of grassland aboveground biomass (gAGB) was conducted using geographic detectors in various climate zones in the study area, so as to conduct a detailed analysis of the driving force of gAGB in each sub-region of the study area based on the Q value, and finally obtain the adaptive key indicators of grassland ecology in each region. This method can provide highly relevant indicators and parameters for grassland ecological monitoring and animal husbandry production.

[0055] (2) Construction of a high-precision prediction model for aboveground biomass of natural grasslands in a large spatial range at a mesoscale:

[0056] In the embodiment of the present invention, it is quite difficult and challenging to complete multiple machine learning prediction models using time-series continuous multi-source remote sensing data and a large amount of ground-measured gAGB data, especially in the context of a large spatial range. A modeling indicator combination that accurately and objectively reflects the survival mechanism of vegetation in each sub-region is screened out based on spatial stratification heterogeneity, and a high-precision gAGB prediction model is constructed using multiple machine learning. The embodiment of the present invention takes the improvement of the existing gAGB prediction model method from the physical mechanism level and the information technology level as the key entry point, and uses a regionalized adaptive indicator data set to construct and optimize four machine learning prediction models, namely random forest (RF), support vector machine (SVM), artificial neural network (ANN) and linear regression (LR), in each sub-region, so as to construct a high-precision natural grassland aboveground biomass prediction model for a large-scale study area at a mesoscale.

[0057] Specifically, the geographic detector model is used to perform regional adaptive screening of biomass driving indicators, including: using geographic detectors to perform spatial stratified heterogeneity analysis of biomass driving indicators in various climate zones within the study area, thereby performing quantitative driving force analysis of biomass driving indicators within each sub-region within the study area based on Q values, and screening out adaptive ecological key indicators within each sub-region. This method can provide highly relevant indicators and parameters for grassland ecological monitoring and animal husbandry production.

[0058] Specifically, the spatially hierarchical heterogeneity of biomass driving indicators was analyzed using geographic probes for each climate zone in the study area, including:

[0059] Optimize the discretization of each continuous driving index;

[0060] The factor detector in the geographic detector was used to analyze the biomass driving force of individual driving indicators of biomass, and the Q value of each driving indicator was obtained;

[0061] The interaction detector in the geographic detector was used to conduct sensitivity analysis of biomass driving force under the interaction of two different driving indicators, and the Q value of each two driving indicators under the interaction was obtained;

[0062] Compare the Q value of each two driving indicators under interaction with the Q value of each corresponding driving indicator to obtain a driving indicator combination sorted by Q value;

[0063] Based on the interaction between each two driving indicators, the ecological detector in the geographic detector is used to determine the significance of the difference in spatial distribution of the driving effects of the two driving indicators on biomass, and to obtain a driving indicator combination that is sensitive to spatial changes in biomass.

[0064] Specifically, Spatial Stratified Heterogeneity (SSH) refers to the different degrees and types of variability or heterogeneity of a specific variable (such as vegetation cover, soil properties, etc.) at different levels or scales in geographic space. SSH mainly describes how the "heterogeneity" in space changes with the change of spatial levels, that is, the heterogeneity from macro to micro scales.

[0065] Spatial heterogeneity: refers to the differences in certain characteristics (such as biodiversity, climate, soil type, etc.) between regions in geographic space. These differences are not evenly distributed and may vary due to environmental factors, land use patterns, ecological processes, etc.

[0066] Spatial stratification refers to dividing spatial areas into different levels or categories according to certain standards or characteristics in order to better understand how spatial heterogeneity manifests itself at different levels.

[0067] Therefore, spatial stratified heterogeneity (SSH) refers to how the spatial distribution characteristics and variability of a variable differ at different spatial levels (such as topography, climate, land use, vegetation type, etc.). Spatial stratified heterogeneity can be classified according to the spatial scale and the level of concern: Spatial heterogeneity at the macro level: At large scales, such as global and regional scales, heterogeneity is usually related to factors such as climate and geographical environment. For example, vegetation differences in different climatic zones. Spatial heterogeneity at the meso level: At medium scales, such as counties, ecosystem units, etc., this level of heterogeneity is usually related to factors such as land use, vegetation cover, and soil type. Spatial heterogeneity at the micro level: At fine scales, such as heterogeneity in small plots of land and local ecosystems, it is mainly related to detailed characteristics such as microclimate, individual plants, and soil properties.

[0068] Specifically, the adaptive ecological key indicators in each sub-area were screened, including:

[0069] The m driving indicators with the highest single Q value, all indicators in the k groups of driving indicator combinations with the highest Q value under interaction, and all indicators with insignificant differences in spatial distribution of driving effects on biomass are taken as key ecological indicators for adaptation in each sub-region.

[0070] Wherein, m and k are both constants. In the embodiment of the present invention, m=k=3.

[0071] Specifically, the calculation method of the Q value of each driving indicator is as follows:

[0072]

[0073] In the formula, h represents the number of categories formed after the driving indicator data is discretized. The number of categories h is spatially reflected as the number of partitions; the meaning of "sub-regions" related to the Q value is the "partition" meaning of h. The entire area is divided into h = 1, ...., 5 layers, that is, 5 sub-regions, by the explanatory variable LAI. N h and N represent the number of units of biomass in sub-region h and the whole region, respectively; and σ 2 are the variances of biomass in sub-region h and the entire region, respectively.

[0074] If Q = 0, there is no correlation between biomass and driving indicators; Q = 1 means that biomass is completely determined by the spatial distribution of driving indicators. Therefore, the larger the Q value, the more significant the impact of driving indicators on biomass change.

[0075] Specifically, the significance of the difference between the two driving indicators in driving biomass in spatial distribution is determined by the F statistic:

[0076]

[0077] Among them, N u and N v is the number of samples of the two indicators, M u and M v is the number of sub-regions of the two indicators, and are the sum of the sub-region variances within the two variables, respectively;

[0078] At a given significance level, set H0: The F distribution table was used to test the significance of the differences in the spatial distribution of biomass driving forces.

[0079] Specifically, the various machine learning models include: random forest, support vector machine, artificial neural network and linear regression.

[0080] Specifically, performing biomass inversion based on the biomass optimal prediction model includes:

[0081] The best remote sensing indicator combination dataset is sampled item by item and input into the optimal biomass prediction model to output the biomass inversion dataset.

[0082] In a specific embodiment of the present invention, a natural grassland biomass remote sensing estimation method based on spatial stratified heterogeneity is provided. Figure 2 As shown, the natural grassland aboveground biomass inversion flow chart of the embodiment of the present invention is generally divided into six parts:

[0083] (1) Collect and process multi-source remote sensing datasets;

[0084] (2) regional adaptive screening of driving indicators of gAGB using the geographic detector model;

[0085] (3) Combine the measured gAGB data to obtain the best remote sensing index combination data set;

[0086] (4) The best remote sensing indicator dataset is divided into training set and test set using stratified random sampling;

[0087] (5) Four machine learning models, namely RF, SVM, ANN, and LR linear regression, were constructed, and the model training was performed using the random search hyperparameter set method combined with the ten-fold cross-validation method. The gAGB prediction model with the highest accuracy was screened and packaged based on the accuracy.

[0088] (6) The remote sensing dataset with the best indicator combination is sampled item by item and input into the gAGB optimal prediction model to produce the gAGB inversion dataset.

[0089] Specifically include:

[0090] Step 1: Multi-source data acquisition and data set organization

[0091] The embodiment of the present invention uses 6 remote sensing data sources, as shown in Table 1. LAI, LST, DEM, and PD data sets can be downloaded from the Google Earth Engine GEE platform, Köppen-Geiger global climate zoning data can be downloaded from the GloH2O platform, precipitation data can be downloaded from the National Center for Environmental Information in the United States, and global GDP grid data can be downloaded from the Gridded global datasets for Gross Domestic Product and Human Development Index over 1990–2015 document. DEM data is used to further produce slope data sets and aspect data sets, and all 9 remote sensing data sets are unified and integrated into Geotiff format. At the same time, all data are batch preprocessed, all data sets are clipped to the scope of the study area, and corresponding resampling, reclassification, mosaic calculation and other operations are performed, and finally all remote sensing data preprocessing is completed, and data preparation work is completed. Furthermore, the embodiment of the present invention needs to collect a certain amount of gAGB field measurement data during the grassland growing season in the study area, including latitude and longitude coordinates and the dry weight of grass samples in the 1m×1m sample box under the coordinates. The nine types of remote sensing data that have been produced are sampled using longitude and latitude coordinates to form a dataset of coordinate-matched gAGB and its driving indicators.

[0092] Table 1 Multi-source remote sensing data set table used in the present invention

[0093]

[0094] Step 2: Optimal discretization

[0095] In order to accurately analyze the spatial stratified heterogeneity of the impact of each indicator on gAGB, it is necessary to optimally discretize the seven numerically continuous indicators (including LAI, LST, Precipitation, DEM, Slope, PD, and GDP) in the embodiment of the present invention. The embodiment of the present invention uses the parameter optimizer in the optimal parameters-based geographic detectors (OPGD) model to discretize the continuous variables and divide them into different discrete intervals to optimize the Q value of the geographic detector. The discretization methods include five methods: equal interval method, natural interval method, quantile interval method, geometric interval method, and standard deviation interval method; for each discretization method, a different number of intervals (number of layers) can be input, and the layering scheme with the largest Q value is usually found between 2 and 10. The OPGD model can freely combine discretization methods and the number of discretization intervals, and all combinations are input into the OPGD model. Among all the combinations of discretization methods and interval numbers, the combination with the highest Q value is selected as the optimal discretization method and interval number for the variable. By optimally discretizing the data set, the reliability of each indicator's interpretation of gAGB is significantly enhanced.

[0096] Step 3: Sensitivity analysis of driving force of gAGB under the action of a single indicator

[0097] The factor detector in the geographic detector was used to analyze the driving force of gAGB on the driving indicators of grassland AGB. As the core part of the geographic detector, the factor detector was used to analyze the explanatory power of a single driving indicator on the spatial stratification heterogeneity of grassland aboveground biomass, which was measured by the Q statistic.

[0098]

[0099] Where: The entire region is divided into h = 1, ..., 5 layers, i.e. 5 sub-regions, by the explanatory variable LAI. h and N represent the number of gAGB units in the sub-region h and the whole region, respectively; and σ 2 are the variances of gAGB in sub-region h and the entire region, respectively. If Q = 0, it means there is no correlation between gAGB and LAI; Q = 1 means that gAGB is completely determined by the spatial distribution of LAI. Therefore, the larger the Q value, the more significant the impact of LAI on the change of gAGB.

[0100] In an embodiment of the present invention, the factor detector can accurately screen the variable sensitivity of a single indicator. For example, the Q values ​​of LAI and LST output by the factor detector are 0.61 and 0.35, indicating that LAI and LST explain 61% and 35% of the spatial distribution of grassland AGB, respectively, indicating that LAI is more sensitive to changes in the spatial distribution of gAGB and has a stronger driving force on gAGB.

[0101] Step 4: Sensitivity analysis of driving force of gAGB under the interaction of indicators

[0102] The embodiment of the present invention uses the interaction detector in the geographic detector to identify the explanatory power of two different indicators X on Y under the interaction. By comparing the Q value obtained under the interaction with the Q value under two single variables, the forms of interaction between indicators can be divided into five categories, including nonlinear weakening of sensitivity, weakening of single variable sensitivity, enhancement of bivariate sensitivity, mutual independence of sensitivity, and nonlinear enhancement of sensitivity. The interaction relationship between the two indicators is described by example, as shown in Table 2.

[0103] Table 2 Interaction forms between indicators

[0104]

[0105] Step 5: Analysis of the driving force differences between indicators on the spatial distribution of gAGB

[0106] The ecological detector in the geographic detector is used to determine whether there is a significant difference in the spatial distribution of the driving effects of the two indicators X1 and X2 on the AGB data Y. The measurement standard is the F statistic.

[0107]

[0108] Among them, N u and N v is the number of samples of the two indicators, M u and M v is the number of sub-regions of the two indicators, and are the sum of the sub-region variances within the two variables, respectively.

[0109] Therefore, at a given significance level, the null hypothesis H0: Through the F distribution table to test, H0 means that there is no significant difference between the two indicators under the spatial distribution of gAGB inversion data.

[0110] If H0 is rejected at the significance level α, it indicates that there is a significant difference between the two indicators under the spatial distribution of gAGB inversion data. The significance level α represents a pre-set probability threshold, which allows the maximum probability of making a type I error (i.e., incorrectly rejecting a null hypothesis that is actually correct) when performing hypothesis testing.

[0111] When the condition for rejecting the null hypothesis H0 at the significance level α is not met, it indicates that the difference in the spatial distribution of the driving effects of the two indicators on gAGB is not significant, which means that the two indicators have similar sensitivities to the spatial changes of gAGB. At this time, the information contained in the two indicators complements each other in the prediction model of gAGB and is suitable as a combination for estimating gAGB.

[0112] Step 6: Remote sensing index screening for regional adaptive natural grassland aboveground biomass prediction model

[0113] Combining steps 2 to 5, the embodiment of the present invention performs optimal discretization processing on each continuous indicator, provides the best discretization method and discretization interval number combination for each indicator, and obtains the maximum Q value for each indicator; performs sensitivity analysis on the driving force of gAGB under the interaction between the indicators, and obtains the indicator combination sorted by Q value; judges the significance of the spatial difference between the driving effects of the two indicators on gAGB, and verifies the indicator combination that is sensitive to the spatial change of gAGB.

[0114] According to the above results, the embodiment of the present invention makes a remote sensing index screening of the prediction model of the aboveground biomass of natural grasslands in a regionalized adaptive manner. The screening method is: the three indexes with the highest Q value of the factor detector, all the indexes in the three groups of index combinations with the highest Q value of each group of interaction detectors, and all the indexes whose driving effects on gAGB are not significantly different in space are all taken as the best indexes reflecting the gAGB survival mechanism in the region, and this index combination is taken as the best remote sensing index combination of the gAGB prediction model in the region.

[0115] By pre-selecting the modeling indicator combination through this method, multicollinearity between variables can be eliminated, model accuracy can be improved, redundant information can be removed, and computational efficiency can be improved. More importantly, the geographic detector provides a clear physical meaning for the method of customizing the screening of variables based on the biomass survival driving mechanism.

[0116] In a specific embodiment of the present invention, a high-precision mesoscale natural grassland aboveground biomass prediction model is constructed and inverted in a large-scale study area. The specific process is as follows:

[0117] Step 1: Construction of a high-precision regional prediction model for aboveground biomass of natural grasslands

[0118] Firstly, the study area is divided into several sub-regions according to the climate zoning raster data according to the Köppen-Geiger climate zoning method, and the eight types of remote sensing data (including LAI, LST, Precipitation, DEM, Aspect, Slope, GDP, and PD) provided in the embodiment of the present invention are combined with the ground measured gAGB data to produce a data set of gAGB and its driving indicators with coordinate matching corresponding to each sub-region.

[0119] Subsequently, all the steps in “Obtaining remote sensing indicator combinations for regionalized adaptive natural grassland aboveground biomass prediction models” were performed for each sub-region to obtain the optimal remote sensing indicator combination for the gAGB prediction model in each sub-region in the study area, and to produce the corresponding gAGB model estimation dataset suitable for input into the machine learning regression prediction model.

[0120] Finally, the gAGB machine learning prediction model for gAGB model parameter optimization is constructed for each sub-region using the best remote sensing indicator dataset of each region. The construction of the gAGB optimal prediction model based on machine learning is divided into the following three parts:

[0121] Part I: Dataset Division. The dataset is divided into 10 layers using the quantile method, and the gAGB model estimation data of each region is randomly divided into training and test datasets at a ratio of 8:2. Stratified spatial random sampling can ensure that the training and test sets have similar spatial distributions to improve the accuracy of model learning and evaluation.

[0122] Part 2: Model training. The training set data is used for model training. The random search hyperparameter set method is combined with a ten-fold cross validation method to train the model. After a large number of model trainings, the coefficient of determination R is used. 2 (The value range is 0-1. The closer the value is to 1, the higher the model accuracy) Evaluate the performance of the model obtained by each set of hyperparameters, and record the R value of each fold of the model obtained by each set of hyperparameters in the 10-fold cross validation process. 2 value:

[0123]

[0124] And calculate 10 R 2 The average value As the prediction accuracy of the model corresponding to this set of hyperparameters:

[0125]

[0126] in accordance with The value is used to score all trained models ( The 0-1 value range of is mapped to 0-100 points). The model corresponding to the hyperparameter set with the highest score is the optimal prediction model for gAGB in this region.

[0127] Part 3: Encapsulate the optimal model for each region. Train the model again on the entire training set data to obtain the final model, combine it with the test set data and use R 2 The root mean square error (RMSE) is used to evaluate the generalization ability and accuracy of the optimal model (the smaller the RMSE value, the more accurate the model). The root mean square error RMSE is calculated as follows:

[0128]

[0129] Specifically, (1) Random Forest (RF) is an ensemble learning method based on decision trees. It improves the generalization ability and accuracy of the model by combining the results of multiple decision trees. When using random forest for gAGB regression prediction analysis, samples are randomly selected to form each decision tree, and the results of all decision trees are averaged to obtain the gAGB prediction result. Therefore, this model can effectively avoid the overfitting problem. The principle formula is as follows:

[0130]

[0131] In the formula, is the final prediction result of gAGB, T i (x) represents the individual prediction result of the i-th decision tree for sample x, and m is the total number of decision trees in the random forest.

[0132] (3) Support Vector Machine

[0133] Support Vector Machine (SVM) is a machine learning model commonly used for classification and regression. SVM regression (SVR) is its application in regression tasks. The goal of SVR is to find an optimal function (regression function) that minimizes the error between the predicted results and the actual observed values, while having good generalization ability. The basic idea of ​​SVR is to map the original data to a high-dimensional space, find a function that is as smooth as possible in the space to fit the data, and control the fitting error. This process is mainly achieved by defining a "margin" and "tolerance" to find an optimal regression hyperplane (or function). Through the kernel technique, SVR can handle nonlinear regression problems and find the best regression function in high-dimensional space.

[0134] (4) Artificial Neural Networks

[0135] Artificial neural network (ANN) consists of input layer, output layer and several hidden layers. The multi-layer feedforward neural network is trained by forward propagation and back propagation algorithm. In the forward propagation stage, the input data P1, P2, ..., P n (CHM, CGF, LAI) after weights ω1, ω2, …, ω n The weighted sum of the gAGB and the bias value b of the neuron is then transformed nonlinearly through the activation function σ to transfer the signal from the input layer to the hidden layer and then to the output layer. The gAGB predicted value of the output layer is compared with the expected value. If the error does not meet the requirements, the weights and biases of each layer are adjusted through the back propagation algorithm until the error reaches the expected value or the set number of training times is reached.

[0136] (5) Linear regression

[0137] Linear regression models include univariate linear and multivariate linear regression models. Let y be the dependent variable, x1, x2, x3…x k Build a linear regression model for grazed grassland biomass as the independent variable:

[0138] y=a1x1+a2x2+a3x3+…+a k x k +a0;

[0139] Wherein, variable y is the grassland AGB predicted by the linear regression model LR, and a1, a2, a3…a k is the regression coefficient of the LRM fitting equation, a0 is the constant term of the fitting equation; when there is only one independent variable whose coefficient is not 0, the equation is a univariate linear regression model.

[0140] Step 2: Inversion of aboveground biomass of natural grassland based on regional optimal model

[0141] According to the best remote sensing indicator combination of each study area, the meta-sampling data sets of each sub-area in the study area are integrated, and the meta-sampling data of each regional indicator combination are input into the optimal prediction model of gAGB in each sub-area obtained in step 1. After the model output, the inversion data sets of each sub-area are obtained. Finally, by mosaicking all the sub-area inversion data sets, the high-precision natural grassland aboveground biomass inversion data set at the mesoscale of the complete study area is completed.

[0142] In the embodiment of the present invention, (1) the regionalized adaptive natural grassland aboveground biomass remote sensing estimation model index screening based on spatial stratified heterogeneity: the embodiment of the present invention uses a geographic detector to perform spatial stratified heterogeneity analysis on the driving index of aboveground biomass of grassland in each climate zone in the study area, obtains the driving index ranking of gAGB quantified by Q-value within each sub-region in the study area, and finally obtains the key index of adaptive response gAGB in each region. This method can provide highly relevant indicators and parameters for grassland ecological monitoring and animal husbandry production.

[0143] (2) Construction of a high-precision mesoscale and large-scale natural grassland aboveground biomass zoning optimization prediction model: The present invention uses nine types of remote sensing data (including: LAI, LST, Precipitation, Climate-zone, DEM, Aspect, Slope, Population-density, GDP) and a large amount of ground-measured gAGB data to optimize the prediction accuracy of four machine learning prediction models (including: RF, SVM, ANN, LR), and optimizes the prediction accuracy of the model with the highest accuracy (R 2 The largest model) was selected as the best prediction model for gAGB. This model construction and model screening method can significantly improve the prediction accuracy of regionalized gAGB.

[0144] In a specific embodiment of the present invention, a natural grassland biomass remote sensing estimation system based on spatial hierarchical heterogeneity includes: a partitioning module for dividing a study area into a plurality of sub-areas based on climate partitioning grid data;

[0145] The dataset generation module is used to collect and process multi-source remote sensing datasets in the study area and generate coordinate-matched datasets of biomass and its driving indicators;

[0146] The adaptive screening module is used to use the geographic detector model to perform regional adaptive screening of biomass driving indicators, and combine the measured biomass data to obtain the best remote sensing indicator combination for each sub-region, thereby obtaining the best remote sensing indicator combination data set;

[0147] The stratified random sampling module is used to divide the best remote sensing indicator data set into training set and test set using stratified random sampling;

[0148] A model packaging module is used to construct a variety of machine learning models, and train the machine learning models respectively by using a random search hyperparameter set method combined with a ten-fold cross validation method, and screen and package to obtain the optimal biomass prediction model;

[0149] The biomass inversion module is used to perform biomass inversion based on the biomass optimal prediction model to obtain a biomass inversion data set.

[0150] The embodiment of the present invention is based on multi-source remote sensing estimation of grassland aboveground biomass with spatial stratified heterogeneity, quantifies the spatial stratified heterogeneity of grassland aboveground biomass driven by multi-source remote sensing data, screens modeling indicators suitable for regionalized grassland aboveground biomass, and establishes a high-precision natural grassland aboveground biomass prediction model. This scheme has important theoretical and practical significance in the estimation of natural grassland aboveground biomass.

[0151] The embodiment of the present invention uses a geographic detector to deeply analyze the spatial stratified heterogeneity of multi-source remote sensing data under different climate zones in a large spatial range of the study area under gAGB, and reveals the gAGB high driving force index under each climate zone based on the Q-value of the obtained single index and the Q-value under the interaction of the index, and combines the obtained high driving force index to produce an index-adaptive gAGB remote sensing inversion dataset under different climate zones. At the same time, this scheme uses four machine learning models, namely random forest (RF), support vector machine (SVM), artificial neural network (ANN) and linear regression (LR), combined with the measured gAGB field measurement dataset and the regionalized adaptive index remote sensing dataset for model screening and accuracy verification, to obtain a high-precision regional optimized gAGB prediction model, and the obtained model can be used to invert the gAGB grid data of each climate zone under a large spatial range, and the high-precision gAGB grid data of the entire study area can be inverted after data splicing and data quality inspection.

[0152] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.

[0153] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A natural grassland biomass remote sensing estimation method based on spatial hierarchical heterogeneity, characterized in that: include: The study area was divided into several sub-areas based on the climate zoning raster data; Collect and process multi-source remote sensing datasets in the study area to generate coordinate-matched datasets of biomass and its driving indicators; The driving indicators of biomass are screened regionally and adaptively, and the best remote sensing indicator combination of each sub-region is obtained by combining the measured biomass data, thus obtaining the best remote sensing indicator combination data set; The best remote sensing indicator dataset is divided into training set and test set using stratified random sampling; Constructing multiple machine learning models, using a random search hyperparameter set method combined with a ten-fold cross validation method to train the machine learning models respectively, and screening and encapsulating the optimal biomass prediction model; Biomass inversion is performed based on the biomass optimal prediction model to obtain a biomass inversion data set.

2. The method for remote sensing estimation of natural grassland biomass based on spatial hierarchical heterogeneity according to claim 1 is characterized in that: Regional adaptive screening of biomass driving indicators includes: using geographic detectors to perform spatial stratified heterogeneity analysis of biomass driving indicators in each sub-region within the study area, and screening out adaptive ecological key indicators in each sub-region.

3. The natural grassland biomass remote sensing estimation method based on spatial hierarchical heterogeneity according to claim 2 is characterized in that: The spatially stratified heterogeneity of biomass driving indicators was analyzed using the Geographic Detector for each sub-region within the study area, including: Perform optimal discretization processing on each driving index with continuous values; The biomass driving force analysis was performed on the individual driving indicators of biomass to obtain the Q value of each driving indicator; The sensitivity analysis of biomass driving force under the interaction of two different driving indicators was carried out to obtain the Q value of each two driving indicators under the interaction; Compare the Q value of each two driving indicators under interaction with the Q value of each corresponding driving indicator to obtain a driving indicator combination sorted by Q value; Based on the interaction between each two driving indicators, the significance of the difference in the spatial distribution of the driving effects of the two driving indicators on biomass was determined, and a combination of driving indicators that was sensitive to spatial changes in biomass was obtained.

4. The method for remote sensing estimation of natural grassland biomass based on spatial hierarchical heterogeneity according to claim 3 is characterized in that: The key ecological indicators of adaptation in each sub-area were screened, including: The m driving indicators with the highest single Q value, all indicators in the k groups of driving indicator combinations with the highest Q value under interaction, and all indicators with insignificant differences in spatial distribution of driving effects on biomass are taken as key ecological indicators for adaptation in each sub-region.

5. The method for remote sensing estimation of natural grassland biomass based on spatial hierarchical heterogeneity according to claim 1 is characterized in that: The Q value of each driving indicator is calculated as follows: In the formula, h represents the number of categories formed after the driving indicator data is discretized. The number of categories h is spatially reflected as the number of partitions; N h and N represent the number of units of biomass in sub-region h and the whole region, respectively; and σ 2 are the variances of biomass in sub-region h and the entire region, respectively.

6. The method for remote sensing estimation of natural grassland biomass based on spatial hierarchical heterogeneity according to claim 5 is characterized in that: The significance of the difference between the two driving indicators in driving biomass in spatial distribution is determined by the F statistic: Among them, N u and N v is the number of samples of two different driving indicators, M u and M v is the number of sub-regions of two different driving indicators, and are the sum of the sub-region variances within the two variables, respectively; At a given significance level, set H0: The F distribution table was used to test the significance of the differences in the spatial distribution of biomass driving forces.

7. The method for remote sensing estimation of natural grassland biomass based on spatial hierarchical heterogeneity according to claim 1 is characterized in that: Multiple machine learning models include: random forest, support vector machine, artificial neural network, and linear regression.

8. The method for remote sensing estimation of natural grassland biomass based on spatial hierarchical heterogeneity according to claim 1 is characterized in that: The biomass inversion is performed based on the biomass optimal prediction model, including: The best remote sensing indicator combination dataset is sampled item by item and input into the optimal biomass prediction model to output the biomass inversion dataset.

9. A natural grassland biomass remote sensing estimation system based on spatial hierarchical heterogeneity, characterized in that: include: Zoning module, which is used to divide the study area into several sub-areas based on climate zoning raster data; The dataset generation module is used to collect and process multi-source remote sensing datasets in the study area and generate coordinate-matched datasets of biomass and its driving indicators; The adaptive screening module is used to perform regional adaptive screening of biomass driving indicators, and obtain the best remote sensing indicator combination for each sub-region by combining the measured biomass data to obtain the best remote sensing indicator combination data set; The stratified random sampling module is used to divide the best remote sensing indicator data set into training set and test set using stratified random sampling; A model packaging module is used to construct a variety of machine learning models, and train the machine learning models respectively by using a random search hyperparameter set method combined with a ten-fold cross validation method, and screen and package to obtain the optimal biomass prediction model; The biomass inversion module is used to perform biomass inversion based on the biomass optimal prediction model to obtain a biomass inversion data set.

Citation Information

Patent Citations

  • Measurement method and measurement system for above-ground biomass of grassland plant population

    CN108663483A

  • A method and system for estimating biomass in alpine grasslands based on a satellite-driven model

    CN113297904B

  • A method for extracting leaf area index and chlorophyll content of crops based on data from a miniature spectrometer.

    CN115372284B

Cited By

  • Shrub biomass evaluation and prediction method based on remote sensing data

    CN120339852A