A method and related device for identifying multi-level management units of soil organic carbon
Through machine learning algorithms and cascade self-organizing mapping models, a multi-level soil organic carbon management unit identification method was constructed, which solved the problem of accurate identification of soil organic carbon management units at the regional scale, achieved efficient and accurate soil organic carbon management, and met dynamic monitoring and management needs.
Patent Information
- Application Number
- CN202510779815.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-06-12
AI Technical Summary
Existing technologies make it difficult to quickly and accurately identify management units with similar problems and solution needs facing soil organic carbon at different regional scales, resulting in insufficient precision and accuracy in soil organic carbon management.
A machine learning algorithm is used based on the reflectivity data, surface temperature data and elevation data of remote sensing monitoring images, combined with the cascade self-organizing map model, to construct a multi-level management unit identification method for soil organic carbon based on hierarchical nesting theory, by calculating environmental variables and constructing the corresponding relationship between management units.
It improves the precision and accuracy of multi-level management unit identification, saves manpower and material resources for field surveys, can cover areas that are difficult to reach with traditional surveys, and ensures the easy updating of results, meeting the needs of regional soil organic carbon dynamic monitoring and management.
Smart Images

Figure CN120298904B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of land parcel division, and in particular to a method for identifying multi-level management units of soil organic carbon and related devices. Background Art
[0002] As the largest carbon reservoir in Earth's surface system, the global soil organic carbon pool is three times that of the atmosphere and 2.5 times that of terrestrial vegetation. Due to its enormous reserves, even small changes in soil organic carbon can have significant impacts on atmospheric carbon dioxide concentrations and climate change. Furthermore, as a key indicator of soil fertility, soil organic carbon profoundly influences food production by affecting soil erosion resistance and water use. Therefore, increasing soil organic carbon is considered a crucial measure for mitigating global climate change and ensuring food security. Identifying multi-level management units is a crucial prerequisite and foundation for improving soil organic carbon management.
[0003] To meet the needs of soil and land resource management, a series of conceptual frameworks and technical systems have been developed both domestically and internationally to identify multi-level soil organic carbon management units. For example, the soil and ecology-based land resource hierarchical framework proposed by the Natural Resources Conservation Service of the United States Department of Agriculture uses soil geography to identify geographic units with similar resource challenges and needs at the regional or landscape scale to support the development of multi-scale projects and management practices. my country has also proposed a variety of zoning approaches, including ecological zoning schemes, eco-geographic zoning, terrestrial ecological basic zoning, and agricultural natural zoning. These approaches identify management zones and units by considering similarities and differences in climate, geomorphology, soils, vegetation, and hydrology at different levels. These frameworks and technical systems provide a solid foundation for the identification of multi-level regional soil organic carbon management units.
[0004] Although the USDA's Land Resource Hierarchy Framework provides different hierarchical structures and unit delineation indicators based on soil geography, it relies on historical soil survey maps for regional-scale management unit delineation. These point-to-surface survey maps are subject to high uncertainty and difficult to update. Furthermore, it does not provide clear delineation variables for regional-scale soil organic carbon management units. my country's zoning schemes, however, primarily focus on large-scale, macroscopic soil and land resource management, with limited consideration of the impacts of regional-scale factors such as human land use and land management on soil organic carbon, which are the primary drivers of regional-scale soil organic carbon dynamics.
[0005] In summary, how to quickly and accurately identify the problems faced by soil organic carbon and management units with similar solution needs at different regional scales to serve regional soil organic carbon management is a key issue that needs to be urgently addressed in this field. Summary of the Invention
[0006] The purpose of this application is to provide a method and related device for identifying multi-level management units of soil organic carbon, which can improve the precision and accuracy of multi-level management unit identification.
[0007] To achieve the above objectives, this application provides the following solutions:
[0008] In a first aspect, the present application provides a method for identifying multi-level management units of soil organic carbon, comprising the following steps:
[0009] Acquire remote sensing monitoring images of the target area; the remote sensing monitoring images include a number of pixels, each pixel represents the reflectance data, surface temperature data and elevation data of several bands at the corresponding position.
[0010] A machine learning algorithm is used to calculate the environmental variables of each pixel based on the reflectivity data, surface temperature data and elevation data of each pixel in the remote sensing monitoring image; the environmental variables include terrain variables, vegetation variables and soil variables.
[0011] A cascade self-organizing map model constructed based on hierarchical nesting theory is used to identify multi-level management units according to the environmental variables of each pixel in the remote sensing monitoring image, and determine the correspondence between each pixel and each management unit at each level; the hierarchical nesting theory shows that the higher the level, the lower the time frequency of changes in the environmental variables of the pixels at that level, and the larger the response time and scale; the cascade self-organizing map model includes several levels, each level includes several management units; each management unit includes several pixels.
[0012] Optionally, terrain variables include altitude, slope, aspect, terrain moisture index, first-class terrain site index, second-class terrain site index, first-class distance from water outlet and second-class distance from water outlet; vegetation variables include cover type, cover function index and vegetation production rate; soil variables include soil organic carbon content, organic carbon density, soil texture type, soil layer thickness and bulk density; the cascade self-organizing map model includes three levels, which are divided into the first level, the second level and the third level from bottom to top; the environmental variables of the first level include vegetation variables, soil organic carbon content, bulk density, altitude, slope, aspect and terrain moisture index; the environmental variables of the second level include vegetation variables, soil variables, altitude, slope, aspect and terrain moisture index; the environmental variables of the third level include terrain moisture index, first-class terrain site index, second-class terrain site index, first-class distance from water outlet and second-class distance from water outlet.
[0013] Optionally, a machine learning algorithm is used to calculate the environmental variables of each pixel based on the reflectance data, surface temperature data, and elevation data of each pixel in the remote sensing monitoring image, specifically including the following steps:
[0014] According to the reflectance value of each pixel in the remote sensing monitoring image, the abundance value of each end member in each pixel and the abundance value image of all end members of the remote sensing monitoring image are obtained based on spectral mixture decomposition; the categories of end members include vegetation, sand, salt and dark matter.
[0015] For any pixel, the abundance values of various end members in the pixel and the abundance values of all end members in the remote sensing monitoring image are input into the pre-trained cover type discrimination model to obtain the corresponding cover type.
[0016] According to the abundance values of various end members in different seasons of the pixel, the corresponding cover function index is calculated.
[0017] According to the change curve of the surface temperature data of the pixel, the corresponding vegetation production rate is determined.
[0018] According to the elevation data of the pixel, the corresponding altitude, slope, aspect, terrain moisture index, first-class terrain site index, second-class terrain site index, first-class distance from the water outlet and second-class distance from the water outlet are determined.
[0019] Based on the abundance values of various end members of the pixel, vegetation variables and terrain variables, the pre-trained soil property prediction model is used to determine the soil organic carbon content, soil texture type, soil layer thickness and bulk density.
[0020] The corresponding organic carbon density is determined based on the soil organic carbon content, soil layer thickness and bulk density of the pixel.
[0021] Optionally, after calculating the environmental variables of each pixel, the method for identifying the soil organic carbon multi-level management unit further includes the following steps:
[0022] The Pearson coefficient was used to evaluate the correlation between the environmental variables in the pixels.
[0023] According to the correlation threshold, redundant environmental variables in each environmental variable of the pixel are screened and eliminated.
[0024] Optionally, several management units at each level form their own self-organizing map model; the number of management units in the self-organizing map model at each level is determined according to the following steps:
[0025] The empirical value of the number of management units corresponding to each level is set based on experience; for the first-level self-organizing map model, the corresponding empirical value of the number of management units is five times the square root of the number of pixels in the remote sensing monitoring image; for the second-level self-organizing map model and the third-level self-organizing map model, the corresponding empirical value of the number of management units is five times the square root of the number of management units in the self-organizing map model of the previous level.
[0026] For the self-organizing map model at any level, several quantity values are selected around the corresponding empirical value of the number of management units to obtain a set of quantity values to be determined.
[0027] For any quantity value in the set of quantity values to be determined, the quantity error and structural error under the quantity value are calculated respectively.
[0028] The number of management units of the hierarchical self-organizing map model is determined according to the quantity error and the structural error under each quantity value.
[0029] Optionally, a soil property prediction model for respectively determining soil organic carbon content, soil texture type, soil layer thickness, and bulk density, and a cover type discrimination model for discriminating the cover type of a pixel are both obtained by training a random forest model using historical data samples of the target area.
[0030] In a second aspect, the present application provides a soil organic carbon multi-level management unit identification system, including the following functional modules:
[0031] The remote sensing monitoring image acquisition module is used to obtain remote sensing monitoring images of the target area; the remote sensing monitoring image includes a number of pixels, each of which represents the reflectivity data, surface temperature data and elevation data of several bands at the corresponding position.
[0032] The pixel environmental variable calculation module is used to use machine learning algorithms to calculate the environmental variables of each pixel based on the reflectance data, surface temperature data and elevation data of each pixel in the remote sensing monitoring image; environmental variables include terrain variables, vegetation variables and soil variables.
[0033] The multi-level management unit identification module is used to identify multi-level management units based on the environmental variables of each pixel in the remote sensing monitoring image using a cascaded self-organizing map model constructed based on the hierarchical nesting theory, and determine the corresponding relationship between each pixel and each management unit at each level; the hierarchical nesting theory shows that the higher the level, the lower the time frequency of changes in the environmental variables of the pixels at that level, and the larger the response time and scale; the cascaded self-organizing map model includes several levels, each level includes several management units; each management unit includes several pixels.
[0034] In a third aspect, the present application provides a computer device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the method for identifying multi-level management units of soil organic carbon as described above.
[0035] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method for identifying multi-level management units of soil organic carbon as described above.
[0036] In a fifth aspect, the present application provides a computer program product, comprising a computer program, which, when executed by a processor, implements the steps of the method for identifying multi-level management units of soil organic carbon as described above.
[0037] According to the specific embodiments provided in this application, this application discloses the following technical effects:
[0038] The present application provides a method and related device for identifying multi-level management units of soil organic carbon. In this method, a machine learning algorithm is used to calculate the terrain variables, vegetation variables and soil variables of each pixel based on the reflectance data, surface temperature data and elevation data of each pixel in the remote sensing monitoring image; then, a cascade self-organizing map model constructed based on hierarchical nesting theory can be used to identify the multi-level management units based on the remote sensing monitoring image, that is, to determine which pixels are included in each management unit at each level; the present application clarifies the multi-level division of management units based on the time frequency, response time and scale of changes in different characteristics of soil organic carbon in the target area under human land use and management, so as to guide soil organic carbon management at different levels such as regional and plot scales, use remote sensing monitoring data and machine learning algorithms to calculate environmental variables, and use a cascade self-organizing map model to automatically identify the correspondence between pixels and management units at each level, which not only saves a lot of manpower and material resources for field surveys, but also can cover areas that are difficult to reach by traditional surveys, thereby improving the precision and accuracy of multi-level management unit identification. In addition, the high temporal and spatial resolution of remote sensing data ensures the easy updating of the results, thus meeting the needs of regional soil organic carbon dynamic monitoring and management. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0040] Figure 1 A flowchart of a method for identifying multi-level soil organic carbon management units provided in one embodiment of the present application.
[0041] Figure 2 This is a flowchart of step A2 in a method for identifying multi-level soil organic carbon management units provided in one embodiment of the present application.
[0042] Figure 3 A two-dimensional scatter plot of principal component analysis in a method for identifying multi-level management units of soil organic carbon provided in one embodiment of the present application.
[0043] Figure 4 A schematic diagram of the TPI and DTP calculation principles in a method for identifying multi-level management units of soil organic carbon provided in one embodiment of the present application.
[0044] Figure 5 This is a flowchart of TPI and DTP calculation in a method for identifying multi-level management units of soil organic carbon provided in one embodiment of the present application.
[0045] Figure 6 A schematic diagram of a technical route for determining soil properties in a method for identifying multi-level management units of soil organic carbon provided in one embodiment of the present application.
[0046] Figure 7 A schematic diagram of a land resource hierarchy framework mentioned in a method for identifying multi-level soil organic carbon management units provided in an embodiment of the present application.
[0047] Figure 8 A schematic diagram of three cascaded self-organizing map models in a method for identifying multi-level management units of soil organic carbon provided in one embodiment of the present application.
[0048] Figure 9 A schematic diagram of the functional modules of a soil organic carbon multi-level management unit identification system provided in one embodiment of the present application.
[0049] Figure 10 A schematic diagram of the structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0050] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0051] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0052] The present application provides a method for identifying a multi-level soil organic carbon management unit. In an exemplary embodiment, Figure 1 As shown, the following steps are included:
[0053] A1. Obtain remote sensing monitoring images of the target area. Remote sensing monitoring images include a number of pixels, each of which represents reflectance data, surface temperature data, and elevation data of several bands at the corresponding location.
[0054] A2. Use machine learning algorithms to calculate the environmental variables of each pixel based on the reflectivity data, surface temperature data, and elevation data of each pixel in the remote sensing monitoring image; environmental variables include terrain variables, vegetation variables, and soil variables.
[0055] Specifically in this embodiment, terrain variables include altitude, slope, aspect, terrain moisture index, first-class terrain site index, second-class terrain site index, first-class distance from water outlet, and second-class distance from water outlet; vegetation variables include cover type, cover function index, and vegetation production rate; soil variables include soil organic carbon content, organic carbon density, soil texture type, soil layer thickness, and bulk density.
[0056] In an exemplary embodiment, Figure 2 As shown, step A2 specifically includes the following steps:
[0057] A21. Based on the reflectance values of each pixel in the remote sensing monitoring image, spectral mixture decomposition is used to obtain the abundance values of each endmember in each pixel and the total endmember abundance value image of the remote sensing monitoring image. Endmember categories include vegetation, sand, salt, and dark matter. Spectral mixture decomposition mainly includes three steps: principal component analysis, endmember extraction based on geometric vertices, and calculation of endmember abundance values for each pixel in the study area.
[0058] First, based on local knowledge of the study area, we determined the months with the highest vegetation cover (e.g., August) and the lowest land cover (e.g., May). We then performed principal component analysis (PCA) on the Landsat surface reflectance data for these months using ENVI software to obtain transformed PCA images. The type and number of endmembers were determined based on existing regional research. In arid regions, monitoring typically involves four endmembers: vegetation (GV), sand (SL), salt (SA), and dark matter (DA). The vegetation (GV) endmember was extracted in August, and the sand (SL), salt (SA), and dark matter (DA) endmembers were extracted in May.
[0059] Then, the end member spectrum curve is extracted by using the two-dimensional scatter plot extraction method. Based on the mathematical mechanism of the linear mixed model, the two-dimensional scatter plot of the principal component analysis should be triangularly distributed, such as Figure 3As shown, this is because the four types of endmembers are in the shape of a tetrahedron. When placed in a two-dimensional scatter plot, they present a triangular distribution. The three vertices of the triangle can be regarded as pure endmembers, and the interior of the triangle is a mixed pixel composed of a linear combination of pure endmembers. In ENVI software, a scatter plot is constructed using the principal component image of the extracted month corresponding to the endmember. 200-400 pixels are selected as endmember pixels at the vertices of the principal component scatter plot, and the average reflectance of the original image at the location of the endmember pixel is used as the endmember spectrum.
[0060] Finally, after obtaining the spectra of various end members in the area, the end member spectra and regional reflectance images were used to perform linear spectral mixing decomposition in ENVI software to obtain the end member abundance values of each end member in the study area.
[0061] Linear spectral mixture decomposition is the main method for abundance estimation, and its expression is:
[0062] .
[0063] in, i is the remote sensing image band, R i is the reflectivity of band i, m is the number of end members, for each end member j (1≤ j ≤ m ) for example, E i,j is the reflectivity of end member j in band i, F j End member j The abundance value, that is, the percentage of the end member in the pixel area, is the residual.
[0064] In practice, in order to make the estimated abundance value have physical meaning, the fully constrained least squares method is generally used to obtain the abundance value of each end member in each pixel and obtain the corresponding abundance value image. j The abundance value satisfies the following formula:
[0065] .
[0066] A22. For any pixel, the abundance values of each end member in the pixel and the abundance values of all end members in the remote sensing monitoring image are input into the pre-trained cover type discrimination model to obtain the corresponding cover type.
[0067] First, according to the land cover classification system (such as forest land, grassland, cultivated land, construction land and water areas), about 1000 pixels were selected for each classification through the high spatial resolution image of Google Earth, and the training sample and validation sample were divided into training samples and validation samples at a ratio of 75% and 25% respectively; then, the training samples and their corresponding end-member abundance values and spectral mixture decomposition were used to obtain the image of all end-member abundance values, which were input into the random forest model for land cover type classification; finally, the accuracy was evaluated based on the land cover type of the validation sample and the corresponding classification results. The evaluation indicators included overall accuracy. p e and Kappa coefficient k , the calculation formula is as follows:
[0068] .
[0069] .
[0070] in, It is the sum of the number of correctly classified samples in each category divided by the total number of samples, that is, the overall classification accuracy. 、 … is the number of true samples of each category, and 、 … is the number of samples in each category in the prediction results, and the total number of samples is n.
[0071] A23. Calculate the corresponding land cover function index based on the abundance values of various end members in different seasons. The land cover function index LCI is calculated according to the following formula:
[0072] .
[0073] in, 、 and are the abundance values of SL, SA and DA end members in the spring vegetation canopy growth phase; It is the end-member abundance value of GV in the summer season when vegetation growth is most vigorous.
[0074] A24. Determine the corresponding vegetation production rate based on the change curve of the pixel's surface temperature data. The vegetation production rate index (VPI) is mainly calculated using MODIS surface temperature data:
[0075] .
[0076] in, is a function of the reference surface temperature curve, is a function of the pixel surface temperature curve, and These are the start and end of the main growing season.
[0077] A25. Based on the pixel elevation data, determine the corresponding altitude, slope, aspect, terrain moisture index, first-class terrain location index, second-class terrain location index, first-class distance from the outlet, and second-class distance from the outlet.
[0078] For example, using elevation data, the slope and aspect can be calculated in the ArcGis hydrology tool module. The slope calculation formula is as follows:
[0079] .
[0080] in, is the slope, is the difference in altitude, is the horizontal distance difference.
[0081] The terrain wetness index TWI can be calculated according to the following formula:
[0082] .
[0083] in, It is the upstream area on the unit contour line through which surface water flows. For the slope.
[0084] like Figure 4 As shown in Figure 1, the terrain position index (TPI) describes the relative terrain position of a pixel within its sub-basin in the vertical direction, effectively improving the incomparability of DEM descriptions of terrain position within different sub-basins. The distance to outlet (DTP) reflects the migration distance of its material and its redistribution by calculating the horizontal distance between the pixel and the sub-basin outlet. The calculation formula is as follows:
[0085] .
[0086] in, is the DEM value of pixel i, is the DEM value of outlet j in the sub-basin where pixel i is located, It is the maximum DEM value in the sub-basin where pixel i is located.
[0087] .
[0088] in, and Pixels i To the outlet of its sub-basin j Distance in the east-west and north-south directions.
[0089] In practical applications, both primary and secondary outlets are considered. Therefore, the TPI and DTP indices are divided into two categories: The first category, within a primary sub-basin, treats each primary outlet as the outlet of each basin, calculating the TPI and DTP for each pixel. The results are labeled TPI1 and DTP1. The second category, within a secondary sub-basin, treats the secondary outlet as the outlet of the basin, calculating the TPI and DTP for each pixel. The results are labeled TPI2 and DTP2. Due to the presence of some real-world terrain (such as karst landforms), the DEM surface contains some concave areas. These areas can lead to illogical or even erroneous flow directions when analyzing water flow directions. Before calculating the direction of water flow, first use the fill tool in the Arcgis hydrological analysis tool to fill the original DEM; then, based on the depression-free DEM generated in the previous step, use the D8 algorithm in the flow direction tool to calculate the water flow direction (i.e., D8 flow direction) of each pixel; next, use the flow tool, input the generated depression-free DEM and flow direction data, and count the flow tool for each pixel; based on the statistical flow, by setting a reasonable empirical value for the minimum surface runoff, use the raster calculator tool to generate river links; secondly, in the raster river network vectorization tool, input the obtained river link raster data and flow direction data to obtain the vector data of the river network, and then use the river network classification tool to classify the obtained river network vector data and determine the outlet; finally, use the watershed analysis tool to extract the watershed based on the flow direction raster data and the river and river network classification data. This process is mainly completed on the basis of ArcGis hydrological tools combined with python programming. The process is as follows Figure 5 shown.
[0090] A26. Based on the abundance values of various end members of the pixel, vegetation variables, and terrain variables, a pre-trained soil property prediction model is used to determine the soil organic carbon content, soil texture type, soil layer thickness, and bulk density.
[0091] In this embodiment, the remote sensing characterization methods for soil organic carbon content, soil texture type, soil layer thickness and bulk density all use digital mapping technology, that is, based on the relationship between soil variables and their characteristic variables (abundance values of various end members, vegetation variables and terrain variables), a random forest model is used to map spatial soil properties.
[0092] like Figure 6 The technical route for determining soil properties is shown in the figure. First, the regional soil sample data is randomly divided into a training dataset and a validation dataset at a ratio of 75% to 25%. Within the training dataset, a grid search method is used to determine the optimal combination of the number of trees, maximum depth, and maximum number of features for the random forest model based on the results of a 5-fold cross-validation. The specific steps are:
[0093] (1) Based on experience, the number of trees in the random forest model is preset to be 100 to 1000, with an interval of 50; the maximum depth is 4 to 10, with an interval of 1; and the maximum number of features is 2 to 5, with an interval of 1.
[0094] (2) Grid retrieval method, which is to combine all possible values of the three parameters within the preset range.
[0095] (3) Then perform 5-fold cross validation on each combination of parameter values:
[0096] The training sample is divided into 5 parts, each part is used as a validation sample, and the remaining four parts are used as training samples. The average R 2 and RMSE, and select the optimal combination of parameters with the highest value to determine the optimal model parameters. Calculate R according to the following formula 2 and RMSE:
[0097] .
[0098] .
[0099] in, and are the predicted value and observed value (true value) of sample point i, is the average value of all sample point observations (true values), r is the number of sample points.
[0100] Then, random forest models were constructed according to their respective optimal model parameter combinations and their accuracy was evaluated.
[0101] The random forest model of the “rpart” package of R language was used to input the observed values of soil properties of each sample and the values of their corresponding characteristic variables. The parameters were set according to the selected optimal model parameters, and the random forest model corresponding to each soil variable in the study area was obtained.
[0102] Next, the land cover type data was used to mask the construction land and water areas to eliminate the interference with the mapping results. The dataset (image) of the feature variables was input into the trained random forest model. In order to ensure the accuracy of the spatial mapping results, the average of the results of running the random forest model 100 times was used as the final prediction result. 2 The accuracy of the mapping results, including the mean soil organic carbon content map, the mean soil layer thickness map, the texture type map, and the mean bulk density map, was evaluated using the RMSE and RSE. The standard deviation of the 100 predictions was also used to assess model uncertainty.
[0103] A27. Determine the corresponding soil organic carbon density based on the soil organic carbon content, soil layer thickness, and bulk density of the pixel. The calculation formula for soil organic carbon density is as follows:
[0104] .
[0105] in, SOCD refers to the soil organic carbon density, SOCC refers to the soil organic carbon content, D represents the thickness of the soil layer, BD Represents bulk density.
[0106] In order to avoid collinearity and redundancy of information between input variables, after calculating the environmental variables of each pixel in step A2, the method for identifying the multi-level management unit of soil organic carbon further includes the following steps:
[0107] B1. Use the Pearson coefficient to evaluate the correlation between the environmental variables of the pixels. In this embodiment, the Pearson coefficient between any two environmental variables is calculated according to the following formula:
[0108] .
[0109] in, and The variables are x and y The values of each sample point, and They are variables x and y The average value of .
[0110] B2. Based on the correlation threshold, filter out the redundant environmental variables in each environmental variable of the pixel. Use the Pearson correlation coefficient re The absolute value of 0.7 was used as the correlation threshold for collinearity. If the correlation coefficient of two variables exceeded the threshold, only one of the variables was retained as input for further analysis.
[0111] A3. Using a cascaded self-organizing map model based on hierarchical nesting theory, we identify multi-level management units based on the environmental variables of each pixel in remote sensing monitoring imagery and determine the corresponding relationship between each pixel and each management unit at each level. Hierarchical nesting theory suggests that higher levels of hierarchy reduce the frequency of changes in the environmental variables of pixels at that level, and increase the response time and scale. The cascaded self-organizing map model includes several levels, each of which contains several management units; each management unit consists of several pixels.
[0112] In this embodiment, the cascade self-organizing map model includes three levels, which are divided into the first level, the second level and the third level from bottom to top; the environmental variables of the first level include vegetation variables, soil organic carbon content, bulk density, altitude, slope, aspect and terrain moisture index; the environmental variables of the second level include vegetation variables, soil variables, altitude, slope, aspect and terrain moisture index; the environmental variables of the third level include terrain moisture index, type I terrain site index, type II terrain site index, type I distance from the water outlet and type II distance from the water outlet.
[0113] like Figure 7 As shown in Figure 1, an existing hierarchical framework for land resources at the regional scale divides the land resources into three layers: the State layer, the Site layer (LS layer), and the Site Group (LSG layer). This hierarchical nesting theory suggests that the higher the level, the lower the temporal frequency of changes in the state variable (the variable that determines state resilience) and the larger the spatial scale of variation. Therefore, a multi-level hierarchical framework for soil organic carbon in the target region was constructed based on this framework and theory. The soil organic carbon content, which has the highest frequency of change (years to decades) and the smallest spatial scale of variation (plots) under human land use and management, was identified as the state variable in the State layer. In addition to soil organic carbon content, soil organic carbon density also considers soil layer thickness, bulk density, and soil texture. These properties vary over decades to centuries under human land use and management, and vary at a much larger spatial scale than plots. Therefore, this was identified as the state variable in the LS layer. The topographic location reflects the differences in parent material and lithology, which determine the upper limit of soil organic carbon. The frequency of change is hundreds to thousands of years, and the spatial scale of the difference is the largest regionally. Therefore, it is used as the state variable of the LSG layer.
[0114] In the constructed multi-level structure framework for soil organic carbon, the three main types of environmental variables (variables influencing soil organic carbon status) play a dominant role in the hierarchical division: soil variables, vegetation variables, and topographic variables. Among the specific variables, five variables related to soil organic carbon content were identified as soil variables: soil organic carbon content, organic carbon density, soil texture type, soil layer thickness, and bulk density. Based on the minimum indicator set established by relevant international sustainable development goals, cover type and vegetation production rate were identified as vegetation variables. Furthermore, to characterize vegetation's response to the natural environment and human management, the cover function index was also identified as a vegetation variable. Among the topographic variables, altitude, slope, and aspect were identified as regional small-scale topographic variables, the topographic wetness index was identified as a regional small-scale topographic variable, and the topographic location index and distance from the water outlet were identified as regional large-scale topographic variables.
[0115] According to the above multi-layer nested structure, at the State layer, all vegetation variables, soil organic carbon content, bulk density and regional small-scale terrain moisture index are determined as its dominant environmental variables; at the LS layer, all vegetation variables, all soil variables and regional small-scale ground moisture index are determined as its dominant environmental variables; at the LSG layer, regional small-scale terrain moisture index, regional large-scale terrain position index and distance from the outlet are determined as the dominant environmental variables.
[0116] Table 1 Dominant environmental variables at each level
[0117]
[0118] In an exemplary embodiment, three cascaded self-organizing map models are used to construct Figure 8 The cascade self-organizing map model shown in Figure 1 is as follows. First, the original value of the State layer environmental variable corresponding to each pixel of the remote sensing monitoring image is converted to x 1. x 2. x 3 and x 4 are combined into an input vector and input into the self-organizing map model 1 to obtain the recognition result of the State layer, which includes the State layer management units and the corresponding relationship between them and the input samples, and the weight vectors of all variables corresponding to each State layer management unit. It should be noted that since the self-organizing map model of each level is in the form of a two-dimensional grid, the management units can also be represented as neurons; then, the weight vectors of the LS layer environmental variables corresponding to each State layer management unit are input into the self-organizing map model 2 to obtain the recognition result of the LS layer, which includes the corresponding relationship between each LS layer management unit and each input State layer management unit, and the weight vectors of all variables corresponding to each LS layer management unit; finally, the weight vectors of the LSG layer environmental variables corresponding to each LS layer management unit are input into the self-organizing map model 3 to obtain the state recognition result of the LSG layer, which includes the corresponding relationship between each LSG layer management unit and each input LS layer management unit, and the weight vectors of all variables corresponding to each LSG layer management unit.
[0119] In an exemplary embodiment, a plurality of management units at each level constitutes a respective self-organizing map model; specifically, the number of management units of the self-organizing map model at each level is determined according to the following steps:
[0120] C1. Set the empirical value of the number of management units corresponding to each level based on experience. For the first-level self-organizing map model, the corresponding empirical value of the number of management units is five times the square root of the number of pixels in the remote sensing monitoring image; for the second-level and third-level self-organizing map models, the corresponding empirical value of the number of management units is five times the square root of the number of management units in the self-organizing map model of the previous level. That is, set the empirical value of the number of management units according to the following formula:
[0121] .
[0122] in, M Experience value for the number of management units, S is the number of pixels in the remote sensing monitoring image or the number of management units in the self-organizing map model of the previous level. For example, assuming the number of pixels in the remote sensing monitoring image is 253, the initial empirical value of the number of management units is 80.
[0123] C2. For any level of the self-organizing map model, select a number of values around the empirical value of the number of management units to obtain a set of values to be determined. Since the self-organizing map model is a two-dimensional network, in this embodiment, based on experience, three values are set around 80. Therefore, the values to be determined are preset to 10×7, 9×8, 11×7, 10×8, 11×8, 10×9, and 12×8.
[0124] C3. For any quantity value in the set of quantity values to be determined, calculate the quantity error and structural error under each quantity value. Calculate the quantity error under each quantity value according to the following formula QE and structural errors TE :
[0125] .
[0126] .
[0127] in, Num is the total number of samples, For the i input samples, It is a sample The corresponding weight vector of the best matching management unit, is the Euclidean distance between the two, and The samples are The first and second neighboring management units, when and When not adjacent, Returns 1, otherwise returns 0. Adjacent here means that in the two-dimensional network of the self-organizing map, the absolute value of the difference between the rows or columns is 1.
[0128] C4. Determine the number of management units of the self-organizing map model at this level based on the quantity error and structural error under each quantity value.
[0129] Specifically in this embodiment, the cover type discrimination model used to discriminate the cover type of the pixel in the above step A22 and the soil property prediction model used to respectively determine the soil organic carbon content, soil texture type, soil layer thickness, and bulk density in step A26 are both obtained by training the random forest model using historical data samples of the target area.
[0130] The above-mentioned method provided in this application, on the one hand, constructs a conceptual framework of the nested hierarchical structure of land systems based on soil organic carbon. Based on the different characteristic rates (frequency of occurrence, response time, and scale) of soil organic carbon at different levels in the region under human land use and management, the state variables and dominant environmental variables at each level are clarified to guide soil organic carbon management at different levels, such as regional and plot scales. On the other hand, the spatial mapping of environmental variables using multi-source remote sensing data and machine learning algorithms not only saves a lot of manpower and material resources for field surveys, but also covers areas that are difficult to reach with traditional surveys, thereby improving the precision and accuracy of the mapping results. In addition, the high temporal and spatial resolution of remote sensing data ensures the easy update of the results, thus meeting the needs of dynamic monitoring and management of regional soil organic carbon.
[0131] In summary, this application combines professional concepts such as soil geography, machine learning algorithms such as random forests, and remote sensing big data to quickly and accurately identify multi-level management units with similar challenges and solutions for soil organic carbon in a region, thereby serving the management of regional soil organic carbon. Furthermore, the management units identified by the methods provided in this application can serve as basic units for monitoring and evaluating regional arable land quality, evaluating and managing ecological quality, and land degradation and restoration.
[0132] Based on the same inventive concept, the present application also provides a system for implementing the above-mentioned method for identifying a multi-level management unit of soil organic carbon. The solution provided by the system is similar to the solution described in the above-mentioned method. In an exemplary embodiment, Figure 9 As shown, the identification system for multi-level soil organic carbon management units includes the following functional modules:
[0133] The remote sensing monitoring image acquisition module is used to obtain remote sensing monitoring images of the target area; the remote sensing monitoring image includes a number of pixels, each of which represents the reflectivity data, surface temperature data and elevation data of several bands at the corresponding position.
[0134] The pixel environmental variable calculation module is used to use machine learning algorithms to calculate the environmental variables of each pixel based on the reflectance data, surface temperature data and elevation data of each pixel in the remote sensing monitoring image; environmental variables include terrain variables, vegetation variables and soil variables.
[0135] The multi-level management unit identification module is used to identify multi-level management units based on the environmental variables of each pixel in the remote sensing monitoring image using a cascaded self-organizing map model constructed based on the hierarchical nesting theory, and determine the corresponding relationship between each pixel and each management unit at each level; the hierarchical nesting theory shows that the higher the level, the lower the time frequency of changes in the environmental variables of the pixels at that level, and the larger the response time and scale; the cascaded self-organizing map model includes several levels, each level includes several management units; each management unit includes several pixels.
[0136] certainly, Figure 9 The architecture shown is only exemplary and can be omitted according to actual needs when implementing different functions. Figure 9 One or at least two components of the system shown.
[0137] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Figure 10 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a method for identifying a multi-level management unit of soil organic carbon provided in the previous embodiment can be implemented.
[0138] Those skilled in the art will understand that Figure 10 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0139] In an exemplary embodiment, a computer device is further provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.
[0140] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0141] In an exemplary embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0142] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0143] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0144] The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may include, but are not limited to, general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic units, data processing logic units based on quantum computing, and the like.
[0145] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0146] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.
Claims
1. A method for identifying multi-level management units of soil organic carbon, characterized in that: include: Acquire a remote sensing monitoring image of the target area; the remote sensing monitoring image includes a plurality of pixels, each pixel representing reflectance data, surface temperature data, and elevation data of a plurality of bands at a corresponding location; A machine learning algorithm is used to calculate the environmental variables of each pixel based on the reflectivity data, surface temperature data, and elevation data of each pixel in the remote sensing monitoring image; the environmental variables include terrain variables, vegetation variables, and soil variables; the terrain variables include altitude, slope, aspect, terrain moisture index, first-class terrain site index, second-class terrain site index, first-class distance to water outlet, and second-class distance to water outlet; the vegetation variables include cover type, cover function index, and vegetation production rate; the soil variables include soil organic carbon content, organic carbon density, soil texture type, soil layer thickness, and bulk density; A cascade self-organizing map model constructed based on hierarchical nesting theory is used to identify multi-level management units according to the environmental variables of each pixel in the remote sensing monitoring image, and determine the corresponding relationship between each pixel and each management unit at each level; the hierarchical nesting theory shows that the higher the level, the lower the time frequency of changes in the environmental variables of the pixels at that level, and the larger the response time and scale; each of the levels includes a number of management units; each of the management units includes a number of pixels; the cascade self-organizing map model includes three levels, which are divided into the first level, the second level and the third level from bottom to top; the environmental variables of the first level include the vegetation variables, soil organic carbon content, bulk density, altitude, slope, aspect and terrain moisture index; the environmental variables of the second level include the vegetation variables, soil variables, altitude, slope, aspect and terrain moisture index; the environmental variables of the third level include terrain moisture index, first-class terrain site index, second-class terrain site index, first-class distance to water outlet and second-class distance to water outlet.
2. The method for identifying multi-level soil organic carbon management units according to claim 1, characterized in that: A machine learning algorithm is used to calculate the environmental variables of each pixel based on the reflectivity data, surface temperature data, and elevation data of each pixel in the remote sensing monitoring image, specifically including: According to the reflectance value of each pixel in the remote sensing monitoring image, the abundance value of each type of end member in each pixel and the image of all end member abundance values of the remote sensing monitoring image are obtained based on spectral mixing decomposition; the categories of the end members include vegetation, sand, salt and dark matter; For any pixel, the abundance values of each end member in the pixel and the abundance value images of all end members in the remote sensing monitoring image are input into the pre-trained cover type discrimination model to obtain the corresponding cover type; According to the abundance values of various end members of the pixel in different seasons, the corresponding cover function index is calculated; Determining a corresponding vegetation production rate according to a change curve of the surface temperature data of the pixel; Determine the corresponding altitude, slope, aspect, terrain moisture index, first-class terrain location index, second-class terrain location index, first-class distance from the water outlet, and second-class distance from the water outlet based on the elevation data of the pixel; Based on the abundance values of various end members of the pixel, vegetation variables and terrain variables, a pre-trained soil property prediction model is used to determine the soil organic carbon content, soil texture type, soil layer thickness and bulk density respectively; The corresponding organic carbon density is determined based on the soil organic carbon content, soil layer thickness and bulk density of the pixel.
3. The method for identifying multi-level soil organic carbon management units according to claim 2, characterized in that: After calculating the environmental variables of each pixel, the method for identifying the soil organic carbon multi-level management unit further includes: The Pearson coefficient was used to evaluate the correlation between environmental variables in pixels; According to the correlation threshold, redundant environmental variables in each environmental variable of the pixel are screened and eliminated.
4. The method for identifying multi-level soil organic carbon management units according to claim 1, characterized in that: Several management units at each level form their own self-organizing map model; the number of management units in the self-organizing map model at each level is determined according to the following steps: An empirical value of the number of management units corresponding to each level is set based on experience; for the first-level self-organizing map model, the corresponding empirical value of the number of management units is five times the square root of the number of pixels in the remote sensing monitoring image; for the second-level self-organizing map model and the third-level self-organizing map model, the corresponding empirical value of the number of management units is five times the square root of the number of management units in the self-organizing map model of the previous level; For any level of the self-organizing map model, a number of quantity values are selected around the corresponding empirical value of the number of management units to obtain a set of quantity values to be determined; For any quantity value in the set of quantity values to be determined, respectively calculating the quantity error and the structural error under the quantity value; The number of management units of the self-organizing map model at the level is determined according to the quantity error and the structural error under each quantity value.
5. The method for identifying multi-level soil organic carbon management units according to claim 2, characterized in that: The soil property prediction model for respectively determining soil organic carbon content, soil texture type, soil layer thickness, and bulk density, and the cover type discrimination model for discriminating the cover type of a pixel, are both obtained by training a random forest model using historical data samples of the target area.
6. A soil organic carbon multi-level management unit identification system, characterized in that: The identification system of the soil organic carbon multi-level management unit is used to implement the identification method of the soil organic carbon multi-level management unit according to any one of claims 1 to 5, and the identification system of the soil organic carbon multi-level management unit includes: A remote sensing monitoring image acquisition module is used to acquire remote sensing monitoring images of the target area; the remote sensing monitoring images include a plurality of pixels, each pixel representing reflectance data, surface temperature data, and elevation data of a plurality of bands at a corresponding position; a pixel environmental variable calculation module, configured to calculate the environmental variables of each pixel based on the reflectivity data, surface temperature data, and elevation data of each pixel in the remote sensing monitoring image using a machine learning algorithm; the environmental variables include terrain variables, vegetation variables, and soil variables; A multi-level management unit identification module is used to use a cascade self-organizing map model constructed based on hierarchical nesting theory to identify multi-level management units according to the environmental variables of each pixel in the remote sensing monitoring image, and determine the corresponding relationship between each pixel and each management unit at each level; the hierarchical nesting theory shows that the higher the level, the lower the time frequency of changes in the environmental variables of the pixels at that level, and the larger the response time and scale; the cascade self-organizing map model includes several levels, each of the levels includes several management units; each management unit includes several pixels.
7. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method for identifying a multi-level management unit of soil organic carbon according to any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for identifying the soil organic carbon multi-level management unit according to any one of claims 1 to 5 is implemented.
9. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method for identifying the soil organic carbon multi-level management unit according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
A soil nutrient classification map generation method based on 3S technology and precision evaluation method thereof
CN108959661A
Soil organic carbon content prediction method based on random forest-ordinary Kriging method
CN109342697A