Identification method of soil organic carbon multi-level management unit and related device

The remote sensing data is processed through machine learning algorithms and cascaded self-organization mapping models to identify the multi-level management unit of soil organic carbon, solving the problems of inaccurate identification and difficulty in updating in traditional methods, and achieving efficient and accurate soil organic carbon management.

CN120298904AActive Publication Date: 2025-07-11CHINA AGRI UNIV
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510779815.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-07-11
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

The prior art is difficult to quickly and accurately identify problems faced by soil organic carbon and solve management units with similar needs at different hierarchical scales in the region. Traditional methods rely on historical soil survey maps with uncertainty and difficulty in updating, and fail to effectively consider the impact of human land use and management factors.

Method used

Machine learning algorithms are used to process remote sensing monitoring image data, combined with cascaded self-organization mapping model and hierarchical nesting theory, multi-level management units for soil organic carbon are identified, and the corresponding relationship of multi-level management units is constructed by calculating environmental variables such as terrain, vegetation and soil variables is considered, and the impact of human land use and management is considered.

Benefits of technology

It improves the accuracy and accuracy of multi-level management unit identification, saves manpower and material resources in field surveys, covers places that are difficult to reach in traditional surveys, and ensures the ease of updating of results, meeting the dynamic monitoring and management needs of regional soil organic carbon.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298904A_ABST
    Figure CN120298904A_ABST
Patent Text Reader

Abstract

The invention discloses an identification method of a soil organic carbon multi-level management unit and a related device, and relates to the technical field of land parcel division. In the method, a machine learning algorithm is adopted to identify a soil organic carbon multi-level management unit according to reflectivity data, surface temperature data and elevation data of each pixel in a remote sensing monitoring image; calculating to obtain a topographic variable, a vegetation variable and a soil variable of each pixel; then, a cascade self-organizing mapping model constructed based on the hierarchical nesting theory is adopted, multi-level management units are recognized according to the remote sensing monitoring images, and pixels included in the management units of each level are determined; the environment variables are calculated through remote sensing monitoring data and a machine learning algorithm, the corresponding relation between the pixels and the management unit of each level is automatically recognized through the cascade self-organizing mapping model, manpower and material resources of a large amount of field investigation are saved, meanwhile, the place where traditional investigation is difficult to reach can be covered, and the field investigation efficiency is improved. And the recognition precision and accuracy of the multi-level management unit are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of plot division, and particularly to a method for identifying multi-level management units of soil organic carbon and related devices. Background Art

[0002] As the largest carbon pool in the Earth's surface system, the global soil organic carbon pool is three times that of the atmospheric carbon pool and 2.5 times that of the terrestrial vegetation carbon pool. Due to the huge soil organic carbon reserves, even a slight change can have a significant impact on the atmospheric carbon dioxide concentration and climate change. At the same time, as an important measure of soil fertility, soil organic carbon can have a profound impact on food production by affecting soil erosion resistance and water use. Therefore, enhancing soil organic carbon is considered an important measure to mitigate global climate change and ensure food security. And how to identify multi-level management units has become an important prerequisite and foundation for improving soil organic carbon management.

[0003] To meet the needs of soil and land resource management, a series of conceptual frameworks and technical systems have been developed at home and abroad for identifying multi-level soil organic carbon management units. For example, the land resource hierarchical framework proposed by the Natural Resources Conservation Service of the US Department of Agriculture identifies geographical units with similar resource problems and solution needs based on soil geography at the regional or landscape scale to support the formulation of multi-scale projects and management practices. China has also proposed various zoning methods including ecological zoning schemes, ecological geographical zoning, terrestrial ecological basic zoning, and agricultural natural zoning, and conducts management zoning and unit identification by considering the similarities and differences of elements such as climate, landform, soil, vegetation, and hydrology at different levels. These frameworks and technical systems provide a solid foundation for the identification of regional multi-level soil organic carbon management units.

[0004] Although the land resource hierarchical framework of the US Department of Agriculture divides different hierarchical structures and unit division indicators based on soil geography, on the one hand, the division of management units at the regional scale in this framework relies on historical soil survey maps, and these point-to-face survey maps have high uncertainties and are difficult to update; on the other hand, for the management units of regional-scale soil organic carbon, this framework does not provide clear division variables. The zoning schemes designed in China mainly target large-scale macro soil and land resource management, and rarely consider the impact analysis of factors such as human land use and land management on soil organic carbon at the regional scale, while these factors are the main driving factors for the dynamics of regional-scale soil organic carbon.

[0005] In summary, how to quickly and accurately identify management units with similar problems and solution needs of soil organic carbon at different hierarchical scales in a region to serve regional soil organic carbon management is a key problem that urgently needs to be solved in this field. Summary of the Invention

[0006] The objective of this application is to provide a method and related device for identifying multi - level management units of soil organic carbon, which can improve the accuracy and precision of identifying multi - level management units.

[0007] To achieve the above objective, this application provides the following solutions: In the first aspect, this application provides a method for identifying multi - level management units of soil organic carbon, including the following steps: Obtain remote sensing monitoring images of the target area; the remote sensing monitoring images include several pixels, and each pixel represents the reflectance data, surface temperature data, and elevation data of several bands at the corresponding position.

[0008] Using a machine learning algorithm, calculate the environmental variables of each pixel based on the reflectance data, surface temperature data, and elevation data of each pixel in the remote sensing monitoring image; the environmental variables include topographic variables, vegetation variables, and soil variables.

[0009] Using a cascaded self - organizing map model constructed based on the hierarchical nesting theory, identify multi - level management units according to the environmental variables of each pixel in the remote sensing monitoring image, and determine the corresponding relationship between each pixel and each management unit at each level; the hierarchical nesting theory indicates that the higher the level, the lower the time frequency of changes in the environmental variables of the pixels at that level, and the longer the response time and larger the scale; the cascaded self - organizing map model includes several levels, and each level includes several management units; each management unit includes several pixels.

[0010] Optionally, the topographic variables include altitude, slope, aspect, topographic wetness index, type - 1 topographic position index, type - 2 topographic position index, type - 1 distance from the water outlet, and type - 2 distance from the water outlet; the vegetation variables include cover type, cover function index, and vegetation production rate; the soil variables include soil organic carbon content, organic carbon density, soil texture type, soil layer thickness, and bulk density; the cascaded self - organizing map model includes three levels, which are divided into the first level, the second level, and the third level from bottom to top; the environmental variables of the first level include vegetation variables, soil organic carbon content, bulk density, altitude, slope, aspect, and topographic wetness index; the environmental variables of the second level include vegetation variables, soil variables, altitude, slope, aspect, and topographic wetness index; the environmental variables of the third level include topographic wetness index, type - 1 topographic position index, type - 2 topographic position index, type - 1 distance from the water outlet, and type - 2 distance from the water outlet.

[0011] Optionally, using a machine learning algorithm to calculate the environmental variables of each pixel based on the reflectance data, surface temperature data, and elevation data of each pixel in the remote sensing monitoring image specifically includes the following steps: Based on the reflectance values of each pixel in the remote sensing monitoring image, the abundance values of various endmembers in each pixel and the entire endmember abundance value image of the remote sensing monitoring image are obtained through spectral mixture analysis; the types of endmembers include vegetation, sand, salt, and dark matter.

[0012] For any pixel, the abundance values of various endmembers in the pixel and the entire endmember abundance value image of the remote sensing monitoring image are input into a pre-trained land cover type discrimination model to obtain the corresponding land cover type.

[0013] Based on the abundance values of various endmembers in the pixel in different seasons, the corresponding land cover function index is calculated.

[0014] Based on the change curve of the surface temperature data of the pixel, the corresponding vegetation production rate is determined.

[0015] Based on the elevation data of the pixel, the corresponding altitude, slope, aspect, terrain wetness index, first-class terrain position index, second-class terrain position index, first-class distance to the water outlet, and second-class distance to the water outlet are determined.

[0016] Based on the abundance values of various endmembers, vegetation variables, and terrain variables of the pixel, a pre-trained soil property prediction model is used to determine the soil organic carbon content, soil texture type, soil layer thickness, and bulk density respectively.

[0017] Based on the soil organic carbon content, soil layer thickness, and bulk density of the pixel, the corresponding organic carbon density is determined.

[0018] Optionally, after calculating the environmental variables of each pixel, the identification method of the soil organic carbon multi-level management unit further includes the following steps: The Pearson coefficient is used to evaluate the correlation between the environmental variables of the pixel.

[0019] According to the correlation threshold, redundant environmental variables among the environmental variables of the pixel are screened and removed.

[0020] Optionally, several management units at each level form their own self-organizing mapping models; the number of management units of the self-organizing mapping models at each level is determined according to the following steps: The empirical values of the number of management units corresponding to each level are set according to experience; for the self-organizing mapping model of the first level, the empirical value of the number of management units is five times the square root of the number of pixels in the remote sensing monitoring image; for the self-organizing mapping models of the second level and the third level, the empirical value of the number of management units is five times the square root of the number of management units of the self-organizing mapping model of the upper level.

[0021] For the self-organizing mapping model at any level, select a number of values on both sides of the empirical value of the corresponding number of management units to obtain a set of values to be determined.

[0022] For any value in the set of values to be determined, calculate the quantity error and structure error under the value respectively.

[0023] Determine the number of management units of the self-organizing mapping model at the level according to the quantity error and structure error under each value.

[0024] Optionally, the soil property prediction model for respectively determining the soil organic carbon content, soil texture type, soil layer thickness, and bulk density, and the cover type discrimination model for discriminating the cover type of pixels are both obtained by training a random forest model with historical data samples of the target area.

[0025] In a second aspect, the present application provides a recognition system for multi-level management units of soil organic carbon, including the following functional modules: A remote sensing monitoring image acquisition module, configured to acquire a remote sensing monitoring image of the target area; the remote sensing monitoring image includes a number of pixels, and each pixel represents reflectance data, surface temperature data, and elevation data of a number of bands at the corresponding position.

[0026] A pixel environmental variable calculation module, configured to calculate the environmental variables of each pixel by using a machine learning algorithm according to the reflectance data, surface temperature data, and elevation data of each pixel in the remote sensing monitoring image; the environmental variables include topographic variables, vegetation variables, and soil variables.

[0027] A multi-level management unit recognition module, configured to use a cascaded self-organizing mapping model constructed based on the hierarchical nesting theory to identify multi-level management units according to the environmental variables of each pixel in the remote sensing monitoring image, and determine the corresponding relationship between each pixel and each management unit at each level; the hierarchical nesting theory indicates that the higher the level, the lower the time frequency of the change of the environmental variables of the pixels at that level, and the longer the response time and the larger the scale; the cascaded self-organizing mapping model includes a number of levels, and each level includes a number of management units; each management unit includes a number of pixels.

[0028] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor executes the computer program to implement the steps of the recognition method for multi-level management units of soil organic carbon described above.

[0029] Fourthly, the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method for identifying the multi-level management unit of soil organic carbon described above are implemented.

[0030] Fifthly, the present application provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the method for identifying the multi-level management unit of soil organic carbon described above are implemented.

[0031] According to the specific embodiments provided by the present application, the following technical effects are disclosed: The present application provides a method for identifying a multi-level management unit of soil organic carbon and related devices. In this method, using a machine learning algorithm, topographic variables, vegetation variables, and soil variables of each pixel are calculated based on the reflectance data, land surface temperature data, and elevation data of each pixel in the remote sensing monitoring image. Subsequently, a cascade self-organizing map model constructed based on the hierarchical nesting theory can be used to identify the multi-level management unit according to the remote sensing monitoring image, that is, to determine which pixels are included in each management unit at each level. The present application clarifies the multi-level division of the management unit according to the time frequency, response time, and scale of the changes in the different characteristics of soil organic carbon in the target area under human land use and management, so as to guide the soil organic carbon management at different levels such as the regional and plot scales. The environmental variables are calculated using remote sensing monitoring data and machine learning algorithms, and the cascade self-organizing map model is used to automatically identify the corresponding relationship between the pixels and the management units at each level, which not only saves a large amount of manpower and material resources for field surveys, but also can cover places that are difficult to reach by traditional surveys, improving the accuracy and precision of the identification of multi-level management units. In addition, the high spatio-temporal resolution of remote sensing data ensures the easy update of the results, thus meeting the requirements of dynamic monitoring and management of regional soil organic carbon. Description of the Drawings

[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0033] Figure 1 It is a flowchart of a method for identifying a multi-level management unit of soil organic carbon provided by an embodiment of the present application.

[0034] Figure 2 It is a flowchart of step A2 in a method for identifying a multi-level management unit of soil organic carbon provided by an embodiment of the present application.

[0035] Figure 3 The two-dimensional scatter plot of the principal component analysis in a method for identifying multi-level management units of soil organic carbon provided by an embodiment of the present application.

[0036] Figure 4 The schematic diagram of the calculation principles of TPI and DTP in a method for identifying multi-level management units of soil organic carbon provided by an embodiment of the present application.

[0037] Figure 5 The calculation flow chart of TPI and DTP in a method for identifying multi-level management units of soil organic carbon provided by an embodiment of the present application.

[0038] Figure 6 The schematic diagram of the technical route for determining soil properties in a method for identifying multi-level management units of soil organic carbon provided by an embodiment of the present application.

[0039] Figure 7 The schematic diagram of a land resource hierarchy framework mentioned in a method for identifying multi-level management units of soil organic carbon provided by an embodiment of the present application.

[0040] Figure 8 The schematic diagram of three cascaded self-organizing mapping models in a method for identifying multi-level management units of soil organic carbon provided by an embodiment of the present application.

[0041] Figure 9 The schematic diagram of the functional modules of a system for identifying multi-level management units of soil organic carbon provided by an embodiment of the present application.

[0042] Figure 10 The schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0043] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0044] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0045] A method for identifying multi-level management units of soil organic carbon provided by an embodiment of the present application, in an exemplary embodiment, as Figure 1 shown, includes the following steps: A1. Obtain remote sensing monitoring images of the target area; the remote sensing monitoring images include a number of pixels, each of which represents reflectivity data, surface temperature data and elevation data of a number of bands at a corresponding position.

[0046] A2. Use machine learning algorithms to calculate the environmental variables of each pixel based on the reflectivity data, surface temperature data and elevation data of each pixel in the remote sensing monitoring image; environmental variables include terrain variables, vegetation variables and soil variables.

[0047] Specifically in this embodiment, terrain variables include altitude, slope, aspect, terrain moisture index, type I terrain site index, type II terrain site index, type I distance from the water outlet, and type II distance from the water outlet; vegetation variables include cover type, cover function index, and vegetation production rate; soil variables include soil organic carbon content, organic carbon density, soil texture type, soil layer thickness, and bulk density.

[0048] In an exemplary embodiment, Figure 2 As shown, step A2 specifically includes the following steps: A21. According to the reflectance value of each pixel in the remote sensing monitoring image, the abundance value of each end member in each pixel and the total end member abundance value image of the remote sensing monitoring image are obtained based on spectral mixture decomposition; the end member categories include vegetation, sand, salt and dark matter. Spectral mixture decomposition mainly includes three steps: principal component analysis, end member extraction based on geometric vertices, and calculation of end member abundance values ​​of each pixel in the study area.

[0049] First, the months with the highest vegetation coverage (such as August) and the lowest surface coverage (such as May) were determined based on local knowledge of the study area. The Landsat surface reflectance data images of these months were subjected to principal component analysis using ENVI software to obtain the transformed principal component images. The types and numbers of end members were determined based on existing research in the region. In arid areas, monitoring is generally carried out on four types of end members: vegetation (GV), sand (SL), salt (SA), and dark matter (DA). Vegetation (GV) end members were extracted in August, and sand (SL), salt (SA), and dark matter (DA) end members were extracted in May.

[0050] Then, the end member spectral curve is extracted by using the two-dimensional scatter plot extraction method. Based on the mathematical mechanism of the linear mixed model, the two-dimensional scatter plot of the principal component analysis should be triangularly distributed, such as Figure 3As shown in the figure, this is because the four types of endmembers are in the shape of a tetrahedron. When placed in a two-dimensional scatter plot, they show a triangular distribution. The three vertices of the triangle can be regarded as pure endmembers, and the interior of the triangle is a mixed pixel, which is linearly combined by pure endmembers. In the ENVI software, a scatter plot is constructed using the principal component image corresponding to the extraction month of the endmembers. 200 - 400 pixels are selected as endmember pixels at the vertices of the principal component scatter plot, and the average reflectance of the original image at the location of the endmember pixels is used as the endmember spectrum.

[0051] Finally, after obtaining the spectra of various endmembers in the area, linear spectral mixture analysis is performed in the ENVI software using the endmember spectra and the regional reflectance image to obtain the endmember abundance values of each endmember in the study area.

[0052] Linear spectral mixture analysis is the main method for estimating abundance values, and its expression is: 。

[0053] Among them, i is the remote sensing image band, R i is the reflectance of band i, m is the number of endmembers. For each endmember j (1 ≤ j ≤ m ), E i,j is the reflectance of endmember j in band i, F j is the endmember j abundance value, that is, the percentage of the endmember occupying the pixel area, is the residual.

[0054] In practice, in order to make the estimated abundance values have physical meanings, generally the fully constrained least squares method can be used to obtain the abundance values of each endmember in each pixel, and the corresponding abundance value image is obtained. The abundance value of endmember j satisfies the following formula: 。

[0055] A22. For any pixel, the abundance values of various endmembers in the pixel and the entire endmember abundance value image of the remote sensing monitoring image are input into the pre-trained cover type discrimination model to obtain the corresponding cover type.

[0056] First, according to the land cover classification system (such as forest land, grassland, cultivated land, construction land, water area, etc.), about 1000 pixels are selected for each classification through the high-spatial-resolution images of Google Earth, and the training samples and validation samples are divided in a ratio of 75% and 25%. Then, the training samples, their corresponding endmember abundance values, and the fully decomposed endmember abundance value images obtained by spectral mixture decomposition are input into the random forest model for land cover type classification. Finally, the accuracy evaluation is carried out according to the land cover types of the validation samples and the corresponding classification results, and the evaluation indicators include overall accuracy p e and Kappa coefficient k , and the calculation formulas are as follows: .

[0057] .

[0058] Among them, is the sum of the number of correctly classified samples in each category divided by the total number of samples, that is, the overall classification accuracy, , … is the number of true samples in each category, while , … is the number of samples in each category in the prediction result, and the total number of samples is n.

[0059] A23. According to the abundance values of various endmembers of the pixel in different seasons, the corresponding land cover function index is calculated. The land cover function index LCI is calculated according to the following formula: .

[0060] Among them, , and are the abundance values of the spring vegetation canopy start growth phenophases SL, SA, and DA endmembers respectively; is the abundance value of the summer vegetation growth most vigorous phenophase GV endmember.

[0061] A24. According to the change curve of the surface temperature data of the pixel, the corresponding vegetation production rate is determined. The vegetation production rate index VPI is mainly calculated using the MODIS surface temperature data: .

[0062] Among them, is the function of the reference surface temperature curve, is the function of the pixel surface temperature curve, and are the start and end times of the main growth season respectively.

[0063] A25. Determine the corresponding altitude, slope, aspect, terrain wetness index, type I terrain position index, type II terrain position index, type I distance from the outlet, and type II distance from the outlet according to the elevation data of the pixels.

[0064] For example, using the elevation data, the slope and aspect can be calculated in the ArcGis hydrological tool module. The calculation formula for the slope is as follows: .

[0065] Where is the slope, is the difference in altitude, is the difference in horizontal distance.

[0066] The terrain wetness index TWI can be calculated according to the following formula: .

[0067] Where is the upstream area of the unit contour line where the surface water flows through, is the slope.

[0068] As Figure 4 shown, the terrain position index TPI describes the relative terrain position of the pixel in its sub-basin in the vertical direction, effectively improving the incomparability of the DEM's description of the terrain position in different sub-basins. The distance from the outlet DTP reflects the migration distance of the material and its redistribution by calculating the distance of the pixel from the outlet point of the sub-basin in the horizontal direction. The calculation formula is as follows: .

[0069] Where is the DEM value of pixel i, is the DEM value of outlet j in the sub-basin where pixel i is located, is the maximum DEM value in the sub-basin where pixel i is located.

[0070] .

[0071] Where and are the distances of pixel i to the outlet j of its sub-basin in the east-west and north-south directions respectively.

[0072] In specific applications, the first - level and second - level outfalls are considered, so the TPI and DTP indices are subdivided into two categories. One category is within the first - level sub - basins, with each first - level outfall as the outfall of each basin, and the TPI and DTP of each pixel are calculated, and the results are marked as TPI1 and DTP1; the other category is within the second - level sub - basins, with the second - level outfall as the outfall of the basin, and the TPI and DTP of each pixel are calculated, and the results are marked as TPI2 and DTP2. Due to the existence of some real terrains (such as karst landforms), there are some sunken areas on the DEM surface. When performing the water - flow direction analysis, the existence of these areas will lead to unreasonable or even incorrect water - flow directions. Before calculating the water - flow direction, first use the fill - depression tool in the Arcgis hydrological analysis tool to perform fill - depression operations on the original DEM; then, based on the DEM without depressions generated in the previous step, use the D8 algorithm in the flow - direction tool to calculate the water - flow direction of each pixel (i.e., D8 flow - direction); next, use the flow - rate tool, input the generated DEM without depressions and the flow - direction data, and count the flow - rate of each pixel; based on the counted flow - rate, by setting a reasonable empirical value of the minimum surface runoff, use the raster calculator tool to generate river links; furthermore, in the raster river - network vectorization tool, input the obtained river - link raster data and flow - direction data to obtain the vector data of the river network, and then use the river - network grading tool to grade the obtained river - network vector data and determine the outfalls; finally, use the water - area analysis tool to extract the basins according to the flow - direction raster data and the river - network grading data. This process is mainly completed by combining python programming on the basis of the ArcGis hydrological tool, and the process is as Figure 5 shown.

[0073] A26. According to the abundance values of various end - members, vegetation variables, and terrain variables of pixels, use the pre - trained soil - property prediction model to determine the soil organic - carbon content, soil - texture type, soil - layer thickness, and bulk density respectively.

[0074] In this embodiment, the remote - sensing characterization methods for soil organic - carbon content, soil - texture type, soil - layer thickness, and bulk density all adopt digital mapping technology, that is, based on the relationship between soil variables and their characteristic variables (abundance values of various end - members, vegetation variables, and terrain variables), use the random - forest model to map the spatial soil properties.

[0075] As Figure 6 shown in the schematic technical route diagram for determining soil properties, first randomly divide the soil - sample data of the region into a training data set and a validation data set according to the ratio of 75% and 25%. In the training data set, adopt the grid - search method to determine the optimal combination of the number of trees, maximum depth, and maximum number of features of the random - forest model according to the results of 5 - fold cross - validation. The specific operation is as follows: (1)Preset the number of trees in the random forest model according to experience to be from 100 to 1000, with an interval of 50; the maximum depth to be from 4 to 10, with an interval of 1; and the maximum number of features to be from 2 to 5, with an interval of 1.

[0076] (2)Grid search method, that is, combine all possible values of the three parameters within the preset range.

[0077] (3)Then perform 5-fold cross-validation on each combination of parameter values: Divide the training samples into 5 parts, successively use each part of the samples as the validation samples, and the remaining four parts as the training samples, and calculate the average R of 5 times 2 and RMSE, and select the optimal combination of parameters with the highest value as the optimal model parameters. Calculate R according to the following formula 2 and RMSE: .

[0078] .

[0079] Among them, and are the predicted value and the observed value (true value) of sample point i respectively, is the average value of the observed values (true values) of all sample points, r is the number of sample points.

[0080] Then construct random forest models according to their respective optimal model parameter combinations and conduct accuracy evaluation.

[0081] Use the random forest model of the "rpart" package in R language, input the observed values of the soil properties of each sample and the values of their corresponding characteristic variables, and set the parameters according to the selected optimal model parameters, and the random forest models corresponding to each soil variable in the study area can be obtained.

[0082] Next, use the land cover type data to mask the construction land and water areas to exclude the interference with the mapping results. Input the dataset (image) of the characteristic variables into the trained random forest model. In order to ensure the accuracy of the spatial mapping results, take the mean value of the results of running the random forest model 100 times as the final prediction result, and use R 2 and RMSE to evaluate the accuracy of the mapping results: the mean map of soil organic carbon content, the mean map of soil layer thickness, the texture type map and the mean map of bulk density. At the same time, use the standard deviation of the 100 predicted values for the evaluation of model uncertainty.

[0083] A27. Determine the corresponding organic carbon density according to the soil organic carbon content, soil layer thickness and bulk density of the pixel. The calculation formula of the density of soil organic carbon is as follows: .

[0084] Among them, SOCD refers to the soil organic carbon density, SOCC refers to the soil organic carbon content, D represents the soil layer thickness, BD represents the bulk density.

[0085] In order to avoid collinearity and redundancy of information among input variables, after calculating the environmental variables of each pixel in step A2, the identification method of the multi-level management unit of soil organic carbon further includes the following steps: B1. Use the Pearson coefficient to evaluate the correlation between the environmental variables of each pixel. In this embodiment, the Pearson coefficient between any two environmental variables is calculated according to the following formula: .

[0086] Among them, and are respectively the values of each sample point of the variables x and y , and are respectively the averages of the variables x and y .

[0087] B2. According to the correlation threshold, screen and eliminate the redundant environmental variables among the environmental variables of each pixel. Use the Pearson correlation coefficient re with an absolute value equal to 0.7 as the correlation threshold for collinearity. If the correlation coefficient of two variables exceeds the threshold, only one of them is retained as the input for further analysis.

[0088] A3. Use the cascade self-organizing map model constructed based on the hierarchical nesting theory to identify the multi-level management unit according to the environmental variables of each pixel in the remote sensing monitoring image, and determine the corresponding relationship between each pixel and each management unit at each level. The hierarchical nesting theory shows that the higher the level, the lower the time frequency of the change of the environmental variables of the pixels at that level, and the longer the response time and the larger the scale; the cascade self-organizing map model includes several levels, and each level includes several management units; each management unit includes several pixels.

[0089] In this embodiment, the cascaded self-organizing mapping model includes three levels, which are divided into the first level, the second level, and the third level from bottom to top; the environmental variables of the first level include vegetation variables, soil organic carbon content, bulk density, altitude, slope, aspect, and topographic wetness index; the environmental variables of the second level include vegetation variables, soil variables, altitude, slope, aspect, and topographic wetness index; the environmental variables of the third level include topographic wetness index, first type of topographic position index, second type of topographic position index, first type of distance from the water outlet, and second type of distance from the water outlet.

[0090] As Figure 7 shown, there is an existing land resource hierarchical framework abroad, which divides three levels of State layer, LS layer, and LSG layer from bottom to top at the regional scale. This hierarchical nesting theory indicates that the higher the level, the lower the time frequency of changes in the state variables (variables determining state resilience) at that level, and the larger the spatial scale of differences. Therefore, a multi-level structure framework of soil organic carbon in the target area is constructed according to the above framework and theory. The soil organic carbon content, which has the highest change frequency (several years - decades) and the smallest spatial scale of differences (plots) under human land use and management, is determined as the state variable of the State layer. In addition to the soil organic carbon content, the density of soil organic carbon also takes into account soil layer thickness, bulk density, and soil texture. These properties change at a frequency ranging from decades to hundreds of years under human land use and management, and the spatial scale of differences is much larger than that of plots. Therefore, it is determined as the state variable of the LS layer. The topographic position reflects the differences in parent material and lithology, which determines the upper limit of soil organic carbon. At the same time, it changes at a frequency ranging from hundreds of years to thousands of years and has the largest spatial scale of differences at the regional level. Therefore, it is used as the state variable of the LSG layer.

[0091] In the constructed multi-level structure framework of soil organic carbon, the types of environmental variables (variables affecting the state) that play a leading role in hierarchical division include three major types: soil variables, vegetation variables, and topographic variables. Among the specific variables, five variables related to the amount of soil organic carbon, namely soil organic carbon content, organic carbon density, soil texture type, soil layer thickness, and bulk density, are determined as soil variables. According to the minimum indicator set determined by international relevant sustainable development goals, the cover type and vegetation production rate are determined as vegetation variables. At the same time, to characterize the response of vegetation to the natural environment and human management, the cover function index is also determined as a vegetation variable. Among the topographic variables, altitude, slope, and aspect are determined as topographic variables at the regional small scale, the topographic wetness index is determined as a topographic variable at the regional medium and small scale, and the topographic position index and the distance from the water outlet are determined as topographic variables at the regional large scale.

[0092] According to the above multi-layer nested structure, in the State layer, all vegetation variables, soil organic carbon content, bulk density, and terrain humidity index at the small and medium scales in the region are determined as its dominant environmental variables; in the LS layer, all vegetation variables, all soil variables, and terrain humidity index at the small and medium scales in the region are determined as its dominant environmental variables; in the LSG layer, the terrain humidity index at the small and medium scales in the region, the terrain position index at the large scale in the region, and the distance from the water outlet are determined as the dominant environmental variables.

[0093] Table 1 Dominant environmental variables at each level

[0094] In an exemplary embodiment, a cascaded self-organizing map model is constructed using three cascaded self-organizing map models as Figure 8 shown. First, the original values of the environmental variables of the State layer corresponding to each pixel of the remote sensing monitoring image x 1, x 2, x 3, and x 4 are combined into an input vector and input into the self-organizing map model 1 to obtain the recognition result of the State layer. This result includes the management units of the State layer, their corresponding relationships with the input samples, and the weight vectors of all variables corresponding to each management unit of the State layer. It should be noted that since the form of the self-organizing map model at each level is a two-dimensional grid form, the management units can also be represented as neurons; then, the weight vectors of the environmental variables of the LS layer corresponding to each management unit of the State layer are input into the self-organizing map model 2 to obtain the recognition result of the LS layer. This result includes the management units of the LS layer, their corresponding relationships with the input management units of the State layer, and the weight vectors of all variables corresponding to each management unit of the LS layer; finally, the weight vectors of the environmental variables of the LSG layer corresponding to each management unit of the LS layer are input into the self-organizing map model 3 to obtain the state recognition result of the LSG layer. This result includes the management units of the LSG layer, their corresponding relationships with the input management units of the LS layer, and the weight vectors of all variables corresponding to each management unit of the LSG layer.

[0095] In an exemplary embodiment, several management units at each level form their respective self-organizing map models; specifically, the number of management units of the self-organizing map model at each level is determined according to the following steps: C1. Set the empirical values of the number of management units corresponding to each level according to experience. For the self-organizing mapping model of the first level, the empirical value of the number of management units is five times the square root of the number of pixels in the remote sensing monitoring image; for the self-organizing mapping models of the second level and the third level, the empirical value of the number of management units is five times the square root of the number of management units of the self-organizing mapping model of the previous level. That is, set the empirical value of the number of management units according to the following formula: .

[0096] Among them, M is the empirical value of the number of management units, S is the number of pixels in the remote sensing monitoring image or the number of management units of the self-organizing mapping model of the previous level. For example, assume that the number of pixels in the remote sensing monitoring image is 253, then the initial empirical value of the number of management units is 80.

[0097] C2. For any self-organizing mapping model of a level, select several numerical values on both sides of the empirical value of the number of management units to obtain a set of numerical values to be determined. Since the self-organizing mapping model is a two-dimensional network, in this embodiment, 3 numerical values are set on both sides of 80 according to experience. Therefore, the numerical values to be determined are preset as 10×7, 9×8, 11×7, 10×8, 11×8, 10×9, 12×8.

[0098] C3. For any numerical value in the set of numerical values to be determined, calculate the numerical error and structural error under the numerical value respectively. Calculate the numerical error under each numerical value according to the following formula QE and the structural error TE : .

[0099] .

[0100] Among them, Num is the total number of samples, is the i th input sample, is the weight vector of the best matching management unit corresponding to the sample , is the Euclidean distance between the two, and are the first and second neighboring management units of the sample respectively. When and are not adjacent, returns 1, otherwise returns 0. Here, adjacent means that in the two-dimensional network of self-organizing mapping, the absolute value of the difference in rows or columns is 1.

[0101] C4. Determine the number of management units of the self-organizing mapping model at this level according to the quantity error and structure error at each quantity value.

[0102] Specifically in this embodiment, the cover type discrimination model for discriminating the cover type of pixels in step A22 and the soil property prediction models for respectively determining soil organic carbon content, soil texture type, soil layer thickness, and bulk density in step A26 are both obtained by training a random forest model with historical data samples of the target area.

[0103] On the one hand, the method provided in this application constructs a conceptual framework of a land system nested hierarchical structure based on soil organic carbon, and clarifies the state variables and dominant environmental variables at each level according to the different characteristic rates (occurrence frequency, response time, and scale) of soil organic carbon at different levels in the region under human land use and management, so as to guide the management of soil organic carbon at different levels such as the regional and plot scales. On the other hand, using multi-source remote sensing data and machine learning algorithms for spatial mapping of environmental variables not only saves a large amount of manpower and material resources for field surveys, but also can cover places that are difficult to reach by traditional surveys, improving the accuracy and precision of the mapping results. In addition, the high spatio-temporal resolution of remote sensing data ensures the easy update of the results, thus meeting the needs of dynamic monitoring and management of regional soil organic carbon.

[0104] In summary, by combining professional concepts such as soil geography, machine learning algorithms such as random forest, and remote sensing big data, this application can quickly and accurately identify the problems faced by soil organic carbon and the multi-level management units with similar solution requirements in the region, serving the management of regional soil organic carbon. In addition, the management units identified by the method provided in this application can be used as the basic units for monitoring and evaluating regional cultivated land quality, ecological quality evaluation and management, and land degradation and restoration.

[0105] Based on the same inventive concept, the embodiment of this application also provides a system for implementing the above-mentioned method for identifying a multi-level management unit of soil organic carbon. The solution provided by this system for solving problems is similar to the solution described in the above method. In an exemplary embodiment, as Figure 9 shown, the above-mentioned system for identifying a multi-level management unit of soil organic carbon includes the following functional modules: A remote sensing monitoring image acquisition module for acquiring remote sensing monitoring images of the target area; the remote sensing monitoring images include a number of pixels, and each pixel represents the reflectance data, surface temperature data, and elevation data of several bands at the corresponding position.

[0106] The pixel environmental variable calculation module is used to calculate the environmental variables of each pixel based on the reflectance data, land surface temperature data, and elevation data of each pixel in the remote sensing monitoring image by using a machine learning algorithm; the environmental variables include topographic variables, vegetation variables, and soil variables.

[0107] The multi-level management unit identification module is used to identify the multi-level management units according to the environmental variables of each pixel in the remote sensing monitoring image by using a cascade self-organizing mapping model constructed based on the hierarchical nesting theory, and determine the corresponding relationship between each pixel and each management unit at each level; the hierarchical nesting theory shows that the higher the level, the lower the time frequency of the change of the environmental variables of the pixels at that level, and the longer the response time and the larger the scale; the cascade self-organizing mapping model includes several levels, and each level includes several management units; each management unit includes several pixels.

[0108] Of course, Figure 9 the architecture shown is only exemplary, and when implementing different functions, one or at least two components in the Figure 9 shown system can be omitted according to actual needs.

[0109] In an exemplary embodiment, a computer device is provided. The computer device can be a server or a terminal, and its internal structure diagram can be as Figure 10 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it can implement a method for identifying a multi-level management unit of soil organic carbon provided in the foregoing embodiments.

[0110] Those skilled in the art can understand that Figure 10 the structure shown in

[0111] In an exemplary embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0112] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0113] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0114] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0115] Those of ordinary skill in the art can understand that all or part of the processes of implementing the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0116] In each of the embodiments provided in this application, the database involved may include at least one of a relational database and a non-relational database. The non-relational database may include a distributed database based on blockchain, etc., without limitation. In each of the embodiments provided in this application, the processor may be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without limitation.

[0117] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0118] Specific examples are used in this article to elaborate on the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.

Claims

1. A method for identifying a multi - level management unit of soil organic carbon, characterized in that, Including: Obtain remote sensing monitoring images of the target area; the remote sensing monitoring images include a number of pixels, and each pixel represents reflectance data, surface temperature data, and elevation data of several bands at the corresponding position; Using a machine learning algorithm, calculate the environmental variables of each pixel based on the reflectance data, surface temperature data, and elevation data of each pixel in the remote sensing monitoring image; the environmental variables include topographic variables, vegetation variables, and soil variables; Using a cascaded self-organizing map model constructed based on the hierarchical nesting theory, identify multi-level management units according to the environmental variables of each pixel in the remote sensing monitoring image, and determine the corresponding relationship between each pixel and each management unit at each level; the hierarchical nesting theory indicates that the higher the level, the lower the time frequency of the change of the environmental variables of the pixels at that level, and the longer the response time and the larger the scale; the cascaded self-organizing map model includes several levels, and each level includes several management units; each management unit includes several pixels.

2. The identification method of the soil organic carbon multi-level management unit according to claim 1, wherein The topographic variables include altitude, slope, aspect, topographic wetness index, first-class topographic position index, second-class topographic position index, first-class distance from the water outlet, and second-class distance from the water outlet; the vegetation variables include cover type, cover function index, and vegetation production rate; the soil variables include soil organic carbon content, organic carbon density, soil texture type, soil layer thickness, and bulk density; the cascaded self-organizing map model includes three levels, which are divided into the first level, the second level, and the third level from bottom to top; the environmental variables of the first level include the vegetation variables, soil organic carbon content, bulk density, altitude, slope, aspect, and topographic wetness index; the environmental variables of the second level include the vegetation variables, the soil variables, altitude, slope, aspect, and topographic wetness index; the environmental variables of the third level include topographic wetness index, first-class topographic position index, second-class topographic position index, first-class distance from the water outlet, and second-class distance from the water outlet.

3. The identification method of the soil organic carbon multi-level management unit according to claim 2, wherein, Using a machine learning algorithm to calculate the environmental variables of each pixel based on the reflectance data, surface temperature data, and elevation data of each pixel in the remote sensing monitoring image specifically includes: Based on the reflectance values of each pixel in the remote sensing monitoring image, obtain the abundance values of various endmembers in each pixel and the full endmember abundance value image of the remote sensing monitoring image through spectral mixture decomposition; the categories of endmembers include vegetation, sand, salt, and dark matter; For any pixel, input the abundance values of various endmembers in the pixel and the full endmember abundance value image of the remote sensing monitoring image into a pre-trained cover type discrimination model to obtain the corresponding cover type; Calculate the corresponding cover function index according to the abundance values of various endmembers of the pixel in different seasons; Determine the corresponding vegetation production rate according to the change curve of the surface temperature data of the pixel; Determine the corresponding altitude, slope, aspect, topographic wetness index, first-class topographic position index, second-class topographic position index, first-class distance from the water outlet, and second-class distance from the water outlet according to the elevation data of the pixel; Based on the abundance values of various endmembers, vegetation variables, and terrain variables of the pixel, a pre-trained soil property prediction model is used to determine the soil organic carbon content, soil texture type, soil layer thickness, and bulk density respectively; Based on the soil organic carbon content, soil layer thickness, and bulk density of the pixel, the corresponding organic carbon density is determined.

4. The identification method of the soil organic carbon multi-level management unit according to claim 3, characterized in that After calculating the environmental variables of each pixel, the method for identifying the multi-level management unit of soil organic carbon further includes: Using the Pearson coefficient to evaluate the correlation between the environmental variables of the pixel; According to the correlation threshold, redundant environmental variables among the environmental variables of the pixel are screened and removed.

5. The identification method of the soil organic carbon multi-level management unit according to claim 2, wherein Several management units at each level form their own self-organizing mapping models; the number of management units of the self-organizing mapping models at each level is determined according to the following steps: Set the empirical value of the number of management units corresponding to each level according to experience; for the self-organizing mapping model of the first level, the corresponding empirical value of the number of management units is five times the square root of the number of pixels of the remote sensing monitoring image; for the self-organizing mapping models of the second level and the third level, the corresponding empirical value of the number of management units is five times the square root of the number of management units of the self-organizing mapping model of the upper level; For any self-organizing mapping model at a certain level, select several numerical values on both sides of the corresponding empirical value of the number of management units to obtain a set of numerical values to be determined; For any numerical value in the set of numerical values to be determined, calculate the numerical error and structural error under the numerical value respectively; Based on the numerical error and structural error under each numerical value, determine the number of management units of the self-organizing mapping model at this level.

6. The identification method of the soil organic carbon multi-level management unit according to claim 3, characterized in that, The soil property prediction model for respectively determining the soil organic carbon content, soil texture type, soil layer thickness, and bulk density and the cover type discrimination model for discriminating the cover type of the pixel are both obtained by training a random forest model with historical data samples of the target area.

7. An identification system for a multi-level management unit of soil organic carbon, characterized in that, Including: A remote sensing monitoring image acquisition module for acquiring a remote sensing monitoring image of the target area; the remote sensing monitoring image includes several pixels, and each pixel represents the reflectance data, surface temperature data, and elevation data of several bands at the corresponding position; A pixel environmental variable calculation module for calculating the environmental variables of each pixel by using a machine learning algorithm based on the reflectance data, surface temperature data, and elevation data of each pixel in the remote sensing monitoring image; the environmental variables include terrain variables, vegetation variables, and soil variables; A multi-level management unit identification module for identifying multi-level management units according to the environmental variables of each pixel in the remote sensing monitoring image by using a cascaded self-organizing mapping model constructed based on the hierarchical nesting theory, and determining the corresponding relationship between each pixel and each management unit at each level; the hierarchical nesting theory indicates that the higher the level, the lower the time frequency of the change of the environmental variables of the pixels at this level, and the longer the response time and the larger the scale; the cascaded self-organizing mapping model includes several levels, and each level includes several management units; each management unit includes several pixels.

8. A computer device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that the processor executes the computer program to implement the method for identifying a multi-level management unit of soil organic carbon according to any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method for identifying a multi-level management unit of soil organic carbon according to any one of claims 1-6.

10. A computer program product comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the method for identifying a multi-level management unit of soil organic carbon according to any one of claims 1-6.

Citation Information

Patent Citations

  • Remote sensing image mixed image element decomposition method based on self-organizing mapping neural network

    CN101221662A

  • A soil nutrient classification map generation method based on 3S technology and precision evaluation method thereof

    CN108959661A

  • Soil organic carbon content prediction method based on random forest-ordinary Kriging method

    CN109342697A

  • Urban infrastructure group settlement evaluation method, electronic equipment and storage medium

    CN114897454A

  • Remote sensing monitoring method, system and device for land desertification in arid region

    CN118258766A