A soil carbon sequestration assessment model construction method and system
Through multi-dimensional data collection and confidence sorting technology, outliers in soil samples are eliminated and the soil carbon sink assessment model is trained, which solves the problems of low efficiency and low precision in existing technologies and achieves high-precision soil carbon sink assessment.
Patent Information
- Application Number
- CN202510981100.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-07-16
AI Technical Summary
Existing soil carbon sequestration assessment methods are inefficient and inaccurate, and are severely affected by accidental errors in experimental equipment.
Multi-dimensional data collection and confidence sorting technology are used. With land use type, stratified soil type, topography, climate and meteorology, and NDVI vegetation index as constraints, the confidence intervals of stratified soil bulk density, gravel volume fraction, and soil carbon content are obtained. Confidence sorting is performed, outliers are eliminated, and a soil carbon sequestration assessment model is trained.
The accuracy and efficiency of soil carbon sequestration assessment were improved, the impact of accidental errors in experimental equipment was reduced, and the generalization ability and prediction accuracy of the model were enhanced.
Smart Images

Figure CN120473031B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of model construction, and in particular to a soil carbon sequestration assessment model construction method and system. Background Art
[0002] Existing soil carbon sequestration assessments rely on laboratory analysis of soil samples. Actual soil sample values are used as application data, and statistical analysis is performed using a set number of experimental values to obtain the required data. However, existing soil carbon sequestration assessments require numerous repeated experiments to obtain statistically valid data, which is inefficient, time-consuming, and labor-intensive. Furthermore, these assessments fail to fully account for the potential for minor, accidental errors in the measurement process. These errors accumulate and amplify during data processing, resulting in poor accuracy in soil carbon sequestration assessments. Summary of the Invention
[0003] The present invention aims to solve the technical problem of low assessment accuracy of soil carbon sink assessment models in the prior art by providing a soil carbon sink assessment model construction method and system.
[0004] The technical solution of the present invention to solve the above technical problems is as follows:
[0005] In the first aspect, the present invention provides a method for constructing a soil carbon sink assessment model, comprising: taking the land use type data of the target area, several layered soil type data, terrain data, climate and meteorological data and NDVI vegetation index as constraints, collecting credible samples and performing statistics to obtain several layered soil capacity reset confidence intervals, several layered gravel volume fraction confidence intervals and several layered soil carbon content confidence intervals; based on the several layered soil capacity reset confidence intervals, the several layered gravel volume fraction confidence intervals and the several layered soil carbon content confidence intervals, the several layered soil sample data are analyzed. Confidence sorting is performed to obtain soil bulk density detection values of several layers, gravel volume fraction detection values of several layers and soil carbon content detection values of several layers; carbon density statistics are performed based on the soil bulk density detection values of several layers, the gravel volume fraction detection values of several layers and the soil carbon content detection values of several layers to obtain carbon density of several layers and total carbon storage; with the carbon density of several layers and the total carbon storage as supervision and the soil bulk density detection values of several layers, the gravel volume fraction detection values of several layers and the soil carbon content detection values of several layers as input, a soil carbon sink assessment model is trained.
[0006] In the second aspect, the present invention provides a soil carbon sink assessment model construction system, including: a sample collection module, which is used to collect credible samples and perform statistics based on the land use type data of the target area, several layered soil type data, terrain data, climate and meteorological data and NDVI vegetation index, to obtain several layered soil capacity reset confidence intervals, several layered gravel volume fraction confidence intervals and several layered soil carbon content confidence intervals; a data screening module, which is used to screen several layered soil sample data based on the several layered soil capacity reset confidence intervals, the several layered gravel volume fraction confidence intervals and the several layered soil carbon content confidence intervals. Confidence sorting is used to obtain the detected values of soil bulk density of several layers, the detected values of gravel volume fraction of several layers and the detected values of soil carbon content of several layers; a carbon density statistics module is used to perform carbon density statistics based on the detected values of soil bulk density of several layers, the detected values of gravel volume fraction of several layers and the detected values of soil carbon content of several layers to obtain the carbon density of several layers and the total carbon storage; a model training module is used to train the soil carbon sink assessment model with the detected values of soil bulk density of several layers, the detected values of gravel volume fraction of several layers and the detected values of soil carbon content of several layers as supervision and the detected values of soil bulk density of several layers, the detected values of gravel volume fraction of several layers and the detected values of soil carbon content of several layers as input.
[0007] The beneficial effects of the present invention are:
[0008] Constrained by land use data for the target area, soil type data for several layers, topography data, climate and meteorological data, and the NDVI vegetation index, a reliable sample was collected and statistically analyzed to obtain confidence intervals for soil bulk density, gravel volume fraction, and soil carbon content for several layers, providing reliable criteria for subsequent data screening. Based on these confidence intervals, the soil sample data were confidence sorted to obtain soil bulk density, gravel volume fraction, and soil carbon content values for several layers. This effectively screened the soil sample data, eliminating outliers that fell outside the confidence intervals and obtaining highly reliable values. This effectively reduced the impact of minor random errors in experimental equipment on model construction. Carbon density statistics were then performed based on the soil bulk density, gravel volume fraction, and soil carbon content values for several layers, resulting in carbon density and total carbon storage for several layers. This high-confidence data-based carbon density analysis improved the accuracy of carbon storage estimates. A soil carbon sequestration assessment model was trained using carbon density and total carbon storage from several layers as supervision, and measured soil bulk density, gravel volume fraction, and soil carbon content from several layers as input. Because the training data underwent rigorous confidence sorting, the model's generalization and prediction accuracy were improved.
[0009] Through the above technical solution, the data distribution interval of the same modal samples collected by big data is used as the confidence interval to verify the credibility of the soil sample detection data, and the reliable data is quickly selected, thereby ensuring the output accuracy of the soil carbon sequestration assessment model. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 A schematic diagram of a process for constructing a soil carbon sequestration assessment model provided by the present invention;
[0011] Figure 2 This is a structural diagram of a soil carbon sequestration assessment model construction system provided by the present invention.
[0012] In the accompanying drawings, the components represented by the reference numerals are as follows:
[0013] Sample collection module 11, data screening module 12, carbon density statistics module 13, model training module 14. DETAILED DESCRIPTION
[0014] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.
[0015] In the description of the present invention, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature specified as "first" or "second" may explicitly or implicitly include one or more of the specified features. In the description of the present invention, "plurality" means two or more, unless otherwise specifically defined.
[0016] In the description of the present invention, the term "for example" is used to mean "used as an example, illustration or illustration". Any embodiment of the present invention described as "for example" is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is given to enable any person skilled in the art to implement and use the present invention. In the following description, details are listed for the purpose of explanation. It should be understood that a person of ordinary skill in the art can recognize that the present invention can be implemented without using these specific details. In other examples, well-known structures and processes are not elaborated in detail to avoid obscuring the description of the present invention with unnecessary details. Therefore, the present invention is not intended to be limited to the embodiments shown, but is consistent with the widest scope consistent with the principles and features disclosed herein.
[0017] Example 1, as Figure 1 As shown, an embodiment of the present invention provides a method for constructing a soil carbon sequestration assessment model, comprising:
[0018] S1. Based on the land use type data, soil type data of several layers, topographic data, climate and meteorological data, and NDVI vegetation index of the target area, reliable samples were collected and statistics were performed to obtain confidence intervals for soil capacity of several layers, confidence intervals for gravel volume fraction of several layers, and confidence intervals for soil carbon content of several layers.
[0019] Specifically, first, multi-dimensional data is collected for the target area, including but not limited to: land use type data, several layered soil type data, terrain data, climate and meteorological data, and NDVI vegetation index. Among them, the target area refers to a specific geographical scope, such as a specific ecosystem area; land use type data refers to classified data reflecting the land resource utilization mode in the target area, usually including categories such as cultivated land, forest land, grassland, wetland, construction land and unused land, and their area proportion information; several stratified soil type data refer to soil type distribution data divided according to different depths (such as 0-20cm, 20-40cm, 40-60cm, etc.), including soil classification, physical and chemical properties and other information; terrain data refers to a data set that characterizes the surface morphological characteristics of the target area, including parameters such as altitude, slope, slope aspect, and terrain relief; climate and meteorological data refers to a data set that reflects the climate characteristics of the target area, including parameters such as annual average temperature, annual precipitation, seasonal temperature changes, and humidity; NDVI vegetation index refers to the normalized difference vegetation index, a quantitative indicator used to characterize the surface vegetation coverage and growth vitality. Its numerical range is usually -1 to 1, and the larger the value, the higher the vegetation coverage.
[0020] Then, the multi-dimensional data collected above is used as a constraint to collect credible samples that meet the constraint. The selection process of these credible samples needs to fully consider the diversity and representativeness of the regional environment to ensure that the selected samples can accurately reflect the soil carbon sequestration characteristics of the target area. Subsequently, statistical analysis is performed on the collected credible samples. For example, through statistical methods such as data clustering and central tendency assessment, several confidence intervals of soil capacity, several confidence intervals of gravel volume fraction, and several confidence intervals of soil carbon content are obtained. These confidence intervals are different from the single mean or fixed threshold in traditional methods. Instead, they are data distribution intervals obtained based on big data statistics, which can more comprehensively reflect the changing patterns and credible ranges of soil parameters.
[0021] Through the above processing, the foundation is laid for the subsequent confidence sorting of soil sample data, thereby effectively solving the problems of low data credibility and low experimental efficiency in the traditional soil carbon sink assessment process.
[0022] S2. Based on the confidence intervals of the soil bulk density of the several layers, the confidence intervals of the gravel volume fraction of the several layers, and the confidence intervals of the soil carbon content of the several layers, confidence sorting is performed on the soil sample data of the several layers to obtain the detected values of the soil bulk density of the several layers, the detected values of the gravel volume fraction of the several layers, and the detected values of the soil carbon content of the several layers.
[0023] Specifically, based on the obtained confidence intervals of soil capacity in several layers, gravel volume fraction in several layers and soil carbon content in several layers, the soil sample data collected in several layers in the target area were systematically confidence sorted.
[0024] First, for each soil sample collected in the target area, obtain the soil bulk density, gravel volume fraction, and soil carbon content values obtained from laboratory testing. Then, compare these test values with the confidence intervals for soil bulk density, gravel volume fraction, and soil carbon content for the corresponding strata. If the test value of a sample falls within the corresponding confidence interval, it is considered statistically reliable and is used as a valid detection value. If the test value falls outside the confidence interval, it indicates that the test value may have been interfered with by slight accidental errors in the experimental equipment or other factors, and is not sufficiently reliable and needs to be discarded or retested.
[0025] Through the above confidence sorting process, several layered soil bulk density detection values, several layered gravel volume fraction detection values, and several layered soil carbon content detection values were obtained. These detection values not only passed conventional laboratory testing, but also underwent confidence interval verification based on big data statistics, with a double guarantee of high credibility, providing a high-quality data foundation for subsequent carbon density statistics and model training. Compared with the traditional method of directly using laboratory test values, the confidence sorting mechanism effectively identifies and eliminates possible outliers and error values, improves the reliability and accuracy of soil parameter data, and lays a solid foundation for building a high-precision soil carbon sequestration assessment model.
[0026] S3. Perform carbon density statistics based on the soil bulk density detection values of the several layers, the gravel volume fraction detection values of the several layers, and the soil carbon content detection values of the several layers to obtain the carbon density and total carbon storage of the several layers.
[0027] Specifically, based on the high-confidence detected values of soil bulk density, gravel volume fraction and soil carbon content in several layers, carbon density statistics are performed to obtain the carbon density and total carbon storage of several layers in the target area.
[0028] First, for each soil layer in the target area, the soil carbon content detection value, soil bulk density detection value and gravel volume fraction detection value of the layer are substituted into the calculation according to the formula: soil carbon density = carbon content × soil bulk density × soil layer thickness × (1-gravel content). Among them, the soil carbon content is expressed as a percentage, the unit of soil bulk density is g / cm³, the unit of soil layer thickness is cm, and the gravel content is expressed as volume fraction. The calculated soil carbon density unit is kg / m². During the calculation process, since there are multiple sampling points in the target area, each layer may have multiple detection values. Therefore, statistical methods such as weighted average or spatial interpolation are used to comprehensively consider the spatial distribution and representativeness of each sampling point, and the average carbon density value of each layer is calculated to obtain the carbon density of several layers. Compared with the simple arithmetic average, this calculation method can more accurately reflect the spatial variation characteristics of carbon density.
[0029] Next, the carbon storage of each layer is calculated based on the carbon density value and corresponding area data of each layer using the formula: Layer carbon storage = layer carbon density × layer area. The carbon storage of each layer is then accumulated to obtain the total carbon storage of the target area, namely: Total carbon storage = ∑(layer carbon density × layer area).
[0030] Through the above carbon density statistics, not only the carbon density distribution of each soil layer in the target area was obtained, but also the total carbon storage of the entire target area was calculated. These data are both the direct results of soil carbon sink assessment and important supervisory data for the subsequent construction of soil carbon sink assessment models. Compared with traditional carbon density calculation methods, calculations based on verified high-confidence detection values effectively avoid calculation errors caused by data anomalies. At the same time, by considering soil stratification characteristics and spatial heterogeneity, more accurate carbon storage assessments are achieved, providing a reliable data basis for the construction of soil carbon sink assessment models.
[0031] S4. Using the carbon density of the plurality of layers and the total carbon storage as supervision, and using the soil bulk density detection values of the plurality of layers, the gravel volume fraction detection values of the plurality of layers, and the soil carbon content detection values of the plurality of layers as input, a soil carbon sink assessment model is trained.
[0032] Specifically, based on the obtained carbon density and total carbon storage of several layers, a soil carbon sink assessment model is constructed and trained to achieve intelligent mapping from basic soil characteristic parameters to carbon sink assessment results, thereby converting the high-quality data obtained in the previous steps into a model with predictive capabilities.
[0033] First, a multi-layer model training architecture was established, using the obtained soil bulk density, gravel volume fraction, and soil carbon content values for several layers as model input features. The calculated carbon density and total carbon storage for several layers served as supervisory signals, i.e., the model's expected output targets. Secondly, a base model for soil carbon sequestration assessment was constructed based on machine learning algorithms such as neural networks and graph convolutional networks. This base model not only considers the numerical characteristics of soil parameters but also incorporates spatial topological relationships through techniques such as graph neural networks, effectively capturing the interactions and correlations between different soil layers. The model structure was designed using a hybrid architecture combining a branch network with a backbone network. Dedicated prediction branches were established for soil bulk density, gravel volume fraction, and soil carbon content, respectively. Feature fusion and comprehensive calculations were then performed within the backbone network. During model training, a combination of layer-by-layer supervision and global supervision was used to simultaneously optimize the model's prediction accuracy for carbon density and total carbon storage in each layer. Specifically, a multi-task learning framework is employed, with a local loss function for stratified carbon density and a global loss function for total carbon storage. These loss functions are weighted and combined to guide the optimization of model parameters. Furthermore, techniques such as regularization and early stopping can be introduced to prevent overfitting and improve generalization.
[0034] Through the above training process, a soil carbon sink assessment model that can accurately assess soil carbon sinks is obtained. The model can calculate carbon density and carbon storage based on existing soil sample data. Based on basic environmental characteristics such as land use type data, soil type data, terrain data, climate and meteorological data, and NDVI vegetation index in the target area, it can directly predict the stratified carbon density and total carbon storage of the region, without the need for tedious soil sample collection and laboratory analysis, thereby improving the efficiency and applicability of soil carbon sink assessment.
[0035] Furthermore, based on the target area's land use type data, several layers of soil type data, topographic data, climate and meteorological data, and NDVI vegetation index as constraints, reliable samples were collected and statistics were performed to obtain several layers of soil capacity confidence intervals, several layers of gravel volume fraction confidence intervals, and several layers of soil carbon content confidence intervals, including:
[0036] S11, extracting first layer soil type data from the plurality of layer soil type data;
[0037] S12. When the candidate credible sample stored on the blockchain meets the requirements of the land use type data, the first layer soil type data, the topographic data, the climate and meteorological data, and the NDVI vegetation index, the candidate credible sample is added to the credible sample;
[0038] S13, extracting the soil bulk density record value set of the credible sample, performing central tendency assessment, obtaining the first layer soil bulk density reset confidence interval, and adding it to the plurality of layer soil bulk density reset confidence intervals;
[0039] S14, extracting a set of recorded gravel volume fraction values of the credible sample, performing a central tendency assessment, obtaining a confidence interval of the gravel volume fraction of the first layer, and adding the confidence intervals of the gravel volume fractions of the plurality of layers;
[0040] S15. Extract the set of soil carbon content record values of the credible sample, perform central tendency assessment, obtain a confidence interval of soil carbon content in the first layer, and add the confidence interval of soil carbon content in the plurality of layers.
[0041] In one feasible implementation, soil type data for several layers of the target area is first analyzed and processed sequentially. Each time, soil type data for a specific depth layer is extracted as the first layer of soil type data. The multiple layers of soil type data include soil type distribution at different depths (e.g., 0-20 cm, 20-40 cm, 40-60 cm, etc.). Soil type data for a specific depth layer (e.g., the surface layer 0-20 cm) is selected as the first layer of soil type data. This first layer of soil type data details the spatial distribution of soil types in the first layer of the target area, including parameters such as soil type, distribution range, and area percentage. The purpose of extracting this first layer of soil type data is to provide precise constraints for subsequent sample screening, ensuring that the selected samples are a good match with the target area in terms of soil type. Blockchain evidence storage technology is then introduced to ensure the credibility and traceability of the selected trusted samples. Specifically, samples that meet the multi-dimensional constraints are selected from the blockchain-verified trusted samples and added to the trusted sample set. The screening process is based on five constraints: land use data, first-layer soil type data, topography data, climate and meteorological data, and the NDVI vegetation index. A sample is considered qualified and added to the credible sample pool only if its environmental characteristics closely match the five constraints for the target area. This rigorous screening mechanism, based on multidimensional environmental characteristics, ensures that the selected samples possess highly consistent environmental backgrounds and soil formation conditions with the target area, thereby ensuring the applicability and accuracy of subsequent statistical analysis results.
[0042] Subsequently, the first-layer soil bulk density records of all samples were extracted from the credible samples obtained through screening to form a set of soil bulk density records. Soil bulk density is a key parameter for calculating soil carbon density and affects the accuracy of carbon storage estimation. A central tendency assessment was performed on the extracted set of soil bulk density records, including but not limited to data distribution test, outlier identification, cluster analysis, and interval estimation. Through these statistical processing, a confidence interval that can accurately reflect the variation characteristics of the soil bulk density of the first layer is obtained, namely, the first-layer soil bulk density reset confidence interval. This confidence interval is usually expressed in the form of [μ-kσ, μ+kσ], where μ is the mean value of the soil bulk density, σ is the standard deviation, and k is an adjustable confidence coefficient. Afterwards, the obtained first-layer soil bulk density reset confidence interval is added to the set of soil bulk density reset confidence intervals of several layers to provide data support for subsequent stratified analysis.
[0043] At the same time, the first-layer gravel volume fraction record values of all samples are extracted from the screened reliable samples to form a set of gravel volume fraction record values. The gravel volume fraction refers to the volume ratio of gravel in the soil and is a key correction factor in the calculation of soil carbon density, because the gravel part usually does not contain organic carbon or has a very low content. The extracted gravel volume fraction record value set is statistically analyzed using a central tendency assessment method similar to step S13. Taking into account the spatial heterogeneity and local concentration characteristics of gravel distribution, spatial autocorrelation analysis and distribution pattern recognition technology can be introduced in the evaluation process to more accurately characterize the change law of gravel volume fraction. Through these processes, the confidence interval of the gravel volume fraction of the first layer is obtained and added to the set of confidence intervals of gravel volume fraction of several layers.
[0044] In addition, the first-layer soil carbon content record values of all samples are extracted from the credible samples obtained by screening to form a set of soil carbon content record values. Soil carbon content is the core parameter for soil carbon sink assessment, which reflects the soil's carbon sequestration capacity and carbon sink potential. In view of the particularity of soil carbon content, when performing the central tendency assessment, not only the numerical distribution characteristics of the carbon content are considered, but also classified statistics are performed in combination with factors such as land use type and vegetation coverage to more finely depict the changing pattern of soil carbon content under different environmental conditions. Preferably, time series analysis technology can also be introduced to consider the seasonal and interannual variation characteristics of soil carbon content, so as to obtain a more representative and adaptive confidence interval. Through the above processing, the confidence interval of the soil carbon content of the first layer is obtained, and it is added to the set of confidence intervals of the soil carbon content of several layers.
[0045] Through the above steps, a systematic processing flow from several stratified soil type data to specific stratification confidence intervals was realized. Through precise environmental constraint screening and statistical analysis, confidence intervals of three key parameters, soil bulk density, gravel volume fraction and soil carbon content, were constructed for each soil stratification, laying a data foundation for subsequent soil sample confidence sorting and carbon density calculation.
[0046] Furthermore, when the candidate credible sample stored on the blockchain satisfies the land use type data, the first layer soil type data, the terrain data, the climate and meteorological data, and the NDVI vegetation index, including:
[0047] S121, extracting sample land use type data, sample soil type data, sample terrain data, sample climate and meteorological data, and sample NDVI vegetation index of the candidate credible sample, wherein the sample layer depth is the same as the first layer depth;
[0048] S122, calculating a first deviation coefficient between the sample land use type data and the land use type data;
[0049] S123, calculating a second deviation coefficient between the sample soil type data and the first layer soil type data;
[0050] S124, calculating a third deviation coefficient between the sample terrain data and the terrain data;
[0051] S125, calculating a fourth deviation coefficient between the sample climate and meteorological data and the climate and meteorological data;
[0052] S126. Calculating a fifth deviation coefficient between the sample NDVI vegetation index and the sample NDVI vegetation index;
[0053] S127. When the first deviation coefficient, the second deviation coefficient, the third deviation coefficient, the fourth deviation coefficient, and the fifth deviation coefficient are all less than or equal to corresponding deviation coefficient thresholds, the candidate credible sample is deemed to meet the requirements of the land use type data, the first layered soil type data, the terrain data, the climate and meteorological data, and the NDVI vegetation index.
[0054] S128. Otherwise, it is deemed unsatisfied.
[0055] In a preferred embodiment, first, feature extraction is performed on the candidate credible samples that have been stored on the blockchain to obtain their environmental feature parameter set. Specifically, key parameters related to environmental constraints are extracted from the metadata records of the candidate credible samples, including sample land use type data, sample soil type data, sample terrain data, sample climate and meteorological data, and sample NDVI vegetation index. These parameters comprehensively describe the environmental background and ecological conditions of the candidate credible samples. In particular, to ensure the consistency of the comparison, ensure that the stratification depth of the extracted sample soil type data is completely consistent with the first stratification depth. For example, if the first stratification depth is 0-20cm, the extracted sample soil type data should also correspond to the 0-20cm depth layer. This depth matching mechanism is an important prerequisite for ensuring the effectiveness of subsequent parameter comparisons.
[0056] Then, a quantitative comparison is performed on the land use type data to calculate the similarity between the candidate credible sample and the target area in terms of land use type data. Specifically, land use types include multiple categories such as cultivated land, forest land, grassland, wetland, and construction land. The area proportion of each land type constitutes the characteristic vector of the land use type data. By calculating the difference in the area proportion of each type of land use between the candidate credible sample and the target area, a first deviation coefficient is obtained. This first deviation coefficient can be calculated using mathematical methods such as weighted Euclidean distance or similarity index. The smaller the value, the more similar the candidate credible sample and the target area are in terms of land use type. In calculating the first deviation coefficient, the differences in the impact of different land use types on soil carbon sequestration can be taken into account, and land types with higher ecological sensitivity (such as wetlands and forest land) can be given higher weights, thereby enhancing the ecological significance of the comparison.
[0057] Simultaneously, a quantitative comparison of soil type data is performed to calculate the similarity of soil type data in the first layer between the candidate credible sample and the target area. Soil type is a key factor in soil carbon sequestration assessment, as different soil types have varying physical and chemical properties and carbon storage capacity. A second coefficient of variation is calculated by comparing the proportions of each soil type between the candidate credible sample and the target area. Given the complexity and regional specificity of soil type classification, a classification comparison method based on soil genetic characteristics can be employed. This considers not only the matching of soil type names but also similarities in soil evolution and the degree of similarity in basic soil properties, leading to a more reasonable assessment of soil type similarity. Furthermore, a quantitative comparison of topographic data is performed to calculate the similarity of topographic characteristics between the candidate credible sample and the target area. Topographic conditions are important environmental factors influencing soil development and carbon cycling, including parameters such as altitude, slope, aspect, and terrain relief. For example, the raw topographic data is first normalized to eliminate differences in the dimensions and numerical ranges of different topographic parameters. Then, the Euclidean distance or cosine similarity is calculated based on the normalized parameter vectors to obtain the third coefficient of variation. The third deviation coefficient comprehensively reflects the degree of match between the selected reliable samples and the target area in terms of terrain characteristics, and takes into account the complex influence mechanism of terrain on soil water and heat conditions and organic matter accumulation.
[0058] At the same time, a quantitative comparison is conducted on climate and meteorological data to calculate the similarity in climatic characteristics between the candidate credible sample and the target area. Climatic conditions are important factors affecting soil carbon input and decomposition rate, including multiple meteorological elements such as temperature, precipitation, humidity, and sunshine. For example, a method combining climate zone matching and weighted comparison of climate elements is adopted, which takes into account both the overall matching of climate types and the quantitative comparison of key meteorological elements. Specifically, it is first determined whether the candidate credible sample and the target area belong to the same climate zone. Then, a weighted distance calculation is performed on key meteorological elements such as annual average temperature, annual precipitation, and dryness and humidity index to comprehensively obtain the fourth deviation coefficient. This fourth deviation coefficient fully considers the different weights of the influence of climate factors on soil carbon dynamics and can accurately reflect the ecological similarity of climate conditions. At the same time, a quantitative comparison is conducted on the NDVI vegetation index to calculate the similarity in vegetation cover between the candidate credible sample and the target area. The NDVI vegetation index is an important remote sensing indicator that represents the coverage and growth status of surface vegetation and is closely related to soil carbon input and surface protection. For example, by comparing the average NDVI values, taking into account the seasonal variation characteristics and spatial distribution patterns of NDVI, and calculating the dynamic similarity of the NDVI time series and the matching degree of spatial distribution characteristics, the fifth deviation coefficient is obtained. This fifth deviation coefficient can comprehensively reflect the similarity between the sample and the target area in terms of vegetation cover and ecosystem function.
[0059] Subsequently, by comparing the relationship between each deviation coefficient and the preset threshold, it is determined whether the selected credible sample meets the constraint conditions. Specifically, corresponding deviation coefficient thresholds are set for the first deviation coefficient, the second deviation coefficient, the third deviation coefficient, the fourth deviation coefficient and the fifth deviation coefficient. These thresholds are determined based on expert knowledge and statistical analysis and can define the similarity boundaries of environmental characteristics. Only when the five deviation coefficients are less than or equal to their respective corresponding deviation coefficient thresholds, the selected credible sample is determined to meet all the constraints, ensuring that the selected credible sample has a high degree of consistency with the target area in terms of multi-dimensional environmental characteristics, providing a reliable data basis for subsequent statistical analysis. When any deviation coefficient exceeds the corresponding deviation coefficient threshold, the selected credible sample is determined to not meet the constraint conditions and is not included in the credible sample, thereby improving the quality and representativeness of the selected credible sample and effectively avoiding errors caused by differences in environmental conditions.
[0060] Through the above steps, a precise sample screening mechanism based on multidimensional environmental characteristics was realized, and quantitative similarity assessments were conducted on factors such as land use type, soil type, topography, climate and vegetation, and multidimensional constraint judgments were made to ensure that the selected credible samples were highly matched with the target area in terms of environmental background, thereby improving the reliability and applicability of subsequent statistical analysis results and laying a data foundation for building a high-precision soil carbon sink assessment model.
[0061] Furthermore, the embodiment of the present application also includes:
[0062] The first deviation coefficient is equal to a normalized parameter of the mean difference in area between the sample land use type data and the same land use type data;
[0063] The second deviation coefficient is equal to a normalized parameter of the mean deviation of the same soil type content ratio between the sample soil type data and the first layered soil type data;
[0064] The third deviation coefficient is equal to the Euclidean distance between the sample terrain data and the normalized value of the characteristic parameter of the terrain data;
[0065] The fourth deviation coefficient is equal to the Euclidean distance between the sample climate and meteorological data and the normalized value of the characteristic parameter of the climate and meteorological data;
[0066] The fifth deviation coefficient is equal to a normalized parameter of the deviation between the sample NDVI vegetation index and the sample NDVI vegetation index.
[0067] In a preferred embodiment, specific calculation methods of the five deviation coefficients are defined in detail to achieve accurate quantitative evaluation of the similarity of environmental features.
[0068] The first deviation coefficient quantifies the degree of difference in land use type distribution between the candidate credible sample and the target area. To calculate this coefficient, land use types are first categorized into multiple categories (such as cultivated land, forest land, grassland, wetland, construction land, and unused land). The difference in the area percentage of each land use type between the candidate credible sample and the target area is then calculated. For example, if the target area has a cultivated land percentage of 40% and the candidate credible sample has a cultivated land percentage of 35%, the difference in cultivated land type area percentage between the two is 5%. This difference in area percentage for all land use types is then calculated and averaged to obtain the mean area difference for each land use type. To eliminate the influence of differences in the number and distribution of land use types across regions, this mean is normalized to obtain the first deviation coefficient. Normalization can use either the Min-Max or Z-score normalization methods to map the raw difference values to a standard interval, facilitating comparison with a preset threshold.
[0069] The second deviation coefficient is used to quantify the degree of difference in the distribution of soil types in the first layer between the candidate credible sample and the target area. The calculation method is similar to the first deviation coefficient. First, all soil types (such as black soil, brown soil, and red soil) present in the candidate credible sample and the target area are identified. Then, the difference in the content percentage of each soil type between the sample area and the target area is calculated. For example, if the black soil content in the first layer of the target area is 60%, while the black soil content at the same depth in the candidate credible sample area is 55%, the difference in the content percentage of the black soil type between the two is 5%. The content percentage deviations of all soil types are calculated in this way, and the average is taken to obtain the mean content percentage deviation of the same soil type. Preferably, to take into account the diversity and complexity of soil types in different regions, a soil type similarity matrix can be introduced into the calculation process to weight the differences between different soil types, thereby more accurately reflecting the overall similarity of soil characteristics. The calculated mean deviation is then normalized to obtain the second deviation coefficient.
[0070] The third deviation coefficient is used to quantify the degree of difference in terrain characteristics between the candidate credible sample and the target area. Topographic data usually contains multiple characteristic parameters, such as average altitude, maximum altitude, minimum altitude, average slope, slope distribution, terrain undulation, etc. Since these characteristic parameters have different dimensions and numerical ranges, in order to achieve comparison, each characteristic parameter is first normalized to eliminate the influence of dimension and numerical range. After normalization, the terrain characteristics of the candidate credible sample and the target area can be represented as vectors in a multidimensional feature space, where the number of dimensions is the number of terrain characteristic parameters. The third deviation coefficient is the Euclidean distance between the two normalized feature vectors in the multidimensional feature space. The Euclidean distance calculation takes into account the square root of the sum of the squares of the differences in each characteristic parameter. The smaller the distance, the more similar the two areas are in terrain characteristics.
[0071] The fourth deviation coefficient is used to quantify the degree of difference in climatic and meteorological characteristics between the candidate credible samples and the target area. Climatic and meteorological data usually contain multiple characteristic parameters, such as average annual temperature, annual precipitation, seasonal temperature changes, relative humidity, sunshine hours, wind speed, etc. Similar to topographic characteristics, these climatic and meteorological parameters also have different dimensions and numerical ranges and need to be normalized. Using the weight distribution principle of climate ecology, weight coefficients are set according to the degree of influence of different climatic factors on the soil carbon cycle, and the normalized characteristic parameters are weighted. For example, in temperate regions, temperature and precipitation have a greater impact on soil carbon dynamics and can be given a higher weight; in tropical regions, humidity and evaporation may be more important. After weight adjustment, the fourth deviation coefficient is calculated as a weighted Euclidean distance, which comprehensively reflects the overall similarity of climatic conditions.
[0072] The fifth deviation coefficient is used to quantify the degree of difference in vegetation coverage between the candidate credible sample and the target area. The NDVI vegetation index is an important remote sensing indicator that characterizes the surface vegetation coverage and growth vitality. Its numerical range is usually from -1 to 1. The larger the value, the higher the vegetation coverage. Taking into account that NDVI has obvious seasonal variation characteristics, it is possible to compare not only the annual average NDVI value, but also the seasonal variation pattern and spatial distribution characteristics of NDVI. Specifically, multi-phase NDVI data is used to calculate the deviation, and the NDVI time series feature vector is constructed. Then, similarity indicators such as the dynamic time warping distance or Pearson correlation coefficient of the NDVI time series between the candidate credible sample and the target area are calculated. Preferably, the spatial distribution characteristics of NDVI can also be considered, such as the spatial uniformity, aggregation and differentiation of NDVI. Taking the time series similarity and spatial distribution similarity into consideration, the original value of the NDVI deviation is calculated, and then it is mapped to the standard interval through normalization to obtain the fifth deviation coefficient.
[0073] By calculating these five deviation coefficients, we achieve a quantitative similarity assessment of multidimensional environmental factors, including land use type, soil type, topography, climate characteristics, and vegetation cover, accurately reflecting the overall similarity of environmental characteristics across regions. Based on these deviation coefficients, combined with preset deviation thresholds, we can efficiently screen for reliable samples that closely match the environmental characteristics of the target region, providing reliable data support for subsequent soil carbon sequestration assessments.
[0074] Furthermore, extracting a set of soil bulk density record values of the credible sample, performing a central tendency assessment, and obtaining a first-layer soil bulk density confidence interval includes:
[0075] S131, performing cluster analysis on the set of soil bulk density record values based on a soil bulk density deviation threshold to obtain multiple clusters of soil bulk density record values;
[0076] S132, deleting clusters whose number of recorded values is less than or equal to the cluster number threshold, to obtain multiple clusters of updated soil bulk density recorded values;
[0077] S133, traversing the multiple clusters to update the soil bulk density record values, performing a central tendency assessment, and obtaining multiple soil bulk density concentration value distribution intervals;
[0078] S134. Take the union of the multiple soil bulk density concentration value distribution intervals to obtain the first layer soil bulk density confidence interval.
[0079] In a preferred embodiment, first, the soil bulk density record value set extracted from the credible sample is preprocessed and data cleaned to eliminate obvious outliers and erroneous records. Then, cluster analysis is performed on the processed record value set based on the soil bulk density deviation threshold. Specifically, a clustering algorithm such as an improved K-means clustering or DBSCAN density clustering is used, and the soil bulk density deviation threshold is used as a clustering parameter to automatically group the data points in the soil bulk density record value set. The soil bulk density deviation threshold is a parameter determined based on professional knowledge and statistical analysis in the field of soil science, and is used to define the boundaries between different soil bulk density clusters. Through cluster analysis, similar soil bulk density record values are classified into the same cluster, and record values with large differences are divided into different clusters. This data-driven clustering method can effectively identify the multimodal distribution characteristics of soil bulk density record values and capture the distribution law of bulk density values under different soil conditions. Through cluster analysis, the original soil bulk density record value set is divided into multiple clusters with internal consistency, obtaining multiple clusters of soil bulk density record values, and the record values within each cluster have similar numerical characteristics, representing the bulk density distribution characteristics under specific soil conditions.
[0080] Subsequently, the multiple clusters of soil bulk density records were screened and optimized, eliminating small clusters with insufficient sample size to improve the reliability of the statistical results. Specifically, a cluster-wide threshold was set as the criterion for quantification. The number of records within each cluster was counted. If the number of records within a cluster was less than or equal to the preset cluster-wide threshold, the cluster was deleted and excluded from subsequent analysis. The cluster-wide threshold is determined based on statistical sample size requirements and empirical knowledge in soil science, typically as a percentage of the total sample size (e.g., 5% or 10%) or as a minimum statistically significant sample size. This cluster-wide sample size-based screening mechanism effectively eliminates small data clusters that may be caused by chance, enhancing the robustness and reliability of the data analysis. By removing small clusters, multiple clusters of updated soil bulk density records were obtained. These updated data clusters have sufficient sample support and statistical representativeness, and more accurately reflect the distribution patterns of bulk density values under different soil conditions.
[0081] Next, a central tendency assessment was performed on each of the multiple clusters of updated soil bulk density data to quantify the distribution characteristics of the data within each cluster and determine the central tendency interval for the soil bulk density data. For example, for each soil bulk density data cluster, the soil bulk density central tendency interval was determined using the mean ± 3 times the standard deviation. Specifically, the mean (μ) and standard deviation (σ) of the soil bulk density data within the cluster were first calculated. Then, the central tendency interval for the soil bulk density data for that cluster was defined as [μ - 3σ, μ + 3σ]. If the data distribution of a cluster deviated from a normal distribution, the central tendency interval was determined using the median ± 2.5 times the interquartile range method, i.e., [median - 2.5 × interquartile range, median + 2.5 × interquartile range] for that cluster. In this way, a central tendency interval for the soil bulk density data was determined for each data cluster that covered the main distribution characteristics of the cluster.
[0082] Subsequently, a union operation was performed on the multiple concentrated value distribution intervals obtained to form an overall confidence interval for the soil bulk density of the first layer. Specifically, the union operation took the minimum lower bound and maximum upper bound of all concentrated value distribution intervals to form a continuous interval containing all valid soil bulk density values. To avoid the problem of excessively wide intervals that may be caused by the union operation, a weighted assignment was first performed on each concentrated value distribution interval before the union operation was performed. Different distribution intervals were assigned different weights based on the sample size, data quality, and representativeness within the cluster. Then, a weighted union operation was performed. In addition, for multiple distribution intervals with significant gaps, a segmented definition approach was adopted to represent the confidence interval for the soil bulk density of the first layer as a combination of multiple subintervals, thereby more accurately characterizing the multimodal distribution characteristics of soil bulk density. Through the above processing, a confidence interval for the soil bulk density of the first layer was obtained. This interval comprehensively considers the multi-cluster distribution characteristics of the soil bulk density data and the statistical confidence level, accurately defining the range of valid soil bulk density values and providing a scientific basis for subsequent confidence sorting of soil samples.
[0083] Through the above steps, the original set of recorded soil bulk density values was transformed into confidence intervals. Cluster analysis was leveraged to identify the multimodal distribution of the data. Combined with strict sample size control and multidimensional statistical analysis, confidence intervals for soil bulk density were constructed. Compared to traditional methods that simply use the mean and standard deviation of the entire sample to determine the interval, this interval construction method more accurately captures the complex distribution characteristics of soil bulk density, effectively addresses the heterogeneity and diversity of soil samples, and provides a more reliable data foundation for subsequent soil carbon sequestration assessments.
[0084] Furthermore, the soil carbon sink assessment model is trained using the carbon density of the plurality of layers and the total carbon storage as supervision, and the detected values of soil bulk density of the plurality of layers, the detected values of gravel volume fraction of the plurality of layers, and the detected values of soil carbon content of the plurality of layers as input, including:
[0085] S41, using the soil bulk density detection values of the multiple layers as supervision, and using the land use type data, the soil type data of the multiple layers, the terrain data, the climate and meteorological data, and the NDVI vegetation index as input, to construct a first branch network;
[0086] S42, using the gravel volume fraction detection values of the plurality of layers as supervision, and using the land use type data, the soil type data of the plurality of layers, the topographic data, the climate and meteorological data, and the NDVI vegetation index as input, to construct a second branch network;
[0087] S43, using the soil carbon content detection values of the plurality of layers as supervision, and using the land use type data, the soil type data of the plurality of layers, the topographic data, the climate and meteorological data, and the NDVI vegetation index as input, to construct a third branch network;
[0088] S44, constructing a hierarchical three-dimensional model of the target area based on the land use type data, the plurality of layered soil type data, and the terrain data, performing topological simulation, and constructing a target area graph neural network;
[0089] S45, using the carbon density of the plurality of layers as the output supervision value of each layer of the target area graph neural network, using the total carbon storage global output supervision value, using the soil bulk density detection values of the plurality of layers, the gravel volume fraction detection values of the plurality of layers, the soil carbon content detection values of the plurality of layers, and the area of the plurality of layers as the input value of each layer of the target area graph neural network, to train the soil carbon sink assessment base model;
[0090] S46. Embed the first branch network, the second branch network and the third branch network into the input nodes corresponding to the soil bulk density detection values, the gravel volume fraction detection values and the soil carbon content detection values of each layer of the target area graph neural network to obtain the soil carbon sink assessment model.
[0091] In a preferred embodiment, a dedicated first-branch network is first constructed for soil bulk density. This first-branch network takes as input the environmental characteristic data of the target area, including land use type data, soil type data of several layers, topographic data, climate and meteorological data, and the NDVI vegetation index. It uses confidence-sorted soil bulk density values of several layers as supervisory signals. The core function of the first-branch network is to establish a mapping relationship between environmental characteristics and soil bulk density, enabling accurate prediction of soil bulk density from environmental data. In terms of network structure design, the first-branch network adopts a hybrid architecture combining a multi-layer perceptron (MLP) and an attention mechanism. The input layer corresponds to five categories of environmental characteristic data, which enter the network after feature encoding and normalization. The hidden layer adopts a multi-layer design, combined with residual connections and batch normalization techniques, to enhance the network's expressive power and training stability. The output layer corresponds to the predicted soil bulk density values of several layers. A loss function is calculated with the supervisory signal (soil bulk density values) to guide network training. During training, cross-validation and early stopping strategies are used to avoid overfitting and improve model generalization. Through the training of the first branch network, the function of directly predicting soil bulk density based on environmental characteristics was realized, providing important support for subsequent comprehensive models.
[0092] A dedicated second-branch network was constructed for gravel volume fraction. This second-branch network has a similar input structure to the first-branch network, also based on five categories of environmental feature data, but uses confidence-sorted gravel volume fraction detection values for several layers as supervisory signals. The core function of the second-branch network is to establish a mapping relationship between environmental features and gravel volume fraction, enabling accurate prediction of gravel content from environmental data. Considering that the distribution characteristics of gravel volume fraction differ from those of soil bulk density, the second-branch network utilizes a convolutional neural network (CNN) architecture, which is more suitable for capturing local features, combined with fully connected layers. For input features with strong spatial correlation (such as topographic data and NDVI data), convolutional layers are used to extract local patterns and spatial associations. For categorical features (such as land use type and soil type), embedding layers are used for feature transformation. After processing these multiple sources of features through their respective modules, they are integrated through a feature fusion layer, ultimately outputting several layers of predicted gravel volume fraction. Through the training of the second-branch network, the ability to directly predict gravel volume fraction based on environmental features is achieved, providing important support for subsequent integrated models.
[0093] In addition, a dedicated third-branch network was constructed for soil carbon content. This third-branch network shares a similar input structure to the first two branches, also based on five categories of environmental characteristic data. However, it utilizes confidence-sorted soil carbon content detection values from several layers as supervisory signals. The core function of the third-branch network is to establish a mapping relationship between environmental characteristics and soil carbon content, enabling accurate prediction of carbon content from environmental data. As a core parameter in carbon sink assessment, soil carbon content is complexly influenced by multiple factors. Therefore, the third-branch network utilizes a more complex deep learning architecture, integrating a long short-term memory (LSTM) network and an attention mechanism. This design not only captures the complex interactions between environmental characteristics but also simulates the cumulative effects of soil carbon accumulation. Specifically, LSTM units are used to extract dynamic vegetation changes from time series data of the NDVI vegetation index. For heterogeneous environmental data from multiple sources, a multi-head self-attention mechanism is employed to dynamically assign weights between features. Through training of the third-branch network, direct prediction of soil carbon content based on environmental characteristics is achieved, providing important support for subsequent integrated models.
[0094] Subsequently, a hierarchical 3D model was constructed based on the spatial data of the target area. Based on this model, a graph neural network suitable for soil carbon sequestration assessment was constructed. First, a hierarchical 3D model of the target area was constructed using land use type data, several layers of soil type data, and terrain data. This model uses a voxel representation method to divide the target area into a 3D grid. Each grid cell contains attribute information such as land use type, soil type, and terrain characteristics. Spatial correlations and material flow exist between adjacent grid cells, forming a complex spatial topological structure. Based on this hierarchical 3D model, a topological simulation was performed to establish the connectivity and influencing mechanisms between grid cells. Specifically, horizontal material migration and vertical infiltration accumulation were considered, and planar adjacency relationships and vertical interlayer relationships were constructed, respectively. The topological simulation process incorporated principles from hydrology, geomorphology, and soil science, such as considering the effects of slope and aspect on material flow and the characteristics of material exchange between different soil layers. Based on this topological structure, a graph neural network for the target area was constructed. Each grid cell was represented as a node in the graph, and the connectivity between nodes was determined based on the topological simulation results. The core of graph neural networks is the message-passing mechanism. By defining node feature update functions and message aggregation functions, this allows for the transfer and integration of information between nodes, thereby capturing complex spatial dependencies. By employing an architecture that combines a graph convolutional network (GCN) with a graph attention network (GAT), it not only considers the local connectivity of nodes but also introduces an attention mechanism for dynamic weight allocation. In this way, the constructed graph neural network can effectively simulate the spatial heterogeneity and inter-layer interactions of the soil carbon cycle.
[0095] Next, a soil carbon sink assessment model was trained based on the constructed graph neural network for the target region. This model employed a training strategy combining hierarchical and global supervision to accurately map basic soil parameters to carbon density and storage. First, confidence-sorted soil bulk density values, gravel volume fraction values, and soil carbon content values for several layers were used as feature inputs for each node of the graph neural network. Combined with the area data for each layer, these served as the complete network input. A two-level supervisory signal was then implemented: on the one hand, the carbon density of each layer served as the output supervisory value for each node in the graph neural network, enabling accurate predictions at the layer level; on the other hand, total carbon storage served as the global output supervisory value to ensure overall prediction accuracy. During model training, a multi-task learning framework was employed to design a comprehensive loss function, simultaneously accounting for the loss of each layer's carbon density prediction and the loss of total carbon storage prediction. Weight coefficients were used to balance the importance of these two supervisory signals. Through end-to-end training, the model's predictive capabilities for both layer-level carbon density and total carbon storage were simultaneously optimized, enabling multi-scale carbon sink assessment.
[0096] Through the above training process, the resulting soil carbon sequestration assessment model can calculate carbon density and carbon storage based on known soil parameter values, but it still relies on input data provided by soil sample collection and laboratory analysis. To further improve the model's automation and applicability, model integration is required. Specifically, the trained first, second, and third branch networks are embedded into the trained target region graph neural network. During the embedding process, the outputs of the three branch networks are connected to the soil bulk density, gravel volume fraction, and soil carbon content input locations at each layer node of the graph neural network, forming a complete deep learning model. On the one hand, the three branch networks can directly predict basic soil parameters from environmental characteristic data; on the other hand, the graph neural network can further calculate carbon density and carbon storage from these predicted soil parameters. Through end-to-end joint optimization, mutual complementation and verification, a highly integrated soil carbon sequestration assessment model is formed. This soil carbon sink assessment model does not require tedious soil sample collection and laboratory analysis. It only needs to input basic environmental data of the target area (land use type, soil type, topography, climate and vegetation index, etc.) to directly predict the stratified carbon density and total carbon storage of the region, thereby improving the automation level and application efficiency of soil carbon sink assessment.
[0097] Furthermore, the embodiment of the present application also includes:
[0098] The land use type data, the soil type data of the several layers, the terrain data, the climate and meteorological data, and the NDVI vegetation index are set as index condition data, and the soil bulk density detection values of the several layers, the gravel volume fraction detection values of the several layers and the soil carbon content detection values of the several layers, the carbon density of the several layers and the total carbon storage are set as index result data to construct a carbon sink assessment database for the target area.
[0099] In one feasible implementation, a dedicated target region carbon sequestration database was constructed to systematically manage and efficiently utilize the multi-source, heterogeneous data generated during soil carbon sequestration assessment. This database utilizes an organizational structure that associates index condition data with index result data, enabling integrated storage and rapid retrieval of environmental characteristics and carbon sequestration assessment results.
[0100] Specifically, five types of environmental characteristic data, including land use type data, soil type data of several layers, terrain data, climate and meteorological data, and NDVI vegetation index, are set as index condition data. These index condition data serve as the primary key and retrieval entry of the database, and each set of index condition data uniquely corresponds to a set of carbon sink assessment results. The index condition data not only contains numerical features, but also contains spatial location information and timestamps, supporting multi-dimensional retrieval and screening operations. At the same time, carbon sink assessment-related data such as soil bulk density detection values of several layers, gravel volume fraction detection values of several layers, soil carbon content detection values of several layers, carbon density and total carbon storage of several layers are set as index result data. These index result data, as the value carrier of the database, record soil parameters and carbon sink assessment results under different environmental conditions, and support multi-level and multi-scale data analysis and application.
[0101] In terms of database structure design, a hybrid architecture combining a relational database and a spatial database can be adopted. The relational database is responsible for managing structured environmental characteristic data and carbon sink assessment results, supporting efficient SQL queries and transaction processing; the spatial database is responsible for managing spatial data with geographic location attributes, supporting spatial indexing and spatial analysis. The two parts of data are linked through unique identifiers to ensure data consistency and integrity.
[0102] By constructing a carbon sequestration assessment database for target regions, we have achieved systematic management and efficient utilization of soil carbon sequestration assessment data, providing a solid data foundation for subsequent carbon sequestration model training, validation, and application. This database not only supports the storage and query of historical data but also enables dynamic updating and optimization as new data accumulates, forming a continuously improving knowledge base.
[0103] Furthermore, the embodiment of the present application also includes:
[0104] Perform visual rendering on the layered three-dimensional model of the target area according to the output value of the soil carbon sequestration assessment model.
[0105] In a feasible implementation, in order to intuitively present the spatial distribution characteristics and three-dimensional structure of soil carbon sinks, based on the output results of the soil carbon sink assessment model, the layered three-dimensional model of the target area is rendered with advanced visualization to achieve a three-dimensional and intuitive display of the soil carbon sink assessment results.
[0106] Specifically, the output data from the soil carbon sequestration assessment model is first obtained, including parameters such as soil carbon content, soil bulk density, gravel volume fraction, and carbon density for each layer, as well as the overall carbon storage distribution characteristics. This data contains both spatial location information and numerical characteristics of various carbon sequestration-related parameters, forming the fundamental data source for 3D visualization. Then, based on the constructed stratified 3D model of the target area, a mapping relationship between parameter values and visualization elements is established. The 3D model uses a voxel structure to represent the spatial stratification characteristics of the target area, with each voxel corresponding to a spatial location and a set of soil parameter values. By establishing mapping rules between parameter values and visualization attributes such as color, transparency, and texture, the abstract numerical data is transformed into intuitive visual elements. For example, a color gradient is used to represent the distribution of high and low carbon density, with red representing high carbon density areas and blue representing low carbon density areas. Transparency variations are used to represent the overlapping relationships between different soil layers, with surface soil appearing semi-transparent and deeper soil layers fully visible. Texture mapping is used to represent the distribution characteristics of different soil types and land use types.
[0107] Next, a multi-layer 3D rendering process is performed. The first is layered rendering, which renders soil layers at different depths separately, supporting multiple visualization modes such as single-layer display, multi-layer combined display, and full-layer perspective display. Next is cross-section rendering, which generates a carbon content distribution map of the soil profile by setting any section, visually demonstrating changes in carbon distribution in the vertical direction. Next is isosurface rendering, which constructs isosurfaces based on parameter values such as carbon density, visually displaying the spatial distribution of the same parameter values. Finally, volume rendering uses techniques such as ray casting and voxel integration to generate a holistic 3D model, fully demonstrating the three-dimensional distribution characteristics of soil carbon sinks.
[0108] By visualizing and rendering the layered 3D model of the target area, the soil carbon sequestration assessment results are transformed into an intuitive and easy-to-understand 3D visualization, improving the interpretability of the assessment results. This 3D visualization method not only supports professional researchers in conducting in-depth scientific analysis, but also provides intuitive information presentation for management decision-makers and the public.
[0109] Example 2, as Figure 2As shown, based on the same inventive concept as the method for constructing a soil carbon sequestration assessment model provided in Example 1, an embodiment of the present invention further provides a soil carbon sequestration assessment model construction system, comprising:
[0110] The sample collection module 11 is used to collect reliable samples and perform statistics based on the land use type data, soil type data of several layers, terrain data, climate and meteorological data, and NDVI vegetation index of the target area to obtain confidence intervals of soil capacity of several layers, confidence intervals of gravel volume fraction of several layers, and confidence intervals of soil carbon content of several layers;
[0111] A data screening module 12 is configured to perform confidence sorting on the soil sample data of the plurality of layers based on the confidence intervals of the soil bulk density of the plurality of layers, the confidence intervals of the gravel volume fraction of the plurality of layers, and the confidence intervals of the soil carbon content of the plurality of layers, to obtain the detected values of the soil bulk density of the plurality of layers, the detected values of the gravel volume fraction of the plurality of layers, and the detected values of the soil carbon content of the plurality of layers;
[0112] A carbon density statistics module 13 is configured to perform carbon density statistics based on the soil bulk density detection values of the plurality of layers, the gravel volume fraction detection values of the plurality of layers, and the soil carbon content detection values of the plurality of layers, to obtain carbon densities and total carbon reserves of the plurality of layers;
[0113] The model training module 14 is used to train the soil carbon sink assessment model using the carbon density of the several layers and the total carbon storage as supervision, and the soil bulk density detection values of the several layers, the gravel volume fraction detection values of the several layers and the soil carbon content detection values of the several layers as input.
[0114] Furthermore, the sample collection module 11 includes the following execution steps:
[0115] Extracting first layer soil type data from the plurality of layer soil type data;
[0116] When the candidate credible sample stored in the blockchain satisfies the land use type data, the first layer soil type data, the terrain data, the climate and meteorological data, and the NDVI vegetation index, it is added to the credible sample;
[0117] Extracting a set of soil bulk density record values of the credible sample, performing a central tendency assessment, obtaining a first-layer soil bulk density reset confidence interval, and adding the reset confidence intervals of the soil bulk density of the plurality of layers;
[0118] Extracting a set of recorded gravel volume fraction values of the credible sample, performing a central tendency assessment, obtaining a confidence interval of gravel volume fraction of a first layer, and adding the confidence intervals of gravel volume fraction of the plurality of layers;
[0119] A set of soil carbon content record values of the credible sample is extracted, and a central tendency assessment is performed to obtain a confidence interval for the soil carbon content of the first layer, which is added to the confidence intervals for the soil carbon content of the plurality of layers.
[0120] Furthermore, the sample collection module 11 further includes the following execution steps:
[0121] Extracting sample land use type data, sample soil type data, sample terrain data, sample climate and meteorological data, and sample NDVI vegetation index of the candidate credible sample, wherein the sample layer depth is the same as the first layer depth;
[0122] Calculating a first deviation coefficient between the sample land use type data and the land use type data;
[0123] Calculating a second deviation coefficient between the sample soil type data and the first layer soil type data;
[0124] calculating a third deviation coefficient between the sample terrain data and the terrain data;
[0125] Calculating a fourth deviation coefficient between the sample climate and meteorological data and the climate and meteorological data;
[0126] Calculating a fifth deviation coefficient between the sample NDVI vegetation index and the sample NDVI vegetation index;
[0127] When the first deviation coefficient, the second deviation coefficient, the third deviation coefficient, the fourth deviation coefficient, and the fifth deviation coefficient are all less than or equal to the corresponding deviation coefficient thresholds, the candidate credible sample is deemed to meet the land use type data, the first layered soil type data, the terrain data, the climate and meteorological data, and the NDVI vegetation index;
[0128] Otherwise, it is deemed unsatisfactory.
[0129] Furthermore, the sample collection module 11 further includes:
[0130] The first deviation coefficient is equal to a normalized parameter of the mean difference in area between the sample land use type data and the same land use type data;
[0131] The second deviation coefficient is equal to a normalized parameter of the mean deviation of the same soil type content ratio between the sample soil type data and the first layered soil type data;
[0132] The third deviation coefficient is equal to the Euclidean distance between the sample terrain data and the normalized value of the characteristic parameter of the terrain data;
[0133] The fourth deviation coefficient is equal to the Euclidean distance between the sample climate and meteorological data and the normalized value of the characteristic parameter of the climate and meteorological data;
[0134] The fifth deviation coefficient is equal to a normalized parameter of the deviation between the sample NDVI vegetation index and the sample NDVI vegetation index.
[0135] Furthermore, the sample collection module 11 further includes the following execution steps:
[0136] performing cluster analysis on the set of soil bulk density record values based on a soil bulk density deviation threshold to obtain multiple clusters of soil bulk density record values;
[0137] The clusters whose number of recorded values is less than or equal to the cluster number threshold are deleted to obtain multiple clusters of updated soil bulk density recorded values;
[0138] Traversing the plurality of clusters to update soil bulk density record values, performing a central tendency assessment, and obtaining a plurality of soil bulk density central value distribution intervals;
[0139] A union of the plurality of soil bulk density centralized value distribution intervals is taken to obtain a confidence interval of the soil bulk density of the first layer.
[0140] Furthermore, the model training module 14 includes the following execution steps:
[0141] Taking the soil bulk density detection values of the plurality of layers as supervision, and taking the land use type data, the soil type data of the plurality of layers, the topographic data, the climate and meteorological data, and the NDVI vegetation index as input, a first branch network is constructed;
[0142] Taking the gravel volume fraction detection values of the plurality of layers as supervision, and taking the land use type data, the soil type data of the plurality of layers, the topographic data, the climate and meteorological data, and the NDVI vegetation index as input, a second branch network is constructed;
[0143] The third branch network is constructed by taking the soil carbon content detection values of the plurality of layers as supervision and taking the land use type data, the soil type data of the plurality of layers, the topographic data, the climate and meteorological data and the NDVI vegetation index as input;
[0144] Based on the land use type data, the plurality of layered soil type data, and the terrain data, a layered three-dimensional model of the target area is constructed, topological simulation is performed, and a graph neural network of the target area is constructed;
[0145] The soil carbon sink assessment base model is trained by using the carbon density of the plurality of layers as the output supervision value of each layer of the target area graph neural network, using the total carbon storage global output supervision value, and using the soil bulk density detection values of the plurality of layers, the gravel volume fraction detection values of the plurality of layers, the soil carbon content detection values of the plurality of layers, and the area of the plurality of layers as the input values of each layer of the target area graph neural network;
[0146] The first branch network, the second branch network and the third branch network are embedded into the input nodes corresponding to the several layered soil bulk density detection values, the several layered gravel volume fraction detection values and the several layered soil carbon content detection values of each layer of the target area graph neural network to obtain the soil carbon sink assessment model.
[0147] Furthermore, the embodiment of the present application further includes a database construction module, which is used to:
[0148] The land use type data, the soil type data of the several layers, the terrain data, the climate and meteorological data, and the NDVI vegetation index are set as index condition data, and the soil bulk density detection values of the several layers, the gravel volume fraction detection values of the several layers and the soil carbon content detection values of the several layers, the carbon density of the several layers and the total carbon storage are set as index result data to construct a carbon sink assessment database for the target area.
[0149] Furthermore, the embodiment of the present application further includes a visualization rendering module, which is used to:
[0150] Perform visual rendering on the layered three-dimensional model of the target area according to the output value of the soil carbon sequestration assessment model.
[0151] It should be noted that, in the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0152] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0153] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded computer, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0154] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0155] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0156] Although preferred embodiments of the present invention have been described, additional changes and modifications to these embodiments may occur to those skilled in the art once the basic inventive concepts become known.
[0157] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the present invention and its equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. A method for constructing a soil carbon sequestration assessment model, characterized in that: include: Based on the land use type data of the target area, soil type data of several layers, topographic data, climate and meteorological data, and NDVI vegetation index, reliable samples were collected and statistics were performed to obtain confidence intervals for soil capacity of several layers, confidence intervals for gravel volume fraction of several layers, and confidence intervals for soil carbon content of several layers, including: Extracting first layer soil type data from the plurality of layer soil type data; When the candidate credible sample stored in the blockchain satisfies the land use type data, the first layer soil type data, the terrain data, the climate and meteorological data, and the NDVI vegetation index, it is added to the credible sample; Extracting a set of soil bulk density record values of the credible sample, performing a central tendency assessment, obtaining a first-layer soil bulk density reset confidence interval, and adding the reset confidence intervals of the soil bulk density of the plurality of layers; Extracting a set of recorded gravel volume fraction values of the credible sample, performing a central tendency assessment, obtaining a confidence interval of gravel volume fraction of a first layer, and adding the confidence intervals of gravel volume fraction of the plurality of layers; Extracting a set of soil carbon content record values of the credible sample, performing a central tendency assessment, obtaining a confidence interval for soil carbon content in a first layer, and adding the confidence intervals for soil carbon content in the plurality of layers; Based on the confidence intervals of the soil bulk density of the several layers, the confidence intervals of the gravel volume fraction of the several layers, and the confidence intervals of the soil carbon content of the several layers, confidence sorting is performed on the soil sample data of the several layers to obtain the detected values of the soil bulk density of the several layers, the detected values of the gravel volume fraction of the several layers, and the detected values of the soil carbon content of the several layers; performing carbon density statistics based on the soil bulk density detection values of the plurality of layers, the gravel volume fraction detection values of the plurality of layers, and the soil carbon content detection values of the plurality of layers to obtain carbon densities and total carbon reserves of the plurality of layers; The soil carbon sink assessment model is trained using the carbon density of the several layers and the total carbon storage as supervision, and the soil bulk density detection values of the several layers, the gravel volume fraction detection values of the several layers and the soil carbon content detection values of the several layers as input.
2. The method for constructing a soil carbon sequestration assessment model according to claim 1, wherein: When the candidate trusted sample stored in the blockchain satisfies the land use type data, the first layer soil type data, the terrain data, the climate and meteorological data, and the NDVI vegetation index, including: Extracting sample land use type data, sample soil type data, sample terrain data, sample climate and meteorological data, and sample NDVI vegetation index of the candidate credible sample, wherein the sample layer depth is the same as the first layer depth; Calculating a first deviation coefficient between the sample land use type data and the land use type data; Calculating a second deviation coefficient between the sample soil type data and the first layer soil type data; calculating a third deviation coefficient between the sample terrain data and the terrain data; Calculating a fourth deviation coefficient between the sample climate and meteorological data and the climate and meteorological data; Calculating a fifth deviation coefficient between the sample NDVI vegetation index and the NDVI vegetation index; When the first deviation coefficient, the second deviation coefficient, the third deviation coefficient, the fourth deviation coefficient, and the fifth deviation coefficient are all less than or equal to the corresponding deviation coefficient thresholds, the candidate credible sample is deemed to meet the land use type data, the first layered soil type data, the terrain data, the climate and meteorological data, and the NDVI vegetation index; Otherwise, it is deemed unsatisfactory.
3. The method for constructing a soil carbon sequestration assessment model according to claim 2, wherein: include: The first deviation coefficient is equal to a normalized parameter of the mean difference in area between the sample land use type data and the same land use type data; The second deviation coefficient is equal to a normalized parameter of the mean deviation of the same soil type content ratio between the sample soil type data and the first layered soil type data; The third deviation coefficient is equal to the Euclidean distance between the sample terrain data and the normalized value of the characteristic parameter of the terrain data; The fourth deviation coefficient is equal to the Euclidean distance between the sample climate and meteorological data and the normalized value of the characteristic parameter of the climate and meteorological data; The fifth deviation coefficient is equal to a normalized parameter of the deviation between the sample NDVI vegetation index and the NDVI vegetation index.
4. The method for constructing a soil carbon sequestration assessment model according to claim 3, wherein: Extracting a set of soil bulk density record values of the credible sample, performing a central tendency assessment, and obtaining a first-layer soil bulk density confidence interval, including: performing cluster analysis on the set of soil bulk density record values based on a soil bulk density deviation threshold to obtain multiple clusters of soil bulk density record values; The clusters whose number of recorded values is less than or equal to the cluster number threshold are deleted to obtain multiple clusters of updated soil bulk density recorded values; Traversing the plurality of clusters to update soil bulk density record values, performing a central tendency assessment, and obtaining a plurality of soil bulk density central value distribution intervals; A union of the plurality of soil bulk density centralized value distribution intervals is taken to obtain a confidence interval of the soil bulk density of the first layer.
5. The method for constructing a soil carbon sequestration assessment model according to claim 1, wherein: The soil carbon sink assessment model is trained using the carbon density of the plurality of layers and the total carbon storage as supervision, and using the detected values of soil bulk density of the plurality of layers, the detected values of gravel volume fraction of the plurality of layers, and the detected values of soil carbon content of the plurality of layers as input, including: Taking the soil bulk density detection values of the plurality of layers as supervision, and taking the land use type data, the soil type data of the plurality of layers, the topographic data, the climate and meteorological data, and the NDVI vegetation index as input, a first branch network is constructed; Taking the gravel volume fraction detection values of the plurality of layers as supervision, and taking the land use type data, the soil type data of the plurality of layers, the topographic data, the climate and meteorological data, and the NDVI vegetation index as input, a second branch network is constructed; The third branch network is constructed by taking the soil carbon content detection values of the plurality of layers as supervision and taking the land use type data, the soil type data of the plurality of layers, the topographic data, the climate and meteorological data and the NDVI vegetation index as input; Based on the land use type data, the plurality of layered soil type data, and the terrain data, a layered three-dimensional model of the target area is constructed, topological simulation is performed, and a graph neural network of the target area is constructed; The soil carbon sink assessment base model is trained by using the carbon density of the plurality of layers as the output supervision value of each layer of the target area graph neural network, using the total carbon storage global output supervision value, and using the soil bulk density detection values of the plurality of layers, the gravel volume fraction detection values of the plurality of layers, the soil carbon content detection values of the plurality of layers, and the area of the plurality of layers as the input values of each layer of the target area graph neural network; The first branch network, the second branch network and the third branch network are embedded into the input nodes corresponding to the several layered soil bulk density detection values, the several layered gravel volume fraction detection values and the several layered soil carbon content detection values of each layer of the target area graph neural network to obtain the soil carbon sink assessment model.
6. The method for constructing a soil carbon sequestration assessment model according to claim 1, wherein: Also includes: The land use type data, the soil type data of the several layers, the terrain data, the climate and meteorological data, and the NDVI vegetation index are set as index condition data, and the soil bulk density detection values of the several layers, the gravel volume fraction detection values of the several layers and the soil carbon content detection values of the several layers, the carbon density of the several layers and the total carbon storage are set as index result data to construct a carbon sink assessment database for the target area.
7. The method for constructing a soil carbon sequestration assessment model according to claim 5, wherein: Also includes: Perform visual rendering on the layered three-dimensional model of the target area according to the output value of the soil carbon sequestration assessment model.
8. A soil carbon sequestration assessment model construction system, characterized in that: For implementing the soil carbon sequestration assessment model construction method according to any one of claims 1 to 7, the system comprises: The sample collection module is used to collect reliable samples and perform statistics based on the land use type data, soil type data of several layers, terrain data, climate and meteorological data, and NDVI vegetation index of the target area to obtain confidence intervals of soil capacity of several layers, confidence intervals of gravel volume fraction of several layers, and confidence intervals of soil carbon content of several layers; a data screening module, configured to perform confidence sorting on the soil sample data of the plurality of layers based on the confidence intervals of the soil bulk density of the plurality of layers, the confidence intervals of the gravel volume fraction of the plurality of layers, and the confidence intervals of the soil carbon content of the plurality of layers, to obtain the detected values of the soil bulk density of the plurality of layers, the detected values of the gravel volume fraction of the plurality of layers, and the detected values of the soil carbon content of the plurality of layers; a carbon density statistics module, configured to perform carbon density statistics based on the soil bulk density detection values of the plurality of layers, the gravel volume fraction detection values of the plurality of layers, and the soil carbon content detection values of the plurality of layers, to obtain carbon densities and total carbon reserves of the plurality of layers; a model training module, configured to train a soil carbon sink assessment model using the carbon density of the plurality of layers and the total carbon storage as supervision, and the detected values of soil bulk density of the plurality of layers, the detected values of gravel volume fraction of the plurality of layers, and the detected values of soil carbon content of the plurality of layers as input; The sample collection module is also used for: Extracting first layer soil type data from the plurality of layer soil type data; When the candidate credible sample stored in the blockchain satisfies the land use type data, the first layer soil type data, the terrain data, the climate and meteorological data, and the NDVI vegetation index, it is added to the credible sample; Extracting a set of soil bulk density record values of the credible sample, performing a central tendency assessment, obtaining a first-layer soil bulk density reset confidence interval, and adding the reset confidence intervals of the soil bulk density of the plurality of layers; Extracting a set of recorded gravel volume fraction values of the credible sample, performing a central tendency assessment, obtaining a confidence interval of gravel volume fraction of a first layer, and adding the confidence intervals of gravel volume fraction of the plurality of layers; A set of soil carbon content record values of the credible sample is extracted, and a central tendency assessment is performed to obtain a confidence interval for the soil carbon content of the first layer, which is added to the confidence intervals for the soil carbon content of the plurality of layers.
Citation Information
Patent Citations
Forest carbon reserve and carbon sink value monitoring system and dynamic evaluation method
CN117350748A
Interactive generation type bridge parameter design method and system and storage medium
CN118296718A