A method for exploring the importance of different factors on soil organic carbon

By combining the SWAT-Carbon model and the random forest model, the problem of unclear changes in soil organic carbon in watershed systems was solved, and quantitative assessment of the impact of various factors was achieved, thereby improving the ability to analyze and predict the global carbon cycle.

CN116994668BActive Publication Date: 2025-11-11CHINESE RES ACAD OF ENVIRONMENTAL SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310767623.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-27
Publication Date
2025-11-11
Estimated Expiration
2043-06-27

AI Technical Summary

Technical Problem

The lack of clarity regarding the baseline and processes of carbon changes at contaminated sites in different regions hinders accurate analysis and prediction of the global carbon cycle. This is particularly true in watershed systems within China's territory, where intense human activity has led to unclear changes in the soil organic carbon pool.

Method used

Using the SWAT-Carbon model combined with the random forest model, watershed boundaries were generated from DEM data, land use and soil data were assigned values, and meteorological data was used for prediction. A watershed-scale soil organic carbon impact analysis method was constructed. Unique heat vector coding was used to process categorical variables, assess variable importance, and quantitatively analyze the impact of each factor on soil organic carbon.

Benefits of technology

This study enables a comprehensive analysis of the impacts of multiple factors on terrestrial and aquatic carbon cycles at the watershed scale, improving our understanding and prediction of carbon balance and dynamic changes under global change, and quantitatively assessing the importance of each variable to soil organic carbon content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116994668B_ABST
    Figure CN116994668B_ABST
Patent Text Reader

Abstract

The application discloses a method for exploring the importance of different factors on soil organic carbon, and comprises the construction of a SWAT-Carbon model and the importance analysis of related variable factors, wherein the construction of the SWAT-Carbon model comprises the generation of a basin based on DEM, the generation of land use data, the generation of soil data, the generation of meteorological data and model running. The importance analysis of the related variable factors adopts a random forest model evaluation, and the random forest model adopts a one-hot vector coding to convert the classification variables into numerical data, which can retain the information and distinguish degree of the type variables. The type variables can be avoided to be sorted or given numerical weight, so as to avoid introducing bias or misunderstanding, and the accuracy and generalization ability of the model can be improved. The exploration method disclosed by the application can comprehensively analyze the influence of various factors on the land and aquatic carbon cycle at the basin scale, and can help us better understand and predict the carbon balance and dynamic change in the basin system under global change.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of organic carbon's impact on soil technology, and more particularly to a method for investigating the importance of different factors influencing soil organic carbon. Background Technology

[0002] Soil is one of the core elements of terrestrial ecosystems. Soil organic carbon (SOC) is a collective term for humus, plant and animal remains, and microorganisms formed through microbial activity; the carbon content within these substances is called soil organic carbon. Soil organic carbon is classified into readily decomposable organic carbon, difficult-to-decompose organic carbon, and inert organic carbon based on its availability to microorganisms.

[0003] Even small fluctuations in soil organic carbon pools can have a significant impact on atmospheric CO2 concentrations and the carbon balance, thus playing a crucial role in regulating the carbon balance of Earth's surface ecosystems and mitigating greenhouse gas emissions. The relationship between soil organic carbon decomposition and climate factors, particularly its sensitivity to temperature and precipitation, is a focal point of academic attention. The combined effects of climate change and human activities will lead to changes in soil organic carbon accumulation and its dynamic balance, causing soil organic carbon pools to potentially function as both carbon sinks and sources, thereby in turn impacting global climate change.

[0004] China has a vast territory with numerous mountains and rivers. Its terrain slopes from west to east, with mountains, plateaus, and hills accounting for approximately 67% of its land area, and basins and plains for about 33%. This diversity has shaped the unique environmental characteristics and agricultural practices of different regions. Due to my country's large population, human activities significantly disrupt natural ecosystems, impacting the carbon cycle. Given my country's vast area, the carbon changes at contaminated sites in different regions are often unclear, with unknown baselines and processes. Therefore, determining the soil carbon storage at contaminated sites in these regions is crucial. Furthermore, accurately analyzing the impact of different factors on terrestrial and aquatic carbon cycles will help us better understand and predict the dynamic changes in carbon balance under global change. Summary of the Invention

[0005] To address the aforementioned problems, this invention aims to provide a method for exploring the importance of different factors in soil organic carbon. This method can comprehensively analyze the impact of multiple factors on terrestrial and aquatic carbon cycles at the watershed scale, helping us to better understand and predict the carbon balance and dynamic changes in watershed systems under global change.

[0006] To achieve the above objectives, the technical solution adopted by this invention is as follows: a method for exploring the importance of different factors on soil organic carbon, characterized by including the construction of a SWAT-Carbon model and importance analysis of relevant variable factors, wherein the construction of the SWAT-Carbon model includes the following steps:

[0007] Step 1, DEM-based watershed generation: Extract watershed boundaries and river networks from DEM data, and divide sub-watersheds and HRUs;

[0008] Step 2, land use data generation: Based on publicly available land use databases, extract the distribution of various land use types within the watershed and assign corresponding parameter values;

[0009] Step 3, Soil Data Generation: Based on publicly available soil databases, extract the spatial distribution of various soil types within the watershed and assign corresponding parameter values;

[0010] Step 4, Meteorological data generation: Using meteorological station or grid data, extract meteorological elements such as rainfall, temperature, wind speed, relative humidity, and solar radiation, and perform quality checks and spatiotemporal forecasts.

[0011] Step 5, Model Run: Convert the above data into the format required by the SWAT model and input it into the SWAT model interface. Adjust the parameters according to the model output and evaluate the final result.

[0012] The importance analysis of the relevant variable factors was evaluated using a random forest model.

[0013] Preferably, the random forest model uses one-hot vector coding to convert categorical variables into numerical data. This coding preserves the information and discriminative power of the categorical variables. It avoids sorting or assigning numerical weights to the categorical variables, thus preventing the introduction of bias or misunderstanding, and improving the model's accuracy and generalization ability.

[0014] Furthermore, the random forest model uses the Gini index as an evaluation metric, and the calculation process for its variable importance is as follows:

[0015] Suppose there are J features X1, X2, X3, ..., X... I Given I decision trees and C categories, we need to calculate the value of each feature X. j Gini Index Score The formula for calculating the Gini index of the i-th tree node q is:

[0016]

[0017] Where C represents the number of categories, P qcThis represents the proportion of category C in node g;

[0018] The change in the Gini exponent before and after the branch at node q is:

[0019]

[0020] Among them, GI l (i) and Gi r (i) These represent the Gini indices of the two new nodes after the branching;

[0021] If feature X j If the nodes appearing in decision tree i are set Q, then X j The importance of the i-th tree is:

[0022]

[0023] Assuming there are I trees in the random forest model, then:

[0024]

[0025] After normalizing the importance of all features, the ranking of the variables' importance to soil organic carbon content was obtained:

[0026]

[0027] The beneficial effects of this invention are: the research method selects variables that may affect soil organic carbon content and quantitatively assesses the importance of each variable to soil organic carbon content, enabling a comprehensive analysis of the impact of multiple factors on terrestrial and aquatic carbon cycles at the watershed scale, which can help us better understand and predict the carbon balance and dynamic changes in watershed systems under global change. Attached Figure Description

[0028] Figure 1 This is a map of the Beijing-Tianjin-Hebei research area of ​​this invention.

[0029] Figure 2 This is a flowchart of the SWAT model of the present invention.

[0030] Figure 3 This is a diagram illustrating the results of the watershed division of the Beijing-Tianjin-Hebei region according to the present invention.

[0031] Figure 4 This is a diagram illustrating the land use type classification results for the Beijing-Tianjin-Hebei region according to the present invention.

[0032] Figure 5 This is a diagram illustrating the soil type classification results of this invention.

[0033] Figure 6This is a diagram illustrating the annual average maximum temperature distribution of the CMADS system according to the present invention.

[0034] Figure 7 This diagram illustrates the distribution of CMADS sites according to the present invention.

[0035] Figure 8 This is a diagram illustrating the temperature data prediction results of the present invention.

[0036] Figure 9 This is a diagram illustrating the meteorological simulation results of the Beijing-Tianjin-Hebei region generated based on CAMDS time-series data according to the present invention.

[0037] Figure 10 This is a diagram illustrating the changes in soil organic carbon in the Beijing-Tianjin-Hebei region from 2000 to 2030.

[0038] Figure 11 This is a graph showing the SOC variation trend of 20 sub-basins from 2000 to 2030 according to the present invention.

[0039] Figure 12 This is a comparison chart of the accuracy of the present invention. (a): SWAT-Carbon simulation results; (b) SoilGrids data; (c) data from the applicant's previous research.

[0040] Figure 13 This is a scatter plot comparing the accuracy of this invention. SOC (MLModel) is the soil organic carbon concentration previously simulated by the applicant based on machine learning methods, and SOC (SWAT-Carbon) is the soil organic carbon result simulated by SWAT-Carbon.

[0041] Figure 14 This is a ranking diagram of the importance of variables in this invention. X0_1-X0_6 are the components of land use type after one-hot vector coding, and X1_4-X1_17 are the components of soil type after one-hot vector coding. Detailed Implementation

[0042] To enable those skilled in the art to better understand the technical solutions of the present invention, the technical solutions of the present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0043] Example

[0044] This application pertains to the Beijing-Tianjin-Hebei region, one of the most important economic areas in northern China, encompassing Beijing, Tianjin, and 11 prefecture-level cities in Hebei Province. Located between the North China Plain and the Taihang Mountains, this region has a warm-temperate, semi-humid continental monsoon climate with distinct seasons and uneven rainfall. As one of my country's most developed regions, the Beijing-Tianjin-Hebei area experiences significant human disturbance to its natural ecosystems, impacting the carbon cycle. However, current understanding of carbon changes at contaminated sites in this region is unclear, lacking a comprehensive picture of the underlying processes. Therefore, determining the soil carbon storage at these contaminated sites is crucial.

[0045] This application provides a method for exploring the importance of different factors on soil organic carbon, including the construction of a SWAT-Carbon model and importance analysis of relevant variable factors. The construction of the SWAT-Carbon model includes the following steps:

[0046] Step 1, DEM-based watershed generation: Extract watershed boundaries and river networks from DEM data, and divide sub-watersheds and HRUs (Hydrological Response Units).

[0047] Sub-basin delineation: This embodiment uses 30m resolution DEM data for each province provided by the Resource and Environmental Science and Data Center of the Institute of Geographic Sciences and Natural Resources Research, Chinese Academy of Sciences. The data is then stitched and projected using ArcGIS software to obtain the DEM data for the study area. Simultaneously, vector network data of major rivers in the Beijing-Tianjin-Hebei region is downloaded from this data center as the river channel input data for the SWAT-Carbon model. Finally, the DEM data and river data are imported into the SWAT-Carbon model, and sub-basin delineation is performed according to certain threshold conditions, generating sub-basin delineation results covering a wide area of ​​the study region. The watershed delineation results for the Beijing-Tianjin-Hebei region are shown below. Figure 3 As shown.

[0048] Step 2, land use data generation: Based on the publicly available land use database, extract the distribution of various land use types within the watershed and assign corresponding parameter values.

[0049] Land use data generation: Land use data is one of the important input data affecting the simulation results of the SWAT-Carbon model. This embodiment uses two types of data to describe the land use situation in the study area: land use distribution map (raster) and land use type index table.

[0050] The land use distribution map was provided by the Resource and Environmental Science and Data Center of the Institute of Geographic Sciences and Natural Resources Research, Chinese Academy of Sciences, and was classified according to the "Classification of Current Land Use" (GB / T21010-2017) standard. In order to match the land cover / vegetation growth database of the SWAT model, this embodiment reclassifies the original land use types into six categories: cultivated land (AGRL), forest land (FRST), grassland (HAY), urban residential land (URML), water area (WATR), and other land uses, and converts them into raster format.

[0051] The land use type index table is a table containing two columns of data. The first column is the value corresponding to each land use type in the raster map, and the second column is the corresponding number for each type in the SWATlandcover / plant database. In this embodiment, the pre-prepared land use type index table is imported using ArcSWAT software, and the values ​​in the raster map are associated with the numbers in the SWATlandcover / plant database.

[0052] The land cover / vegetation growth database stores the parameter information required for SWAT simulation of specific land cover growth, including plant growth under ideal conditions and the impact of stress factors on plant growth. This embodiment uses the default parameter values ​​provided with the SWAT-Carbon model without modification or calibration. The land use type classification results are as follows: Figure 4 As shown.

[0053] Step 3, Soil Data Generation: Based on publicly available soil databases, extract the spatial distribution of various soil types within the watershed and assign corresponding parameter values.

[0054] Soil data generation: Soil data is one of the important input data affecting the simulation results of the SWAT-Carbon model. This embodiment uses the Harmonized World Soil Database (HWSD) constructed by the Food and Agriculture Organization of the United Nations (FAO) and the International Institute for Applied Systems Science in Vienna (IIASA) as the main data source, and makes appropriate adjustments to the 1:1,000,000 soil data in China provided by the Nanjing Institute of Soil Science during the Second National Land Survey.

[0055] The soil data consists of three parts: a soil type distribution map, a soil type index table, and soil physical property files. The soil type distribution map uses a grid raster format with WGS84 projection and is classified according to the FAO-90 standard. The soil type index table is a table that connects the values ​​in the raster map with the numbers in the soilsol database included with the SWAT model.

[0056] This embodiment uses ArcSWAT software to import a pre-prepared index table for association. The soilsol database stores parameters such as hydrological and water conduction properties of various soils, including parameters input by soil type and by soil layer. Due to the lack or inaccuracy of some parameters in HWSD, this embodiment uses multiple methods to obtain or calculate relevant parameter values. Organic carbon content is obtained by converting organic matter content; clay, silt, sand, and gravel content are obtained from measured data; saturated hydraulic conductivity and effective field capacity are calculated using SPAW software; other parameters such as hydrological unit groups and surface reflectance are calculated using formulas; a few parameters, such as electrical conductivity, use the default values ​​of the SWAT model. Soil parameters are defined in Table 1 below, and the soil type classification structure is as follows: Figure 5 As shown.

[0057] Table 1 Soil parameters

[0058] Variable name Model definition Notes SNAM Soil Name NLAYERS Soil stratification HYDGRP Soil Hydrology Group SOL_ZMX Maximum root depth in soil profile ANION_EXCL Anion exchange porosity Default 0.5 SOL_CRK maximum compressibility of soil Default 0.5 TEXTURE Soil layer structure SOL_Z Depth (mm) of each soil layer from the bottom layer to the top layer The sum of the depths of the first few layers SOL_BD <![CDATA[Soil wet density (mg / m 3 or g / cm 3 )]]> SOL_AWC Effective water holding capacity of soil layer (mm) SOL_K Saturated hydraulic conductivity / saturated hydraulic transfer coefficient (mm / h) SOL_CBN Organic carbon content in soil layer Organic matter content*0.58 CLAY Clay content, soil particle composition with a diameter <0.0002 mm SILT Loam content, soil particle composition with a diameter of 0.0002-0.05 mm. SAND Gravel content, soil particle composition with a diameter >2.0 mm. ROCK Gravel content, composition of soil particles with a diameter >2.0 mm SOL_ALB Surface reflectance (wet) Default 0.01 USLE_K Soil erosivity factor in the USLE equation SOL_EC Soil electrical conductivity (dS / m) The default value is 0.

[0059] Step 4, Meteorological data generation: Using meteorological station or grid data, extract meteorological elements such as rainfall, temperature, wind speed, relative humidity, and solar radiation, and perform quality checks and spatiotemporal forecasts.

[0060] Meteorological data generation: Meteorological data is an essential input for building the SWAT-Carbon model, and it has a significant impact on crop growth and water balance in the watershed. The meteorological features included in the SWAT-Carbon model are daily precipitation, daily maximum / minimum temperature, relative humidity, and average wind speed, etc., which are organized according to the SWAT-Carbon model driving data format and units.

[0061] In the SWAT-Carbon model, the weather generator simulates daily meteorological data based on multi-year monthly meteorological data. Its main input data include monthly average maximum temperature, monthly average minimum temperature, standard deviation of maximum temperature, monthly average precipitation, standard deviation of precipitation, number of dry days in a month, dew point temperature, and monthly average solar radiation. The calculation formulas are shown in Table 2. Precipitation and temperature data are important factors directly affecting runoff, so measured daily-scale data should be used whenever possible. This study uses 33 meteorological stations in the Beijing-Tianjin-Hebei region provided by the National Climate Data Center (NCDC) as the basic meteorological data.

[0062] Table 2 Formulas for Calculating Weather Generator Parameters

[0063]

[0064] This embodiment aims to establish a refined model of soil organic carbon distribution in the Beijing-Tianjin-Hebei region in recent years using input data with higher spatiotemporal resolution. Therefore, this embodiment selects the CMADS (China Meteorological Assimilation Driving Datasets for the SWAT model) dataset as the meteorological data source. The CMADS dataset is developed based on the China Meteorological Administration's Atmospheric Assimilation System (CLD AS) technology, employing various techniques such as data loop nesting, resampling, model extrapolation, and bilinear interpolation. It has also been formatted and corrected according to the SWAT model input-driven data format. The CMADS dataset includes the following aspects:

[0065] (1) Temperature, air pressure, specific humidity and wind speed driving data: The data are derived from the NCEP / GFS background field and hourly observation data of basic meteorological elements on the ground since January 2009 from 2,421 national automatic stations and 29,452 regional automatic stations.

[0066] (2) Precipitation: It is a combination of precipitation data from multiple satellites and ground automatic stations.

[0067] The CMADS dataset contains 205 stations in the Beijing-Tianjin-Hebei region, far exceeding the 33 stations included in the U.S. National Climate Center dataset. Therefore, this dataset provides a more accurate characterization of meteorological elements.

[0068] Meteorological Data Imputation Based on LSTM Algorithm: Since the currently available CMADS dataset only covers data up to 2018, this embodiment employs a time-series forecasting method based on historical CMADS data to estimate meteorological data after 2018 in order to achieve a fine-grained characterization of the SOC in the Beijing-Tianjin-Hebei region. Among numerous time-series forecasting models, this embodiment selects the LSTM model as the forecasting tool. The LSTM model is a recurrent neural network capable of capturing long-term dependencies, exhibiting high accuracy and robustness. By training and testing the CMADS dataset using the LSTM model, this embodiment obtains meteorological data that meets the input requirements of the SWAT model and uses it for modeling and simulating the SOC distribution in the Beijing-Tianjin-Hebei region. In terms of simulation accuracy for maximum and minimum temperatures, the MSE values ​​in the test set are only 0.11 and 0.08, respectively. Figure 8-9 As shown, a relatively good prediction result was achieved.

[0069] Hydrological Response Unit (HRU) Delineation: The SWAT model divides sub-basins into different HRUs based on land use type, soil type, and slope, assuming that HRUs of the same type have the same hydrological behavior. During model calculation, the hydrological processes of various HRUs are superimposed at the sub-basin outlet to obtain the sub-basin's hydrological output. The number of HRUs affects the model's speed and accuracy. This embodiment uses a 10% threshold for HRU delineation; land use, soil distribution, slope type, etc., below these thresholds will be merged into other types, thus ensuring a moderate number of HRUs within each sub-basin without affecting the reliability of the calculation results.

[0070] Step 5, Model Run: Convert the above data into the format required by the SWAT model and input it into the SWAT model interface. Adjust the parameters according to the model output results. The simulation time range is from January 1, 2000 to December 31, 2030. Evaluate the final results.

[0071] Visualization of implementation results

[0072] 1. Spatial distribution of SOC from 2000 to 2030

[0073] Based on soil organic carbon data from 2000 to 2030, this study plotted a spatial distribution map of soil organic carbon content in the Beijing-Tianjin-Hebei region (e.g., Figure 10 (As shown in the image). The results indicate that the soil organic carbon content in the Beijing-Tianjin-Hebei region exhibits a pattern of higher levels in the south and lower levels in the north, consistent with the region's topographical characteristics. The southern region has a higher altitude, more complex terrain, and predominantly forest and grassland land use, which are conducive to the accumulation and stabilization of soil organic carbon. In contrast, the northern region has a lower altitude, flatter terrain, and predominantly arable land land use, resulting in lower soil organic carbon content.

[0074] To analyze the dynamic changes in soil organic carbon in the Beijing-Tianjin-Hebei region, this embodiment selected 20 typical sub-basins and plotted the change curves of soil organic carbon content from 2000 to 2030 (e.g., Figure 11 (As shown in the figure). The results show that the soil organic carbon content in the Beijing-Tianjin-Hebei region generally showed an upward trend over the 31 years, but the growth rates varied. During 2000-2019, the soil organic carbon content in most sub-basins increased rapidly, with a period of accelerated growth around 2006 and 2015. During 2020-2030, the growth rate of soil organic carbon content in most sub-basins tended to stabilize or slow down, and some sub-basins even showed a fluctuation of first decreasing and then increasing.

[0075] 2. Accuracy Comparison (Comparison with Machine Learning Results)

[0076] To evaluate the accuracy of the SWAT-Carbon model in simulating soil organic carbon content in the Beijing-Tianjin-Hebei region, this embodiment employs two comparative analysis methods. One method compares measured values ​​from monitoring stations with simulated values. However, since the SWAT-Carbon model simulates sub-basins, and each sub-basin covers a large area, soil organic carbon content exhibits significant spatial variability (e.g., between 1-10 kg / m²), this method may have limitations and be unreasonable. The other method compares publicly available online spatial distribution data of soil organic carbon (SoilGrids™ (hereinafter referred to as SoilGrids) is a global digital soil mapping system that uses advanced machine learning methods to map the spatial distribution of soil properties globally) with the average values ​​of previously obtained spatial distribution data and simulation results of soil organic carbon in the Beijing-Tianjin-Hebei region within the watershed. Figure 12 As shown, this is to verify the reliability of the SWAT-Carbon model.

[0077] It can be observed that all three datasets exhibit the same spatial distribution pattern, demonstrating the good reliability of the spatial distribution of soil organic carbon simulated based on the SWAT-Carbon model. To further verify the specific accuracy, the SOC values ​​of the sub-basins were compared with the average soil organic carbon concentration in the Beijing-Tianjin-Hebei region simulated by the research group using machine learning methods within the sub-basins. The verification results are shown in the figure below. A strong positive correlation exists between the two datasets, and the overall fitting accuracy reached 0.68. Figure 13 As shown, it has a high degree of credibility.

[0078] Results Analysis

[0079] The results analysis focused on the importance of relevant variable factors. This application employed a random forest model to evaluate and analyze 12 variables that may affect soil organic carbon content, including slope, altitude, population density, GDP, annual precipitation, daily relative humidity, daily average solar radiation, daily average maximum temperature, daily average minimum temperature, river density, soil type, and land use type. The importance of each variable to soil organic carbon content was quantitatively assessed. Feature importance can be used for feature selection, i.e., removing irrelevant or redundant features, thereby improving the efficiency and effectiveness of the model.

[0080] Random forest algorithms are ensemble learning methods based on decision trees. Decision trees require numerical comparisons of features at each node to select the optimal split point. Therefore, features that are categorical variables cannot be directly used for modeling, such as soil type and land use type variables mentioned above. Categorical variables need to be transformed before they can be used in random forest algorithms. One-hot vector coding is a commonly used transformation method. It converts each categorical variable into multiple binary variables representing whether it belongs to a certain category. For example, if a categorical variable has three possible values ​​(A, B, C), it can be converted into three binary variables (A, B, C), where A = 1 indicates belonging to category A, otherwise 0; B = 1 indicates belonging to category B, otherwise 0; and C = 1 indicates belonging to category C, otherwise 0. This converts categorical variables into numerical data, making them suitable for use in random forest algorithms.

[0081] Using one-hot vector encoding has the following advantages: it preserves the information and discriminative power of type variables; it avoids sorting or assigning numerical weights to type variables, thus preventing the introduction of bias or misinterpretation; and it improves the accuracy and generalization ability of the model.

[0082] The random forest model uses the Gini index as an evaluation metric, and the calculation process for its variable importance is as follows:

[0083] Suppose there are J features X1, X2, X3, ..., X... I Given I decision trees and C categories, we need to calculate the value of each feature X. j Gini Index Score That is, the average change in node splitting impurity of the j-th feature across all decision trees in the random forest model; the formula for calculating the Gini index of node q in the i-th tree is:

[0084]

[0085] Where C represents the number of categories, P qc This represents the proportion of class C in node g; that is, the probability that two samples randomly drawn from node q will have different class labels:

[0086] Feature X j The importance of node q in the i-th tree, i.e., the change in the Gini index before and after node q branches, is:

[0087]

[0088] Among them, GI l (i) and Gi r (i) These represent the Gini indices of the two new nodes after the branching;

[0089] If feature X j If the nodes appearing in decision tree i are set Q, then X j The importance of the i-th tree is:

[0090]

[0091] Assuming there are I trees in the random forest model, then:

[0092]

[0093] After normalizing the importance of all features, the ranking of the importance of each variable to soil organic carbon content was obtained.

[0094]

[0095] The result of the variable importance ranking based on the random forest algorithm is as follows: Figure 14 As shown, among all variables, slope has the greatest impact on SOC, followed by altitude and population density. This is mainly because slope affects soil properties such as soil erosion, moisture, and temperature, leading to a general increase in soil organic carbon content with increasing slope. At the same time, altitude affects ecological and environmental conditions such as climate, vegetation, and soil type, causing soil organic carbon content to gradually increase with increasing altitude. Similarly, population density affects socio-economic factors such as land use patterns, vegetation cover, and management intensity, causing soil organic carbon content to change with population density.

[0096] However, because land use type and soil type variables were split through one-hot vector coding, while this improved the overall modeling accuracy, it negatively impacted the importance of these two types of variables, resulting in generally lower importance. Nevertheless, among the numerous decomposed variables after one-hot vector coding, the importance of cultivated land and forest land remained far higher than other independent variables. Cultivated land and forest land represent two distinct land use patterns, and their impacts on soil organic carbon content differ significantly. The influence of cultivated land and forest land on soil organic carbon is mainly reflected in three aspects: input, transformation, and output of soil organic carbon.

[0097] Farmland reduces soil organic carbon input because it has less vegetation cover; at the same time, it accelerates soil organic carbon conversion and output because it damages soil structure, alters soil microorganisms, and reduces soil moisture, making soil organic carbon more susceptible to oxidation and decomposition or erosion by wind and water. Conversely, forest land increases soil organic carbon input because it has more plant residues; at the same time, it slows down soil organic carbon conversion and output because it maintains higher soil moisture and temperature, forms complex root networks, and has multi-layered vegetation structures, making it easier for soil organic carbon to be converted into recalcitrant forms or preserved in the soil.

[0098] In summary, this embodiment uses the research method disclosed in this application to investigate organic carbon in the Beijing-Tianjin-Hebei region and analyzes the impact of various factors on terrestrial and aquatic carbon cycles, which can help us better understand and predict the carbon balance and dynamic changes in this watershed system under global change.

[0099] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Various changes and modifications can be made to the present invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention.

Claims

1. A method for investigating the importance of different factors on soil organic carbon, characterized by: This includes the construction of a SWAT-Carbon model and importance analysis of related variable factors. The construction of the SWAT-Carbon model includes the following steps: Step 1, DEM-based watershed generation: Extract watershed boundaries and river networks from DEM data, and divide sub-watersheds and HRUs; Step 2, land use data generation: Based on publicly available land use databases, extract the distribution of various land use types within the watershed and assign corresponding parameter values; Step 3, Soil Data Generation: Based on publicly available soil databases, extract the spatial distribution of various soil types within the watershed and assign corresponding parameter values; Step 4, Meteorological data generation: Using meteorological station or grid data, extract meteorological elements such as rainfall, temperature, wind speed, relative humidity, and solar radiation, and perform quality checks and spatiotemporal forecasts. Step 5, Model Run: Convert the above data into the format required by the SWAT model and input it into the SWAT model interface. Adjust the parameters according to the model output and evaluate the final result. The importance analysis of the relevant variable factors was evaluated using a random forest model; The random forest model uses one-hot vector coding to convert categorical variables into numerical data. The random forest model uses the Gini index as an evaluation metric.

2. The research method according to claim 1, characterized in that: The process of calculating variable importance when using the Gini index as an evaluation metric in the random forest model is as follows: Assume there is J Features , I A decision tree, C For each category, we need to calculate each feature. X j of Gini Index Score , No. i Tree nodes q of Gini The formula for calculating the index is: (1) in, C Indicates that there is C Categories P qc Represents a node g Medium category C The proportion; node q Before and after branching Gini The change in the index is: (2) in, GI l (i) and Gi r (i) These represent the two new nodes after the branch. Gini index; After normalizing the importance of all features, the ranking of the variables' importance to soil organic carbon content was obtained: (3)。 3. The research method according to claim 2, characterized in that: If the characteristic in equation (2) is X j In decision tree i The nodes appearing in the set Q ,So X j In the i The importance of each tree is: (4) Assume there are a total of I A tree, then: (5)。

Citation Information

Patent Citations

  • Method for analyzing influence of land utilization change on runoff process based on SWAT model

    CN116167193A

  • Soil loss evaluation method based GIS

    KR1020180000619A