A method and system for predicting cadmium content distribution in cultivated soil based on machine learning

By constructing a soil cadmium content distribution prediction model based on machine learning, the problem of insufficient nonlinear interaction and dynamic simulation of soil heavy metal distribution prediction in the existing technology is solved, the prediction accuracy and adaptability are improved, and decision support for environmental management is provided.

CN119989183BActive Publication Date: 2025-08-26TECH CENT FOR SOIL AGRI & RURAL ECOLOGY & ENVIRONMENT MINIST OF ECOLOGY & ENVIRONMENT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510481124.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-08-26
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

The prior art is difficult to capture the nonlinear interactions of complex environmental factors in soil heavy metal distribution prediction, and lacks dynamic simulation of pollutant migration and transformation process, resulting in insufficient prediction accuracy and cannot meet the needs of high-precision pollution warning and prevention and control.

Method used

A machine learning-based method is adopted, combined with multi-dimensional indicators (pollution source, migration process, soil attributes) to calculate characteristic factors, and a random forest model is constructed, taking into account the impact of the input and output process of soil cadmium element, and a prediction model for soil cadmium content distribution in cultivated land is constructed.

Benefits of technology

The accuracy of soil cadmium content distribution prediction is improved, dynamic simulation of the impact of pollution sources and nonlinear relationships are captured, the dynamic adaptability and prediction accuracy of the model are enhanced, and reliable decision-making indicators are provided for environmental management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989183B_ABST
    Figure CN119989183B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of soil data analysis technology, and discloses a method and system for predicting the distribution of cadmium content in cultivated land soil based on machine learning. The method comprises: obtaining historical data on multi-dimensional indicators of cultivated land soil in a target area and cadmium content data at known soil locations; calculating characteristic factors based on the historical data of the multi-dimensional indicators; constructing a dataset based on the historical data of the multi-dimensional indicators, the characteristic factors, and the cadmium content data at known soil locations; training a random forest model using the dataset to obtain a model for predicting the distribution of cadmium content in cultivated land soil; obtaining multi-dimensional indicator data of a target time series, calculating the characteristic factors, and inputting them into the model for predicting the distribution of cadmium content in cultivated land soil, to obtain a predicted result for the distribution of cadmium content in cultivated land soil. This solution can consider the impact of the input and output processes of the soil cadmium element on the soil cadmium content, thereby improving the accuracy of the model's prediction results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of soil data analysis, and in particular relates to a method and system for predicting the distribution of cadmium (Cd) content in cultivated soil based on machine learning. Background Art

[0002] Spatial prediction of heavy metals in soil is an important foundation for understanding regional soil pollution and carrying out soil pollution prevention and control. Spatial prediction methods based on machine learning models are currently an important research direction for predicting the spatial distribution of heavy metals in soil.

[0003] Existing technologies for predicting heavy metal distribution primarily rely on traditional statistical models and analysis of single environmental factors. These typically involve constructing regression models based on limited monitoring data or integrating them with geographic information systems for spatial interpolation (e.g., kriging). However, these methods struggle to capture the nonlinear interactions of complex environmental factors. Some studies have incorporated machine learning algorithms (e.g., random forests and neural networks) to optimize prediction accuracy, but these approaches are often limited to static data-driven models and lack dynamic simulation of pollutant migration and transformation processes. Traditional methods generally suffer from limitations such as a single data dimension, insufficient spatiotemporal dynamics, and weak explanatory power, making them incapable of meeting the demands for high-precision pollution early warning and prevention.

[0004] Therefore, there is an urgent need to develop a method and system for predicting the distribution of cadmium content in cultivated soil based on machine learning, which can consider the impact of the input and output processes of soil cadmium elements on the soil cadmium content and improve the accuracy of the model prediction results. Summary of the Invention

[0005] In order to solve the above technical problems, the present invention provides a method and system for predicting the distribution of cadmium content in cultivated soil based on machine learning, which can consider the impact of the input and output process of soil cadmium elements on the soil cadmium content and improve the accuracy of the model prediction results.

[0006] The present invention provides a method for predicting the distribution of cadmium content in cultivated soil based on machine learning, the method comprising the following steps:

[0007] S1. Obtain historical data on multi-dimensional indicators of cultivated land in the target area and soil cadmium content data at known locations; multi-dimensional indicators include pollution source indicators, migration process indicators, and soil property indicators;

[0008] S2. Calculate characteristic factors based on historical data of multi-dimensional indicators; the characteristic factors include pollution source attenuation factor, dynamic migration factor, and soil interaction factor;

[0009] S3. Construct a data set based on historical data of multi-dimensional indicators, characteristic factors, and soil cadmium content data at known locations;

[0010] S4. Use the dataset to train the random forest model to obtain a prediction model for the distribution of cadmium content in cultivated soil;

[0011] S5. Obtain multi-dimensional indicator data of the target time series, calculate characteristic factors, and input them into the cultivated land soil cadmium content distribution prediction model to obtain the cultivated land soil cadmium content distribution prediction results.

[0012] Furthermore, in S1, pollution source indicators include:

[0013] PM2.5 content, average distance to non-ferrous metal smelting enterprises, average distance to non-ferrous metal mining and dressing enterprises, average distance to battery manufacturing enterprises, average distance to stone building material manufacturing enterprises, average distance to waste resource utilization enterprises and night light intensity.

[0014] Furthermore, in S1, the migration process indicators include:

[0015] River density, road density, net primary production, soil erosion rate, wind speed, air temperature, air pressure, relative humidity, precipitation, and digital elevation models.

[0016] Furthermore, in S1, soil property indicators include:

[0017] pH, organic matter content, cation exchange capacity and soil moisture content.

[0018] Furthermore, in S2, the pollution source attenuation factor is calculated based on the historical data of pollution source indicators. The calculation formula is as follows:

[0019] ;

[0020] in, represents the pollution source attenuation factor, m represents the mth type of pollution source enterprise, M represents the total number of pollution source enterprises, d m represents the average distance to the mth type of pollution source enterprise, λ m represents the diffusion rate of the mth type of pollution source enterprise, ω m represents the pollution weight of the mth type of pollution source enterprise, C PM2.5 represents the PM2.5 content in the air, α represents the pollution weight of PM2.5 content, L night represents the night light intensity, and β represents the pollution weight of the night light intensity.

[0021] Furthermore, in S2, the dynamic migration factor is calculated based on the historical data of the migration process indicators. The calculation formula is as follows:

[0022] ;

[0023] in, represents the dynamic migration factor, γ, δ, and ϵ represent the influence weights of hydraulic migration, atmospheric diffusion, and vegetation retardation, respectively. 降水 Indicates precipitation, D 河流 Indicates river density, DEM T represents the terrain driving factor, V 风速 represents wind speed, T represents temperature, T base represents the base temperature, NPP represents the net primary productivity of vegetation, R 土壤 represents the soil erosion rate, and R represents the erosion threshold.

[0024] Furthermore, the calculation formula of the terrain driving factor is as follows:

[0025] ;

[0026] Among them, S norm represents the normalized slope factor, FA log represents the logarithmic runoff accumulation, C curv represents the normalized curvature, and λ represents the curvature contribution coefficient.

[0027] Furthermore, in S2, the soil interaction factor is calculated based on the historical data of soil property indicators. The calculation formula is as follows:

[0028] ;

[0029] in, represents the soil interaction factor, pH represents the pH value, k1 represents the nonlinear effect coefficient of pH value, Y represents the organic matter content, η represents the adjustment coefficient, CEC represents the cation exchange capacity, lp represents the clay content, W represents the soil moisture content, and k2 represents the moisture influence coefficient.

[0030] The present invention also provides a machine learning-based arable land soil cadmium content distribution prediction system, which is used to implement the above-mentioned machine learning-based arable land soil cadmium content distribution prediction method, characterized in that the system includes the following modules:

[0031] The data acquisition module is used to obtain historical data on multi-dimensional indicators of cultivated land in the target area and soil cadmium content data at known locations; the multi-dimensional indicators include pollution source indicators, migration process indicators, and soil property indicators;

[0032] A characteristic factor calculation module is connected to the data acquisition module and is used to calculate characteristic factors based on historical data of multi-dimensional indicators; wherein the characteristic factors include pollution source attenuation factor, dynamic migration factor and soil interaction factor;

[0033] The data set construction module is connected to the data acquisition module and the characteristic factor calculation module, and is used to construct a data set based on the historical data of multi-dimensional indicators, characteristic factors, and soil cadmium content data at known locations;

[0034] The model building module is connected to the dataset building module and is used to train the random forest model using the dataset to obtain a prediction model for the distribution of cadmium content in cultivated soil;

[0035] The output module is connected to the model building module and is used to obtain multi-dimensional indicator data of the target time series, calculate characteristic factors, and input them into the cultivated land soil cadmium content distribution prediction model to obtain the cultivated land soil cadmium content distribution prediction results.

[0036] The embodiments of the present invention have the following technical effects:

[0037] The present invention considers the influencing factors of soil cadmium content from multiple dimensions, combines multidimensional indicators, and introduces characteristic factors, so that the model can capture the impact of the input and output processes of soil cadmium elements on soil cadmium content, potential nonlinear relationships, and more realistically simulate the change in the intensity of the impact of pollution sources on the target area with distance, thereby improving the accuracy of prediction; the pollution source attenuation factor integrates distance, atmospheric deposition and human activities, the dynamic migration factor couples meteorological-topographic-ecological multiple processes, and the soil interaction factor quantifies the chemical synergistic effect, solving the characterization limitations of a single feature; by integrating the three characteristic factors constructed by environmental mechanisms, a deep integration of domain knowledge-driven and data-driven is achieved, which improves the dynamic adaptability and prediction accuracy of the model, provides a reliable tool for predicting the distribution of heavy metals in soil, and provides directly operational decision-making indicators for environmental management. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0039] Figure 1 This is a flow chart of a method for predicting cadmium content distribution in cultivated soil based on machine learning provided by an embodiment of the present invention;

[0040] Figure 2 is a schematic diagram of a digital soil mapping study area provided by an embodiment of the present invention;

[0041] Figure 3 This is a schematic diagram of a random forest model provided by an embodiment of the present invention;

[0042] Figure 4 This is an analysis chart of a comparison between predicted values ​​and observed values ​​of a cadmium content distribution prediction model for cultivated land soil and an analysis chart of the prediction interval provided by an embodiment of the present invention;

[0043] Figure 5 This is a schematic diagram of a model for predicting soil cadmium content distribution in cultivated land provided by an embodiment of the present invention for predicting soil cadmium content and mapping;

[0044] Figure 6 This is a schematic diagram of a Kriging interpolation model for predicting soil cadmium content and mapping provided by an embodiment of the present invention;

[0045] Figure 7 This is a structural diagram of a system for predicting cadmium content distribution in cultivated soil based on machine learning provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0046] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention are described clearly and completely below. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are also within the scope of protection of the present invention.

[0047] The present invention provides a method for predicting the distribution of cadmium content in cultivated soil based on machine learning. Figure 1 This is a flow chart of a method for predicting cadmium content distribution in cultivated soil based on machine learning provided by an embodiment of the present invention. Figure 2 This is a schematic diagram of a digital soil mapping study area provided by an embodiment of the present invention, see Figure 1 and Figure 2 , the method comprises the following steps:

[0048] S1. Obtain historical data on multi-dimensional indicators of cultivated land in the target area and soil cadmium content data at known locations.

[0049] The data used in this study primarily included soil cadmium (Cd) concentrations and multi-dimensional indicators for soil Cd content prediction. Sampling adhered to HJ / T 166-2004, the "Technical Specification for Soil Environmental Monitoring," using a grid-based sampling method. Each sample consisted of five subsamples collected using a five-point sampling method. All soil samples were air-dried indoors, crushed, and then passed through a nylon sieve (0.25-0.4 mm mesh size) and mixed thoroughly. Finally, soil Cd concentrations were determined using atomic absorption spectrophotometry.

[0050] Among them, the multidimensional indicators include pollution source indicators, migration process indicators and soil property indicators. The variables included in each indicator selected in this embodiment are shown in Table 1.

[0051] Table 1 Multi-dimensional indicator data for spatial distribution prediction of Cd content in regional cultivated land soil

[0052]

[0053] Among the pollution source indicators, PM2.5 content refers to the concentration of airborne particles with a diameter of 2.5 microns or less. These particles may carry heavy metals such as cadmium, contaminating the soil. The average distance from various pollution source enterprises reflects the likelihood that cultivated land will be affected by pollution from these enterprises. The closer the distance, the higher the risk of contamination. Nighttime light intensity can, to a certain extent, indicate the intensity of human activity. Areas with more frequent human activity are likely to generate more pollution sources and, therefore, have a greater impact on soil cadmium levels.

[0054] These indicators were chosen as pollution source indicators because non-ferrous metal smelting and mining companies, battery manufacturers, stone building material manufacturers, and waste resource utilization companies can all generate cadmium-containing pollutants during their production processes, contaminating surrounding farmland soil through atmospheric diffusion and wastewater discharge. PM2.5, a major component of atmospheric pollutants, can be transported over long distances and deposited in the soil, increasing soil cadmium levels. Nighttime light intensity is closely related to human activity. High levels of human activity are often accompanied by increased industrial production, transportation, and other activities, which can lead to higher cadmium emissions.

[0055] By comprehensively considering these pollution source indicators, we can more accurately assess the potential risk of cadmium contamination in farmland soils within a target area. For example, the closer the average distance to nonferrous metal smelting enterprises, the greater the likelihood that farmland in that area will be contaminated by cadmium from those enterprises; and higher PM2.5 levels indicate a greater potential for cadmium to enter the soil through atmospheric deposition. Combining these indicators provides rich and accurate pollution source information for subsequent characteristic factor calculations and model predictions, improving the reliability of prediction results.

[0056] Among the migration process indicators: river density represents the length of rivers per unit area, reflecting the hydrological characteristics of the region. As a key carrier of pollutant migration, river density affects the migration and diffusion of cadmium in the soil. Road density reflects the development of the regional transportation network. Traffic activities may cause cadmium pollution, and roads may also affect the transmission paths of pollutants. Net primary production refers to the amount of organic carbon fixed by plants through photosynthesis per unit time and per unit area, minus the amount consumed by plants through respiration. It reflects the growth of vegetation, which has a certain absorption and barrier effect on pollutants such as cadmium. Soil erosion rate indicates the rate of soil erosion. Soil erosion can cause cadmium to spread to other areas along with the migration of soil particles. Meteorological factors such as wind speed, temperature, air pressure, relative humidity, and precipitation can affect the diffusion and transport of pollutants in the atmosphere. Digital elevation models, through digital representation of terrain, can reflect the topography of a region, which affects the flow of water and pollutants.

[0057] The selection of migration process indicators is based on the migration and transformation patterns of cadmium in the environment. Rivers and roads serve as channels for material transmission, and their density affects the range and speed of cadmium migration. Meteorological factors affect the movement and diffusion capacity of the atmosphere, thereby affecting the transmission distance and deposition distribution of cadmium in the atmosphere. The net primary production of vegetation and the soil erosion rate are closely related to the migration and accumulation of cadmium in the soil-vegetation system. Digital elevation models can help analyze the impact of topography on cadmium migration. For example, in areas with large terrain undulations, water flow may be faster, thereby accelerating cadmium migration.

[0058] Considering these migration process indicators provides a more comprehensive understanding of cadmium's environmental transport behavior. For example, in areas with a high river density, cadmium may be more easily transported through surface runoff; in areas with high wind speeds, cadmium may diffuse more widely in the atmosphere. Analysis of these indicators can provide a more accurate basis for predicting cadmium's distribution in soil, helping to develop more targeted pollution prevention and control measures.

[0059] Among them, among the soil property indicators: pH value is an important indicator to measure the acidity and alkalinity of the soil, which affects the existence form and biological effectiveness of cadmium in the soil. Organic matter content refers to the total amount of organic matter contained in the soil. Organic matter has an adsorption effect on cadmium and can affect the migration and accumulation of cadmium in the soil. Cation exchange capacity refers to the ability of soil colloids to adsorb cations. It is closely related to the soil's fertilizer retention capacity and pollutant adsorption capacity. Heavy metal ions such as cadmium can be adsorbed by soil colloids through cation exchange. Soil moisture content refers to the mass fraction of water in the soil. It affects the physical, chemical and biological properties of the soil, and thus affects the migration and transformation of cadmium.

[0060] Soil property indicators play a key role in the behavior of cadmium in the soil. Changes in pH value will change the solubility and chemical form of cadmium. For example, under acidic conditions, the solubility of cadmium is higher and it is more easily absorbed by plants, thereby increasing the safety risk of agricultural products. Soils with high organic matter content can adsorb more cadmium, reducing its mobility and bioavailability in the soil. Soils with high cation exchange capacity have a strong adsorption capacity for cadmium, which helps to reduce the migration of cadmium. The moisture content of the soil affects the diffusion rate and chemical reaction rate of cadmium in the soil. For example, when the moisture content is high, cadmium may be more likely to migrate with water.

[0061] By monitoring and analyzing these soil property indicators, we can gain a deeper understanding of the soil's ability to adsorb, desorb, and migrate cadmium, providing an important basis for predicting the distribution of cadmium content in soil. For example, in acidic soils with low pH values, cadmium has a higher bioavailability, necessitating special attention to the quality and safety of agricultural products. In soils with high organic matter content, cadmium has lower mobility and a relatively lower risk of contamination spread.

[0062] S2. Calculate characteristic factors based on historical data of multi-dimensional indicators.

[0063] Among them, characteristic factors include pollution source attenuation factor, dynamic migration factor and soil interaction factor.

[0064] In some embodiments, the pollution source attenuation factor is calculated based on historical data of pollution source indicators, and the calculation formula is as follows:

[0065] ;

[0066] in, represents the pollution source attenuation factor, m represents the mth type of pollution source enterprise, M represents the total number of pollution source enterprises, d m represents the average distance to the mth type of pollution source enterprise, λ m represents the diffusion rate of the mth type of pollution source enterprise, ω m represents the pollution weight of the mth type of pollution source enterprise, C PM2.5 represents the PM2.5 content in the air, α represents the pollution weight of PM2.5 content, L night represents the night light intensity, and β represents the pollution weight of the night light intensity.

[0067] Based on the fundamental principle of pollutant diffusion, a pollution source attenuation factor was constructed. For each pollution source enterprise, its impact on the cadmium content in the surrounding soil gradually decreases with increasing distance. Different pollution source enterprises have different diffusion rates and pollution weights due to their production processes and pollutant emission intensities. PM2.5 levels and nighttime light intensity were used as additional pollution source indicators to further improve the assessment of pollution source impacts.

[0068] In some embodiments, the dynamic migration factor is calculated based on historical data of migration process indicators, and the calculation formula is as follows:

[0069] ;

[0070] in, represents the dynamic migration factor, γ, δ, and ϵ represent the influence weights of hydraulic migration, atmospheric diffusion, and vegetation retardation, respectively. 降水 Indicates precipitation, D 河流 Indicates river density, DEM T represents the terrain driving factor, V 风速 represents wind speed, T represents temperature, T base represents the base temperature, NPP represents the net primary productivity of vegetation, R 土壤 represents the soil erosion rate, and R represents the erosion threshold.

[0071] Each term in the formula corresponds to a different migration mechanism. The hydraulic migration term Taking into account the role of precipitation and rivers in the transport of cadmium, the greater the precipitation and the higher the river density, the stronger the ability of hydraulic migration. The topographic driving factor reflects the influence of topography on the speed and direction of water flow, which in turn affects the migration of cadmium. The effects of wind speed and temperature on the diffusion of cadmium in the atmosphere are taken into account. The greater the wind speed and the greater the difference between the temperature and the reference temperature, the stronger the atmospheric diffusion effect. Taking into account the blocking effect of vegetation on cadmium migration, the higher the net primary productivity of vegetation, the stronger its blocking ability on cadmium. When the soil erosion rate exceeds the erosion threshold, the blocking effect of vegetation will weaken.

[0072] The calculation of the dynamic migration factor can comprehensively reflect the migration ability of cadmium in the environment, enabling the model to capture the impact of the input and output processes of soil Cd elements on the soil Cd content and the potential nonlinear relationship, providing key information for predicting the distribution of soil cadmium content.

[0073] Furthermore, the calculation formula of the terrain driving factor is as follows:

[0074] ;

[0075] Among them, S norm represents the normalized slope factor, FA log represents the logarithmic runoff accumulation, C curv represents the normalized curvature, and λ represents the curvature contribution coefficient.

[0076] Among them, the normalized slope factor S norm That is, the ratio of the actual slope to the maximum slope in the area, the logarithmic runoff accumulation FA logThe normalized curvature C is obtained by calculating the runoff accumulation and taking the natural logarithm. curv It means that the curvature is standardized to a mean of 0 and a standard deviation of 1. λ represents the curvature contribution coefficient, that is, the regulation intensity of curvature on the terrain driving factor, which needs to be calibrated according to the regional terrain characteristics.

[0077] The terrain driving factor is obtained by integrating terrain factors such as slope, runoff accumulation and curvature in the digital elevation model. The normalized slope factor and logarithmic runoff accumulation reflect the impact of terrain on water flow from the perspective of slope and runoff respectively, while the standardized curvature takes into account the changes in the direction and speed of water flow caused by the curvature of the terrain. The curvature contribution coefficient is used to adjust the degree of influence of curvature on the terrain driving factor, so that the factor can more accurately reflect the effect of actual terrain on pollutant migration. The introduction of terrain driving factors can more accurately analyze the impact of terrain on cadmium migration. For example, in areas with large slopes, high runoff accumulation and large curvature, the water flow speed is fast and the path is complex. Cadmium may be more likely to migrate to farther places with the water flow, thereby increasing the risk of cadmium contamination in the soil of the area. By accurately calculating the terrain driving factor, a more reliable basis can be provided for the calculation of the dynamic migration factor, thereby improving the accuracy of the entire prediction model.

[0078] In some embodiments, the soil interaction factor is calculated based on historical data of soil property indicators using the following formula:

[0079] ;

[0080] in, represents the soil interaction factor, pH represents the pH value, k1 represents the nonlinear effect coefficient of the pH value, exemplarily, k1 can be set to 1.5, Y represents the organic matter content, η represents the adjustment coefficient, CEC represents the cation exchange capacity, lp represents the clay content, W represents the soil moisture content, and k2 represents the moisture influence coefficient, exemplarily, k2 can be set to 0.3.

[0081] Considering the mechanisms by which different soil properties influence cadmium, the nonlinear effect coefficient of pH accounts for the nonlinear influence of pH on cadmium's form and adsorption / desorption processes. Organic matter content (Y) is calculated after natural logarithmic transformation, reflecting the contribution of organic matter to cadmium adsorption capacity. Cation exchange capacity (CEC) and clay content (lp) are closely related to the soil's cadmium adsorption capacity; their product represents the comprehensive cadmium adsorption capacity of soil colloids. The power function term for soil moisture content (W) accounts for the effect of moisture content on cadmium migration, and the moisture effect coefficient adjusts the degree of influence of moisture content on soil interaction factors.

[0082] S3. Construct a data set based on historical data of multi-dimensional indicators, characteristic factors, and soil cadmium content data at known locations.

[0083] Specifically, the original monitoring data is combined with characteristic factors processed by domain knowledge (pollution source attenuation factor, dynamic migration factor, soil interaction factor) to form a complete data chain that includes pollution sources, migration process, soil properties and cadmium content results, breaking through the limitations of traditional single-dimensional data and providing more comprehensive input information for the model.

[0084] S4. Use the dataset to train the random forest model to obtain a prediction model for the distribution of cadmium content in cultivated soil.

[0085] In some embodiments, Figure 3 This is a schematic diagram of a random forest model provided by an embodiment of the present invention. The random forest model is trained using a data set to obtain a prediction model for the distribution of cadmium content in cultivated soil. Figure 4 This is a diagram showing the comparison of predicted values ​​and observed values ​​of a prediction model for cadmium content distribution in cultivated soil and an analysis of the prediction interval provided by the embodiment of the present invention. Figure 4 The data show the performance of the cultivated soil cadmium content distribution prediction model trained according to the present invention on soil Cd content in both the training and test sets, demonstrating the model's ability to fit and generalize across different datasets. In the training set, the model's predicted values ​​(dashed red line) closely match the observed values ​​(solid black line), demonstrating the model's excellent performance on the training data and its ability to accurately capture the changing trends in Cd content. The model's predictions are particularly accurate in areas with low Cd content. This indicates that the model has effectively learned the patterns in the training data and can accurately reflect fluctuations in Cd content within the samples. At several key peaks (such as the 25th and 75th samples), the model also reliably predicts these high-content areas, demonstrating its high adaptability to the training data. Furthermore, most predicted values ​​fall within the 90% prediction interval, further demonstrating the model's stability and reliability on the training set. The cultivated soil cadmium content distribution prediction model trained according to the present invention achieved a goodness-of-fit R² of 0.96 for the training data and 0.81 for the test data, demonstrating the model's strong predictive ability on the test data.

[0086] S5. Obtain multi-dimensional indicator data of the target time series, calculate characteristic factors, and input them into the cultivated land soil cadmium content distribution prediction model to obtain the cultivated land soil cadmium content distribution prediction results.

[0087] In some embodiments, Figure 5 This is a schematic diagram of a model for predicting soil cadmium content distribution in cultivated land provided by an embodiment of the present invention, for predicting soil cadmium content and mapping the soil cadmium content. Figure 6 This is a schematic diagram of a Kriging interpolation model for predicting soil cadmium content mapping provided by an embodiment of the present invention, see Figure 5 and Figure 6, showing the comparison of the prediction results of soil Cd content by the cultivated land soil cadmium content distribution prediction model and the Kriging interpolation model provided by the embodiment of the present invention. Figure 5 The prediction results based on the cadmium content distribution prediction model of cultivated land soil show that the spatial distribution of soil Cd content has significant heterogeneity, and the high-concentration area is concentrated in the north-central position of the map, showing an obvious hot spot area, with the soil Cd content as high as 9.5 mg / kg. This concentration reflects that the cadmium content distribution prediction model of cultivated land soil can capture the local characteristics of soil Cd content, which may be highly correlated with specific pollution sources or regional factors, such as the influence of industrial activities or mining areas. In addition, the low-concentration areas in the map are widely distributed, indicating that the degree of cadmium Cd pollution in other areas is relatively light, and the model can distinguish between high and low pollution areas well. In contrast, Figure 6 The Kriging interpolation model in the presents a smoother distribution of soil Cd content. Figure 5 The concentration areas of Cd are roughly the same, but the boundaries of the concentration distribution are more blurred, and the high-Cd concentration areas show a more diffuse trend. This smoothing process highlights the Kriging interpolation model's focus on spatial autocorrelation when processing spatial data, emphasizing the continuity of Cd concentrations across regions. Unlike the prediction model for cadmium content distribution in cultivated soil, the Kriging model tends to weaken the extreme values ​​of local hotspots and instead makes the pollution appear more uniform across space, resulting in a slightly larger expansion of the high-concentration area but a gentler concentration gradient.

[0088] The present invention considers the influencing factors of soil cadmium content from multiple dimensions, combines multidimensional indicators, and introduces characteristic factors, so that the model can capture the impact of the input and output processes of soil cadmium elements on soil cadmium content, potential nonlinear relationships, and more realistically simulate the change in the intensity of the impact of pollution sources on the target area with distance, thereby improving the accuracy of prediction; the pollution source attenuation factor integrates distance, atmospheric deposition and human activities, the dynamic migration factor couples meteorological-topographic-ecological multiple processes, and the soil interaction factor quantifies the chemical synergistic effect, solving the characterization limitations of a single feature; by integrating the three characteristic factors constructed by environmental mechanisms, a deep integration of domain knowledge-driven and data-driven is achieved, which improves the dynamic adaptability and prediction accuracy of the model, provides a reliable tool for predicting the distribution of heavy metals in soil, and provides directly operational decision-making indicators for environmental management.

[0089] The embodiment of the present invention also provides a system for predicting the distribution of cadmium content in cultivated soil based on machine learning, which is used to execute the above-mentioned method for predicting the distribution of cadmium content in cultivated soil based on machine learning. Figure 7 This is a schematic diagram of a system for predicting cadmium content distribution in cultivated soil based on machine learning provided by an embodiment of the present invention. Figure 7 , the system includes the following modules:

[0090] The data acquisition module is used to obtain historical data on multi-dimensional indicators of cultivated land in the target area and soil cadmium content data at known locations; the multi-dimensional indicators include pollution source indicators, migration process indicators, and soil property indicators;

[0091] A characteristic factor calculation module is connected to the data acquisition module and is used to calculate characteristic factors based on historical data of multi-dimensional indicators; wherein the characteristic factors include pollution source attenuation factor, dynamic migration factor and soil interaction factor;

[0092] The data set construction module is connected to the data acquisition module and the characteristic factor calculation module, and is used to construct a data set based on the historical data of multi-dimensional indicators, characteristic factors, and soil cadmium content data at known locations;

[0093] The model building module is connected to the dataset building module and is used to train the random forest model using the dataset to obtain a prediction model for the distribution of cadmium content in cultivated soil;

[0094] The output module is connected to the model building module and is used to obtain multi-dimensional indicator data of the target time series, calculate characteristic factors, and input them into the cultivated land soil cadmium content distribution prediction model to obtain the cultivated land soil cadmium content distribution prediction results.

[0095] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the technical solutions of the embodiments of the present invention.

Claims

1. A method for predicting the distribution of cadmium content in cultivated soil based on machine learning, characterized in that: The method comprises the following steps: S1. Obtain historical data on multi-dimensional indicators of cultivated land in the target area and soil cadmium content data at known locations; multi-dimensional indicators include pollution source indicators, migration process indicators, and soil property indicators; S2. Calculate characteristic factors based on historical data of multi-dimensional indicators; the characteristic factors include pollution source attenuation factor, dynamic migration factor, and soil interaction factor; The calculation formula for the pollution source attenuation factor is as follows: ; in, represents the pollution source attenuation factor, m represents the mth type of pollution source enterprise, M represents the total number of pollution source enterprises, d m represents the average distance to the mth type of pollution source enterprise, λ m represents the diffusion rate of the mth type of pollution source enterprise, ω m represents the pollution weight of the mth type of pollution source enterprise, C PM2.5 represents the PM2.5 content in the air, α represents the pollution weight of PM2.5 content, L night represents the nighttime light intensity, and β represents the pollution weight of the nighttime light intensity; The dynamic migration factor is calculated as follows: ; in, Represents dynamic migration factors, γ, δ, They represent the impact weights of hydraulic migration, atmospheric diffusion, and vegetation retardation, respectively. 降水 Indicates precipitation, D 河流 Indicates river density, DEM T represents the terrain driving factor, V 风速 represents wind speed, T represents temperature, T base represents the base temperature, NPP represents the net primary productivity of vegetation, R 土壤 represents the soil erosion rate, R represents the erosion threshold; The soil interaction factor calculation formula is as follows: ; in, represents soil interaction factor, pH represents pH value, k1 represents nonlinear effect coefficient of pH value, Y represents organic matter content, η represents adjustment coefficient, CEC represents cation exchange capacity, lp represents clay content, W represents soil moisture content, and k2 represents moisture influence coefficient; S3. Construct a data set based on historical data of multi-dimensional indicators, characteristic factors, and soil cadmium content data at known locations; S4. Use the dataset to train the random forest model to obtain a prediction model for the distribution of cadmium content in cultivated soil; S5. Obtain multi-dimensional indicator data of the target time series, calculate characteristic factors, and input them into the cultivated land soil cadmium content distribution prediction model to obtain the cultivated land soil cadmium content distribution prediction results.

2. The method for predicting the distribution of cadmium content in cultivated soil based on machine learning according to claim 1, wherein: In S1, the pollution source indicators include: PM2.5 content, average distance to non-ferrous metal smelting enterprises, average distance to non-ferrous metal mining and dressing enterprises, average distance to battery manufacturing enterprises, average distance to stone building material manufacturing enterprises, average distance to waste resource utilization enterprises and night light intensity.

3. The method for predicting the distribution of cadmium content in cultivated soil based on machine learning according to claim 1, wherein: In S1, the migration process indicators include: River density, road density, net primary production, soil erosion rate, wind speed, air temperature, air pressure, relative humidity, precipitation, and digital elevation models.

4. The method for predicting the distribution of cadmium content in cultivated soil based on machine learning according to claim 1, wherein: In S1, the soil property indicators include: pH, organic matter content, cation exchange capacity and soil moisture content.

5. The method for predicting the distribution of cadmium content in cultivated soil based on machine learning according to claim 1, wherein: The calculation formula of the terrain driving factor is as follows: ; Among them, S norm represents the normalized slope factor, FA log represents the logarithmic runoff accumulation, C curv represents the normalized curvature, and λ represents the curvature contribution coefficient.

6. A system for predicting the distribution of cadmium content in cultivated soil based on machine learning, for executing the method for predicting the distribution of cadmium content in cultivated soil based on machine learning according to any one of claims 1 to 5, characterized in that: The system includes the following modules: The data acquisition module is used to obtain historical data on multi-dimensional indicators of cultivated land in the target area and soil cadmium content data at known locations; the multi-dimensional indicators include pollution source indicators, migration process indicators, and soil property indicators; A characteristic factor calculation module, connected to the data acquisition module, is used to calculate characteristic factors based on historical data of multi-dimensional indicators; wherein the characteristic factors include pollution source attenuation factor, dynamic migration factor and soil interaction factor; A data set construction module, connected to the data acquisition module and the characteristic factor calculation module, is used to construct a data set based on the historical data of multi-dimensional indicators, characteristic factors and soil cadmium content data at known points; A model construction module, connected to the data set construction module, is used to train a random forest model using the data set to obtain a prediction model for the distribution of cadmium content in cultivated soil; The output module is connected to the model building module and is used to obtain multi-dimensional indicator data of the target time series, calculate characteristic factors, and input them into the cultivated land soil cadmium content distribution prediction model to obtain the cultivated land soil cadmium content distribution prediction results.

Citation Information

Patent Citations

  • Agricultural non-point source pollution risk diagnosis method suitable for irrigation area

    CN111882182A

  • Soil cadmium pollution risk identification method based on spatial information enhancement algorithm

    CN119721698A