Cultivated land soil cadmium content distribution prediction method and system based on machine learning
Through the machine learning-based prediction method of cadmium content distribution in cultivated soil, combined with multi-dimensional indicators and characteristic factors, a random forest model is constructed, which solves the problem of low prediction accuracy of soil heavy metals in the existing technology, achieves higher prediction accuracy and dynamic adaptability, and provides reliable decision support for environmental management.
Patent Information
- Application Number
- CN202510481124.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-17
AI Technical Summary
The prior art is difficult to capture the nonlinear interactions of complex environmental factors in the prediction of spatial distribution of heavy metals in soil, and lacks dynamic simulation of pollutant migration and transformation process, resulting in low prediction accuracy and difficult to meet the needs of high-precision pollution warning and prevention and control.
Using machine learning-based prediction method for cadmium content distribution in cultivated soil, we use the historical data of multi-dimensional indicators to calculate the pollution source attenuation factor, dynamic migration factor and soil interaction factor, and build a random forest model and train it to obtain the prediction model.
It improves the accuracy of soil cadmium content distribution prediction, can more realistically simulate the impact of pollution sources on the target area, captures the nonlinear relationship in the input and output of soil cadmium elements, enhances the dynamic adaptability of the model, and provides direct operational decision-making indicators for environmental management.
Smart Images

Figure CN119989183A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of soil data analysis, and in particular relates to a method and system for predicting the distribution of cadmium (Cd) content in cultivated soil based on machine learning. Background Art
[0002] Spatial prediction of soil heavy metals is an important basis for understanding regional soil pollution and carrying out soil pollution prevention and control. Spatial prediction methods based on machine learning models are an important research direction for the spatial distribution prediction of soil heavy metals.
[0003] Existing technologies for predicting heavy metal distribution mainly rely on traditional statistical models and single environmental factor analysis. They usually build regression models based on limited monitoring data, or combine geographic information systems for spatial interpolation (such as Kriging interpolation). However, such methods are difficult to capture the nonlinear interactions of complex environmental factors. Some studies have introduced machine learning algorithms (such as random forests and neural networks) to optimize prediction accuracy, but they are mostly limited to static data-driven models and lack dynamic simulation of pollutant migration and transformation processes. Traditional methods generally have problems such as single data dimension, insufficient spatiotemporal dynamics, and weak mechanism explanatory power, making it difficult to meet the needs of high-precision pollution warning and prevention and control.
[0004] Therefore, there is an urgent need to develop a method and system for predicting the distribution of cadmium content in cultivated soil based on machine learning, which can consider the impact of the input and output process of soil cadmium elements on the soil cadmium content and improve the accuracy of model prediction results. Summary of the invention
[0005] In order to solve the above technical problems, the present invention provides a method and system for predicting the distribution of cadmium content in cultivated soil based on machine learning, which can consider the influence of the input and output process of soil cadmium elements on the soil cadmium content and improve the accuracy of model prediction results.
[0006] The present invention provides a method for predicting the distribution of cadmium content in cultivated soil based on machine learning, the method comprising the following steps: S1. Obtain historical data of multi-dimensional indicators of cultivated soil in the target area and soil cadmium content data at known locations; the multi-dimensional indicators include pollution source indicators, migration process indicators and soil property indicators; S2. Calculate characteristic factors based on historical data of multi-dimensional indicators; the characteristic factors include pollution source attenuation factor, dynamic migration factor and soil interaction factor; S3. Construct a data set based on historical data of multi-dimensional indicators, characteristic factors and soil cadmium content data at known locations; S4. Use the data set to train the random forest model to obtain a prediction model for the distribution of cadmium content in cultivated soil; S5. Obtain multi-dimensional indicator data of the target time series, calculate characteristic factors, and input them into the cultivated land soil cadmium content distribution prediction model to obtain the cultivated land soil cadmium content distribution prediction results.
[0007] Furthermore, in S1, the pollution source indicators include: PM2.5 content, average distance to non-ferrous metal smelting enterprises, average distance to non-ferrous metal mining and selection enterprises, average distance to battery manufacturing enterprises, average distance to stone building material manufacturing enterprises, average distance to waste resource utilization enterprises and night light intensity.
[0008] Furthermore, in S1, the migration process indicators include: River density, road density, net primary production, soil erosion rate, wind speed, air temperature, air pressure, relative humidity, precipitation, and digital elevation models.
[0009] Furthermore, in S1, soil property indicators include: pH, organic matter content, cation exchange capacity and soil moisture content.
[0010] Furthermore, in S2, the pollution source attenuation factor is calculated based on the historical data of the pollution source indicators. The calculation formula is as follows: ; in, represents the pollution source attenuation factor, m represents the mth type of pollution source enterprise, M represents the total number of pollution source enterprises, d m represents the average distance to the mth type of pollution source enterprise, λ m represents the diffusion rate of the mth type of pollution source enterprise, ω m represents the pollution weight of the mth type of pollution source enterprise, C PM2.5 represents the PM2.5 content in the air, α represents the pollution weight of PM2.5 content, L night represents the night light intensity, and β represents the pollution weight of the night light intensity.
[0011] Furthermore, in S2, the dynamic migration factor is calculated based on the historical data of the migration process indicators. The calculation formula is as follows: ; in, represents the dynamic migration factor, γ, δ, and ϵ represent the influence weights of hydraulic migration, atmospheric diffusion, and vegetation retardation, respectively. 降水 Indicates precipitation, D 河流 Indicates river density, DEM T represents the terrain driving factor, V 风速 represents wind speed, T represents temperature, T baserepresents the base temperature, NPP represents the net primary productivity of vegetation, R 土壤 represents the soil erosion rate, and R represents the erosion threshold.
[0012] Furthermore, the calculation formula of terrain driving factor is as follows: ; Among them, S norm represents the normalized slope factor, FA log represents the logarithmic runoff accumulation, C curv represents the normalized curvature, and λ represents the curvature contribution coefficient.
[0013] Furthermore, in S2, the soil interaction factor is calculated based on the historical data of soil property indicators, and the calculation formula is as follows: ; in, represents soil interaction factor, pH represents pH value, k1 represents nonlinear effect coefficient of pH value, Y represents organic matter content, η represents adjustment coefficient, CEC represents cation exchange capacity, lp represents clay content, W represents soil moisture content, and k2 represents moisture influence coefficient.
[0014] The present invention also provides a system for predicting the distribution of cadmium content in cultivated soil based on machine learning, which is used to execute the above-mentioned method for predicting the distribution of cadmium content in cultivated soil based on machine learning, and is characterized in that the system includes the following modules: The data acquisition module is used to obtain historical data of multi-dimensional indicators of cultivated soil in the target area and soil cadmium content data at known points; the multi-dimensional indicators include pollution source indicators, migration process indicators and soil property indicators; A characteristic factor calculation module is connected to the data acquisition module and is used to calculate characteristic factors based on historical data of multi-dimensional indicators; wherein the characteristic factors include pollution source attenuation factors, dynamic migration factors and soil interaction factors; A data set construction module, connected with the data acquisition module and the characteristic factor calculation module, is used to construct a data set based on historical data of multi-dimensional indicators, characteristic factors and soil cadmium content data at known points; A model building module, connected to the data set building module, is used to train the random forest model using the data set to obtain a prediction model for the distribution of cadmium content in cultivated soil; The output module is connected to the model building module and is used to obtain multi-dimensional indicator data of the target time series, calculate characteristic factors, and input them into the cultivated soil cadmium content distribution prediction model to obtain the cultivated soil cadmium content distribution prediction results.
[0015] The embodiments of the present invention have the following technical effects: The present invention considers the influencing factors of soil cadmium content from multiple dimensions, combines multi-dimensional indicators, and introduces characteristic factors, so that the model can capture the impact of soil cadmium element input and output processes on soil cadmium content, potential nonlinear relationships, and more realistically simulate the change in the intensity of the impact of pollution sources on the target area with distance, thereby improving the accuracy of prediction; the pollution source attenuation factor integrates distance, atmospheric deposition and human activities, the dynamic migration factor couples meteorological-topographic-ecological multiple processes, and the soil interaction factor quantifies the chemical synergistic effect, solving the characterization limitations of a single feature; by integrating the three characteristic factors constructed by the environmental mechanism, a deep integration of domain knowledge-driven and data-driven is achieved, the dynamic adaptability and prediction accuracy of the model are improved, a reliable tool is provided for the prediction of soil heavy metal distribution, and a directly operable decision-making indicator is provided for environmental management. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0017] Figure 1 It is a flow chart of a method for predicting the distribution of cadmium content in cultivated soil based on machine learning provided by an embodiment of the present invention; Figure 2 is a schematic diagram of a digital soil mapping study area provided by an embodiment of the present invention; Figure 3 is a schematic diagram of a random forest model provided by an embodiment of the present invention; Figure 4 It is an analysis diagram of a comparison between predicted values and observed values of a prediction model for cadmium content distribution in cultivated soil and a prediction interval provided by an embodiment of the present invention; Figure 5 It is a schematic diagram of a cultivated land soil cadmium content distribution prediction model provided by an embodiment of the present invention for predicting soil cadmium content and mapping; Figure 6 It is a schematic diagram of a Kriging interpolation model for predicting soil cadmium content mapping provided by an embodiment of the present invention; Figure 7 It is a structural schematic diagram of a system for predicting the distribution of cadmium content in cultivated soil based on machine learning provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0018] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be described clearly and completely below. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work belong to the scope of protection of the present invention.
[0019] The embodiment of the present invention provides a method for predicting the distribution of cadmium content in cultivated soil based on machine learning. Figure 1 is a flow chart of a method for predicting the distribution of cadmium content in cultivated soil based on machine learning provided by an embodiment of the present invention. Figure 2 is a schematic diagram of a digital soil mapping study area provided by an embodiment of the present invention, see Figure 1 and Figure 2 , the method comprises the following steps: S1. Obtain historical data on multi-dimensional indicators of cultivated soil in the target area and soil cadmium content data at known locations.
[0020] The data used in this study mainly include soil cadmium (Cd) content data and multi-dimensional indicator data for soil Cd content prediction. The sampling process follows HJ / T 166-2004 "Technical Specifications for Soil Environmental Monitoring", and each sample is composed of 5 sub-samples collected by the five-point sampling method according to grid sampling. All soil samples were naturally air-dried indoors, crushed and ground, and then passed through a nylon sieve (sieve aperture 0.25-0.4mm), mixed evenly, and finally the content of soil heavy metal Cd was determined by atomic absorption spectrophotometry.
[0021] Among them, the multi-dimensional indicators include pollution source indicators, migration process indicators and soil property indicators. The variables included in each indicator selected in this embodiment are shown in Table 1.
[0022] Table 1 Multi-dimensional index data for spatial distribution prediction of Cd content in regional cultivated soil
[0023] Among the pollution source indicators: PM2.5 content refers to the concentration of particles with a diameter less than or equal to 2.5 microns in the air. These particles may carry heavy metals such as cadmium and pollute the soil. The average distance from various pollution source enterprises reflects the possibility that cultivated land will be affected by pollution from such enterprises. The closer the distance, the higher the risk of pollution. The intensity of night lights can, to a certain extent, represent the intensity of human activities. The more frequent human activities are in an area, the more pollution sources may be generated, and the greater the impact on the cadmium content in the soil may be.
[0024] These indicators are selected as pollution source indicators because non-ferrous metal smelting, mining and selection enterprises, battery manufacturing enterprises, stone building material manufacturing enterprises and waste resource utilization enterprises may produce cadmium-containing pollutants in the production process, which may pollute the surrounding cultivated land soil through atmospheric diffusion, wastewater discharge and other channels. PM2.5, as an important component of atmospheric pollutants, can be transmitted over long distances and settle in the soil, increasing the cadmium content in the soil. The intensity of night lights is closely related to human activities. High-intensity human activities are often accompanied by more industrial production, transportation and other activities, which may lead to more cadmium pollution emissions.
[0025] By comprehensively considering these pollution source indicators, the potential risk of cadmium contamination in the cultivated land soil of the target area can be more accurately assessed. For example, the closer the average distance to the non-ferrous metal smelting enterprise, the greater the possibility that the cultivated land in the area is contaminated by cadmium from the enterprise; the higher the PM2.5 content, the more cadmium may enter the soil through atmospheric deposition. Combining these indicators can provide rich and accurate pollution source-related information for subsequent characteristic factor calculations and model predictions, and improve the reliability of prediction results.
[0026] Among them, in the migration process indicators: river density refers to the length of the river per unit area, which reflects the hydrological characteristics of the region. As an important carrier of pollutant migration, the density of the river will affect the migration and diffusion of cadmium in the soil. Road density reflects the degree of development of the regional transportation network. Traffic activities may bring cadmium pollution, and roads may also affect the transmission path of pollutants. Net primary production refers to the residual amount of organic carbon fixed by plants through photosynthesis per unit time and per unit area after deducting their own respiratory consumption. It reflects the growth status of vegetation. Vegetation has a certain absorption and blocking effect on pollutants such as cadmium. Soil erosion rate indicates the speed at which soil is eroded. Soil erosion will cause cadmium to spread to other areas with the migration of soil particles. Meteorological factors such as wind speed, temperature, air pressure, relative humidity and precipitation will affect the diffusion and transmission of pollutants in the atmosphere. Digital elevation model is a digital representation of terrain that can reflect the terrain undulation of the region. The terrain will affect the flow direction of water and pollutants.
[0027] The selection of migration process indicators is based on the migration and transformation laws of cadmium in the environment. As channels for material transmission, the density of rivers and roads will affect the migration range and speed of cadmium. Meteorological factors affect the transmission distance and deposition distribution of cadmium in the atmosphere by affecting the movement and diffusion capacity of the atmosphere. The net primary production of vegetation and the soil erosion rate are closely related to the migration and accumulation of cadmium in the soil-vegetation system. Digital elevation models can help analyze the impact of terrain on cadmium migration. For example, in areas with large terrain undulations, the water flow rate may be faster, thereby accelerating the migration of cadmium.
[0028] Considering these migration process indicators can provide a more comprehensive understanding of the migration behavior of cadmium in the environment. For example, in areas with high river density, cadmium may be more likely to migrate to other places through surface runoff; in areas with high wind speeds, cadmium diffuses more widely in the atmosphere. By analyzing these indicators, a more accurate basis can be provided for predicting the distribution of cadmium in the soil, which will help to formulate more targeted pollution prevention and control measures.
[0029] Among them, among the soil property indicators: pH value is an important indicator to measure the acidity and alkalinity of the soil, which will affect the existence form and biological effectiveness of cadmium in the soil. Organic matter content refers to the total amount of organic matter contained in the soil. Organic matter has an adsorption effect on cadmium and can affect the migration and accumulation of cadmium in the soil. Cation exchange capacity refers to the ability of soil colloids to adsorb cations. It is closely related to the soil's fertilizer retention capacity and pollutant adsorption capacity. Heavy metal ions such as cadmium can be adsorbed by soil colloids through cation exchange. Soil moisture content refers to the mass fraction of water in the soil. It affects the physical, chemical and biological properties of the soil, and thus affects the migration and transformation of cadmium.
[0030] Soil property indicators play a key role in the behavior of cadmium in the soil. Changes in pH value will change the solubility and chemical form of cadmium. For example, under acidic conditions, the solubility of cadmium is higher and it is more easily absorbed by plants, thereby increasing the safety risk of agricultural products. Soils with high organic matter content can adsorb more cadmium, reducing its mobility and bioavailability in the soil. Soils with high cation exchange capacity have a strong adsorption capacity for cadmium, which helps to reduce the migration of cadmium. The moisture content of the soil affects the diffusion rate and chemical reaction rate of cadmium in the soil. For example, when the moisture content is high, cadmium may be more likely to migrate with water.
[0031] By monitoring and analyzing these soil property indicators, we can gain a deeper understanding of the soil's ability to adsorb, desorb and migrate cadmium, providing an important basis for predicting the distribution of soil cadmium content. For example, in acidic soils with low pH values, cadmium has a higher bioavailability, and special attention should be paid to the quality and safety of agricultural products; in soils with high organic matter content, cadmium has a lower mobility and the risk of pollution spread is relatively small.
[0032] S2. Calculate characteristic factors based on historical data of multi-dimensional indicators.
[0033] Among them, the characteristic factors include pollution source attenuation factor, dynamic migration factor and soil interaction factor.
[0034] In some embodiments, the pollution source attenuation factor is calculated based on historical data of pollution source indicators, and the calculation formula is as follows: ; in, represents the pollution source attenuation factor, m represents the mth type of pollution source enterprise, M represents the total number of pollution source enterprises, d m represents the average distance to the mth type of pollution source enterprise, λ m represents the diffusion rate of the mth type of pollution source enterprise, ω m represents the pollution weight of the mth type of pollution source enterprise, C PM2.5 represents the PM2.5 content in the air, α represents the pollution weight of PM2.5 content, L night represents the night light intensity, and β represents the pollution weight of the night light intensity.
[0035] Based on the basic principle of pollutant diffusion, the pollution source attenuation factor is constructed. For each pollution source enterprise, as the distance increases, its impact on the cadmium content in the surrounding soil will gradually weaken. Different types of pollution source enterprises have different diffusion rates and pollution weights due to their different production processes, pollutant emission intensities, etc. PM2.5 content and night light intensity are used as additional pollution source indicators to further improve the assessment of the impact of pollution sources.
[0036] In some embodiments, the dynamic migration factor is calculated based on historical data of the migration process indicator, and the calculation formula is as follows: ; in, represents the dynamic migration factor, γ, δ, and ϵ represent the influence weights of hydraulic migration, atmospheric diffusion, and vegetation retardation, respectively. 降水 Indicates precipitation, D 河流 Indicates river density, DEM T represents the terrain driving factor, V 风速 represents wind speed, T represents temperature, T base represents the base temperature, NPP represents the net primary productivity of vegetation, R 土壤 represents the soil erosion rate, and R represents the erosion threshold.
[0037] Each term in the formula corresponds to a different migration mechanism. The hydraulic migration term Taking into account the role of precipitation and rivers in the transport of cadmium, the greater the precipitation and the higher the river density, the stronger the ability of hydraulic migration. The topographic driving factor reflects the influence of topography on the speed and direction of water flow, which in turn affects the migration of cadmium. The effects of wind speed and temperature on the diffusion of cadmium in the atmosphere are taken into account. The greater the wind speed and the greater the difference between the temperature and the reference temperature, the stronger the atmospheric diffusion effect. Taking into account the blocking effect of vegetation on cadmium migration, the higher the net primary productivity of vegetation, the stronger its blocking ability on cadmium. When the soil erosion rate exceeds the erosion threshold, the blocking effect of vegetation will weaken.
[0038] The calculation of the dynamic migration factor can comprehensively reflect the migration capacity of cadmium in the environment, enabling the model to capture the impact of the input and output process of soil Cd elements on the soil Cd content and the potential nonlinear relationship, providing key information for predicting the distribution of soil cadmium content.
[0039] Furthermore, the calculation formula of terrain driving factor is as follows: ; Among them, S norm represents the normalized slope factor, FA log represents the logarithmic runoff accumulation, C curv represents the normalized curvature, and λ represents the curvature contribution coefficient.
[0040] Among them, the normalized slope factor S norm That is, the ratio of the actual slope to the maximum slope in the area, the logarithmic runoff accumulation FA log The normalized curvature C is obtained by calculating the runoff accumulation and taking the natural logarithm. curv It means that the curvature is standardized to a mean of 0 and a standard deviation of 1. λ represents the curvature contribution coefficient, that is, the regulation intensity of curvature on the terrain driving factor, which needs to be calibrated according to the regional terrain characteristics.
[0041] The terrain driving factor is obtained by integrating terrain factors such as slope, runoff accumulation and curvature in the digital elevation model. The normalized slope factor and logarithmic runoff accumulation reflect the impact of terrain on water flow from the perspective of slope and runoff, respectively, while the standardized curvature takes into account the change of water flow direction and speed caused by terrain curvature. The curvature contribution coefficient is used to adjust the degree of influence of curvature on the terrain driving factor, so that the factor can more accurately reflect the effect of actual terrain on pollutant migration. The introduction of terrain driving factors can more accurately analyze the impact of terrain on cadmium migration. For example, in areas with large slopes, high runoff accumulation and large curvature, the water flow speed is fast and the path is complex. Cadmium may be more likely to migrate to farther places with the water flow, thereby increasing the risk of soil cadmium pollution in the area. By accurately calculating the terrain driving factor, a more reliable basis can be provided for the calculation of the dynamic migration factor, thereby improving the accuracy of the entire prediction model.
[0042] In some embodiments, the soil interaction factor is calculated based on historical data of soil property indicators, and the calculation formula is as follows: ; in, represents the soil interaction factor, pH represents the pH value, k1 represents the nonlinear effect coefficient of the pH value, and k1 can be set to 1.5 by way of example, Y represents the organic matter content, η represents the adjustment coefficient, CEC represents the cation exchange capacity, lp represents the clay content, W represents the soil moisture content, and k2 represents the moisture influence coefficient, and k2 can be set to 0.3 by way of example.
[0043] Considering the influence mechanism of different soil properties on cadmium, the nonlinear effect coefficient of pH value takes into account the nonlinear effect of pH value on the existence form and adsorption and desorption process of cadmium. The organic matter content Y is calculated after natural logarithm transformation, reflecting the contribution of organic matter to the adsorption capacity of cadmium. The cation exchange capacity CEC and clay content lp are closely related to the adsorption capacity of soil for cadmium, and their product represents the comprehensive adsorption capacity of soil colloids for cadmium. The power function term of soil moisture content W takes into account the effect of moisture content on cadmium migration, and the moisture influence coefficient is used to adjust the degree of effect of moisture content on soil interaction factors.
[0044] S3. Construct a data set based on historical data of multi-dimensional indicators, characteristic factors and soil cadmium content data at known locations.
[0045] Specifically, the original monitoring data is combined with characteristic factors processed by domain knowledge (pollution source attenuation factor, dynamic migration factor, soil interaction factor) to form a complete data chain that includes pollution sources, migration process, soil properties and cadmium content results, breaking through the limitations of traditional single-dimensional data and providing more comprehensive input information for the model.
[0046] S4. Use the dataset to train the random forest model and obtain a prediction model for the distribution of cadmium content in cultivated soil.
[0047] In some embodiments, Figure 3 This is a schematic diagram of a random forest model provided by an embodiment of the present invention. The random forest model is trained using a data set to obtain a prediction model for the distribution of cadmium content in cultivated soil. Figure 4 This is a comparison chart of the predicted value and observed value of a prediction model for cadmium content distribution in cultivated soil and an analysis chart of the prediction interval provided by the embodiment of the present invention, see Figure 4, showing the prediction performance of the cultivated soil cadmium content distribution prediction model trained in the embodiment of the present invention on the soil Cd content in the training set and the test set, reflecting the model's fitting and generalization capabilities on different data sets. In the training set part, the model's predicted value (red dotted line) is basically consistent with the actual observed value (black solid line), indicating that the model performs well in processing training data and can accurately capture the changing trend of Cd content, especially in low-content areas, where the model's prediction is very accurate. This shows that the model has learned the patterns in the training data very well and can better reflect the fluctuations in Cd content in the samples. At several key peak points (such as the 25th and 75th samples), the model can also predict these high-content areas well, showing a high degree of adaptability to the training data. At the same time, most of the predicted values are within the 90% prediction interval, further demonstrating the stability and reliability of the model on the training set. The cultivated soil cadmium content distribution prediction model trained in the embodiment of the present invention has a goodness of fit R² of 0.96 for the training data and a goodness of fit R² of 0.81 for the test data, indicating that the model has good prediction capabilities for the test data.
[0048] S5. Obtain multi-dimensional indicator data of the target time series, calculate characteristic factors, and input them into the cultivated land soil cadmium content distribution prediction model to obtain the cultivated land soil cadmium content distribution prediction results.
[0049] In some embodiments, Figure 5 is a schematic diagram of a model for predicting soil cadmium content distribution in cultivated land provided by an embodiment of the present invention for predicting soil cadmium content and mapping. Figure 6 is a schematic diagram of a Kriging interpolation model for predicting soil cadmium content mapping provided by an embodiment of the present invention, see Figure 5 and Figure 6 , showing the comparison of the prediction results of soil Cd content by the cultivated land soil cadmium content distribution prediction model and the Kriging interpolation model provided in the embodiment of the present invention. Figure 5 The prediction results based on the prediction model for the distribution of cadmium content in cultivated soil show that the spatial distribution of soil Cd content has significant heterogeneity, and the high-concentration area is concentrated in the north-central position of the map, showing an obvious hot spot area, with the highest soil Cd content reaching 9.5 mg / kg. This concentration reflects that the prediction model for the distribution of cadmium content in cultivated soil is able to capture the local characteristics of soil Cd content, which may be highly correlated with specific pollution sources or regional factors, such as the influence of industrial activities or mining areas. In addition, the low-concentration areas in the figure are widely distributed, indicating that the degree of cadmium Cd pollution in other areas is relatively light, and the model can distinguish between high and low pollution areas well. In contrast, Figure 6 The Kriging interpolation model in the above shows a smoother distribution of soil Cd content. Figure 5The concentration areas of Cd are roughly the same, but the boundaries of the content distribution are more blurred, and the high-content areas of Cd show a more diffuse trend. This smoothing process highlights that the Kriging interpolation model focuses on spatial autocorrelation when processing spatial data, and emphasizes the continuity of Cd content between regions. Unlike the cadmium content distribution prediction model in cultivated soil, the Kriging model tends to weaken the extreme values of local hotspots and instead makes the pollution present a more uniform transition in space, resulting in a slightly larger expansion range of high-concentration areas, but a slower concentration gradient.
[0050] The present invention considers the influencing factors of soil cadmium content from multiple dimensions, combines multi-dimensional indicators, and introduces characteristic factors, so that the model can capture the impact of soil cadmium element input and output processes on soil cadmium content, potential nonlinear relationships, and more realistically simulate the change in the intensity of the impact of pollution sources on the target area with distance, thereby improving the accuracy of prediction; the pollution source attenuation factor integrates distance, atmospheric deposition and human activities, the dynamic migration factor couples meteorological-topographic-ecological multiple processes, and the soil interaction factor quantifies the chemical synergistic effect, solving the characterization limitations of a single feature; by integrating the three characteristic factors constructed by the environmental mechanism, a deep integration of domain knowledge-driven and data-driven is achieved, the dynamic adaptability and prediction accuracy of the model are improved, a reliable tool is provided for the prediction of soil heavy metal distribution, and a directly operable decision-making indicator is provided for environmental management.
[0051] The embodiment of the present invention also provides a system for predicting the distribution of cadmium content in cultivated soil based on machine learning, which is used to execute the above-mentioned method for predicting the distribution of cadmium content in cultivated soil based on machine learning. Figure 7 is a schematic diagram of a system for predicting the distribution of cadmium content in cultivated soil based on machine learning provided by an embodiment of the present invention, see Figure 7 , the system includes the following modules: The data acquisition module is used to obtain historical data of multi-dimensional indicators of cultivated soil in the target area and soil cadmium content data at known points; the multi-dimensional indicators include pollution source indicators, migration process indicators and soil property indicators; A characteristic factor calculation module is connected to the data acquisition module and is used to calculate characteristic factors based on historical data of multi-dimensional indicators; wherein the characteristic factors include pollution source attenuation factors, dynamic migration factors and soil interaction factors; A data set construction module, connected with the data acquisition module and the characteristic factor calculation module, is used to construct a data set based on historical data of multi-dimensional indicators, characteristic factors and soil cadmium content data at known points; A model building module, connected to the data set building module, is used to train the random forest model using the data set to obtain a prediction model for the distribution of cadmium content in cultivated soil; The output module is connected to the model building module and is used to obtain multi-dimensional indicator data of the target time series, calculate characteristic factors, and input them into the cultivated soil cadmium content distribution prediction model to obtain the cultivated soil cadmium content distribution prediction results.
[0052] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the technical solutions of the embodiments of the present invention.
Claims
1. A method for predicting the distribution of cadmium content in cultivated soil based on machine learning, characterized in that: The method comprises the following steps: S1. Obtain historical data of multi-dimensional indicators of cultivated soil in the target area and soil cadmium content data at known locations; the multi-dimensional indicators include pollution source indicators, migration process indicators and soil property indicators; S2. Calculate characteristic factors based on historical data of multi-dimensional indicators; the characteristic factors include pollution source attenuation factor, dynamic migration factor and soil interaction factor; S3. Construct a data set based on historical data of multi-dimensional indicators, characteristic factors and soil cadmium content data at known locations; S4. Use the data set to train the random forest model to obtain a prediction model for the distribution of cadmium content in cultivated soil; S5. Obtain multi-dimensional indicator data of the target time series, calculate characteristic factors, and input them into the cultivated land soil cadmium content distribution prediction model to obtain the cultivated land soil cadmium content distribution prediction results.
2. The method for predicting the distribution of cadmium content in cultivated soil based on machine learning according to claim 1, characterized in that: In S1, the pollution source indicators include: PM2.5 content, average distance to non-ferrous metal smelting enterprises, average distance to non-ferrous metal mining and selection enterprises, average distance to battery manufacturing enterprises, average distance to stone building material manufacturing enterprises, average distance to waste resource utilization enterprises and night light intensity.
3. The method for predicting the distribution of cadmium content in cultivated soil based on machine learning according to claim 1, characterized in that: In S1, the migration process indicators include: River density, road density, net primary production, soil erosion rate, wind speed, air temperature, air pressure, relative humidity, precipitation, and digital elevation models.
4. The method for predicting the distribution of cadmium content in cultivated soil based on machine learning according to claim 1, characterized in that: In S1, the soil property indicators include: pH, organic matter content, cation exchange capacity and soil moisture content.
5. The method for predicting the distribution of cadmium content in cultivated soil based on machine learning according to claim 2, characterized in that: In S2, the pollution source attenuation factor is calculated based on the historical data of the pollution source index. The calculation formula of the pollution source attenuation factor is as follows: ; in, represents the pollution source attenuation factor, m represents the mth type of pollution source enterprise, M represents the total number of pollution source enterprises, d m represents the average distance to the mth type of pollution source enterprise, λ m represents the diffusion rate of the mth type of pollution source enterprise, ω m represents the pollution weight of the mth type of pollution source enterprise, C PM2.5 represents the PM2.5 content in the air, α represents the pollution weight of PM2.5 content, L night represents the night light intensity, and β represents the pollution weight of the night light intensity.
6. The method for predicting the distribution of cadmium content in cultivated soil based on machine learning according to claim 3, characterized in that: In S2, the dynamic migration factor is calculated based on the historical data of the migration process indicator. The dynamic migration factor calculation formula is as follows: ; in, represents the dynamic migration factor, γ, δ, and ϵ represent the influence weights of hydraulic migration, atmospheric diffusion, and vegetation retardation, respectively. 降水 Indicates precipitation, D 河流 Indicates river density, DEM T represents the terrain driving factor, V 风速 represents wind speed, T represents temperature, T base represents the base temperature, NPP represents the net primary productivity of vegetation, R 土壤 represents the soil erosion rate, and R represents the erosion threshold.
7. The method for predicting the distribution of cadmium content in cultivated soil based on machine learning according to claim 6, characterized in that: The calculation formula of the terrain driving factor is as follows: ; Among them, S norm represents the normalized slope factor, FA log represents the logarithmic runoff accumulation, C curv represents the normalized curvature, and λ represents the curvature contribution coefficient.
8. The method for predicting the distribution of cadmium content in cultivated soil based on machine learning according to claim 4, characterized in that: In S2, the soil interaction factor is calculated based on the historical data of soil property indicators, and the soil interaction factor calculation formula is as follows: ; in, represents soil interaction factor, pH represents pH value, k1 represents nonlinear effect coefficient of pH value, Y represents organic matter content, η represents adjustment coefficient, CEC represents cation exchange capacity, lp represents clay content, W represents soil moisture content, and k2 represents moisture influence coefficient.
9. A system for predicting the distribution of cadmium content in cultivated soil based on machine learning, used to execute the method for predicting the distribution of cadmium content in cultivated soil based on machine learning as described in any one of claims 1 to 8, characterized in that: The system includes the following modules: The data acquisition module is used to obtain historical data of multi-dimensional indicators of cultivated soil in the target area and soil cadmium content data at known points; the multi-dimensional indicators include pollution source indicators, migration process indicators and soil property indicators; A characteristic factor calculation module, connected to the data acquisition module, is used to calculate characteristic factors based on historical data of multi-dimensional indicators; wherein the characteristic factors include pollution source attenuation factors, dynamic migration factors and soil interaction factors; A data set construction module, connected to the data acquisition module and the characteristic factor calculation module, for constructing a data set based on historical data of multi-dimensional indicators, characteristic factors and soil cadmium content data at known points; A model building module, connected to the data set building module, is used to train the random forest model using the data set to obtain a prediction model for the distribution of cadmium content in cultivated soil; The output module is connected to the model building module and is used to obtain multi-dimensional indicator data of the target time series, calculate characteristic factors, and input the cadmium content distribution prediction model of cultivated soil to obtain the prediction result of the cadmium content distribution of cultivated soil.
Citation Information
Patent Citations
Method for simulating heavy metal behaviours in drainage basin dynamically and quantitatively
CN105046043A
Agricultural non-point source pollution risk diagnosis method suitable for irrigation area
CN111882182A
Regional soil heavy metal pollution risk zoning and control method based on machine learning
CN117292768A
Soil cadmium pollution risk identification method based on spatial information enhancement algorithm
CN119721698A
Method of monitoring, evaluating and early-warning soil and groundwater pollution in industrial park
US20240135389A1
Cited By
Slope cropland ridge side cultivation segmented corrosion control and acid reduction and soil fertility improvement method
CN120814382A
Prediction method for effective-state cadmium and methane emission of soil under dry-wet alternation condition and related equipment
CN122045818A
Prediction method and related equipment for soil available cadmium and methane emission under wet-dry alternating conditions
CN122045818B