Soil heavy metal content prediction method for data sample sparse region
By adopting a fully connected neural network model in soil heavy metal content prediction, integrating spatial location and environmental variable information, and using data augmentation and transfer learning technology, the problem of data scarcity in soil heavy metal content prediction is solved, achieving higher prediction accuracy and model generalization ability.
Patent Information
- Application Number
- CN202510086212.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to effectively utilize scarce sampling data in soil heavy metal content prediction, resulting in limited prediction accuracy.
The fully connected neural network model is adopted, and the model is pre-trained and fine-tuned to improve prediction performance by fusing spatial location and multiple environment variable information as input features, and using data augmentation and transfer learning techniques.
It significantly improves the prediction accuracy of soil heavy metal content in sparse samples, solves the problem of insufficient marking data in target domains, and improves the generalization ability of the model.
Smart Images

Figure CN119990443A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of soil heavy metal pollution monitoring, and relates to a method for predicting soil heavy metal content in an area with sparse data samples. Background Art
[0002] Accurately predicting the continuous spatial distribution of regional soil heavy metal content is crucial for monitoring farmland pollution and ensuring the sustainable development of ecological agriculture. Traditional soil heavy metal pollution surveys mainly rely on on-site actual sampling, which is costly, inefficient, and limited by natural conditions, resulting in a very limited number of sampling points. How to predict the continuous spatial distribution of soil heavy metal content based on limited observation samples to better assist soil pollution investigations has become a long-term research hotspot in the field of ecological environment. At present, existing studies have focused on predicting the spatial distribution of soil heavy metals using spatial position relationships and attribute similarity relationships.
[0003] The method based on spatial position relationship is based on the first and second laws of geography. This type of method is based on the spatial autocorrelation of the distribution of heavy metal content in soil to predict the concentration of heavy metals at unsampled locations, such as Inverse Distance Weighting (IDW), Ordinary Kriging (OK) and spline function interpolation. Although this type of method is simple and efficient, it ignores the impact of environmental variables on the spatial heterogeneity of soil heavy metals. When the content of heavy metals in regional soil is affected by the complex influence of soil inherent properties and external environmental factors and shows significant spatial heterogeneity, the prediction results obtained by using this type of interpolation method often have certain deviations.
[0004] The attribute similarity-based method predicts the content of heavy metals in soil by constructing a quantitative association between environmental variables and the content of heavy metals in soil. Among these methods, machine learning (ML) and its higher-order branch deep learning (DL) can flexibly model the nonlinear relationship between environmental variables and soil properties, and have become a widely used method in soil heavy metal prediction. Commonly used machine learning and deep learning models include random forest (RF), extreme gradient boosting (XGBoost), support vector machine (SVM), deep neural network (DNN), and convolutional neural network (CNN). Although machine learning and deep learning models have been widely used in predicting the content of heavy metals in soil, the effective training of these models is highly dependent on the support of large-scale data sets. However, the scarcity of soil sampling data in reality usually cannot meet the needs of machine learning and deep learning model training, which limits the accuracy of prediction. Summary of the invention
[0005] In response to the current technical problem of inaccurate prediction of soil heavy metal content, the present invention provides a method for predicting soil heavy metal content in areas with sparse data samples. The method integrates spatial position and multiple environmental variable information as input features of the model to achieve accurate prediction of soil heavy metal content in areas with sparse samples. The method also utilizes data enhancement and transfer learning techniques to effectively solve the problem of insufficient labeled data in the target domain and significantly improve the model performance.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] The present invention provides a method for predicting soil heavy metal content in areas with sparse data samples, comprising the following steps:
[0008] S1: Select the area to be predicted and arrange virtual sampling points, obtain the locations of original sampling points and virtual sampling points in the area to be predicted, soil heavy metal concentrations and environmental variable information, and construct the original sampling point data set and virtual sampling point data set of the area to be predicted;
[0009] S2: construct a fully connected neural network model, and use the virtual sampling point data set to pre-train the fully connected neural network model to obtain a pre-trained fully connected neural network model;
[0010] S3: Freeze the network layer parameters of the pre-trained fully connected neural network model except the last learnable layer, and use the original sampling point dataset to fine-tune the parameters of the last learnable layer of the pre-trained fully connected neural network model to obtain a transfer learning model based on fine-tuning;
[0011] S4: Select a point to be predicted from the area to be predicted, input the location and environmental variable information of the point to be predicted into the transfer learning model based on fine-tuning, and output the predicted value of soil heavy metal concentration of the point to be predicted.
[0012] In the above technical solution, step S1 specifically includes the following steps:
[0013] S11: Obtain the publicly available soil heavy metal content data and environmental variable distribution information, extract the location of the original sampling points in the area to be predicted, the soil heavy metal concentration and environmental variable information, and construct the original sampling point data set of the area to be predicted;
[0014] S12: Multiple virtual sampling points are arranged at equal intervals in the area to be predicted. According to the positions of the original sampling points and the corresponding heavy metal concentrations, the Kriging interpolation method is used to obtain the soil heavy metal concentrations of the virtual sampling points. According to the positions of the virtual sampling points, the environmental variable information of the virtual sampling points is extracted from the environmental variable distribution information to construct a virtual sampling point data set for the area to be predicted.
[0015] In the above technical solution, the environmental variables referred to in the environmental variable distribution information in step S11 include: total nitrogen (TN), total phosphorus (TP), total potassium (TK), soil acidity (pH), soil organic matter (SOM), cation exchange capacity (CEC), soil bulk density (BD), sand content (Sand), silt content (Silt), clay content (Clay), altitude (DEM), ground slope (Slope), lithology (Lithology), normalized difference vegetation index (NDVI), vegetation coverage (FVC), drought index (DI), annual precipitation (AP), relative humidity (RH), average annual temperature (AT), evapotranspiration (EVA), gross domestic product (GDP), road network density (DRN), population density (PD), electronics factory density (ETF), chemical plant density (CP), machinery manufacturing enterprise density (MME), metal manufacturing plant density (MMP), mining enterprise density (ME), livestock farm density (LF), printing and dyeing factory density (PDF).
[0016] In the above technical solution, the fully connected neural network model in step S2 includes an input layer, multiple hidden layers and an output layer, wherein each hidden layer has multiple neurons.
[0017] Compared with the prior art, the present invention has the following beneficial effects:
[0018] The present invention integrates spatial location and multiple environmental variable information as input features of the model to achieve accurate prediction of soil heavy metal content in areas with sparse samples; uses data enhancement technology to expand the scale of training sets and increase data diversity to improve the training effect and generalization ability of deep learning models, and then uses transfer learning technology to transfer the knowledge learned from the virtual data set generated by data enhancement in the source domain to the target domain based on the measured data set, effectively solving the problem of insufficient labeled data in the target domain and significantly improving model performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is a schematic diagram of the flow of the method for predicting soil heavy metal content in data sparse areas of the present invention.
[0020] Figure 2 Schematic diagram of the transfer learning process of the present invention.
[0021] Figure 3 Comparison of the predicted values of heavy metal concentrations of the present invention with those predicted by other methods. DETAILED DESCRIPTION
[0022] The following examples are used to illustrate the present invention, but are not intended to limit the scope of protection of the present invention. Unless otherwise specified, the technical means used in the examples are conventional means well known to those skilled in the art. The test methods in the following examples are conventional methods unless otherwise specified.
[0023] Embodiment 1
[0024] like Figure 1 As shown in the flowchart, the method for predicting soil heavy metal content in areas with sparse data samples in this embodiment specifically includes the following steps:
[0025] (1) Data Acquisition
[0026] Obtain the published soil heavy metal content data by investigating literature and books, and extract the location of the original sampling point in the area to be predicted. raw , the corresponding soil heavy metal concentration c raw ; Obtain the distribution information of environmental variables in the area to be predicted from multiple data platforms, and extract the environmental variable information at the original sampling points raw This example obtains 203 original soil Zn sampling data in the Yangtze River Delta region from the heavy metal data attached to the document "Distribution of Heavy Metal Pollution in Surface Soil Samples in China: A Graphical Review", including the location and the corresponding soil Zn concentration data k=1,2,…,203; obtain the distribution information of various environmental variables in the Yangtze River Delta region from multiple environmental data platforms, and calculate the distribution information of various environmental variables based on the location of each original sampling point. Extract the environmental variable information of each original sampling point from the environmental variable distribution information Thus constructing the original sampling point data set
[0027] The locations of the original sampling points are represented by longitude and latitude coordinates. There are 30 environmental variables, including total nitrogen (TN), total phosphorus (TP), total potassium (TK), soil acidity and alkalinity (pH), soil organic matter (SOM), cation exchange capacity (CEC), soil bulk density (BD), sand content (Sand), silt content (Silt), clay content (Clay), altitude (DEM), ground slope (Slope), lithology (Lithology), normalized difference vegetation index (NDVI), vegetation coverage (FVC), drought index (DI), annual precipitation (AP), relative humidity (RH), average annual temperature (AT), evapotranspiration (EVA), gross domestic product (GDP), road network density (DRN), population density (PD), electronics factory density (ETF), chemical plant density (CP), machinery manufacturing enterprise density (MME), metal manufacturing plant density (MMP), mining enterprise density (ME), livestock farm density (LF), and printing and dyeing factory density (PDF). Soil organic matter (SOM), total nitrogen (TN), total phosphorus (TP), and total potassium (TK) were obtained from the National Earth System Science Data Center (https: / / www.geodata.cn). Soil pH, cation exchange capacity (CEC), bulk density (BD), sand, silt, and clay content were obtained from the World Soil Database. Soil property data were mainly obtained from the Second National Soil Survey of China. Altitude (DEM), slope, drought index (DI), average annual precipitation (AP), relative humidity (RH), average annual temperature (AT), evapotranspiration (EVA), and gross domestic product (GDP) were obtained from the Resources and Environmental Science Data Center (http: / / www.resdc.cn). Lithology data were obtained from the National Geological Survey. Normalized difference vegetation index (NDVI) and fractional vegetation cover (FVC) were obtained from NASA Earth Data (https: / / www.earthdata.nasa.gov). Population density (PD) data were obtained from the LandScan program initiated by Oak Ridge National Laboratory (https: / / landscan.ornl.gov / ). Road network data were downloaded from OpenStreetMap (https: / / www.openstreetmap.org), and the road network density (DRN) was calculated using ArcGIS 10.8. Considering that the impact of industrial activities on the accumulation of heavy metals in soil cannot be ignored, the application programming interface in Amap (https: / / lbs.amap.com) was used to locate several points of interest and calculate the density of various factories. These points of interest include electronics factories, chemical plants, machinery manufacturing enterprises, metal manufacturing plants, mining enterprises, livestock farms, and printing and dyeing factories. The Kernel Density function in ArcGIS 10.8 was used to calculate the density of various factories in the predicted area.
[0028] It is worth noting that, due to the limited data samples obtained, and in order to facilitate the verification of the prediction effect of the present invention, the original sampling point data obtained can be randomly divided into two parts during implementation, namely, a training set and a validation set, wherein: the training set is used to obtain the soil heavy metal concentration of the virtual sampling point and to fine-tune the pre-trained fully connected neural network model, and the validation set is used to demonstrate and evaluate the prediction accuracy of the transfer learning model. In this embodiment, the original sampling point data obtained is randomly divided into a training set and a validation set in a ratio of 3:1. and validation set
[0029]
[0030] (2) Data enhancement
[0031] Multiple virtual sampling points are evenly spaced in the area to be predicted. Based on the training set data information of the original sampling point data set, the Kriging interpolation method is used to obtain the soil heavy metal concentration of the virtual sampling point, and the environmental variable information of the virtual sampling point is extracted from the environmental variable distribution information according to the location of the virtual sampling point. In this embodiment, 6,400 virtual sampling points are evenly spaced in the Yangtze River Delta region. v=1,2,3,…,6400; according to D train The locations of the 153 original sampling points and the corresponding heavy metal concentrations Information, using Kriging interpolation method, interpolation to obtain the location of each virtual sampling point Zn concentration on With the help of ArcGIS 10.8 software, according to the location of each virtual sampling point Extract the environmental variable information of each virtual sampling point from the environmental variable distribution information Thus, a virtual sampling point data set is constructed
[0032]
[0033] (3) Model construction and pre-training
[0034] A fully connected neural network (FCNN) model was constructed. The FCNN model integrated the longitude and latitude coordinates of the sampling point location with 30 environmental variables as input features, and used the heavy metal concentration of the sampling point as the output result. Pre-train the FCNN model to obtain a pre-trained fully connected neural network model FCNN pre .
[0035] In this embodiment, the FCNN includes 1 input layer, 4 hidden layers and 1 output layer. The number of neurons in the 4 hidden layers is 80, 40, 20 and 1 respectively. The parameters of the first layer of the FCNN include the weight matrix and the deviation vector All parameters of FCNN can be expressed as:
[0036] W={W1,W2,W3,W4},
[0037] b={b1,b2,b3,b4}.
[0038] Given a d-dimensional (d=32) input vector x∈R d , k-dimensional (k=1) output That is, the FCNN approximation of the function u(x) takes the following form:
[0039]
[0040] in, represents the approximation of FCNN, for which the following formula applies:
[0041]
[0042] Among them, y l is the output of layer l, g l is the activation function of layer l, and {W,b} can be estimated by minimizing the following loss function:
[0043]
[0044] Where l(W,b) is the loss function, usually the mean square error between the predicted value and the true value.
[0045] (5) Fine-tune the pre-trained model
[0046] Freezing FCNN pre The network layer parameters except the last learnable layer are then taken as the training set D of the original sampling point data set. train For the target data set, pre Fine-tune the parameters of the last learnable layer of the FCNN to obtain the fine-tuning-based transfer learning model FTL The specific fine-tuning process is as follows: Figure 2 shown.
[0047] In this embodiment, FCNN pre The last learnable layer is the fourth hidden layer, with the training set D of the original sampling point data set train Fine-tune the hidden layer parameters for the target data set to adapt the network to the target data set and obtain the fine-tuning-based transfer learning model FCNNFTL In the process of fine-tuning the last hidden layer, in order to retain the common features of the pre-trained model, avoid overfitting, and ensure the stability of the training process, the learning rate of this layer is set to a smaller value of 5×10 -5 .
[0048] (6) Prediction
[0049] Select the points to be predicted from the area to be predicted, and input the location and environmental variable information of the points to be predicted into the transfer learning model FCNN FTL , then the predicted value of soil heavy metal concentration at the predicted point is output. test The 50 points (n test =50) as the point to be predicted, and its location and environmental variable information j=1,2,3,…,n test Enter FCNN FTL , output the predicted value of soil Zn concentration at the predicted point
[0050] like Figure 3 As shown, the Kriging interpolation is used to obtain ( Figure 3 -a) Single use D train Training FCNN output ( Figure 3 -b) and D train and D VSG Joint training FCNN output ( Figure 3 -c) and FCNN of the present invention FTL Output ( Figure 3 -d), compared with the actual soil Zn concentration value c in the validation set of the original sampling point dataset test Compare the fitting results. It can be seen that compared with the Kriging interpolation method and other models trained based on different data sets, FCNN FTL The predicted Zn concentration value at the predicted point is closer to the actual value, and its R 2 It can be seen that the present invention can accurately predict the content of heavy metals in regional soil.
[0051] The embodiments described above are only preferred embodiments of the present invention and are only used to explain the present invention, not to limit the scope of implementation of the present invention. For those skilled in the art, other implementation methods can certainly be easily made by replacement or modification based on the technical contents disclosed in this specification. Therefore, all changes and improvements made on the principles of the present invention should be included in the scope of the patent application of the present invention.
Claims
1. A method for predicting soil heavy metal content in areas with sparse data samples, characterized in that: The following steps are involved: S1: Select the area to be predicted and arrange virtual sampling points, obtain the locations of original sampling points and virtual sampling points in the area to be predicted, soil heavy metal concentrations and environmental variable information, and construct the original sampling point data set and virtual sampling point data set of the area to be predicted; S2: construct a fully connected neural network model, and use the virtual sampling point data set to pre-train the fully connected neural network model to obtain a pre-trained fully connected neural network model; S3: Freeze the network layer parameters of the pre-trained fully connected neural network model except the last learnable layer, and use the original sampling point dataset to fine-tune the parameters of the last learnable layer of the pre-trained fully connected neural network model to obtain a transfer learning model based on fine-tuning; S4: Select a point to be predicted from the area to be predicted, input the location and environmental variable information of the point to be predicted into the transfer learning model based on fine-tuning, and output the predicted value of soil heavy metal concentration of the point to be predicted.
2. The method for predicting soil heavy metal content in areas with sparse data samples according to claim 1 is characterized in that: Step S1 specifically includes the following steps: S11: Obtain the publicly available soil heavy metal content data and environmental variable distribution information, extract the location of the original sampling points in the area to be predicted, the soil heavy metal concentration and environmental variable information, and construct the original sampling point data set of the area to be predicted; S12: Multiple virtual sampling points are arranged at equal intervals in the area to be predicted. According to the positions of the original sampling points and the corresponding heavy metal concentrations, the Kriging interpolation method is used to obtain the soil heavy metal concentrations of the virtual sampling points. According to the positions of the virtual sampling points, the environmental variable information of the virtual sampling points is extracted from the environmental variable distribution information to construct a virtual sampling point data set for the area to be predicted.
3. The method for predicting soil heavy metal content in areas with sparse data samples according to claim 2 is characterized in that: The environmental variables referred to in the environmental variable distribution information in step S11 include: total nitrogen, total phosphorus, total potassium, soil acidity and alkalinity, soil organic matter, cation exchange capacity, soil bulk density, sand content, silt content, clay content, altitude, ground slope, lithology, normalized vegetation difference index, vegetation coverage, drought index, annual precipitation, relative humidity, average annual temperature, evapotranspiration, gross domestic product, road network density, population density, electronics factory density, chemical plant density, machinery manufacturing enterprise density, metal manufacturing plant density, mining enterprise density, livestock farm density and printing and dyeing factory density.
4. The method for predicting soil heavy metal content in areas with sparse data samples according to claim 1, characterized in that: The fully connected neural network model in step S2 includes an input layer, multiple hidden layers and an output layer, wherein each hidden layer has multiple neurons.