Plateau lake basin cultivated land soil pollution multi-index synchronous monitoring method
By constructing a multi-task deep learning model and combining CNN and LSTM branches, the problem of simultaneous monitoring of multiple indicators of farmland soil pollution in plateau lake basins was solved, realizing efficient and low-cost simultaneous monitoring and risk assessment of multiple indicators, and improving the accuracy and timeliness of monitoring.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- KUNMING UNIV OF SCI & TECH
- Filing Date
- 2026-01-04
- Publication Date
- 2026-05-05
AI Technical Summary
Existing technologies are insufficient for simultaneous monitoring of multiple indicators of farmland soil pollution in plateau lake basins. They suffer from high costs, time-consuming and labor-intensive processes, limited representativeness, poor timeliness, insufficient accuracy, and poor model adaptability. In particular, under the complex terrain and cloudy climate conditions of plateau lake basins, it is difficult to achieve long-term, large-scale, and accurate monitoring.
A multi-task deep learning framework was adopted, combining CNN and LSTM branches, and introducing channel attention and time step attention mechanisms. Using multi-source remote sensing data and soil formation factors, a multi-task prediction model for soil pollution detection indicators in cultivated land of plateau lake watershed was constructed to achieve synchronous monitoring of soil pollution indicators.
It enables efficient and low-cost simultaneous monitoring of multiple indicators, improves data availability and estimation accuracy, accurately characterizes the spatial heterogeneity and temporal variation of soil pollution, provides pollution distribution maps and risk assessments with high spatial and temporal resolution, and supports the comprehensive assessment of agricultural non-point source pollution.
Smart Images

Figure FT_1 
Figure FT_2 
Figure SMS_1
Abstract
Description
Technical Field
[0001] This invention relates to the field of environmental monitoring technology, specifically a method for simultaneous monitoring of multiple indicators of soil pollution in cultivated land in plateau lake basins. Background Technology
[0002] The Importance and Current Technological Bottlenecks of Monitoring Farmland Soil Pollution Indicators: Ensuring farmland soil health and preventing agricultural environmental pollution. Agricultural pollution mainly includes point source pollution and non-point source pollution. Among them, non-point source pollution, such as eutrophication caused by fertilizer and pesticide runoff, is much more difficult to monitor and control than point source pollution due to its characteristics of dispersed emissions, strong concealment, and high randomness.
[0003] Currently, the monitoring of soil pollution indicators in arable land and the resulting pollution risk assessment mainly rely on field sampling and laboratory chemical analysis. While this method is highly accurate, it has the following inherent limitations: 1) High cost and time-consuming: It requires a lot of manpower, material resources and time for sample collection and laboratory analysis, and it is difficult to carry out it frequently and on a large scale; 2) Representing the whole by a few points: Discrete sampling points cannot fully and continuously reflect the spatial heterogeneity of soil nutrients and pollutants in the whole area, and cannot accurately depict the distribution range and degree of agricultural non-point source pollution. 3) Poor timeliness: The long cycle from sampling to obtaining results results in serious delays in monitoring data, making it impossible to provide real-time data support for rapid early warning and precise policy implementation for environmental pollution; The Development and Existing Challenges of Remote Sensing Technology: To overcome the limitations of traditional methods, remote sensing technology, due to its advantages of macroscopic, rapid, and dynamic monitoring, has gradually become the mainstream means for long-term, large-scale monitoring of soil physicochemical properties. However, it still has significant shortcomings: 1) Traditional methods often rely on single-spectral-band information from a single satellite platform, making it difficult to capture the deep, nonlinear relationship between soil physicochemical properties and complex spectral characteristics, thus limiting their accuracy and applicability; 2) Optical remote sensing, such as Landsat and Sentinel-2, is easily affected by clouds and rain. Cloudy climates in plateau lake basins lead to serious data loss. 3) Radar remote sensing, such as Sentinel-1, although it has penetrating power, cannot independently characterize the biochemical properties of soil carbon. 4) Insufficient model adaptability: Geostatistical methods, such as Kriging interpolation, depend on the density of sampling points, and sparse samples lead to amplified errors; Mechanistic models, such as carbon cycle process models, have complex input parameters, such as net primary productivity of vegetation and basic respiration of soil, which are difficult to promote in heterogeneous terrain. 5) Special characteristics of the plateau region: The plateau lake basin has fragmented topography, scattered farmland, and diverse soil types, such as brown calcareous soil and chestnut calcareous soil. Existing models are not very suitable for this region. Advances and limitations of deep learning applications: The development of deep learning has improved the accuracy of remote sensing estimation of soil physicochemical properties, but multi-method collaborative optimization is still insufficient. 1) Single model dominance: Although Random Forest (RF) and Support Vector Machine (SVM) are widely used, such as SOC prediction in the black soil area of Jilin Province, feature extraction relies on manual design and is difficult to capture deep correlations in multi-source data. 2) Insufficient fusion of multi-source data: The fusion of visible-near infrared (VNIR) and hyperspectral imaging (HSI) can improve accuracy, but it does not integrate carbon cycle mechanism parameters, such as soil respiration rate, resulting in weak explanatory power of ecological processes. Its suitability in plateau lake basins also needs to be explored. 3) Lack of dynamic monitoring capabilities: Existing methods mostly focus on static estimation and lack time series modeling capabilities. In particular, they are difficult to predict multiple indicators simultaneously and cannot directly serve the comprehensive assessment of pollution risks. Unique Challenges in Plateau Lake Basins: Long-term, large-scale monitoring of soil pollution indicators in arable land in plateau lake basins faces unique challenges. 1) Difficulty in data acquisition: Large differences in terrain elevation and variable climate lead to unstable quality of optical and radar data; 2) Complex ecological processes: Plateau lake basins, such as the Erhai Lake basin, are sensitive to human activities. The complex interaction between soil and soil-forming environment brings many uncertainties that are difficult to quantify to soil monitoring. It is necessary to consider the soil formation process under special conditions, coupled environmental variables, and long-term monitoring. 3) Poor model generalization: Existing deep learning models show a sharp drop in accuracy in regions with few training samples and lack cross-regional transfer capabilities; Therefore, there is an urgent need in this field for a comprehensive and innovative technical solution that can coordinate multi-source data, multi-model deep learning frameworks and soil formation process mechanisms, be applicable to plateau lake basins, achieve simultaneous and accurate monitoring of multiple indicators of farmland soil pollution over a long period and on a large scale, and serve agricultural non-point source pollution assessment, so as to overcome the shortcomings of existing technologies. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention provides a method for simultaneous monitoring of multiple indicators of soil pollution in cultivated land in plateau lake basins, which has the advantages of low cost and high efficiency, and solves the aforementioned technical problems.
[0005] To achieve the above objectives, the present invention provides the following technical solution: a method for simultaneous monitoring of multiple indicators of soil pollution in cultivated land in plateau lake basins, comprising the following steps: S1: Obtain measured data of soil pollution detection indicators in cultivated land in plateau lake basins; S2: Obtain soil-forming factor data synchronized with the measured data of soil pollution detection indicators in cultivated land of plateau lake basins; S3: Obtain farmland extraction data for plateau lake basins; S4: Based on the measured data of soil pollution detection indicators in cultivated land in the plateau lake basin and the synchronous soil formation factor data, a multi-task prediction model for soil pollution detection indicators in cultivated land in the plateau lake basin is constructed. S5: Input the extracted data of cultivated land in the plateau lake basin into the multi-task prediction model of soil pollution detection indicators of cultivated land in the plateau lake basin, and obtain the spatial distribution map of key soil pollution indicators of cultivated land in the plateau lake basin, the interannual variation map of key soil pollution indicators of cultivated land in the plateau lake basin, and the distribution map of ground source pollution risk level of cultivated land in the plateau lake basin.
[0006] As a preferred technical solution of the present invention, the measured data of soil pollution detection indicators in the plateau lake basin in step S1 is obtained by collecting soil samples from the topsoil, and after the soil samples are air-dried, measuring the measured data of soil pollution detection indicators, which include the content of soil organic matter, nitrogen, phosphorus and potassium.
[0007] As a preferred technical solution of the present invention, in step S2, the soil formation factor data obtained are synchronized with the measured data of soil pollution detection indicators in cultivated land of plateau lake basin, including soil parent material data, soil property data, annual average surface temperature data, annual average precipitation data, topographic variable data, and vegetation information. The expressions for the soil parent material data and soil property data are as follows: Among them, NDVI is the Normalized Difference Vegetation Index; NBR2 is the Normalized Burning Ratio Index; BSI is the Bare Soil Index; SBI is the Surface Brightness Index; NDMI is the Normalized Difference Water Index; CMSI is the Crop Water Stress Index; CI is the Clay Index; the near-infrared band in the NIR multispectral remote sensing image data; Red is the red light band; SWIR1 is the shortwave infrared band 1, SWIR2 is the shortwave infrared band 2, Blue is the blue light band, and Green is the green light band; soil property data includes NDVI, NBR2, BSI, SBI, NDMI, and CMSI; soil parent material data includes CI; The annual average surface temperature data is obtained by first acquiring the annual average surface temperature, processing it into orthorectified surface temperature, filtering images with cloud cover less than a set value throughout the year, converting them into surface temperature data in degrees Celsius, and then obtaining the annual average surface temperature data. The annual average precipitation data is obtained by first acquiring the monthly precipitation dataset for the entire year of the sampling year, and then deriving the annual average precipitation data. The terrain variable data includes slope, vector ruggedness, water flow power index, terrain ruggedness index, composite terrain index, roughness, terrain location index, and elevation; The vegetation information includes EVI and LSP, where EVI is the annual average enhanced vegetation index and LSP represents vegetation phenology index.
[0008] As a preferred technical solution of the present invention, the multi-task prediction model for soil pollution detection indicators in the plateau lake basin of S4 includes a CNN branch and an LSTM branch. The outputs of the two branches are fused to predict the predicted value. The CNN branch introduces a channel attention mechanism and the LSTM branch introduces a time step attention mechanism. The multi-task prediction function of the multi-task prediction model for soil pollution detection indicators in cultivated land of plateau lake basins is as follows: in, This represents any indicator of soil pollution in arable land. () It is a function representing the detection index of soil pollution in any arable land, composed of seven types of soil-forming factors. Represents soil property data, Representing climate, including average annual surface temperature data and average annual precipitation data, Represents vegetation information. Represents terrain variable data, Represents soil parent material data, Represents time, The above seven categories of soil-forming factors are used to comprehensively characterize the soil-forming environment of the study area. Input x_cnn_common into the CNN branch, which outputs f_cnn. Input x_ts_evi_lsp into the LSTM branch, which outputs f_lstm. Concatenate the output f_cnn from the CNN branch with the output f_lstm from the LSTM branch to form a 32-dimensional fused feature. The specific expression is as follows: f=[f cnn ; f lstm ]∈R 32 in, Represents 32-dimensional fused features; [;] represents vector concatenation operation; f cnn f is the output of the CNN branch; lstm The output of the LSTM branch; The system then enters the final fully connected regression layer and outputs the predicted values. .
[0009] As a preferred embodiment of the present invention, the CNN branch extracts soil property data, soil parent material data, annual average surface temperature data, annual average precipitation data, and topographic variable data, which are used as x_cnn_common. x_cnn_common is then input into the CNN branch to obtain a 16-dimensional representation of the output data f_cnn. The specific steps are as follows: S4a.1:x_cnn_common is taken as input, and after two layers of convolution and ReLU operation, it goes through a 2×2 max pooling operation; S4a.2: Connect to SEBlock channel attention; S4a.3: After flattening, a fully connected layer is used to obtain a 16-dimensional representation f_cnn∈R. 16 ; The shape of x_cnn_common is: (B, E, H, W) Where E is the number of image channels; B is the batch size; and H and W are the height and width of the feature map, respectively.
[0010] As a preferred embodiment of the present invention, a channel attention mechanism is introduced into the CNN branch, first performing a global average pooling operation on the channels in the CNN branch: In the formula, For channel The global average, and It refers to the spatial height and width of the feature map. Indicates channel In position Pixel values; For feature maps Summation operation is performed on each height; For feature maps Summation operation is performed on each height; Two fully connected layers are used to model the dependencies between channels and output the weights: In the formula, ∈ R C It is the descriptor for all channels. and It is the weight matrix of the fully connected layer. It is the ReLU activation function. It is the Sigmoid activation function. ∈(0,1) C This represents the attention weight for each channel; Scaling the feature map of each channel by weights: In the formula, Representative Channel The reweighted feature map Representative Channel Updated attention weights; For channel Feature map before weighting.
[0011] As a preferred embodiment of the present invention, the LSTM branch extracts the time features of the time series EVI and LSP dynamic variables as x_ts_evi_lsp, and inputs x_ts_evi_lsp into the LSTM branch to obtain the 16-dimensional representation f_lstm. The specific steps are as follows: S4b.1: Input the time series data of EVI and LSP into the LSTM; S4b.2: Following the attention at each time step, the hidden states at each time step are weighted and summed to obtain the context vector; S4b.3: A 16-dimensional representation f_lstm∈R is obtained through a fully connected layer. 16 ; The shape of x_ts_evi_lsp is: (B, T, D) in, For time steps, B is the size of the input feature dimension; B is the batch size. = + in, The input EVI feature dimension; is the input LSP feature dimension.
[0012] As a preferred embodiment of the present invention, the LSTM branch scores the time steps: In the formula, Represents time step The score, Represents the LSTM at time steps The hidden state vector. Represents the number of time steps. Represents a learnable attention weight vector; Softmax normalization transforms the original scores into a probability distribution. In the formula, Represents time step t Attention weights Represents time step t The score; Represents the time step index; For all time steps The sum; It is an exponential function; Weighted convergence into a context vector: In the formula, Represents the context vector. Represents time step Attention weights Represents LSTM at time steps t The hidden state vector; Represents the number of time steps; For all time steps from 1 to T The sum of .
[0013] Compared with existing technologies, this invention provides a method for simultaneous monitoring of multiple indicators of soil pollution in cultivated land in plateau lake basins, which has the following beneficial effects: This invention addresses the data gap problem in cloudy areas by synthesizing time-series data, significantly improving data availability and estimation accuracy. By introducing parameters derived from the Digital Elevation Model (DEM) (vector ruggedness, hydrodynamic power index, topographic ruggedness index, etc.) and combining them with a CNN model to extract spatial features of soil-forming factors, it effectively utilizes geospatial background information to accurately characterize the spatial heterogeneity of organic matter, nitrogen, phosphorus, and potassium in arable land soil. The use of LSTM to extract temporal features from time-series vegetation data significantly improves the model's prediction accuracy and effectively represents historical accumulation using long-term crop growth information. The soil-forming model acquires seven types of soil-forming factors that describe the soil-forming process in arable land, enhancing the interpretability of ecological processes, and the integration of a dual attention mechanism improves the model's physical interpretability. Through the constructed multi-task prediction method for arable land soil pollution detection indicators in plateau lake basins, and the resulting high spatial and temporal resolution spatial distribution maps, interannual variation maps, and non-point source pollution risk level distribution maps of various key soil pollution indicators, it efficiently and cost-effectively achieves long-term, large-scale, multi-indicator synchronous monitoring and risk control. Attached Figure Description
[0014] Figure 1 This is a schematic diagram of the overall process of the present invention; Figure 2 This is a schematic diagram of the overall framework of the present invention. Detailed Implementation
[0015] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0016] Please see Figure 1 and Figure 2 A method for simultaneous monitoring of multiple indicators of soil pollution in cultivated land in plateau lake basins includes the following steps: S1: Obtain measured data of soil pollution detection indicators in cultivated land in plateau lake basins; In S1, the measured data of soil pollution detection indicators in cultivated land in plateau lake basins are obtained by collecting soil samples from the topsoil and measuring the soil pollution detection indicators after the soil samples have been air-dried. When collecting soil samples from the topsoil, a combination of expert knowledge sampling, stratified stepwise sampling, and heuristic uncertainty-oriented sampling methods are used. During the bare soil period after crop harvest, generally from March to May and from September to November each year, soil samples are collected from the topsoil at a depth of 0-20 cm. Each sample is composed of 5-6 subsamples selected from a 30m×30m quadrat and mixed evenly. The coordinates of the soil samples are recorded by a portable GPS. After the soil samples were air-dried, the measured values of soil pollution indicators were determined, including the content of soil organic matter, nitrogen, phosphorus, and potassium. The unit for soil organic matter was g / kg. 1 The nitrogen, phosphorus, and potassium content is mg / kg. 1 ; S2: Obtain soil-forming factor data synchronized with the measured data of soil pollution detection indicators in cultivated land of plateau lake basins; S2 acquires soil-forming factor data synchronized with the measured data of cultivated land soil pollution detection indicators, including soil parent material data, soil property data, annual average surface temperature data, annual average precipitation data, topographic variable data, and vegetation information; Sentinel-2 MSI remote sensing reflectance data were downloaded from the GEE platform to obtain soil parent material and soil property data. Bare soil pixels were acquired from images with cloud cover less than 10% by setting three thresholds: NDVI less than 0.35, NBR2 less than 0.125, and BSI greater than 0.021. The data acquisition time range was broadened to October of the year preceding the soil sampling period to April of the following year to ensure comprehensive coverage of the study area. A median composite method was used to calculate the median of multiple observations at the same location within each time period, and bilinear interpolation was used to resample to 30m to obtain a composite bare soil image for the sampling year. The relevant expressions are as follows: Among them, NDVI is the Normalized Difference Vegetation Index; NBR2 is the Normalized Burning Ratio Index; BSI is the Bare Soil Index; SBI is the Surface Brightness Index; NDMI is the Normalized Difference Water Index; CMSI is the Crop Water Stress Index; CI is the Clay Index; the near-infrared band in the NIR multispectral remote sensing image data; Red is the red light band; SWIR1 is the shortwave infrared band 1, SWIR2 is the shortwave infrared band 2, Blue is the blue light band, and Green is the green light band; soil property data includes NDVI, NBR2, BSI, SBI, NDMI, and CMSI; soil parent material data includes CI; Landsat 8 T1 L2 remote sensing reflectance data were downloaded from the GEE platform to obtain the annual average land surface temperature. This dataset contains a thermal infrared (TIR) band and was processed into orthorectified land surface temperature. After screening images with cloud cover of less than 10% throughout the year, cloud masking, shading saturation, band extraction, and radiometric calibration were performed to convert the data into land surface temperature data in degrees Celsius. The mean synthesis method and bilinear interpolation were used to resample to 30m to obtain the annual average land surface temperature data. The monthly precipitation dataset for the entire year of the sampling year was obtained through the 1km monthly precipitation dataset of China from 1901 to 2022 via the official website of the National Qinghai-Tibet Plateau Scientific Data Center. The annual average precipitation data was obtained by resampling to 30m using bilinear interpolation and then accumulating the data. Terrain variables were obtained by downloading the Geomorpho90m and SRTM 30m datasets from the GEE platform. From Geomorpho90m, slope (Slp), vector ruggedness (VRM), hydrodynamic power index (SPI), terrain ruggedness index (TRI), composite terrain index (CTI), roughness, and terrain location index (TPI) were obtained and resampled to 30m using bilinear interpolation. Elevation was obtained from SRTM30m. Vegetation information for the time series was obtained from the MOD13Q1 and MCD12Q2 datasets downloaded from the GEE platform, with the time series from 2000 to the sampling year. EVI (Enhanced Vegetation Index) was obtained from the MOD13Q1 dataset, and resampled to 30m using bilinear interpolation. The annual average EVI was calculated for each year. Eleven vegetation phenological indices (Greenup, Mid-Greenup, Maturity, Peak, Senescence, Mid-Greendown, Dormancy, Num-Cycles, EVI-Amplitude, EVI-Area, and EVI-Minimum) were obtained from the MCD12Q2 dataset for each year, and resampled to 30m using bilinear interpolation. S3: Download the GLC_FCS30 land use change dataset from the official website of the International Research Center for Big Data for Sustainable Development as the basic data source for farmland extraction, obtain farmland extraction data of the study area and resample to 30m using bilinear interpolation; S4: Based on the measured data of soil pollution detection indicators in cultivated land in the plateau lake basin and the synchronous soil formation factor data, a multi-task prediction model for soil pollution detection indicators in cultivated land in the plateau lake basin is constructed. The S4 multi-task prediction model for soil pollution detection indicators in cultivated land in the plateau lake basin includes a CNN branch and an LSTM branch. The CNN branch introduces a channel attention mechanism, and the LSTM branch introduces a time step attention mechanism. The multi-task prediction function of the multi-task prediction model for soil pollution detection indicators in cultivated land in plateau lake basins is as follows: in, This represents any indicator of soil pollution in arable land. () It is a function representing the detection index of soil pollution in any arable land, composed of seven types of soil-forming factors. Represents soil property data, Representing climate, including average annual surface temperature data and average annual precipitation data, Represents vegetation information. Represents terrain variable data, Represents soil parent material data, Represents time, The above seven categories of soil-forming factors are used to comprehensively characterize the soil-forming environment of the study area. In the CNN branch, channel-wide average pooling introduces a channel attention mechanism. This mechanism adaptively weights channel features by modeling the correlation between channels, making the model focus more on discriminative spectral and spatial information. In the CNN branch, channel-wide average pooling is performed on each channel to obtain a channel descriptor. Its function is to compress a two-dimensional feature map into a scalar while preserving the global statistical information of that channel. In the formula, For channel The global average, and It refers to the spatial height and width of the feature map. Indicates channel In position Pixel values; For feature maps Summation operation is performed on each height; For feature maps Summation operation is performed on each height; Two fully connected layers model the dependencies between channels and output weights. Their purpose is to learn the importance scores of different channels. The output weight expression is as follows: In the formula, ∈ R C It is the descriptor for all channels. and It is the weight matrix of the fully connected layer. It is the ReLU activation function. It is the Sigmoid activation function. ∈(0,1) C This represents the attention weight for each channel; The feature map of each channel is scaled according to its weights to suppress interfering channels and highlight discriminative channels. The gradient updates both the convolutional kernel and the channel weights simultaneously during end-to-end training. The expression for scaling the feature map of each channel according to its weights is as follows: In the formula, Representative Channel The reweighted feature map This represents the attention weight after the channel update; For channel Feature map before weighting.
[0017] A time step attention mechanism is introduced into the LSTM branch, which automatically assigns different weights according to the importance of each time step in the input sequence. The context vector is generated by weighted summation of the LSTM output, thereby highlighting key time features. Its function is to assign learnable weights to different time steps. In the LSTM branch, time steps are scored to measure their importance to the prediction: In the formula, Represents time step The score, Represents the LSTM at time steps The hidden state vector. Represents the number of time steps. Represents a learnable attention weight vector; Softmax normalization transforms the original scores into a probability distribution. In the formula, Represents time step t Attention weights Represents time step t The score, Represents the time step index; For all time steps The sum; It is an exponential function; The weighted convergence is a context vector that comprehensively considers all time steps, but the contribution of different time steps is controlled by attention weights. The expression for the context vector is as follows: In the formula, Represents the context vector. Represents time step Attention weights Represents LSTM at time steps t The hidden state vector; Represents the number of time steps; For all time steps from 1 to T The sum of .
[0018] Compared to simply taking the final time step or averaging, attention amplifies relatively important time steps and weakens irrelevant fluctuations. The channel attention weights s_c∈(0,1) directly reflect the importance of CNN channels and time steps. t The attention weight αt reflects the marginal contribution of each time step to the prediction; The CNN branch extracts soil property data, soil parent material data, annual average surface temperature data, annual average precipitation data, and topographic variable data as x_cnn_common. x_cnn_common is then input into the CNN branch to obtain the 16-dimensional representation f_cnn of the output data. The specific steps are as follows: S4a.1:x_cnn_common is taken as input, and after two layers of convolution and ReLU operation, it goes through a 2×2 max pooling operation; S4a.2: Connect to SEBlock channel attention; S4a.3: After flattening, a fully connected layer is used to obtain a 16-dimensional representation f_cnn∈R. 16 ; The shape of x_cnn_common is: (B, E, H, W) Where E is the number of image / raster channels; B is the batch size; H and W are the height and width of the feature map, respectively. This branch extracts spatial context information to enhance the predictive model's ability to map detailed information, using information from multiple dimensions to jointly represent adjacent soil samples, thus providing richer prior information for the predictive model. The LSTM branch extracts the temporal features of the dynamic variables of time series EVI and vegetation phenology indicators as x_ts_evi_lsp, with the time series from 2000 to the sampling year. The specific steps of inputting x_ts_evi_lsp into the LSTM branch to obtain the 16-dimensional representation f_lstm are as follows: S4b.1: Input the time series of EVI into the LSTM; S4b.2: Following the attention at each time step, the hidden states at each time step are weighted and summed to obtain the context vector; S4b.3: A 16-dimensional representation f_lstm∈R is obtained through a fully connected layer. 16 ; The shape of x_ts_evi_lsp is: (B, T, D) in, For time steps, B is the size of the input feature dimension; B is the batch size. = + in, The input EVI feature dimension; is the input LSP feature dimension.
[0019] This branch can improve the accuracy of prediction by extracting features from historical data that can characterize the accumulation and changing trends of soil organic matter, nitrogen, phosphorus, potassium and other contents. However, if only the data of the input variable in a single period or the mean in a period of one vegetation growing season is considered as input, the prediction error is likely to increase. This is because the lack of time series data will greatly reduce the dimensionality and complexity of the input data and lose historical information, that is, the long-term change characteristics of the variable. The obtained f_cnn and f_lstm are concatenated to form a 32-dimensional fusion feature f, the specific expression of which is as follows: f=[f cnn ; f lstm ]∈R 32 f represents a 32-dimensional fusion feature; [;] represents a vector concatenation operation; f cnn f is the output of the CNN branch; lstm The output of the LSTM branch; The system then enters the final fully connected regression layer and outputs the predicted values. ; Using a multi-task prediction model for soil pollution detection indicators in cultivated land in plateau lake basins, the spatiotemporal spectral characteristics of seven types of soil-forming factors were extracted. Combined with measured soil data, the relationship between measured values of soil pollution detection indicators and input variables was established, and a multi-task prediction model for soil pollution detection indicators in cultivated land in plateau lake basins was constructed. S5: Input the extracted data of cultivated land in the plateau lake basin into the multi-task prediction model of soil pollution detection indicators of cultivated land in the plateau lake basin, and obtain the spatial distribution map of key soil pollution indicators of cultivated land in the plateau lake basin, the interannual variation map of key soil pollution indicators of cultivated land in the plateau lake basin, and the distribution map of ground source pollution risk level of cultivated land in the plateau lake basin. Based on S1-S4 and combined with the farmland extraction data of the corresponding years in the plateau lake watershed, the spatial distribution maps of soil organic matter, nitrogen, phosphorus, and potassium content over multiple years can be obtained, and then the interannual variation map of key soil pollution indicators can be obtained. According to the actual situation of the target area and the interannual variation of key soil pollution indicators, thresholds are set to determine the risk level, and finally the distribution map of ground source pollution risk level of farmland in the plateau lake watershed is obtained, sampled to 30m. The attention weight of the input variables is calculated by using a dual attention mechanism, and the driving factors of the spatiotemporal heterogeneity of soil organic matter, nitrogen, phosphorus, and potassium content in farmland in the plateau lake watershed are identified by correlation analysis.
[0020] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for simultaneous monitoring of multiple indicators of soil pollution in cultivated land in plateau lake basins, characterized in that... Includes the following steps: S1: Obtain measured data of soil pollution detection indicators in cultivated land in plateau lake basins; S2: Obtain soil-forming factor data synchronized with the measured data of soil pollution detection indicators in cultivated land of plateau lake basins; S3: Obtain farmland extraction data for plateau lake basins; S4: Based on the measured data of soil pollution detection indicators in cultivated land in the plateau lake basin and the synchronous soil formation factor data, a multi-task prediction model for soil pollution detection indicators in cultivated land in the plateau lake basin is constructed. S5: Input the extracted data of cultivated land in the plateau lake basin into the multi-task prediction model of soil pollution detection indicators of cultivated land in the plateau lake basin, and obtain the spatial distribution map of key soil pollution indicators of cultivated land in the plateau lake basin, the interannual variation map of key soil pollution indicators of cultivated land in the plateau lake basin, and the distribution map of ground source pollution risk level of cultivated land in the plateau lake basin.
2. The method for simultaneous monitoring of multiple indicators of soil pollution in cultivated land in plateau lake basins according to claim 1, characterized in that: The measured data of soil pollution detection indicators in the plateau lake basin obtained in S1 are obtained by collecting soil samples from the topsoil, and after the soil samples are air-dried, the measured data of soil pollution detection indicators are determined. Soil pollution detection indicators include soil organic matter, nitrogen, phosphorus and potassium content.
3. The method for simultaneous monitoring of multiple indicators of soil pollution in cultivated land in plateau lake basins according to claim 1, characterized in that: In S2, soil-forming factor data are obtained synchronously with the measured data of soil pollution detection indicators in cultivated land of plateau lake basins, including soil parent material data, soil property data, annual average surface temperature data, annual average precipitation data, topographic variable data, and vegetation information. The expressions for the soil parent material data and soil property data are as follows: Among them, NDVI is the Normalized Difference Vegetation Index; NBR2 is the Normalized Burning Ratio Index; BSI is the Bare Soil Index; SBI is the Surface Brightness Index; NDMI is the Normalized Difference Water Index; CMSI is the Crop Water Stress Index; CI is the Clay Index; the near-infrared band in the NIR multispectral remote sensing image data; Red is the red light band; SWIR1 is the shortwave infrared band 1, SWIR2 is the shortwave infrared band 2, Blue is the blue light band, and Green is the green light band; soil property data includes NDVI, NBR2, BSI, SBI, NDMI, and CMSI; soil parent material data includes CI; The annual average surface temperature data is obtained by first acquiring the annual average surface temperature, processing it into orthorectified surface temperature, filtering images with cloud cover less than a set value throughout the year, converting them into surface temperature data in degrees Celsius, and then obtaining the annual average surface temperature data. The annual average precipitation data is obtained by first acquiring the monthly precipitation dataset for the entire year of the sampling year, and then deriving the annual average precipitation data. The terrain variable data includes slope, vector ruggedness, water flow power index, terrain ruggedness index, composite terrain index, roughness, terrain location index, and elevation; The vegetation information includes EVI and LSP, where EVI is the annual average enhanced vegetation index and LSP represents vegetation phenology index.
4. The method for simultaneous monitoring of multiple indicators of soil pollution in cultivated land in plateau lake basins according to claim 1, characterized in that: The S4 multi-task prediction model for soil pollution detection indicators in cultivated land in the plateau lake basin includes a CNN branch and an LSTM branch. The outputs of the two branches are fused to obtain the predicted value. The CNN branch introduces a channel attention mechanism and the LSTM branch introduces a time step attention mechanism. The multi-task prediction function of the multi-task prediction model for soil pollution detection indicators in cultivated land of plateau lake basins is as follows: in, This represents any indicator of soil pollution in arable land. () It is a function representing the detection index of soil pollution in any arable land, composed of seven types of soil-forming factors. Represents soil property data, Representing climate, including average annual surface temperature data and average annual precipitation data, Represents vegetation information. Represents terrain variable data, Represents soil parent material data, Represents time, The above seven categories of soil-forming factors are used to comprehensively characterize the soil-forming environment of the study area. Input x_cnn_common into the CNN branch, which outputs f_cnn. Input x_ts_evi_lsp into the LSTM branch, which outputs f_lstm. Concatenate the output f_cnn from the CNN branch with the output f_lstm from the LSTM branch to form a 32-dimensional fused feature. The specific expression is as follows: f=[f cnn ; f lstm ]∈R 32 in, Represents 32-dimensional fused features; [;] represents vector concatenation operation; f cnn f is the output of the CNN branch; lstm The output of the LSTM branch; The system then enters the final fully connected regression layer and outputs the predicted values. .
5. The method for simultaneous monitoring of multiple indicators of soil pollution in cultivated land in plateau lake basins according to claim 4, characterized in that: The CNN branch extracts soil property data, soil parent material data, annual average surface temperature data, annual average precipitation data, and topographic variable data as x_cnn_common. x_cnn_common is then input into the CNN branch to obtain the 16-dimensional representation f_cnn of the output data. The specific steps are as follows: S4a.1:x_cnn_common is taken as input, and after two layers of convolution and ReLU operation, it goes through a 2×2 max pooling operation; S4a.2: Connect to SEBlock channel attention; S4a.3: After flattening, a fully connected layer is used to obtain a 16-dimensional representation f_cnn∈R. 16 ; The shape of x_cnn_common is: (B, E, H, W) Where E is the number of image channels; B is the batch size; and H and W are the height and width of the feature map, respectively.
6. The method for simultaneous monitoring of multiple indicators of soil pollution in cultivated land in plateau lake basins according to claim 4, characterized in that: The CNN branch introduces a channel attention mechanism, first performing global average pooling on the channels within the CNN branch: In the formula, For channel The global average, and It refers to the spatial height and width of the feature map. Indicates channel In position Pixel values; For feature maps Summation operation is performed on each height; For feature maps Summation operation is performed on each height; Two fully connected layers are used to model the dependencies between channels and output the weights: In the formula, ∈ R C It is the descriptor for all channels. and It is the weight matrix of the fully connected layer. It is the ReLU activation function. It is the Sigmoid activation function. ∈(0,1) C This represents the attention weight for each channel; Scaling the feature map of each channel by weights: In the formula, Representative Channel The reweighted feature map Representative Channel Updated attention weights; For channel Feature map before weighting.
7. The method for simultaneous monitoring of multiple indicators of soil pollution in cultivated land in plateau lake basins according to claim 4, characterized in that: The LSTM branch extracts the temporal features of the time series EVI and LSP dynamic variables as x_ts_evi_lsp. The specific steps of inputting x_ts_evi_lsp into the LSTM branch to obtain the 16-dimensional representation f_lstm are as follows: S4b.1: Input the time series data of EVI and LSP into the LSTM; S4b.2: Following the attention at each time step, the hidden states at each time step are weighted and summed to obtain the context vector; S4b.3: A 16-dimensional representation f_lstm∈R is obtained through a fully connected layer. 16 ; The shape of x_ts_evi_lsp is: (B, T, D) in, For time steps, B is the size of the input feature dimension; B is the batch size. = + in, The input EVI feature dimension; is the input LSP feature dimension.
8. The method for simultaneous monitoring of multiple indicators of soil pollution in cultivated land in plateau lake basins according to claim 4, characterized in that: The LSTM branch scores the time steps: In the formula, Represents time step The score, Represents the LSTM at time steps The hidden state vector. Represents a learnable attention weight vector; Softmax normalization transforms the original scores into a probability distribution. In the formula, Represents time step t Attention weights Represents time step t The score; Represents the time step index; For all time steps The sum; It is an exponential function; Weighted convergence into a context vector: In the formula, Represents the context vector. Represents time step Attention weights Represents LSTM at time steps t The hidden state vector; Represents the number of time steps; For all time steps from 1 to T The sum of .