Soil heavy metal cadmium runoff output flux prediction method based on small sample learning
By integrating physical reinforcement transfer learning with a self-attention mechanism, a few-shot learning model was developed to solve the problem of predicting soil heavy metal cadmium runoff output flux in small-shot scenarios, achieving high-precision, low-cost, and fast-response prediction results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-23
- Publication Date
- 2026-03-10
AI Technical Summary
Existing technologies struggle to accurately predict soil cadmium runoff flux in small sample scenarios. Traditional methods require extensive monitoring data, are costly and time-consuming, and have poor adaptability to complex terrain.
A few-sample learning model that integrates physical reinforcement transfer learning and self-attention mechanism is adopted. The model is trained with a few-sample data, and key features are extracted by combining physical constraints and self-attention mechanism to predict the runoff flux of cadmium heavy metal in soil.
High-precision predictions were achieved under small sample conditions, reducing monitoring costs and time requirements, improving the applicability and stability of the model, and enabling rapid identification of high-risk areas to reduce the risk of pollution spread.
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of soil environment monitoring and prediction, in particular to a soil heavy metal cadmium runoff output flux prediction method based on small sample learning. BACKGROUND
[0002] Soil heavy metal cadmium pollution is one of the serious environmental problems currently faced. Its runoff output not only leads to soil quality decline, but also pollutes the surrounding environment such as water bodies, posing a serious threat to the ecological system and human health. Accurate prediction of soil heavy metal cadmium runoff output flux is of great significance for formulating effective pollution prevention and control measures and protecting the ecological environment.
[0003] Traditional prediction methods often rely on a large amount of monitoring data to establish statistical models or physical models for prediction. There are currently a variety of soil heavy metal runoff prediction schemes proposed in related patents, for example:
[0004] Patent CN202310215678.9 (A soil heavy metal cadmium runoff prediction method based on BP neural network): This method uses BP neural network as the core model, collects more than 10 index data such as soil cadmium content, rainfall, soil texture, and constructs a data-driven prediction model. However, it explicitly requires that the training sample size should not be less than 200 groups, and at least 3 complete hydrological years of monitoring data should be covered. In the small sample scenario (such as sample size less than 50 groups), the model is prone to overfitting, the prediction RMSE exceeds 0.2, and the R² is less than 0.6, which is difficult to meet the accuracy requirements.
[0005] However, in actual situations, it is difficult to obtain a large amount of soil heavy metal cadmium runoff output related data due to high monitoring cost (a single complete monitoring requires an investment of 3000-5000 yuan), long monitoring period (covering different periods such as rainy season and dry season), and complex terrain in some areas (such as mountainous and hilly areas where it is difficult to set up monitoring points). Small sample data scenarios (less than 50 monitoring points in the target area, less than 30 effective sample groups) are common.
[0006] Small sample learning technology can achieve good learning and prediction effect under limited data samples, providing a new idea to solve the above problems. Transfer learning can use the knowledge trained from a large amount of data in other related fields to migrate to the target task of small samples, improve the performance of the model, and the self-attention mechanism can focus on key features and enhance the model's ability to capture important information. By integrating transfer learning and self-attention mechanism into the small sample learning model, it is expected to improve the prediction accuracy of soil heavy metal cadmium runoff output flux. SUMMARY
[0007] To achieve the above objectives, this invention provides the following technical solution: a method for predicting soil heavy metal cadmium runoff flux based on small sample learning, comprising the following steps:
[0008] Data collection and preprocessing steps: Collect small sample data related to the heavy metal cadmium in soil, including but not limited to soil cadmium content data, soil physicochemical property data, meteorological data, topographic data, and runoff monitoring data; clean the collected data, remove outliers and missing values, and standardize the data to make different types of data comparable;
[0009] Constructing a few-shot learning model: A few-shot learning model architecture that integrates physical augmentation transfer learning and self-attention mechanism is adopted;
[0010] Model parameter initialization: The pre-trained parameters of the base model remain unchanged. The parameters of the physical constraint layer are initialized according to the theoretical values of the physical equations. The parameters of the adaptation layer, self-attention mechanism module and output layer are randomly initialized using the Xavier initialization method. The initial learning rate is set to 1e-4 to 1e-3, and the number of iterations is 500-1000.
[0011] Model training steps: Divide the preprocessed data into training set, validation set and test set; use the training set to train the small sample learning model. During the training process, adjust the model parameters according to the model's performance on the validation set to prevent overfitting and improve the model's generalization ability. Through continuous iterative training, the model learns the relationship between soil heavy metal cadmium runoff output flux and various related factors.
[0012] Feature extraction and selection steps: During model training, the model's feature extraction capability is used to extract features closely related to soil heavy metal cadmium runoff output flux from the input data; feature selection algorithms, such as correlation analysis-based methods or machine learning-based feature importance assessment methods, are used to further screen out the most representative features in order to reduce the computational complexity of the model and improve prediction efficiency.
[0013] Model prediction and evaluation steps: Use the trained model to predict the soil heavy metal cadmium runoff flux using data from the test set; evaluate the model's predictive performance by comparing it with actual runoff flux data and using appropriate evaluation indicators such as root mean square error, mean absolute error, and coefficient of determination; optimize and adjust the model based on the evaluation results to improve prediction accuracy.
[0014] Preferably, in the data collection and preprocessing steps, the soil physicochemical property data includes soil pH value, soil organic matter content, and soil texture data, which are obtained through laboratory analysis methods; the meteorological data includes rainfall, temperature, and wind speed data, which are obtained from meteorological monitoring stations; and the topographic data includes slope and aspect data, which are obtained through a geographic information system.
[0015] Preferably, in the step of constructing a few-shot learning model, the pre-training dataset of the basic model needs to contain at least 50 soil environment-related records, covering more than 5 soil types and more than 8 environmental variables, and the number of hidden layer nodes of the fully connected network of the adaptation layer is 64-256, which is dynamically adjusted according to the input feature dimension.
[0016] Preferably, in the model training step, cross-validation is used to divide the training set, validation set, and test set to make full use of the limited small sample data. During the training process, the stochastic gradient descent algorithm or its variants, such as Adagrad, Adadelta, Adam, etc., are used to update the model parameters.
[0017] Preferably, in the model prediction and evaluation step, when the RMSE is less than the set error threshold, the MAE is within an acceptable range, and the R² is greater than the set goodness-of-fit threshold, the model is considered to have good predictive performance and can be used for actual prediction of soil heavy metal cadmium runoff output flux. If the model performance does not meet the requirements, return to the model training step or adjust the parameters in the feature extraction and selection step, and re-optimize the model.
[0018] It has the following beneficial effects:
[0019] This method for predicting soil cadmium runoff flux based on few-sample learning relies on this technique. It requires only 30-50 sets of valid monitoring samples to complete model training and achieve stable predictions. The measured root mean square error (RMSE) is as low as 1.5, and the coefficient of determination (R²) is as high as 0.76. It accurately captures the variation patterns of cadmium runoff flux and avoids overfitting. By leveraging transfer learning to fully utilize pre-trained knowledge from large-scale soil environmental datasets, it only requires collecting readily available indicators such as soil cadmium content, pH value, rainfall, and slope. The cost of a single complete monitoring session is controlled to around 800 yuan. Furthermore, the conventional monitoring equipment is small in size and easy to operate, and its deployment success rate is high in complex terrain areas such as mountains and hills. With an accuracy rate exceeding 90%, the model significantly reduces costs and enhances applicability. The self-attention mechanism embedded in the model can automatically identify core influencing features, with attention weights for soil cadmium content, rainfall, and slope reaching 38%, 32%, and 20%, respectively. The weight of secondary features is less than 10%, ensuring that even if some secondary feature data fluctuate slightly, the increase in prediction error can be controlled within 3%-5%, guaranteeing prediction stability. At the same time, it does not require long-term data accumulation; training and prediction can be completed by covering monitoring data for only one hydrological season (approximately 3 months), significantly shortening the data acquisition cycle. This enables rapid support for pollution prevention and control, helping staff to identify high-risk areas in a timely manner, formulate targeted measures, and reduce the risk of pollution spread. Detailed Implementation
[0020] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0021] This invention provides a technical solution: a method for predicting soil heavy metal cadmium runoff flux based on small sample learning, comprising the following steps:
[0022] Data collection and preprocessing steps: Collect small sample data related to the heavy metal cadmium in soil, including but not limited to soil cadmium content data, soil physicochemical property data, meteorological data, topographic data, and runoff monitoring data; clean the collected data, remove outliers and missing values, and standardize the data to make different types of data comparable;
[0023] Building a few-shot learning model: A few-shot learning model architecture that integrates physical reinforcement transfer learning and self-attention mechanism is adopted. The specific construction steps are as follows:
[0024] Building a basic model for physical-enhanced transfer learning:
[0025] A deep neural network pre-trained on a large-scale soil environment dataset was selected as the core of the basic model. The deep neural network is a Transformer encoder. The large-scale soil environment dataset contains more than 100 soil environment-related records, covering more than 7 soil types and more than 10 environmental variables. The specific variables and pre-training tasks are as follows:
[0026] Core input variables include: basic soil physicochemical properties (pH value, organic matter content, bulk density, clay content), basic soil heavy metal data (content of common heavy metals such as lead, copper, zinc, and cadmium), meteorological parameters (daily rainfall, average daily temperature, humidity, sunshine duration), topographic attributes (altitude, slope, aspect, curvature), and vegetation cover (NDVI index).
[0027] Pre-training objective task: The core task is "prediction of the comprehensive environmental risk index of soil". This index is calculated by weighting the input variables (the weights are determined by expert experience and principal component analysis) and is used to characterize the comprehensive risk level of soil affected by environmental factors.
[0028] Embedded physical mechanism constraint module: Based on the physical equations of cadmium migration and transformation in soil, a physical constraint layer is constructed. This constraint layer acts on the basic model in the following ways. The physical equations of cadmium migration and transformation in soil involved in the physical mechanism constraint module include:
[0029] Convection-diffusion equation: Where C is the cadmium concentration, t is time, D is the dispersion coefficient, v is the pore water flow velocity, and x is the spatial coordinate;
[0030] Runoff scour model: F = k × C × R × S, where F is the cadmium runoff output flux, k is the scour coefficient, R is the runoff volume, and S is the soil surface area.
[0031] Apply a physical dimension consistency check to the intermediate features output by the base model, and remove features that do not conform to the dimension rules. The physical dimension consistency check is specifically as follows:
[0032] Dimensional analysis was performed on the intermediate features output by the model to ensure that the dimensions of soil cadmium concentration (mg / kg), water flow velocity (m / s), and runoff (m³ / s) conformed to the International System of Units (SI). Linear transformation was performed to correct any features that did not conform to the dimensions.
[0033] By introducing a physical parameter regularization term, a penalty term for deviation of physical parameters is added to the model loss function, as shown in the formula: ,in To predict losses, The physical parameter deviation loss is represented by λ, which is the weighting coefficient.
[0034] Model adaptation layer design: An adaptation layer is added to the output of the physical augmentation transfer learning basic model. The adaptation layer consists of 2-3 fully connected networks. Each network layer is followed by a Batch-Normalization layer and a ReLU activation function to convert the general features after physical augmentation into specific features related to the output flux of soil heavy metal cadmium runoff.
[0035] Self-attention mechanism module embedding: A self-attention mechanism module is embedded between the adaptation layer and the output layer. By calculating the attention weights between feature vectors, the influence weight of key features is strengthened, and the interference of irrelevant features is weakened. The calculation process of the self-attention mechanism module is as follows:
[0036] A linear transformation is performed on the feature matrix output by the adaptation layer to generate a query matrix Q, a key matrix K, and a value matrix V.
[0037] Calculate the attention score: Perform a dot product operation on the transposes of Q and K, and then scale by dividing by the square root of the feature dimension to obtain the original attention score matrix;
[0038] The original attention score matrix is normalized using the Softmax function to obtain the attention weight matrix;
[0039] The attention weight matrix and the value matrix V are weighted and summed to output the feature matrix optimized by the attention mechanism.
[0040] The combination of the physical enhancement transfer learning model and the self-attention mechanism module forms a closed loop of "prior knowledge constraints + key feature reinforcement":
[0041] The physical-enhanced transfer learning model builds a "basic logical framework" for the model through physical mechanisms and transfer knowledge, ensuring that all learned patterns conform to the physical nature of Cd transfer and avoiding "spurious associations" under small sample conditions (such as incorrectly associating accidental low Cd flux with irrelevant high sunshine duration).
[0042] The self-attention mechanism further focuses on the association of core features within this framework, maximizing the use of effective information in limited samples and solving the problem of "key patterns being masked by noise" in small samples.
[0043] The two work together to achieve an improvement from "reasonable to accurate": physical constraints ensure that the prediction results are "not out of line" (in line with the basic laws of Cd migration), while the self-attention mechanism "focuses on key points" (strengthens the influence of core factors) on this basis, and finally achieves Cd flux prediction with both physical rationality and prediction accuracy in small samples;
[0044] Output layer construction: A single-layer fully connected network is used as the output layer, and the feature matrix optimized by the attention mechanism is used as the input to output the predicted value of soil heavy metal cadmium runoff output flux.
[0045] Model parameter initialization: The pre-trained parameters of the base model remain unchanged. The parameters of the physical constraint layer are initialized according to the theoretical values of the physical equations. The parameters of the adaptation layer, self-attention mechanism module and output layer are randomly initialized using the Xavier initialization method. The initial learning rate is set to 1e-4 to 1e-3, and the number of iterations is 500-1000.
[0046] Model training steps: Divide the preprocessed data into training set, validation set and test set; use the training set to train the small sample learning model. During the training process, adjust the model parameters according to the model's performance on the validation set to prevent overfitting and improve the model's generalization ability. Through continuous iterative training, the model learns the relationship between soil heavy metal cadmium runoff output flux and various related factors.
[0047] Feature extraction and selection steps: During model training, the model's feature extraction capability is used to extract features closely related to soil heavy metal cadmium runoff output flux from the input data; feature selection algorithms, such as correlation analysis-based methods or machine learning-based feature importance assessment methods, are used to further screen out the most representative features in order to reduce the computational complexity of the model and improve prediction efficiency.
[0048] Model prediction and evaluation steps: Use the trained model to predict the soil heavy metal cadmium runoff flux using data in the test set. Compare the model with the actual runoff flux data and use appropriate evaluation indicators, such as root mean square error (RMSE), mean absolute error (MAE), and coefficient of determination (R²), to evaluate the model's predictive performance. Based on the evaluation results, optimize and adjust the model to improve the accuracy of the prediction.
[0049] The present invention will be further described in detail below with reference to specific embodiments:
[0050] Data collection and preprocessing:
[0051] Small sample data related to the heavy metal cadmium in soil were collected. These included 50 data points on soil cadmium content obtained through laboratory testing; 50 data points each on soil physicochemical properties (pH, organic matter content, and texture) obtained through laboratory analysis; 50 data points each on meteorological data (rainfall, temperature, and wind speed) obtained from local meteorological monitoring stations; 50 data points each on topography (slope and aspect) obtained through a Geographic Information System (GIS); and 50 data points on runoff monitoring.
[0052] The collected data were cleaned, and three soil cadmium content data were found to be outliers and two rainfall data were missing. The outliers were deleted, and the missing values were filled with the mean. Then, all data were standardized to be within the range of [0,1] to make different types of data comparable.
[0053] Building a few-shot learning model:
[0054] Selection of base model for transfer learning: ResNet, which was pre-trained on a dataset containing 100 soil environment-related records (covering 7 soil types and 10 environmental variables), was selected as the base model;
[0055] Model adaptation layer design: Two fully connected networks are added to the output of ResNet as adaptation layers. The first fully connected network has 128 hidden nodes and the second has 64. Each layer is followed by a BatchNormalization layer and a ReLU activation function.
[0056] Embedding of self-attention mechanism module: Construct the module according to the above calculation process of self-attention mechanism module and embed it between the adaptation layer and the output layer;
[0057] Output layer construction: A single-layer fully connected network is used as the output layer, with 1 output node, corresponding to the predicted value of soil heavy metal cadmium runoff output flux;
[0058] Model parameter initialization: The pre-trained parameters of ResNet remain unchanged. The parameters of the adaptation layer, self-attention mechanism module and output layer are randomly initialized using the Xavier initialization method. The initial learning rate is set to 5e-4 and the number of iterations is 800.
[0059] Model training:
[0060] The preprocessed data was divided into training, validation, and test sets in a ratio of 7:1:2. The model was trained using the training set, and the Adam algorithm was used to update the model parameters. During the training process, the model performance was evaluated using the validation set every 100 iterations, and the model parameters were adjusted based on the evaluation results to prevent overfitting.
[0061] Feature extraction and selection:
[0062] During model training, relevant features were extracted using the model's feature extraction capabilities. A correlation analysis-based method was used to calculate the Pearson correlation coefficient between each feature and the soil heavy metal cadmium runoff output flux. Features with an absolute value of correlation coefficient greater than 0.5 were selected, and finally, soil cadmium content, rainfall, and slope were selected as key features.
[0063] Model prediction and evaluation:
[0064] The trained model was used to predict the test set data, and the prediction results were compared with the actual runoff output flux data. The calculated RMSE was 1.5, MAE was 1.3, and R² was 0.76, indicating that the model has good predictive performance.
[0065] The reliability of this method is fully demonstrated through three dimensions: small sample adaptability, accuracy stability, and the effectiveness of the core mechanism. In different types of regions, using typical small sample sizes to train the model, even with low sample sizes, the model can still converge within a reasonable number of iterations. Furthermore, the prediction accuracy shows a reasonable upward trend with increasing sample size, without abnormal fluctuations, proving that it can stably learn the changing patterns of soil heavy metal cadmium runoff flux within a small sample range. For common interference situations in real-world scenarios, such as missing secondary features, outliers, and cross-seasonal data migration, the fluctuation range of the model's prediction accuracy is far below the error tolerance threshold in practical applications, indicating that it can maintain stable prediction results even when the data is incomplete or contains interference. Through the control variable method, removing the transfer learning module or the self-attention mechanism module significantly reduced the model's accuracy, proving that transfer learning is the core for ensuring accuracy under small sample conditions, and the self-attention mechanism is the key to improving anti-interference capabilities. The synergistic effect of these two mechanisms ensures the reliability of the method's performance.
[0066] From a practical application perspective, the examples highlight the superiority of this method through comparative experiments: In the same area, this method only requires monitoring common and easily obtainable indicators, eliminating the need for high-cost physical parameter monitoring. The monitoring equipment is small in size and easy to operate, and has a high success rate in complex terrain areas such as mountains and hills, effectively solving the problems of high cost and poor terrain adaptability of traditional methods. In terms of data acquisition and model training cycle, training and prediction can be completed with only a short period of monitoring data, significantly shortening the overall cycle from data collection to output prediction results. This can quickly provide data support for the prevention and control of cadmium pollution in soil, avoiding the risk of pollution spread due to excessively long data accumulation cycles. In terms of prediction accuracy, this method meets the preset high-performance standards. Compared with simplified models without core modules and traditional small-sample prediction methods, it has a significant accuracy advantage and can more accurately capture the changing patterns of cadmium runoff output flux in soil.
[0067] Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. All other embodiments obtained by those skilled in the art and related fields based on the embodiments of the present invention without inventive effort should fall within the scope of protection of the present invention. Structures, devices, and operating methods not specifically described and explained in the present invention, unless otherwise specified or limited, shall be implemented according to conventional means in the art.
Claims
1. A method for predicting soil heavy metal cadmium runoff output flux based on small sample learning, characterized in that, Comprising the following steps: Data collection and preprocessing step: Collect small sample data related to soil heavy metal cadmium, including but not limited to soil cadmium content data, soil physicochemical property data, meteorological data, topographic data and runoff monitoring data, clean up the collected data, remove outliers and missing values, and standardize the data to make different types of data comparable; Constructing a small sample learning model: a small sample learning model architecture is adopted, which combines physical enhancement transfer learning and self-attention mechanism; Model parameter initialization: the pre-training parameters of the basic model remain unchanged, the physical constraint layer parameters are initialized according to the theoretical value of the physical equation, the parameters of the adaptive layer, the self-attention mechanism module and the output layer are randomly initialized using the Xavier initialization method, the initial learning rate is set to 1e-4 to 1e-3, and the number of iterations is 500-1000 times; Model training step: divide the preprocessed data into training set, validation set and test set, use the training set to train the small sample learning model, in the training process, according to the performance of the model on the validation set, adjust the parameters of the model to prevent overfitting and improve the generalization ability of the model, through continuous iteration training, the model learns the relationship between soil heavy metal cadmium runoff output flux and each related factor; Feature extraction and selection step: in the model training process, the model's feature extraction capability is used to extract features closely related to soil heavy metal cadmium runoff output flux from input data, and feature selection algorithms such as correlation analysis-based methods or machine learning-based feature importance evaluation methods are used to further filter out the most representative features to reduce the computational complexity of the model and improve the prediction efficiency; Model prediction and evaluation step: use the trained model to predict the soil heavy metal cadmium runoff output flux in the test set, compare it with the real runoff output flux data, use appropriate evaluation indicators to evaluate the prediction performance of the model, and optimize and adjust the model according to the evaluation results to improve the prediction accuracy.
2. The method for predicting soil heavy metal cadmium runoff output flux based on small sample learning according to claim 1, characterized in that: In the data collection and preprocessing step, the soil physicochemical property data includes soil pH value, soil organic matter content and soil texture data, which are obtained by laboratory analysis method, the meteorological data includes rainfall, temperature and wind speed data, which are obtained from meteorological monitoring station, and the topographic data includes slope and slope direction data, which are obtained by geographic information system.
3. The method according to claim 1, wherein the method is characterized by: In the step of constructing a small sample learning model, the pre-training data set of the basic model needs to contain at least 50 soil environment related records, covering more than 5 soil types and more than 8 environmental variables, and the number of hidden layer nodes of the fully connected network of the adaptive layer is 64-256, which is dynamically adjusted according to the input feature dimension.
4. The method according to claim 1, wherein the method is characterized by: In the model training step, the training set, validation set and test set are divided by cross-validation to make full use of the limited small sample data, and the random gradient descent algorithm or its variants are used to update the model parameters during the training process.
5. The method according to claim 1, wherein the method is characterized by: In the model prediction and evaluation step, when the RMSE is less than the set error threshold, and the MAE is within the acceptable range, and the R² is greater than the set goodness-of-fit threshold, the prediction performance of the model is considered good, and can be used for actual soil heavy metal cadmium runoff output flux prediction. If the model performance does not meet the requirements, return to the model training step or adjust the parameters in the feature extraction and selection step to re-optimize the model.
Citation Information
Patent Citations
Small sample learning and LSTM (Long Short Term Memory)-based runoff prediction method for areas lacking data
CN114372631A
Space-time attention air pollution prediction method based on physical diffusion process
CN118862656A
Mineral resource intelligent prediction method and system based on multi-source heterogeneous data fusion and deep learning
CN121303465A
System and Method for Multi-Modal Hyperspectral Image Generation with Cross-Modal Attention and Adaptive Quality Assurance
US20250315932A1