Rapid detection method for soil heavy metal pollution based on hyperspectral imaging
By using spectral preprocessing and nonlinear mapping models of hyperspectral imaging technology, a multidimensional spectral feature set of heavy metals in soil is extracted, which solves the problem of long detection time in traditional detection methods and realizes rapid, full-domain spatial detection of heavy metal pollution in soil.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- YANTAI TONGWO MECHANICAL EQUIP CO LTD
- Filing Date
- 2026-04-24
- Publication Date
- 2026-05-29
AI Technical Summary
Traditional methods for detecting heavy metal pollution in soil are cumbersome and time-consuming. Existing hyperspectral imaging technology has failed to effectively combine spatial dimension information for full-domain calculations, resulting in limited improvement in spectral data quality and making it difficult to meet the needs of rapid detection.
By acquiring hyperspectral data of the soil surface, performing spectral preprocessing, extracting multidimensional spectral feature sets, and constructing a nonlinear mapping model for point-by-point inversion calculation, a spatial distribution map of heavy metal content is generated.
It enables rapid and accurate spatial distribution detection of heavy metal pollution in soil, simplifies the detection process, and is suitable for rapid detection application scenarios.
Smart Images

Figure CN122108980A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of hyperspectral technology for soil pollution detection, specifically a rapid detection method for heavy metal pollution in soil based on hyperspectral imaging. Background Technology
[0002] Traditional methods for detecting heavy metal pollution in soil often rely on laboratory chemical analysis, requiring on-site sampling followed by digestion and titration, resulting in a cumbersome and time-consuming process. Existing hyperspectral imaging technologies often directly analyze full-spectrum data or extract only single spectral features to characterize soil heavy metal properties. Some technologies only retrieve heavy metal content from a single soil sample, failing to incorporate hyperspectral spatial information for comprehensive computation. Spectral preprocessing often involves only noise reduction, without simultaneously integrating noise removal, optical path scattering correction, and smoothing, resulting in limited improvement in spectral data quality. Full-spectrum features suffer from data redundancy, and single spectral features cannot fully reflect the spectral response patterns of heavy metals. Single-point retrieval methods cannot present the spatial distribution of heavy metal pollution in soil, making it difficult to meet the demand for rapid spatial detection of heavy metal pollution in soil. Summary of the Invention
[0003] The present invention aims to solve the technical problems existing in the prior art; Therefore, this invention proposes a rapid detection method for heavy metal pollution in soil based on hyperspectral imaging, comprising: Obtain raw hyperspectral data of the soil surface in the target area, wherein the raw hyperspectral data includes spatial pixel coordinate information and spectral continuous reflectance information; The original hyperspectral data is subjected to a spectral preprocessing process, which includes noise removal, optical path scattering correction, and spectral smoothing and denoising, to generate standard hyperspectral data with improved spectral quality. The spectral response range reflecting the characteristics of heavy metals in the soil is extracted from the standard hyperspectral data. The spectral response range is divided into multiple spectral sub-bands, and the spectral absorption depth, absorption peak symmetry and spectral first-order differential eigenvalue are calculated for each spectral sub-band to form a multidimensional spectral feature set. A nonlinear mapping model is constructed to associate the multidimensional spectral feature set with the soil heavy metal content. The nonlinear mapping model is trained using spectral data of soil samples with known heavy metal concentrations to generate a fully trained heavy metal content prediction model. Based on the heavy metal content prediction model, the multidimensional spectral feature set of each pixel in the standard hyperspectral data is inverted point by point to calculate and output a spatial distribution map of heavy metal content in the soil of the target area.
[0004] Further, the spectral preprocessing procedure performed on the original hyperspectral data includes noise removal, optical path scattering correction, and spectral smoothing and denoising to generate standard hyperspectral data with improved spectral quality, including: The original hyperspectral data is processed using a noise removal algorithm based on statistical thresholds to remove abnormal spectral bands whose reflectance values exceed the normal soil reflectance range, thus completing the noise removal process. The optical path scattering correction process is completed by applying a multivariate scattering correction algorithm, using the average reflectance spectrum of the entire hyperspectral image as the reference spectrum, and performing linear transformation correction on the spectral curve of each pixel to eliminate the scattering effect caused by uneven distribution of soil particles and surface roughness. A convolutional smoothing algorithm is used to filter the spectral curve of each pixel after correction to suppress the high-frequency jitter of the spectrum caused by random noise of the instrument, while preserving the absorption and reflection characteristics of the soil spectrum, thus completing the spectral smoothing and denoising process. The hyperspectral data, after noise removal, optical path scattering correction, and spectral smoothing and denoising, are integrated and reconstructed into the standard hyperspectral data with spatial-spectral information alignment.
[0005] Furthermore, the step of extracting spectral response intervals reflecting soil heavy metal characteristics from the standard hyperspectral data, dividing the spectral response intervals into multiple spectral sub-bands, and calculating the spectral absorption depth, absorption peak symmetry, and first-order differential spectral characteristic values for each spectral sub-band to form a multidimensional spectral feature set, including: Based on the characteristic absorption bands of typical heavy metal elements in soil, one or more continuous spectral band ranges are defined as the spectral response intervals within the full spectrum of the standard hyperspectral data. For each defined spectral response interval, it is divided into multiple spectral subbands with narrower band ranges using equal intervals or non-equal intervals based on absorption characteristics. For each spectral sub-band that has been divided, perform the following calculation process: Calculate the difference between the lowest reflectance point within the spectral sub-band and the average reflectance of the two shoulder areas, and normalize the difference to obtain the spectral absorption depth characteristic value of the spectral sub-band. Calculate the similarity measure of the spectral curves on the left and right sides with the point of lowest reflectance as the center, and use the similarity measure as the absorption peak symmetry characteristic value of the spectral subband. Calculate the average value and variance of the first derivative of the spectral reflectance curve within the spectral sub-band, and combine the average value and variance as the spectral first-order differential eigenvalue of the spectral sub-band. The spectral absorption depth characteristic values, absorption peak symmetry characteristic values, and first-order differential characteristic values of all spectral sub-bands are collected and arranged in ascending order of wavelength to form a multidimensional spectral feature set characterizing the spectral morphology of the corresponding pixels.
[0006] Further, the step of calculating the difference between the lowest reflectance point within the spectral sub-band and the average reflectance of the two shoulder areas, and then normalizing the difference to obtain the spectral absorption depth characteristic value of the spectral sub-band, includes: In the spectral curve of the spectral sub-band, locate the wavelength position with the minimum reflectance value and denote it as the absorption center point; Using the absorption center point as the boundary, select multiple consecutive bands in both the short-wave and long-wave directions. Calculate the average reflectance of the selected bands in the short-wave direction as the reflectance of the left shoulder, and calculate the average reflectance of the selected bands in the long-wave direction as the reflectance of the right shoulder. Calculate the arithmetic mean of the reflectivity of the left shoulder and the reflectivity of the right shoulder as the comprehensive shoulder reflectivity; The difference between the overall shoulder reflectivity and the reflectivity at the absorption center point is calculated to obtain the original absorption depth; Dividing the original absorption depth by the comprehensive shoulder reflectance yields the normalized spectral absorption depth characteristic value.
[0007] Furthermore, the construction of the nonlinear mapping model for relating the multidimensional spectral feature set to soil heavy metal content includes: An artificial neural network with a multi-layer hidden layer structure is selected as the basic architecture of the nonlinear mapping model, and the number of input layer nodes of the artificial neural network is the same as the dimension of the multidimensional spectral feature set. The output layer of the artificial neural network is designed with the same number of nodes as the types of heavy metal elements to be detected, and each output node corresponds to the predicted content of one heavy metal element. The hidden layers of the artificial neural network employ nonlinear activation functions, and a random deactivation mechanism is introduced between the hidden layers to prevent model overfitting. A loss function is defined with the objective of minimizing the error between the predicted content and the laboratory measured content, and an optimization algorithm based on the gradient descent principle is selected to train the nonlinear mapping model.
[0008] Furthermore, the step of training the nonlinear mapping model using spectral data of soil samples with known heavy metal concentrations to generate a fully trained heavy metal content prediction model includes: Collect soil samples with laboratory chemical analysis results, which include the exact concentration values of various heavy metal elements; Hyperspectral data of the soil sample surface was acquired using a hyperspectral imaging device. After the spectral preprocessing process, the multidimensional spectral feature set corresponding to each soil sample was extracted to form a training sample dataset. Each sample in the training sample dataset consists of a multidimensional spectral feature vector and a corresponding heavy metal concentration true value vector. The training sample dataset is divided into a training subset and a validation subset according to a preset ratio; Using the data from the training subset, the nonlinear mapping model is trained iteratively. In each iteration, the multidimensional spectral feature vector is input into the nonlinear mapping model to obtain the predicted heavy metal concentration. The loss between the predicted heavy metal concentration and the true heavy metal concentration is calculated, and the weight parameters of the artificial neural network are adjusted using the optimization algorithm. After each complete training iteration, the performance of the nonlinear mapping model is evaluated using the data from the validation subset. The training process is terminated when the prediction error of the nonlinear mapping model on the validation subset no longer decreases after multiple consecutive iterations. The model parameters obtained at the end of training are fixed and saved as the fully trained heavy metal content prediction model.
[0009] Further, dividing the training sample dataset into a training subset and a validation subset according to a preset ratio includes: Read the complete training sample dataset consisting of multidimensional spectral feature vectors and corresponding heavy metal concentration true value vectors, and randomly sort the samples in the training sample dataset to eliminate systematic bias caused by sample collection order or numbering. Set a first proportion of the training subset to the total number of samples and a second proportion of the validation subset to the total number of samples. The sum of the first proportion and the second proportion is 1, and the value of the first proportion is greater than the value of the second proportion. Based on the set first ratio value, the first N samples are selected sequentially from the randomly sorted training sample dataset and assigned to the training subset, where N is equal to the total number of samples multiplied by the first ratio value and rounded down. The remaining samples that were not selected into the training subset are automatically assigned to the validation subset to ensure that there is no sample overlap between the training subset and the validation subset, and that they together cover the complete training sample dataset. The samples in the training subset are labeled according to the type of heavy metal element, while the samples in the validation subset retain the corresponding ground truth labels of heavy metal concentrations, so as to independently evaluate the model performance during the model training process.
[0010] Furthermore, based on the heavy metal content prediction model, the multidimensional spectral feature set of each pixel in the standard hyperspectral data is inverted point by point to calculate and output a spatial distribution map of heavy metal content in the soil of the target area, including: Traverse every spatial pixel in the standard hyperspectral data; For the spatial pixel currently being traversed, a multidimensional spectral feature set is extracted from its corresponding spectral curve using the same method as in the training phase, forming the feature vector of the spatial pixel. The feature vectors of the spatial pixels are input into the input layer of the fully trained heavy metal content prediction model; The fully trained heavy metal content prediction model performs forward propagation calculations on the feature vector of the input spatial pixel, and obtains the predicted content values of various heavy metal elements corresponding to the current spatial pixel at the output layer. The calculated predicted content values of various heavy metal elements are associated and stored with the spatial coordinate information of the spatial pixels; After all spatial pixels have been traversed, the predicted content value of each heavy metal element is rendered and a spatial distribution map of the corresponding heavy metal element in the target area is generated based on its associated spatial coordinate information.
[0011] Furthermore, the method also includes a detection result optimization step based on spatial autocorrelation: After generating the spatial distribution map of heavy metal content, the local spatial autocorrelation index of the predicted content of each pixel in the spatial distribution map of heavy metal content and the predicted content of pixels in its spatial neighborhood is calculated. Pixels whose local spatial autocorrelation index is lower than a preset threshold are identified and marked as spatial outliers; Obtain the original hyperspectral reflectance curve corresponding to the spatial anomaly point, and calculate the spectral angle cosine distance between it and the average spectral curve of the surrounding neighboring pixels. When the spectral angle cosine distance is greater than the preset spectral difference threshold, the spatial anomaly is determined to be a spectral anomaly, and the median value of the predicted content of its spatial neighboring pixels is used to replace the original predicted content value of the spectral anomaly. The spatial distribution map of heavy metal content is updated using the replaced predicted content values to generate a spatially continuous optimized distribution map of heavy metal content.
[0012] Furthermore, before acquiring the raw hyperspectral data of the soil surface in the target area, a standard whiteboard calibration step is also included: Before collecting soil data with the hyperspectral imaging device, the hyperspectral imaging device is aimed at a standard diffuse reflectance white board to obtain the hyperspectral reflectance data of the standard white board under the same lighting conditions. Calculate the average reflectance value of each band from the hyperspectral reflectance data of the standard whiteboard; When collecting raw hyperspectral data of the soil in the subsequent process, the raw digital signal value received by each pixel in each band is divided by the average reflectance value of the standard white board in the corresponding band to achieve relative standardization of reflectance and obtain the raw hyperspectral data.
[0013] Compared with the prior art, the beneficial effects of the present invention are: The spectral response intervals reflecting soil heavy metal characteristics are divided into multiple spectral sub-bands. Within each sub-band, the spectral absorption depth, absorption peak symmetry, and first-order differential eigenvalues are calculated independently, and these are combined to form a multidimensional spectral feature set. Screening the spectral response intervals eliminates irrelevant spectral interference, sub-band division refines the granularity of spectral characterization, and simultaneous calculation of multiple features captures the spectral variation patterns corresponding to heavy metals from different dimensions. This avoids analytical interference caused by redundancy in full-spectrum features, compensates for the incomplete characterization of single spectral features, weakens the impact of spectral noise and optical path scattering on feature extraction, and strengthens the correlation between spectral features and soil heavy metal content.
[0014] Based on a well-trained nonlinear mapping model, point-by-point inversion calculations are performed on the multidimensional spectral feature set of each pixel in standard hyperspectral data, directly outputting a spatial distribution map of heavy metal content in the target area's soil. Pixel-level point-by-point inversion covers all spatial dimensions of hyperspectral data, extending traditional single-point inversion to full-domain spatial computation. The nonlinear mapping model adapts to the complex correlation between spectral features and heavy metal content, and the point-by-point inversion results are directly converted into spatial distribution maps, intuitively presenting the spatial distribution pattern of soil heavy metal content. This simplifies the process of obtaining spatial distribution information on soil heavy metal pollution and is suitable for rapid detection applications. Attached Figure Description
[0015] Figure 1 This is a flowchart illustrating the steps of the rapid detection method for heavy metal pollution in soil based on hyperspectral imaging as described in this invention. Figure 2 A flowchart for calculating spectral absorption depth characteristic values; Figure 3 Loss curve for training neural network models; Figure 4 To verify the statistical distribution box plots of the contents of the three heavy metals cadmium, lead, and arsenic within the set; Figure 5 The training performance curve of the heavy metal content prediction model. Detailed Implementation
[0016] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0017] Example See Figure 1 This invention provides a rapid detection method for heavy metal pollution in soil based on hyperspectral imaging, the method comprising: The process begins by acquiring raw hyperspectral data of the soil surface in the target area. This raw hyperspectral data includes spatial pixel coordinates and continuous reflectance information in the spectral dimension. A spectral preprocessing procedure is then performed on the raw hyperspectral data, including noise removal, optical path scattering correction, and spectral smoothing and denoising, to generate standard hyperspectral data with improved spectral quality. Spectral response intervals reflecting soil heavy metal characteristics are extracted from the standard hyperspectral data. These intervals are divided into multiple spectral sub-bands, and the spectral absorption depth, absorption peak symmetry, and first-order differential eigenvalues are calculated for each sub-band, forming a multidimensional spectral feature set. A nonlinear mapping model is constructed to correlate the multidimensional spectral feature set with soil heavy metal content. This model is trained using spectral data from soil samples with known heavy metal concentrations, generating a fully trained heavy metal content prediction model. Based on this prediction model, a point-by-point inversion calculation is performed on the multidimensional spectral feature set of each pixel in the standard hyperspectral data, outputting a spatial distribution map of heavy metal content in the target area's soil.
[0018] In one embodiment of the present invention, before the hyperspectral imaging device acquires soil data, a standard whiteboard calibration step is performed. The hyperspectral imaging device is aligned with a standard diffuse reflectance whiteboard to acquire hyperspectral reflectance data of the standard whiteboard under the same illumination conditions. The average reflectance value of each band is calculated from the hyperspectral reflectance data of the standard whiteboard. When subsequently acquiring raw hyperspectral data of the soil, the raw digital signal value received by each pixel in each band is divided by the average reflectance value of the corresponding band of the standard whiteboard to achieve relative standardization of reflectance, thus obtaining the raw hyperspectral data.
[0019] The raw hyperspectral data undergoes a spectral preprocessing procedure to generate standard hyperspectral data. This includes processing the raw hyperspectral data using a statistical threshold-based noise removal algorithm to remove abnormal spectral bands with reflectance values exceeding the normal soil reflectance range, thus completing noise removal. A multivariate scattering correction algorithm is applied, using the average reflectance spectrum of the entire hyperspectral image as a reference spectrum, to perform linear transformation correction on the spectral curve of each pixel, eliminating scattering effects caused by uneven soil particle distribution and surface roughness, thus completing optical path scattering correction. A convolutional smoothing algorithm is then used to filter the spectral curve of each corrected pixel, suppressing high-frequency spectral jitter caused by instrument random noise while preserving the absorption and reflection characteristics of the soil spectrum, thus completing spectral smoothing and denoising. Finally, the hyperspectral data processed by noise removal, optical path scattering correction, and spectral smoothing and denoising are integrated to reconstruct standard hyperspectral data with spatially and spectrally aligned information.
[0020] In practice, heavy metal pollution testing was conducted on farmland soil near an industrial area. A hyperspectral imaging device scanned the target area under fixed height and illumination conditions to acquire raw hyperspectral data. A standard whiteboard calibration step was performed before the hyperspectral imaging device acquired soil data. The device was aligned with a standard diffuse reflectance whiteboard to obtain hyperspectral reflectance data under the same illumination conditions. The average reflectance value for each band was calculated from the standard whiteboard's hyperspectral reflectance data. When subsequently acquiring raw hyperspectral data of the soil, the raw digital signal value received by each pixel in each band was divided by the corresponding band's average reflectance value from the standard whiteboard, achieving relative standardization of reflectance and obtaining the raw hyperspectral data. Reflectance standardization was achieved using the following formula:
[0021] in: This represents the soil reflectance of band b. This represents the original digital signal value of the soil in band b. The standard whiteboard average reflectance value for band b is represented by the standard whiteboard calibration step, which makes the data collected at different times or under different lighting conditions comparable. The reflectance value in the original hyperspectral data reflects the true spectral characteristics of the soil surface.
[0022] In some embodiments, a spectral preprocessing procedure is performed on the raw hyperspectral data to generate standard hyperspectral data. This includes processing the raw hyperspectral data using a statistical threshold-based noise removal algorithm to remove abnormal spectral bands whose reflectance values exceed the normal soil reflectance range, thus completing noise removal processing; applying a multivariate scattering correction algorithm, using the average reflectance spectrum of the entire hyperspectral image as a reference spectrum, to perform linear transformation correction on the spectral curve of each pixel, eliminating scattering effects caused by uneven soil particle distribution and surface roughness, thus completing optical path scattering correction processing; using a convolutional smoothing algorithm to filter the spectral curve of each pixel after correction, suppressing high-frequency spectral jitter caused by instrument random noise, while preserving the absorption and reflection characteristics of the soil spectrum, thus completing spectral smoothing and denoising processing; and then integrating the hyperspectral data after noise removal processing, optical path scattering correction processing, and spectral smoothing and denoising processing to reconstruct standard hyperspectral data with spatial-spectral information alignment.
[0023] In the data comparison, after noise removal processing, the abnormal bands of the original hyperspectral data were removed, the outliers in the spectral curves were reduced, and after optical path scattering correction processing, the spectral curves of different pixels were corrected in terms of baseline shift and amplification effect, and the spectral shape tended to be consistent. After spectral smoothing and denoising processing, the high-frequency fluctuations of the spectral curves were smoothed, and the characteristic absorption valleys and reflection peaks were more clearly visible.
[0024] In one embodiment of the present invention, a spectral response interval reflecting the characteristics of heavy metals in soil is extracted from standard hyperspectral data. This spectral response interval is divided into multiple spectral sub-bands, and characteristic values are calculated for each sub-band to form a multidimensional spectral feature set. This process includes defining one or more continuous spectral band ranges as spectral response intervals within the full spectrum of standard hyperspectral data, based on the characteristic absorption bands of typical heavy metal elements in the soil. For each defined spectral response interval, it is divided into multiple photonic bands with narrower band ranges using either equal or non-equal intervals based on absorption characteristics. For each defined spectral sub-band, the difference between the lowest reflectance point within the sub-band and the average reflectance of the two shoulder areas is calculated. This difference is normalized and used as the spectral absorption depth characteristic value of the sub-band. A similarity measure of the spectral curve shapes on both sides of the lowest reflectance point is calculated and used as the absorption peak symmetry characteristic value of the sub-band. The average and variance of the first derivative of the spectral reflectance curve within the sub-band are calculated, and the combination of this average and variance is used as the spectral first-order differential characteristic value of the sub-band. The spectral absorption depth characteristic values, absorption peak symmetry characteristic values, and first-order differential characteristic values of all spectral sub-bands are collected and arranged in ascending order of wavelength to form a multidimensional spectral feature set characterizing the spectral morphology of the corresponding pixels.
[0025] See Figure 2The calculation of the spectral absorption depth feature value involves locating the wavelength position with the minimum reflectance value in the spectral curve of the spectral sub-band, denoted as the absorption center point. Multiple consecutive bands are selected in both the short-wave and long-wave directions, using the absorption center point as the boundary. The average reflectance of the selected bands in the short-wave direction is calculated as the left shoulder reflectance, and the average reflectance of the selected bands in the long-wave direction is calculated as the right shoulder reflectance. The arithmetic mean of the left and right shoulder reflectances is calculated as the composite shoulder reflectance. The difference between the composite shoulder reflectance and the reflectance at the absorption center point is calculated to obtain the original absorption depth. The original absorption depth is divided by the composite shoulder reflectance to obtain the normalized spectral absorption depth feature value. In data comparison, the original continuous spectral curve, after feature extraction, is transformed into a fixed-dimensional multidimensional spectral feature set vector. This multidimensional spectral feature set vector contains multidimensional information such as spectral absorption depth, absorption peak symmetry, and the first-order differential of the spectrum.
[0026] Heavy metal pollution detection was conducted on farmland soil near an industrial area. Preprocessed standard hyperspectral data was obtained, which includes spatial pixel coordinate information and spectral continuous reflectance information. In practice, spectral response intervals reflecting the characteristics of heavy metals in the soil were extracted from the standard hyperspectral data. The spectral response intervals were divided into multiple spectral sub-bands, and feature values were calculated for each spectral sub-band to form a multidimensional spectral feature set. Based on the characteristic absorption bands of typical heavy metal elements in the soil, one or more continuous spectral band ranges were defined as spectral response intervals within the full spectrum of the standard hyperspectral data. In practice, each defined spectral response interval was divided into multiple spectral sub-bands with narrower band ranges using equal intervals or non-equal intervals based on absorption characteristics.
[0027] In some embodiments, for each divided spectral sub-band, a feature calculation process is performed to calculate the difference between the lowest reflectance point within the spectral sub-band and the average reflectance of the two shoulder areas. This difference is normalized and used as the spectral absorption depth feature value of the spectral sub-band. The similarity measure of the spectral curve shapes on the left and right sides, centered on the lowest reflectance point, is calculated and used as the absorption peak symmetry feature value of the spectral sub-band. The absorption peak symmetry feature value is calculated using the following formula:
[0028] in: Indicates the symmetry eigenvalue of the absorption peak. This represents the vector consisting of the reflectance values of the i-th band to the left of the absorption center point. Let represent the vector consisting of the reflectance values of the i-th band to the right of the absorption center point. This represents the average value of the reflectivity vector on the left. The mean value of the reflectance vector on the right is represented by , and n represents the number of selected bands. The mean value and variance of the first derivative of the spectral reflectance curve within the spectral sub-band are calculated. The mean value and variance are combined as the first differential spectral characteristic value of the spectral sub-band. The spectral absorption depth characteristic value, absorption peak symmetry characteristic value and first differential spectral characteristic value calculated for all spectral sub-bands are collected and arranged in ascending order of wavelength of the spectral sub-bands to form a multidimensional spectral feature set characterizing the spectral morphology of the corresponding pixel.
[0029] In one embodiment of the present invention, a nonlinear mapping model for correlating a multidimensional spectral feature set with soil heavy metal content is constructed. This process includes selecting an artificial neural network with a multi-layer hidden layer structure as the basic architecture of the nonlinear mapping model. The number of nodes in the input layer of the artificial neural network is the same as the dimension of the multidimensional spectral feature set. The output layer of the artificial neural network is designed with the same number of nodes as the number of heavy metal elements to be detected, with each output node corresponding to the predicted content of one heavy metal element. The hidden layers of the artificial neural network employ nonlinear activation functions, and a random deactivation mechanism is introduced between the hidden layers to prevent model overfitting. A loss function is defined with the objective of minimizing the error between the predicted content and the laboratory measured content, and an optimization algorithm based on the gradient descent principle is selected to train the nonlinear mapping model.
[0030] To detect heavy metal pollution in farmland soil near an industrial area, a model needs to be built to predict the concentrations of cadmium, lead, and arsenic in the soil from an extracted set of multidimensional spectral features. This involves constructing a nonlinear mapping model to correlate the multidimensional spectral feature set with the heavy metal content in the soil. Specifically, this includes selecting an artificial neural network with multiple hidden layers as the basic architecture of the nonlinear mapping model. The number of nodes in the input layer of the artificial neural network is the same as the dimension of the multidimensional spectral feature set. In practice, the input layer of the artificial neural network receives a fixed-length multidimensional spectral feature vector, where each dimension corresponds to a specific feature value, such as spectral absorption depth, absorption peak symmetry, or the first-order differential of the spectrum. The output layer of the artificial neural network is designed with the same number of nodes as the number of heavy metal elements to be detected. In the example scenario, the heavy metal elements to be detected include cadmium, lead, and arsenic; therefore, the output layer of the nonlinear mapping model is designed with three nodes, each corresponding to the predicted content of one heavy metal element.
[0031] In some embodiments, the hidden layers of the artificial neural network employ nonlinear activation functions, and a random deactivation mechanism is introduced between the hidden layers to prevent model overfitting. The nonlinear activation function enables the artificial neural network to learn and simulate the complex nonlinear relationship between a multidimensional spectral feature set and soil heavy metal content. The random deactivation mechanism randomly ignores a portion of the hidden layer nodes during model training, thereby reducing the model's over-reliance on noise or specific samples in the training data. A loss function is defined with the objective of minimizing the error between the predicted content and the laboratory measured content. The loss function is defined by the following formula:
[0032] in: This represents the value of the loss function. Indicates the number of training samples. This indicates the number of different types of heavy metal elements to be predicted. Indicates the first In the training samples, the th Laboratory measured content values of several heavy metal elements. Represents the nonlinear mapping model for the first... In the training samples, the th The predicted content values of a certain heavy metal element can be understood as follows: the loss function measures the overall difference between the model's predicted value and the actual value; and an optimization algorithm based on the gradient descent principle is selected to train the nonlinear mapping model. The optimization algorithm calculates the gradient of the loss function with respect to each weight parameter of the artificial neural network and updates the parameters in the opposite direction of the gradient in order to gradually minimize the value of the loss function.
[0033] See Figure 3In the training process of a neural network model for detecting heavy metal pollution in soil based on hyperspectral imaging, the loss curves intuitively reflect the model's learning convergence and generalization performance. The figure uses a logarithmic scale (log scale) on the vertical axis and the training epochs on the horizontal axis, plotting the training set loss (solid line) and validation set loss (dashed line) curves respectively: In the early training phase (0-10 epochs), both types of losses decrease rapidly, indicating that the model efficiently learns the nonlinear mapping relationship between spectral features and heavy metal content, and gradient updates effectively reduce prediction errors. In the middle training phase (10-40 epochs), the loss curves continue to converge, with the training set loss consistently lower than the validation set loss, and both maintaining a similar decreasing trend without significant divergence or separation, indicating that the model has not suffered severe overfitting and has good generalization ability. In the later training phase (40-50 epochs), the training set loss continues to decrease steadily to near the order of 1, while the validation set loss fluctuates slightly at a low level, reflecting the model's continued fitting ability on the training set. Simultaneously, the fluctuation in the validation set also reflects the combined effect of sample noise and model regularization (such as random inactivation). The overall shape of the loss curve conforms to the typical convergence law of deep learning training: rapid initial descent, steady convergence in the middle stage, and stabilization in the later stage. This verifies the effectiveness of the constructed multi-layer neural network architecture with random deactivation mechanism and the gradient descent-based optimization algorithm in the task of predicting soil heavy metal content, and provides a reliable model foundation for subsequent spatial distribution inversion of heavy metal content.
[0034] In one embodiment of the present invention, a nonlinear mapping model is trained using spectral data of soil samples with known heavy metal concentrations to generate a fully trained heavy metal content prediction model. This process includes collecting soil samples with laboratory chemical analysis results, which contain the exact concentration values of various heavy metal elements. Hyperspectral data of the soil sample surface is acquired using a hyperspectral imaging device, and after spectral preprocessing, a multidimensional spectral feature set corresponding to each soil sample is extracted to form a training sample dataset. Each sample in the training sample dataset consists of a multidimensional spectral feature vector and a corresponding ground truth vector of heavy metal concentration. The training sample dataset is divided into a training subset and a validation subset according to a preset ratio. The nonlinear mapping model is trained iteratively using the data from the training subset. In each iteration, the multidimensional spectral feature vector is input into the nonlinear mapping model to obtain the predicted heavy metal concentration. The loss between the predicted heavy metal concentration and the ground truth heavy metal concentration is calculated, and the weight parameters of the artificial neural network are adjusted using an optimization algorithm. After each complete training iteration, the performance of the nonlinear mapping model is evaluated using the data from the validation subset. The training process is terminated when the prediction error of the nonlinear mapping model on the validation subset no longer decreases after multiple consecutive iterations. The model parameters obtained at the end of training are fixed and saved as a fully trained heavy metal content prediction model.
[0035] A nonlinear mapping model is trained using spectral data from soil samples with known heavy metal concentrations to generate a fully trained heavy metal content prediction model. In the example scenario, the goal is to predict the content of three heavy metals—cadmium, lead, and arsenic—from farmland soil samples. Training the nonlinear mapping model using spectral data from soil samples with known heavy metal concentrations to generate a fully trained heavy metal content prediction model involves collecting soil samples with laboratory chemical analysis results containing the exact concentration values of multiple heavy metals. Hyperspectral data of the soil sample surface is acquired using a hyperspectral imaging device, and after spectral preprocessing, a multidimensional spectral feature set corresponding to each soil sample is extracted to form a training sample dataset. Each sample in the training sample dataset consists of a multidimensional spectral feature vector and a corresponding ground truth vector of heavy metal concentration. In specific implementation, one hundred soil samples are collected from the target area and its surroundings. The exact concentration values of cadmium, lead, and arsenic in each soil sample are obtained in the laboratory using atomic absorption spectrometry, forming a ground truth vector of heavy metal concentration. At the same time, each soil sample undergoes hyperspectral imaging scanning, spectral preprocessing, and feature extraction to generate a multidimensional spectral feature vector containing forty-five feature dimensions.
[0036] The training sample dataset is divided into a training subset and a validation subset according to a preset ratio. In practice, a complete training sample dataset consisting of multidimensional spectral feature vectors and corresponding heavy metal concentration truth vectors is read. The samples in the training sample dataset are randomly sorted to eliminate systematic biases caused by the sample collection order or numbering. A first ratio of the training subset to the total number of samples and a second ratio of the validation subset to the total number of samples are set. The sum of the first and second ratios is one, and the first ratio is greater than the second ratio. Based on the set first ratio, the first subset is selected sequentially from the randomly sorted training sample dataset. N samples are assigned to the training subset, where N equals the total number of samples multiplied by the first proportion value and rounded down. The remaining samples not selected for the training subset are automatically assigned to the validation subset. The samples in the training subset are labeled according to the type of heavy metal element, while the samples in the validation subset retain their corresponding ground truth labels for heavy metal concentrations, which are used to independently evaluate model performance during model training. Refer to Table 1. In the example scenario, the first proportion value is set to 0.8 and the second proportion value is set to 0.2. From one hundred randomly sorted samples, the first eighty samples are selected as the training subset, and the remaining twenty samples are selected as the validation subset. The partitioning results are shown in Table 1.
[0037] Table 1: Training Sample Dataset Partition Table
[0038] In some embodiments, for adjusting the weight parameters of an artificial neural network using an optimization algorithm, the optimization algorithm updates the weights and bias parameters in the network according to the gradient direction of the loss function. After each complete training iteration, the model performance is evaluated using data from a validation subset, and the validation error is calculated using the following formula:
[0039] in: This represents the root mean square error of the current model on the validation subset. This indicates the number of samples in the validation subset. This indicates the number of different types of heavy metal elements. The model represents the first The first verification sample Predicted content of each element Indicates the first The first verification sample The true values of the laboratory measured content of each element are used to terminate the training process when the prediction error of the model on the validation subset no longer decreases after multiple iterations. It can be understood that the validation subset provides an independent dataset to evaluate the model's ability to generalize to unseen data and prevents the model from overfitting on the training subset.
[0040] See Figure 4In the statistical analysis phase of the sample data for the rapid detection method of heavy metal pollution in soil based on hyperspectral imaging, the distribution characteristics of the three target heavy metal elements, cadmium (Cd), lead (Pb), and arsenic (As), within the validation set can be quantitatively analyzed using box plots. Statistically, the validation set samples for the three heavy metal elements all exhibit similar distribution patterns: Median level: The median content of cadmium (Cd), lead (Pb), and arsenic (As) is stable within the range of 1.4–1.5 mg / kg, indicating a high degree of consistency in the central trends of the three heavy metals in the validation set samples, providing a balanced benchmark reference for model training. Interquartile Range (IQR): The box heights (i.e., interquartile ranges) of the three elements are similar, indicating a similar degree of dispersion in their core data distribution. The box ranges for lead (Pb) and arsenic (As) are slightly wider than those for cadmium (Cd), reflecting slightly larger sample fluctuations for both in the medium content range. Extreme Value Range: The whisker concentrations of all three elements extend to approximately 0.2–3.0 mg / kg, covering the complete range of concentration fluctuations. The validation set includes data across the entire gradient from low to high concentrations, effectively verifying the model's generalization prediction ability for samples with different pollution levels. Distribution Symmetry: The relative positions of the box plots and the median show that the concentration distributions of the three heavy metals exhibit a slight right skewness, meaning that the tails of high-concentration samples are slightly longer, consistent with the actual distribution patterns of heavy metal pollution in farmland soils. Box plots visually present the statistical dispersion and distribution characteristics of the validation set samples, providing a reliable sample distribution basis for subsequent training error evaluation and generalization ability verification of the nonlinear mapping model. They also serve as a benchmark for comparing the consistency between the model's predictions and the measured true values.
[0041] In one embodiment of the present invention, based on a heavy metal content prediction model, a point-by-point inversion calculation is performed on the multidimensional spectral feature set of each pixel in the standard hyperspectral data to output a spatial distribution map of heavy metal content in the soil of the target area. This process includes traversing each spatial pixel in the standard hyperspectral data. For the currently traversed spatial pixel, a multidimensional spectral feature set is extracted from its corresponding spectral curve using the same method as in the training phase to form a feature vector for the spatial pixel. The feature vector of the spatial pixel is input into the input layer of the fully trained heavy metal content prediction model. The fully trained heavy metal content prediction model performs forward propagation calculation on the feature vector of the input spatial pixel, and obtains the predicted content values of multiple heavy metal elements corresponding to the current spatial pixel at the output layer. The calculated predicted content values of multiple heavy metal elements are associated and stored with the spatial coordinate information of the spatial pixel. After all spatial pixels have been traversed, the predicted content value of each heavy metal element is rendered according to its associated spatial coordinate information to generate a spatial distribution map of the content of the corresponding heavy metal element in the target area.
[0042] After generating the spatial distribution map of heavy metal content, a detection result optimization step based on spatial autocorrelation is performed. This involves calculating the local spatial autocorrelation index between the predicted content of each pixel in the spatial distribution map and the predicted content of pixels in its spatial neighborhood. Pixels with local spatial autocorrelation indices below a preset threshold are identified and marked as spatial outliers. The original hyperspectral reflectance curve corresponding to each spatial outlier is obtained, and the spectral angle cosine distance between it and the average spectral curve of its surrounding neighboring pixels is calculated. When the spectral angle cosine distance is greater than a preset spectral difference threshold, the spatial outlier is determined to be a spectral outlier, and its original predicted content value is replaced with the median value of the predicted content of its spatial neighboring pixels. The spatial distribution map of heavy metal content is updated using the replaced predicted content value, generating a spatially continuous optimized heavy metal content distribution map.
[0043] Heavy metal pollution detection was conducted on farmland soil near an industrial area. Standard hyperspectral data and a well-trained heavy metal content prediction model were obtained. Based on the heavy metal content prediction model, the multidimensional spectral feature set of each pixel in the standard hyperspectral data was inverted point by point to calculate and output a spatial distribution map of heavy metal content in the target area soil. This process involves traversing every spatial pixel in the standard hyperspectral data. In the specific implementation, the standard hyperspectral data contains 100,000 spatial pixels. The program code uses a double loop structure to traverse each row and column of the image to access each pixel. For the currently traversed spatial pixel, the multidimensional spectral feature set is extracted from its corresponding spectral curve using the same method as in the training phase to form the feature vector of the spatial pixel. It can be understood that the extraction method of this feature vector includes defining the spectral response interval, dividing sub-bands, and calculating the spectral absorption depth feature value, absorption peak symmetry feature value, and spectral first-order differential feature value to ensure complete consistency with the feature dimensions and calculation logic used during model training.
[0044] The feature vectors of spatial pixels are input into the input layer of a well-trained heavy metal content prediction model. The model performs forward propagation calculations on the input feature vectors and obtains the predicted content values of various heavy metal elements corresponding to the current spatial pixel at the output layer. In data comparison, for a spatial pixel, the input 45-dimensional feature vector undergoes multi-layer nonlinear calculations by an artificial neural network model, ultimately outputting a vector containing the predicted concentration values of three heavy metals. The calculated predicted content values of various heavy metal elements are associated with the spatial coordinate information of the spatial pixel and stored. After all spatial pixels have been traversed, the predicted content value of each heavy metal element is rendered based on its associated spatial coordinate information to generate a spatial distribution map of the corresponding heavy metal element in the target area. During the rendering process, each spatial coordinate point is assigned a color, and the color intensity corresponds to the level of the predicted heavy metal content value at that point, thus forming a visual spatial distribution map.
[0045] In some embodiments, after generating the spatial distribution map of heavy metal content, a detection result optimization step based on spatial autocorrelation is performed to calculate the local spatial autocorrelation index between the predicted content of each pixel in the spatial distribution map of heavy metal content and the predicted content of pixels in its spatial neighborhood. The local spatial autocorrelation index is calculated using the following formula:
[0046] in: The local spatial autocorrelation index represents the local spatial autocorrelation index of spatial cell i. This represents the predicted heavy metal content of spatial pixel i. This represents the predicted heavy metal content of spatial pixel j. This represents the average predicted heavy metal content of all pixels in the entire study area. This represents the variance of the predicted heavy metal content across the entire study area. This represents the spatial weight between spatial pixel i and spatial pixel j. This represents the total number of pixels in the neighborhood of spatial pixel i. Pixels with a local spatial autocorrelation index lower than a preset threshold are identified and marked as spatial outliers. In the example scenario, the preset threshold is set to -1.5. Pixels with a local spatial autocorrelation index lower than -1.5 are considered outliers with extremely weak spatial correlation with their surrounding pixels.
[0047] Obtain the original hyperspectral reflectance curve corresponding to the spatial anomaly point, and calculate the spectral angle cosine distance between it and the average spectral curve of the surrounding neighboring pixels. In some embodiments, the surrounding neighborhood can be defined as all pixels within a 3×3 or 5×5 window centered on the spatial anomaly point. When the spectral angle cosine distance is greater than a preset spectral difference threshold, the spatial anomaly point is determined to be a spectral anomaly point, and the median value of the predicted content of its spatial neighboring pixels is used to replace the original predicted content value of the spectral anomaly point. It can be understood that the median is not sensitive to extreme values and can better represent the concentration trend of the neighborhood content. The spatial distribution map of heavy metal content is updated using the replaced predicted content value to generate a spatially continuous optimized heavy metal content distribution map.
[0048] Optional, spatial weights The threshold can be defined based on the reciprocal of the Euclidean distance between spatial pixels i and j, or based on their binary relationship of adjacency. In the data comparison, the unoptimized distribution map contained a few isolated, abnormally high-brightness points with significantly different content values from their surrounding areas. After optimization, the values of these points were replaced by the median content of their neighborhoods, making the final distribution map more spatially continuous and smooth, better reflecting the continuous spatial diffusion pattern of soil pollutants. Optionally, the preset spectral difference threshold can be set based on the average difference range between normal pixels and their neighborhood spectra in historical data. In the data comparison, the spatial distribution map of heavy metal content, after optimization based on spatial autocorrelation, exhibits better spatial continuity in its spatial statistical characteristics, such as the variogram, and spatially anomalous noise is effectively suppressed.
[0049] See Figure 5 In evaluating the training performance of the heavy metal content prediction model, the convergence trend and generalization ability of the model are characterized by joint analysis of the training and validation set loss curves. Specifically, a 45-dimensional hyperspectral feature vector is used as the model input, which is mapped through a multi-layer nonlinear neural network to output predicted values for three heavy metal concentrations. The training process uses mean squared error as the basic loss function, while a random deactivation mechanism is introduced to suppress overfitting. The training set loss (dashed line) and validation set loss (solid line) decrease synchronously with the number of iterations, reflecting that the model maintains good generalization while learning the feature-concentration mapping relationship: the training set loss continuously decreases from approximately 0.30 to below 0.03, while the validation set loss converges synchronously to around 0.06, and the two curves do not show significant separation, indicating that the model has not overfitted. When the loss value decreases to the preset convergence threshold (0.05), the training set loss continues to decrease slowly, while the validation set loss tends to stabilize, meeting the convergence criteria for the heavy metal content prediction task. During parameter configuration, the maximum number of training iterations is set to 50, and the early stopping trigger condition is that the validation set loss does not decrease significantly after 5 consecutive iterations, ensuring that the model can fully learn the feature patterns while avoiding redundant calculations.
Claims
1. A rapid detection method for heavy metal pollution in soil based on hyperspectral imaging, characterized in that, The method includes: Obtain raw hyperspectral data of the soil surface in the target area, wherein the raw hyperspectral data includes spatial pixel coordinate information and spectral continuous reflectance information; The original hyperspectral data is subjected to a spectral preprocessing process, which includes noise removal, optical path scattering correction, and spectral smoothing and denoising, to generate standard hyperspectral data with improved spectral quality. The spectral response range reflecting the characteristics of heavy metals in the soil is extracted from the standard hyperspectral data. The spectral response range is divided into multiple spectral sub-bands, and the spectral absorption depth, absorption peak symmetry and spectral first-order differential eigenvalue are calculated for each spectral sub-band to form a multidimensional spectral feature set. A nonlinear mapping model is constructed to associate the multidimensional spectral feature set with the soil heavy metal content. The nonlinear mapping model is trained using spectral data of soil samples with known heavy metal concentrations to generate a fully trained heavy metal content prediction model. Based on the heavy metal content prediction model, the multidimensional spectral feature set of each pixel in the standard hyperspectral data is inverted point by point to calculate and output a spatial distribution map of heavy metal content in the soil of the target area.
2. The rapid detection method for soil heavy metal pollution based on hyperspectral imaging according to claim 1, characterized in that, The process of performing spectral preprocessing on the original hyperspectral data includes noise removal, optical path scattering correction, and spectral smoothing and denoising to generate standard hyperspectral data with improved spectral quality. The original hyperspectral data is processed using a noise removal algorithm based on statistical thresholds to remove abnormal spectral bands whose reflectance values exceed the normal soil reflectance range, thus completing the noise removal process. The optical path scattering correction process is completed by applying a multivariate scattering correction algorithm, using the average reflectance spectrum of the entire hyperspectral image as the reference spectrum, and performing linear transformation correction on the spectral curve of each pixel to eliminate the scattering effect caused by uneven distribution of soil particles and surface roughness. A convolutional smoothing algorithm is used to filter the spectral curve of each pixel after correction, suppressing the high-frequency spectral jitter caused by instrument random noise, while preserving the absorption and reflection characteristics of the soil spectrum, thus completing the spectral smoothing and denoising process. The hyperspectral data, after noise removal, optical path scattering correction, and spectral smoothing and denoising, are integrated and reconstructed into the standard hyperspectral data with spatial-spectral information alignment.
3. The rapid detection method for soil heavy metal pollution based on hyperspectral imaging according to claim 2, characterized in that, The process involves extracting spectral response intervals reflecting soil heavy metal characteristics from the standard hyperspectral data, dividing these intervals into multiple spectral sub-bands, and calculating the spectral absorption depth, absorption peak symmetry, and first-order differential spectral eigenvalues for each sub-band, thus constructing a multidimensional spectral feature set, including: Based on the characteristic absorption bands of typical heavy metal elements in soil, one or more continuous spectral band ranges are defined as the spectral response intervals within the full spectrum of the standard hyperspectral data. For each defined spectral response interval, it is divided into multiple spectral subbands with narrower band ranges using equal intervals or non-equal intervals based on absorption characteristics. For each spectral sub-band that has been divided, perform the following calculation process: Calculate the difference between the lowest reflectance point within the spectral sub-band and the average reflectance of the two shoulder areas, and normalize the difference to obtain the spectral absorption depth characteristic value of the spectral sub-band. Calculate the similarity measure of the spectral curves on the left and right sides with the point of lowest reflectance as the center, and use the similarity measure as the absorption peak symmetry characteristic value of the spectral subband. Calculate the average value and variance of the first derivative of the spectral reflectance curve within the spectral sub-band, and combine the average value and variance as the spectral first-order differential eigenvalue of the spectral sub-band. The spectral absorption depth characteristic values, absorption peak symmetry characteristic values, and first-order differential characteristic values of all spectral sub-bands are collected and arranged in ascending order of wavelength to form a multidimensional spectral feature set characterizing the spectral morphology of the corresponding pixels.
4. The rapid detection method for soil heavy metal pollution based on hyperspectral imaging according to claim 3, characterized in that, The step of calculating the difference between the lowest reflectance point within the spectral sub-band and the average reflectance of the two shoulder areas, and then normalizing the difference to obtain the spectral absorption depth characteristic value of the spectral sub-band, includes: In the spectral curve of the spectral sub-band, locate the wavelength position with the minimum reflectance value and denote it as the absorption center point; Using the absorption center point as the boundary, select multiple consecutive bands in both the short-wave and long-wave directions. Calculate the average reflectance of the selected bands in the short-wave direction as the reflectance of the left shoulder, and calculate the average reflectance of the selected bands in the long-wave direction as the reflectance of the right shoulder. Calculate the arithmetic mean of the reflectivity of the left shoulder and the reflectivity of the right shoulder as the comprehensive shoulder reflectivity; The difference between the overall shoulder reflectivity and the reflectivity at the absorption center point is calculated to obtain the original absorption depth; Dividing the original absorption depth by the comprehensive shoulder reflectance yields the normalized spectral absorption depth characteristic value.
5. The rapid detection method for soil heavy metal pollution based on hyperspectral imaging according to claim 4, characterized in that, The construction of the nonlinear mapping model for relating the multidimensional spectral feature set to soil heavy metal content includes: An artificial neural network with a multi-layer hidden layer structure is selected as the basic architecture of the nonlinear mapping model, and the number of input layer nodes of the artificial neural network is the same as the dimension of the multidimensional spectral feature set. The output layer of the artificial neural network is designed with the same number of nodes as the types of heavy metal elements to be detected, and each output node corresponds to the predicted content of one heavy metal element. The hidden layers of the artificial neural network employ nonlinear activation functions, and a random deactivation mechanism is introduced between the hidden layers to prevent model overfitting. A loss function is defined with the objective of minimizing the error between the predicted content and the laboratory measured content, and an optimization algorithm based on the gradient descent principle is selected to train the nonlinear mapping model.
6. The rapid detection method for soil heavy metal pollution based on hyperspectral imaging according to claim 5, characterized in that, The process of training the nonlinear mapping model using spectral data from soil samples with known heavy metal concentrations to generate a fully trained heavy metal content prediction model includes: Collect soil samples with laboratory chemical analysis results, which include the exact concentration values of various heavy metal elements; Hyperspectral data of the soil sample surface was acquired using a hyperspectral imaging device. After the spectral preprocessing process, the multidimensional spectral feature set corresponding to each soil sample was extracted to form a training sample dataset. Each sample in the training sample dataset consists of a multidimensional spectral feature vector and a corresponding heavy metal concentration true value vector. The training sample dataset is divided into a training subset and a validation subset according to a preset ratio; Using the data from the training subset, the nonlinear mapping model is trained iteratively. In each iteration, the multidimensional spectral feature vector is input into the nonlinear mapping model to obtain the predicted heavy metal concentration. The loss between the predicted heavy metal concentration and the true heavy metal concentration is calculated, and the weight parameters of the artificial neural network are adjusted using the optimization algorithm. After each complete training iteration, the performance of the nonlinear mapping model is evaluated using the data from the validation subset. The training process is terminated when the prediction error of the nonlinear mapping model on the validation subset no longer decreases after multiple consecutive iterations. The model parameters obtained at the end of training are fixed and saved as the fully trained heavy metal content prediction model.
7. The rapid detection method for soil heavy metal pollution based on hyperspectral imaging according to claim 6, characterized in that, The step of dividing the training sample dataset into a training subset and a validation subset according to a preset ratio includes: Read the complete training sample dataset consisting of multidimensional spectral feature vectors and corresponding heavy metal concentration true value vectors, and randomly sort the samples in the training sample dataset to eliminate systematic bias caused by sample collection order or numbering. Set a first proportion of the training subset to the total number of samples and a second proportion of the validation subset to the total number of samples. The sum of the first proportion and the second proportion is 1, and the value of the first proportion is greater than the value of the second proportion. Based on the set first ratio value, the first N samples are selected sequentially from the randomly sorted training sample dataset and assigned to the training subset, where N is equal to the total number of samples multiplied by the first ratio value and rounded down. The remaining samples that were not selected into the training subset are automatically assigned to the validation subset to ensure that there is no sample overlap between the training subset and the validation subset, and that they together cover the complete training sample dataset. The samples in the training subset are labeled according to the type of heavy metal element, while the samples in the validation subset retain the corresponding ground truth labels of heavy metal concentrations, so as to independently evaluate the model performance during the model training process.
8. The rapid detection method for soil heavy metal pollution based on hyperspectral imaging according to claim 7, characterized in that, The method based on the heavy metal content prediction model performs point-by-point inversion calculations on the multidimensional spectral feature set of each pixel in the standard hyperspectral data, outputting a spatial distribution map of heavy metal content in the soil of the target area, including: Traverse every spatial pixel in the standard hyperspectral data; For the spatial pixel currently being traversed, a multidimensional spectral feature set is extracted from its corresponding spectral curve using the same method as in the training phase, forming the feature vector of the spatial pixel. The feature vectors of the spatial pixels are input into the input layer of the fully trained heavy metal content prediction model; The fully trained heavy metal content prediction model performs forward propagation calculation on the feature vector of the input spatial pixel, and obtains the predicted content values of various heavy metal elements corresponding to the current spatial pixel in the output layer. The calculated predicted content values of various heavy metal elements are associated and stored with the spatial coordinate information of the spatial pixels; After all spatial pixels have been traversed, the predicted content value of each heavy metal element is used to render a spatial distribution map of the corresponding heavy metal element in the target area based on its associated spatial coordinate information.
9. The rapid detection method for soil heavy metal pollution based on hyperspectral imaging according to claim 8, characterized in that, The method also includes a detection result optimization step based on spatial autocorrelation: After generating the spatial distribution map of heavy metal content, the local spatial autocorrelation index of the predicted content of each pixel in the spatial distribution map of heavy metal content and the predicted content of pixels in its spatial neighborhood is calculated. Pixels whose local spatial autocorrelation index is lower than a preset threshold are identified and marked as spatial outliers; Obtain the original hyperspectral reflectance curve corresponding to the spatial anomaly point, and calculate the spectral angle cosine distance between it and the average spectral curve of the surrounding neighboring pixels. When the spectral angle cosine distance is greater than the preset spectral difference threshold, the spatial anomaly is determined to be a spectral anomaly, and the median value of the predicted content of its spatial neighboring pixels is used to replace the original predicted content value of the spectral anomaly. The spatial distribution map of heavy metal content is updated using the replaced predicted content values to generate a spatially continuous optimized distribution map of heavy metal content.
10. The rapid detection method for soil heavy metal pollution based on hyperspectral imaging according to claim 9, characterized in that, Before acquiring the raw hyperspectral data of the soil surface in the target area, a standard whiteboard calibration step is also included: Before collecting soil data with the hyperspectral imaging device, the hyperspectral imaging device is aimed at a standard diffuse reflectance white board to obtain the hyperspectral reflectance data of the standard white board under the same lighting conditions. Calculate the average reflectance value of each band from the hyperspectral reflectance data of the standard whiteboard; When collecting raw hyperspectral data of the soil in the subsequent process, the raw digital signal value received by each pixel in each band is divided by the average reflectance value of the standard white board in the corresponding band to achieve relative standardization of reflectance and obtain the raw hyperspectral data.