A soil heavy metal pollution detection method and system based on spectral data

By constructing a nonlinear mapping characteristic factor group and a credibility factor group and dynamically adjusting the model weights, the problem of insufficient accuracy in soil heavy metal concentration estimation in existing technologies is solved, and high-precision detection and system robustness in the medium and low concentration ranges are achieved, which is suitable for complex soil environments.

CN120369649BActive Publication Date: 2025-09-30WEIFANG XINBO PHYSICAL & CHEM TESTING CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510863508.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-09-30
Estimated Expiration
2045-06-26

AI Technical Summary

Technical Problem

The existing technology estimates soil heavy metal concentrations based on linear relationships, ignoring the nonlinear spectral response of the complex component structure of the soil and the occurrence state of heavy metals, resulting in low prediction accuracy in the medium and low concentration ranges.

Method used

A soil heavy metal pollution detection method based on spectral data is adopted. By obtaining the spectral reflectance image of the soil surface, a hyperspectral feature sequence is constructed, the nonlinear mapping feature factor group and the credibility factor group are calculated, the credibility representation value is constructed, the reliability level labels are divided, the model weights are dynamically adjusted, the local fitting sub-model is called, and the final predicted concentration results are output.

Benefits of technology

It improves the accuracy of heavy metal identification and fitting in the low and medium concentration ranges, enhances the robustness and adaptability of the detection system, can improve the stability and interpretability of prediction results in complex environments, has self-correction capabilities, and is suitable for the detection needs of actual soil environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120369649B_ABST
    Figure CN120369649B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of soil detection technology, and discloses a method and system for detecting heavy metal contamination in soil based on spectral data. The detection method comprises: obtaining a spectral reflectance image of the soil surface in the test area, extracting a hyperspectral feature sequence to construct a spectral response curve for the target soil; calculating a nonlinear mapping feature factor group and a credibility factor group based on the spectral response curve; constructing a credibility representation value based on the two factors and assigning reliability level labels; and determining whether there is model fitting distortion, whether to call a local fitting sub-model, and whether to adjust the model weights based on the labels, and outputting a final predicted concentration result. The present invention has high detection accuracy for relatively low levels of heavy metals in soil.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of soil detection technology, and in particular to a method and system for detecting soil heavy metal pollution based on spectral data. Background Art

[0002] In recent years, with the development of spectral imaging technology, multispectral and hyperspectral sensing technologies have been gradually applied to the identification and estimation of heavy metal content in soil. Existing techniques often compare collected soil spectra with established spectral libraries, using spectral angle matching and linear regression models to predict concentrations.

[0003] Chinese patent publication number: CN113791040B discloses a soil heavy metal detection method and system. This detection method combines near-infrared spectral analysis with hyperspectral imaging technology. It first obtains a photo and spectral image of the soil, then extracts spectral feature data from the spectral image and calculates the spectral angle with the spectral feature data in the spectral library. The similarity between the two is determined by the size of the spectral angle, thereby qualitatively analyzing the type of heavy metal in the soil. The concentration of heavy metals in the soil is then quantitatively calculated using a preset linear regression equation calculation program between the spectral feature data and the absolute concentration. The detection system includes a flight device, an analysis system, and an optical camera, a multispectral sensor, and a GPS positioning system mounted on the flight device; the spectral library is located in the analysis system.

[0004] Although the above scheme can be used for on-site and rapid analysis of the types and concentrations of heavy metals in soil, the following problems still exist:

[0005] Concentration estimation based on linear relationships ignores the nonlinear spectral response caused by the complex component structure of the soil and the occurrence status of heavy metals, resulting in low prediction accuracy in the medium and low concentration ranges. Summary of the Invention

[0006] In view of this, the present invention proposes a soil heavy metal pollution detection method and system based on spectral data to solve the problems existing in the prior art.

[0007] On the one hand, the present invention proposes a method for detecting heavy metal pollution in soil based on spectral data, comprising:

[0008] Obtain the spectral reflectance image of the soil surface in the test area, extract the corresponding hyperspectral feature sequence and construct the spectral response curve of the target soil;

[0009] Based on the spectral response curve, calculating a nonlinear mapping characteristic factor group and a credibility factor group of the soil sample;

[0010] Constructing a credibility representation value according to the credibility factor group and the nonlinear mapping characteristic factor group;

[0011] Based on the credibility characterization value, reliability level labels are divided, and the following steps are performed according to the reliability level labels: based on the change of the nonlinear mapping characteristic factor group of the soil sample, whether there is model fitting distortion is determined to determine whether to call a local fitting sub-model;

[0012] and, based on the credibility representation value, determining whether to perform model weight adjustment;

[0013] Output the final predicted concentration result;

[0014] The nonlinear mapping characteristic factor group is calculated based on the multi-order nonlinear function fitting residual of the hyperspectral principal component, and the credibility factor group is calculated based on the residual amplitude and the spectral structure similarity.

[0015] Furthermore, the calculation process of the nonlinear mapping characteristic factor group includes:

[0016] Performing multi-order polynomial fitting on the hyperspectral feature sequence to obtain a predicted concentration sequence;

[0017] Calculating the residual sum of squares between the predicted concentration sequence and the original concentration label, and determining a first nonlinear characteristic factor based on the ratio of the residual sum of squares to a reference residual threshold; the original concentration label is the actual heavy metal concentration value measured in the laboratory;

[0018] The deviation between the spectral response curve and its kernel principal component projection curve is taken as the second nonlinear characteristic factor;

[0019] performing weighted summation of the first nonlinear characteristic factor and the second nonlinear characteristic factor to form a nonlinear mapping characteristic factor group of the sample;

[0020] All soil samples are stored after establishing association relationships with the nonlinear mapping characteristic factor group.

[0021] Furthermore, the calculation process of the credibility factor group includes:

[0022] The residual between the spectral response curve of each soil sample and the model-predicted concentration was extracted as the first credibility factor;

[0023] Calculating the concentration deviation dispersion of the current soil sample under multiple model prediction paths as a second credibility factor; the multiple models include a global fitting model, a local fitting sub-model, a multi-fitting path model, a regularized fitting model, and a kernel function regression model;

[0024] Calculating the spectral structure cosine similarity between the current spectral response curve and the most similar sample in the spectral library as a third credibility factor; the spectral library is a database that stores heavy metal types, corresponding contents, and corresponding spectral response curves;

[0025] Taking a weighted average of the first credibility factor, the second credibility factor, and the third credibility factor to construct the credibility factor group;

[0026] An index table is established for each soil sample and credibility factor group.

[0027] Furthermore, the calculation process of the credibility representation value includes:

[0028] Analyze the credibility factor groups of different soil samples in the same test area, and obtain the difference between the maximum credibility factor group and the minimum credibility factor group as the credibility fluctuation;

[0029] The ratio of the credibility fluctuation to the preset credibility fluctuation threshold is used as the credibility representation value.

[0030] Furthermore, the process of classifying reliability level labels includes:

[0031] If the credibility representation value is less than the preset credibility low risk threshold, it is marked as a high credibility label;

[0032] If the credibility representation value is higher than the preset credibility high risk threshold, it is marked as a low credibility label;

[0033] If the credibility representation value is between the credibility low risk threshold and the credibility high risk threshold, it is marked as a medium credibility label.

[0034] Furthermore, if the detection result is a low confidence label, the following process is triggered:

[0035] Based on the changes in the nonlinear mapping characteristic factor group of the current soil sample, the degree of model fitting distortion is judged. If it is higher than the fitting deviation threshold, the local fitting sub-model is called, and the local fitting sub-model is used to predict the heavy metal concentration in the current test area.

[0036] Furthermore, the process of determining the degree of model fitting distortion includes:

[0037] Calculate the degree of deviation between the nonlinear mapping characteristic factor group of the current soil sample and the historical average characteristic factor group;

[0038] If the degree of deviation is higher than a preset fitting deviation threshold, it is determined that there is fitting distortion, and the local sub-model is called to refit the path.

[0039] Furthermore, the model weight adjustment includes:

[0040] When the reliability level label is a medium credibility label, weight adjustment is performed on the multiple models based on the first credibility factor, the second credibility factor, and the third credibility factor;

[0041] If the first credibility factor of any model is greater than the preset residual threshold, the weight of the current model is reduced;

[0042] If the second credibility factor of any model is greater than the preset discreteness threshold, the weight of the current model is reduced;

[0043] If the third credibility factor of any model is less than the similarity threshold, the weight of the current model is reduced;

[0044] After the model weights are adjusted, the weight adjustment results are used to perform weighted summation on the prediction results of the multiple models to output the final predicted concentration result.

[0045] Furthermore, the process of constructing the hyperspectral feature sequence includes:

[0046] Based on the collected spectral response curve, the main component characteristics of the spectrum, the band ratio characteristics and the slope change characteristics of the continuous spectrum band in the specified band of 450-2450nm are extracted; and a multidimensional combined feature vector is constructed as the hyperspectral feature sequence of the soil sample;

[0047] The process of obtaining soil spectral reflectance images includes:

[0048] Obtain spectral reflectance images of different soil surface layers in the test area; divide the spectral reflectance images into multiple sub-areas and extract the average spectral response curve of each sub-area; and extract the spectral response curves of all sub-areas in turn as input for subsequent processing.

[0049] Compared with the prior art, the present invention has the following beneficial effects:

[0050] The present invention improves the ability of spectral information to express the occurrence status of heavy metals in soil by introducing the construction of hyperspectral feature sequences and multidimensional combination feature extraction, and can more comprehensively reflect the spectral response characteristics of soil components at different concentration levels.

[0051] During the model construction process, the present invention adopts a nonlinear mapping characteristic factor group to replace the traditional linear fitting method. By fusing multi-order polynomial residuals with kernel principal component deviations, the modeling ability in the nonlinear response region is improved, which is particularly suitable for the identification and fitting of medium and low concentrations of heavy metals.

[0052] By constructing the first, second, and third credibility factors and weighting them to form a credibility factor group, a multi-angle comprehensive evaluation of the stability of the prediction results and the adaptability of the model can be conducted, focusing not only on the predicted values ​​but also on the ability to judge the credibility of the prediction results themselves.

[0053] Introducing a division mechanism of credibility characterization values ​​and reliability level labels into the detection process enables the system to flexibly call local fitting sub-models or adjust model weights according to different credibility levels, thereby enhancing local adaptability while maintaining the versatility of the overall model.

[0054] When the system determines that the detection result is of medium or low confidence, it can actively trigger the model path adjustment or refitting process based on the deviation between historical data and current features, thereby improving the robustness of the prediction results in scenarios with environmental disturbances or strong sample specificity.

[0055] Through the dynamic adjustment mechanism of model weights and combining the performance of different models in residuals, discreteness and similarity, the impact of incompatible models can be effectively weakened, the interpretability and rationality of the overall prediction results can be enhanced, and the risk of misjudgment under the dominance of a single model can be avoided.

[0056] In addition, the linkage mechanism between the constructed spectral library and the model enables the system to have a certain self-correction ability when facing samples with heterogeneous spectral characteristics. It can form a closed loop of "judgment-adjustment-re-prediction" during the detection process, thereby improving the flexibility and intelligence level of the detection process.

[0057] On the other hand, the present invention proposes a soil heavy metal pollution detection system based on spectral data, comprising:

[0058] An acquisition module is configured to obtain a spectral reflectance image of the soil surface in the test area, extract a corresponding hyperspectral feature sequence, and construct a spectral response curve of the target soil;

[0059] a calculation module configured to calculate a nonlinear mapping characteristic factor group and a credibility factor group of the soil sample based on the spectral response curve;

[0060] a processing module configured to classify reliability level labels based on the credibility characterization values, and adjust a process for obtaining the predicted concentration according to the reliability level labels;

[0061] The output module is configured to output the final predicted concentration result.

[0062] It should be noted that the soil heavy metal pollution detection method and system based on spectral data of the present invention have the same beneficial effects and will not be described in detail here. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present invention. The same reference symbols are used throughout the drawings to represent the same components. In the drawings:

[0064] Figure 1 A flow chart of a method for detecting heavy metal pollution in soil based on spectral data provided by an embodiment of the present invention.

[0065] Figure 2 This is a functional block diagram of a soil heavy metal pollution detection system based on spectral data provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0066] The exemplary embodiments disclosed in the present application will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art. It should be noted that, unless there is a conflict, the embodiments of the present invention and the features described in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0067] See Figure 1 As shown, an embodiment of the present invention provides a method for detecting heavy metal pollution in soil based on spectral data, comprising:

[0068] S1: Obtain the spectral reflectance image of the soil surface in the test area, extract the corresponding hyperspectral feature sequence and construct the spectral response curve of the target soil;

[0069] S2: Based on the spectral response curve, calculate the nonlinear mapping characteristic factor group and credibility factor group of the soil sample;

[0070] S3: Constructing a credibility representation value based on the credibility factor group and the nonlinear mapping characteristic factor group;

[0071] S4: dividing the reliability level labels based on the credibility characterization values, and performing the following steps according to the reliability level labels: judging whether there is model fitting distortion based on the changes in the nonlinear mapping characteristic factor group of the soil sample, so as to determine whether to call the local fitting sub-model;

[0072] and, based on the credibility representation value, determining whether to adjust the model weight;

[0073] S5: Output the final predicted concentration result;

[0074] Among them, the nonlinear mapping characteristic factor group is calculated based on the residuals of the multi-order nonlinear function fitting of the hyperspectral principal component, and the credibility factor group is calculated based on the similarity between the residual amplitude and the spectral structure.

[0075] It can be understood that this method retains the important spectral information of soil reflectance in the 450-2450nm band through the systematic extraction of hyperspectral feature sequences in step S1, and constructs a more detailed spectral response curve through multidimensional features (such as principal components, ratio features, spectral slope, etc.), providing higher information dimensions and data resolution for subsequent modeling and error evaluation.

[0076] In step S2, a breakthrough in traditional linear modeling was achieved by calculating a set of nonlinear mapping characteristic factors and a set of credibility factors. The introduction of nonlinear mapping characteristic factors enabled the model to identify the nonlinear relationship between soil heavy metal content and spectral response, improving prediction accuracy in the low and medium concentration ranges. The introduction of credibility factors enabled self-diagnosis, identifying the stability and applicability of the current prediction results.

[0077] In step S3, a credibility representation value is constructed, and multiple credibility factors are further summarized into a single indicator, so that the system can identify the overall reliability of sample predictions in a structured manner, avoid misleading overall judgments due to local fluctuations of a certain factor, and help improve the robustness and responsiveness of the overall prediction system.

[0078] The reliability level labeling and response strategy design in step S4 ensure that the system not only outputs concentration values ​​but also provides results screening and risk warning capabilities. If the reliability is determined to be low or medium, the system proactively determines whether model fitting distortion exists and triggers adjustments to local fitting sub-models or model weights. This dynamic feedback and local adaptation mechanism enhances versatility and interference resistance in complex soil conditions.

[0079] Finally, S5 completes model path adjustments or prediction optimization before outputting results, ensuring the output of predicted concentration results with high reliability, stability, and interpretability, suitable for the diverse and complex testing needs of actual soil environments. The overall process forms an intelligent testing system with a closed loop of "feature construction - factor extraction - trusted diagnosis - model regulation - result output."

[0080] In some embodiments of the present application, the calculation process of the nonlinear mapping characteristic factor group includes:

[0081] Perform multi-order polynomial fitting on the hyperspectral feature sequence to obtain the predicted concentration sequence;

[0082] The residual sum of squares between the predicted concentration series and the original concentration labels is calculated, and the first nonlinear characteristic factor is determined based on the ratio of the residual sum of squares to a reference residual threshold; the original concentration labels are the actual heavy metal concentration values ​​measured in the laboratory;

[0083] The deviation between the spectral response curve and its kernel principal component projection curve is taken as the second nonlinear characteristic factor;

[0084] The first nonlinear characteristic factor and the second nonlinear characteristic factor are weighted and summed to form a nonlinear mapping characteristic factor group of the sample;

[0085] All soil samples are stored after establishing association relationships with nonlinear mapping characteristic factor groups.

[0086] It is understandable that applying multi-order polynomial regression analysis to the hyperspectral feature sequence and establishing a predictive model for soil spectral features and heavy metal concentrations yields a predicted concentration sequence. Multi-order polynomial fitting can better capture the potential nonlinear relationship between spectra and concentrations, avoiding complex phenomena that cannot be described by simple linear models.

[0087] The magnitude of the fitting error is measured by calculating the sum of squared residuals between the predicted concentration series and the actual concentrations measured in the laboratory (raw concentration labels). The sum of squared residuals is then compared with a preset reference residual threshold to obtain the first nonlinear characteristic factor, which reflects the relative magnitude of the model fitting error under the reference standard. A larger value indicates a more significant degree of nonlinearity in the fitting.

[0088] Calculate the deviation between the spectral response curve and its kernel principal component projection curve. Kernel principal component projection is a representation of high-dimensional spectral data mapped into a low-dimensional principal component space using kernel principal component analysis. The deviation reflects the difference between the original spectral data and its principal component projection. Larger values ​​indicate greater complexity or anomalies in the sample's spectral structure, suggesting the presence of nonlinear mapping characteristics.

[0089] The first nonlinear characteristic factor and the second nonlinear characteristic factor are weighted and summed according to the preset weights to obtain the overall nonlinear mapping characteristic factor group of the sample, which serves as an important basis for subsequent model judgment and processing.

[0090] The nonlinear mapping feature factor groups corresponding to all samples are associated with the sample information and stored to facilitate subsequent rapid query and model training updates.

[0091] It should be noted that the degree of deviation represents the degree of difference between two data sequences. Specifically, the original hyperspectral response curve is mapped to a low-dimensional space using kernel principal component analysis to obtain the principal component projection curve of the sample. The difference between the original spectral curve and the projected curve is calculated using a certain metric, such as Euclidean distance, cosine distance, or other similarity metrics. This difference is the "deviation," which reflects the degree of deviation of the original spectral data from its low-dimensional principal component representation.

[0092] To construct the nonlinear mapping characteristic factor set, the first and second nonlinear characteristic factors are first numerically normalized to eliminate differences in their dimensionality and numerical range. This normalization method can employ linear normalization, mapping the numerical value of each factor to the range of 0 to 1. Specifically, the minimum value of each factor across all samples is subtracted from the factor value, and then divided by the difference between the maximum and minimum values ​​across all samples, ensuring that all factor values ​​fall within the range of 0 to 1 after normalization. After normalization, correlation metrics, such as correlation coefficients or goodness-of-fit statistics, are calculated between the two normalized factors and the actual test accuracy based on a large amount of historical sample data. Based on this correlation metric, a weighted ratio is automatically determined between the two factors. The weights reflect the contribution of each factor to the accuracy of the test results, and the sum of the weights is guaranteed to be 1. Subsequently, the calculated weights are used to perform a weighted summation of the normalized first and second nonlinear characteristic factors to form the comprehensive value of the nonlinear mapping characteristic factor set for the soil sample. This comprehensive value can more objectively reflect the nonlinear characteristics of the sample, which is helpful for subsequent model fitting distortion judgment and credibility analysis.

[0093] In some embodiments of the present application, the calculation process of the credibility factor group includes:

[0094] The residual between the spectral response curve of each soil sample and the model-predicted concentration was extracted as the first credibility factor;

[0095] Calculate the concentration deviation dispersion of the current soil sample under multiple model prediction paths as the second credibility factor; the multiple models include global fitting model, local fitting sub-model, multi-fitting path model, regularized fitting model and kernel function regression model;

[0096] Calculate the spectral structure cosine similarity between the current spectral response curve and the most similar sample in the spectral library as the third credibility factor; the spectral library is a database that stores heavy metal types, corresponding contents and corresponding spectral response curves;

[0097] Taking a weighted average of the first credibility factor, the second credibility factor, and the third credibility factor to construct a credibility factor group;

[0098] An index table is established for each soil sample and credibility factor group.

[0099] It should be noted that the credibility factor group is intended to comprehensively reflect the consistency and reliability between the soil sample spectral data and the model prediction results. The specific calculation process includes three aspects:

[0100] First, the confidence factor characterizes the magnitude of the prediction error by extracting the residual (i.e., the difference between the spectral response curve of the soil sample and the corresponding model-predicted concentration). The smaller the residual, the more accurate the model prediction and the higher the confidence level. Conversely, the smaller the residual, the lower the confidence level.

[0101] The second credibility factor (concentration deviation dispersion): For the same soil sample, multiple model prediction paths (including global fitting models, local fitting sub-models, multi-fitting path models, regularized fitting models, and kernel function regression models) are used to calculate the predicted heavy metal concentration values. Concentration deviation dispersion is the deviation dispersion between these different predicted values. It is obtained by calculating the statistical dispersion index of all predicted concentration values. The common method is to calculate the standard deviation or variance of the predicted concentration values. The smaller the standard deviation, the more consistent the prediction results of the models, and the higher the credibility. The larger the standard deviation, the more dispersed the prediction results, the greater the model uncertainty, and the lower the credibility.

[0102] The third confidence factor calculates the cosine similarity between the spectral response curve of the current soil sample and the spectral structure of the most similar sample in the spectral library. Cosine similarity measures the cosine of the angle between two spectral vectors. Values ​​closer to 1 indicate more similar spectral structures, indicating a high degree of match between the sample's spectral characteristics and the existing sample library data, and a more reliable model prediction.

[0103] After normalizing each of these three factors, we then take a weighted average based on pre-determined weights to construct a comprehensive credibility factor group value, which is used to comprehensively assess the reliability of the sample prediction. The credibility factor group values ​​corresponding to all samples are indexed with the sample to facilitate subsequent query and analysis.

[0104] Method for obtaining concentration deviation dispersion: For a particular soil sample, multiple models are used to independently predict its heavy metal concentration, resulting in a set of predicted concentration values. The standard deviation of the predicted concentration value set is the concentration deviation dispersion of the sample.

[0105] In some embodiments of the present application, the calculation process of the credibility representation value includes:

[0106] Analyze the credibility factor groups of different soil samples in the same test area, and obtain the difference between the maximum credibility factor group and the minimum credibility factor group as the credibility fluctuation;

[0107] The ratio of the credibility fluctuation to the preset credibility fluctuation threshold is used as the credibility representation value.

[0108] It can be understood that this method effectively reflects the stability and reliability of prediction results by analyzing the fluctuations in the credibility factors of multiple soil samples within the same area. Using the difference between the maximum and minimum credibility factors as a fluctuation indicator, the degree of consistency of data within a region can be identified. The credibility representation value is then calculated by comparing it with a preset threshold, achieving a quantitative assessment of the credibility of soil heavy metal pollution detection results.

[0109] In some embodiments of the present application, the process of classifying reliability level labels includes:

[0110] If the credibility representation value is less than the preset credibility low risk threshold, it is marked as a high credibility label;

[0111] If the credibility representation value is higher than the preset credibility high risk threshold, it is marked as a low credibility label;

[0112] If the credibility representation value is between the credibility low risk threshold and the credibility high risk threshold, it is marked as a medium credibility label.

[0113] It should be noted that the method for obtaining the preset low-risk and high-risk confidence thresholds is to collect a large amount of historical soil sample test data, calculate the corresponding distribution of confidence characterization values, and then determine the distribution range of the confidence characterization values ​​through statistical analysis (quantile analysis or cluster analysis).

[0114] In this embodiment, the lower interval in the distribution (the first 20%) is used as the preset credibility low-risk threshold, indicating a data range with higher credibility; the higher interval in the distribution (the last 20%) is used as the high-risk threshold, indicating a data range with lower credibility.

[0115] In some embodiments of the present application, if the detection result is a low confidence label, the following process is triggered:

[0116] Based on the changes in the nonlinear mapping characteristic factor group of the current soil sample, the degree of model fitting distortion is judged. If it is higher than the fitting deviation threshold, the local fitting sub-model is called, and the local fitting sub-model is used to predict the heavy metal concentration in the current test area.

[0117] It should be noted that when the detection result is marked as a low-confidence label, the nonlinear mapping characteristic factor group of the current soil sample is first analyzed in detail to determine the degree of distortion of the model fitting. The specific steps include: comparing the nonlinear mapping characteristic factor group of the current soil sample with the average nonlinear mapping characteristic factor group of the same or similar area in the historical database, and calculating the deviation between the two. The deviation is the cosine similarity difference. The calculated deviation is compared with the preset fitting deviation threshold. The fitting deviation threshold is determined by historical data statistics and model performance testing to reflect the maximum tolerance of the normal model fitting range. If the deviation exceeds the fitting deviation threshold, it is judged that the current model has significant fitting distortion for the sample and cannot accurately reflect the distribution of soil heavy metal concentrations. At this time, the local fitting sub-model is started to re-fit the soil spectral characteristics and nonlinear relationships specifically for the current area to be tested.

[0118] It should be noted that the local fitting sub-model is a regional nonlinear prediction model based on Support Vector Regression (SVR), which is specifically used to supplement the prediction of spectral fitting distortion samples. The specific training method of this model is as follows:

[0119] First, based on the historical spectral data of the target area and the laboratory-measured heavy metal concentration labels, a local training sample set was selected that met the following criteria: First, the principal component Euclidean distance between the sample's spectral response curve and the current sample in the 450-2450nm range was less than a set similarity threshold; second, the weighted deviation between the nonlinear mapping feature factor group and the current sample was less than a set fitting tolerance threshold. Samples that met these conditions were defined as the local sample set.

[0120] Then, using the hyperspectral feature sequence from the local sample set as input and its corresponding laboratory concentration values ​​as output, an SVR model was constructed, using the radial basis kernel (RBF) as the kernel function. A grid search method was used to determine the optimal penalty coefficient C and kernel width γ under cross-validation. This SVR model is the local fitting submodel.

[0121] After the model training is completed, in the detection process, if the current sample is marked as a low-confidence label and the deviation between its nonlinear mapping feature factor group and the historical average feature factor group is higher than the preset fitting deviation threshold, the local SVR sub-model is automatically called, the spectral feature sequence of the sample is input, and the corrected predicted concentration result is output, which is then fused and weighted with the global model prediction result.

[0122] In some embodiments of the present application, the process of determining the degree of model fitting distortion includes:

[0123] Calculate the degree of deviation between the nonlinear mapping characteristic factor group of the current soil sample and the historical average characteristic factor group;

[0124] If the degree of deviation is higher than the preset threshold of fitting deviation, it is determined that there is fitting distortion, and the local sub-model is called to refit the path.

[0125] Specifically, the process of determining the degree of model fitting distortion includes the following steps:

[0126] The first step is to obtain the historical average characteristic factor group. The system selects historically measured samples with similar soil types and compositions to the current test area, screens out samples with "high confidence labels," and extracts two indicators from their nonlinear mapping characteristic factor group: the first nonlinear characteristic factor (calculated by the ratio of the multi-order polynomial fitting residual to the threshold) and the second nonlinear characteristic factor (calculated by the deviation between the spectral response curve and the kernel principal component projection curve). The mean of these two factors is calculated and recorded as the historical average characteristic factor group.

[0127] The second step is to extract the nonlinear mapping characteristic factor group of the current soil sample to be tested.

[0128] The third step is to calculate the feature deviation between the current sample and the historical average sample. This deviation is obtained using the weighted Euclidean distance calculation formula.

[0129] The fourth step is to compare the calculated deviation with the preset fitting deviation threshold. If the deviation is greater than the preset fitting deviation threshold, it is considered that there is a significant difference in the nonlinear feature mapping between the current sample and the historical credible samples, which may cause distortion in the global fitting model prediction.

[0130] In the fifth step, if it is determined that there is model fitting distortion, the local fitting sub-model corresponding to the current area is called to replace the global model to predict the concentration of the current sample, thereby improving the adaptability and accuracy of the model.

[0131] In some embodiments of the present application, model weight adjustment includes:

[0132] When the reliability level label is a medium credibility label, weight adjustment is performed on the multiple models based on the first credibility factor, the second credibility factor, and the third credibility factor;

[0133] If the first credibility factor of any model is greater than the preset residual threshold, the weight of the current model is reduced;

[0134] If the second credibility factor of any model is greater than the preset discreteness threshold, the weight of the current model is reduced;

[0135] If the third credibility factor of any model is less than the similarity threshold, the weight of the current model is reduced;

[0136] After the model weights are adjusted, the weight adjustment results are used to perform weighted summation on the prediction results of multiple models to output the final predicted concentration result.

[0137] It should be noted that the purpose of the model weight adjustment process is to improve the stability and accuracy of model combination predictions by introducing a dynamic model scoring and weight distribution mechanism in situations where there is a certain degree of uncertainty (i.e., when the test results are marked as "medium-trusted"). The following is a step-by-step explanation of this process:

[0138] When a soil sample's credibility representation falls between the "low credibility risk threshold" and the "high credibility risk threshold," the system labels the sample as "medium credibility." While the prediction for this sample has some uncertainty, it's not enough to trigger the local fitting sub-model. In this case, you can't rely solely on any one model; instead, you need to dynamically adjust the model weights based on their performance on the current soil sample.

[0139] The first credibility factor (residual value): indicates the error between the model-predicted concentration and the actual laboratory concentration. If the residual value of a model on the current sample is too large, it means that the model has poor fitting ability for the sample. The second credibility factor (multi-model prediction dispersion): indicates the degree of deviation between the predicted value of the model and the predicted value of other models. If the deviation is large, it means that the model output is not representative or lacks consistency with other models. The third credibility factor (spectral structure similarity): indicates whether the sample on which the model is based is close to similar samples in the spectral library. If the similarity is low, it means that the model is not representative enough for the current sample.

[0140] If the first credibility factor is greater than the residual threshold (i.e., the prediction error is significantly larger), the model is judged to be inaccurately fitted on the sample and its weight should be reduced; if the second credibility factor is greater than the dispersion threshold (i.e., the prediction results of the model are too different from those of other models), it means that the model prediction lacks robustness and its weight should also be reduced; if the third credibility factor is less than the set similarity threshold (i.e., the sample is significantly different from the model training sample), the reliability of the model inference is limited and the weight should also be lowered.

[0141] These conditions are judged independently. Once the model triggers any threshold, it can be regarded as "poor credibility performance" and its weight will be adjusted accordingly. The amount of weight adjustment is the preset value.

[0142] After all model weights are updated, the predictions from multiple models are weighted averaged using the updated weights to produce the final predicted concentration. This approach avoids a single low-confidence model dominating the output while leveraging information from high-performing models to improve overall prediction stability.

[0143] This mechanism doesn't simply exclude certain models; instead, it meticulously regulates the predictive contribution of each model through the quantitative calculation of its credibility factor. This weighting adjustment ensures the system maintains robust adaptability and predictive robustness, especially when credibility fluctuates but isn't severe enough to warrant a model rebuild. This mechanism offers a more realistic strategy for practical testing applications in current conditions with unstable sample sizes and complex, heterogeneous soil compositions.

[0144] In some embodiments of the present application, the process of constructing a hyperspectral feature sequence includes:

[0145] Based on the collected spectral response curve, the main component characteristics, band ratio characteristics and continuous spectral slope change characteristics of the spectral band within the specified band of 450-2450nm are extracted; a multidimensional combined feature vector is constructed as the hyperspectral feature sequence of the soil sample;

[0146] The process of obtaining soil spectral reflectance images includes:

[0147] Obtain spectral reflectance images of different soil surface layers in the test area; divide the spectral reflectance images into multiple sub-areas and extract the average spectral response curve of each sub-area; and extract the spectral response curves of all sub-areas in turn as input for subsequent processing.

[0148] It should be noted that the construction process of the hyperspectral feature sequence includes: The hyperspectral feature sequence is a multidimensional feature vector set used to represent the spectral response characteristics of the soil, which has the advantages of high information density and obvious feature differences. The specific construction process is as follows:

[0149] Principal component analysis (PCA) was performed on the original spectral response curves in the 450-2450nm band (i.e., the curve showing how the reflectance of each soil sub-region changes with wavelength). Principal component vectors with a cumulative contribution rate of more than 95% were selected, and the corresponding principal component scores were extracted as "spectral principal component features." These features are used to characterize the main change trends of the samples in high-dimensional spectral space, which helps to reduce dimensionality and remove redundancy.

[0150] Set several band pairs with diagnostic significance (such as 680 / 550nm, 950 / 850nm, etc.), and calculate the reflectance ratio between these band pairs; the band ratio can reveal the absorption difference between specific bands, the relative relationship between spectral distortion and components such as water / organic matter; the obtained ratios constitute the "band ratio feature" vector.

[0151] The entire spectral curve is divided into several continuous band windows (e.g., each 50 nm is a segment), and a linear fit is performed on each segment to extract the slope of the fitted line. At the same time, the slope change rate between adjacent spectral segments is calculated to obtain the "slope difference feature." This feature is used to identify the spectral change rate and local absorption edge information, which helps to identify abnormal absorption behavior of trace elements in the soil.

[0152] The spectral principal component characteristics, band ratio characteristics and continuous spectral slope change characteristics are spliced ​​and combined to form a multidimensional combined feature vector of unified dimension, that is, the hyperspectral feature sequence of the soil sample; the feature sequences of all sub-region samples are summarized together for subsequent feature factor calculation and model training.

[0153] In addition, the acquisition and processing process of the spectral reflectance image includes: the soil spectral reflectance image is the physical basis for obtaining hyperspectral features, and its acquisition and processing steps are as follows:

[0154] Use hyperspectral imaging equipment (such as ground-based hyperspectral cameras and airborne imaging spectrometers) to scan the area to be measured. Obtain a spectral reflectance image data cube (a three-dimensional matrix of x, y, and λ), where the value of λ ranges from 450 to 2450 nm. Each pixel records the reflectance of the soil at that location at different wavelengths.

[0155] The collected soil image is divided into several sub-regions according to the spatial dimension (such as 10×10 pixels or 50×50 pixels); the reflectance values ​​of all pixels in each sub-region at each wavelength are averaged to obtain the average spectral response curve of the region; the averaged curve removes local noise and spatial stray interference, reflecting the representative soil spectral characteristics of the region.

[0156] The spectral response curves (reflectance curves with wavelengths of 450-2450nm) extracted from each sub-region are numbered in sequence; all sub-region curves are input into the subsequent hyperspectral feature extraction process as a group of samples; and the hyperspectral feature sequence construction of the samples in the entire test area is completed.

[0157] See Figure 2 As shown, an embodiment of the present invention provides a soil heavy metal pollution detection system based on spectral data, comprising:

[0158] An acquisition module is configured to obtain a spectral reflectance image of the soil surface in the test area, extract a corresponding hyperspectral feature sequence, and construct a spectral response curve of the target soil;

[0159] a calculation module configured to calculate a nonlinear mapping characteristic factor group and a credibility factor group of the soil sample based on the spectral response curve;

[0160] a processing module configured to classify reliability level labels based on the credibility representation values ​​and adjust a process for obtaining the predicted concentration according to the reliability level labels;

[0161] The output module is configured to output the final predicted concentration result.

[0162] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the scope of protection of the claims of the present invention.

Claims

1. A method for detecting heavy metal pollution in soil based on spectral data, characterized in that: include: Obtain the spectral reflectance image of the soil surface in the test area, extract the corresponding hyperspectral feature sequence and construct the spectral response curve of the target soil; Based on the spectral response curve, calculating a nonlinear mapping characteristic factor group and a credibility factor group of the soil sample; Constructing a credibility representation value according to the credibility factor group and the nonlinear mapping characteristic factor group; Dividing reliability level labels based on the credibility characterization values, and performing the following steps according to the reliability level labels: judging whether there is model fitting distortion based on changes in the nonlinear mapping characteristic factor group of the soil sample, so as to judge whether to call a local fitting sub-model; and, based on the credibility representation value, determining whether to perform model weight adjustment; Output the final predicted concentration result; The nonlinear mapping characteristic factor group is calculated based on the multi-order nonlinear function fitting residual of the hyperspectral principal component, and the credibility factor group is calculated based on the similarity between the residual amplitude and the spectral structure; The calculation process of the nonlinear mapping characteristic factor group includes: Performing multi-order polynomial fitting on the hyperspectral feature sequence to obtain a predicted concentration sequence; Calculating the residual sum of squares between the predicted concentration sequence and the original concentration label, and determining a first nonlinear characteristic factor based on the ratio of the residual sum of squares to a reference residual threshold; the original concentration label is the actual heavy metal concentration value measured in the laboratory; The deviation between the spectral response curve and its kernel principal component projection curve is taken as the second nonlinear characteristic factor; performing weighted summation of the first nonlinear characteristic factor and the second nonlinear characteristic factor to form a nonlinear mapping characteristic factor group of the sample; All soil samples are stored after establishing association relationships with the nonlinear mapping characteristic factor group.

2. The method for detecting heavy metal pollution in soil based on spectral data according to claim 1, characterized in that: The calculation process of the credibility factor group includes: The residual between the spectral response curve of each soil sample and the model-predicted concentration was extracted as the first credibility factor; Calculating the concentration deviation dispersion of the current soil sample under multiple model prediction paths as a second credibility factor; the multiple models include a global fitting model, a local fitting sub-model, a multi-fitting path model, a regularized fitting model, and a kernel function regression model; Calculating the spectral structure cosine similarity between the current spectral response curve and the most similar sample in the spectral library as a third credibility factor; the spectral library is a database that stores heavy metal types, corresponding contents, and corresponding spectral response curves; Taking a weighted average of the first credibility factor, the second credibility factor, and the third credibility factor to construct the credibility factor group; An index table is established for each soil sample and credibility factor group.

3. The method for detecting heavy metal pollution in soil based on spectral data according to claim 2, characterized in that: The calculation process of the credibility characterization value includes: Analyze the credibility factor groups of different soil samples in the same test area, and obtain the difference between the maximum credibility factor group and the minimum credibility factor group as the credibility fluctuation; The ratio of the credibility fluctuation to the preset credibility fluctuation threshold is used as the credibility representation value.

4. The method for detecting heavy metal pollution in soil based on spectral data according to claim 3, characterized in that: The process of assigning reliability level labels includes: If the credibility representation value is less than the preset credibility low risk threshold, it is marked as a high credibility label; If the credibility representation value is higher than the preset credibility high risk threshold, it is marked as a low credibility label; If the credibility representation value is between the credibility low risk threshold and the credibility high risk threshold, it is marked as a medium credibility label.

5. The method for detecting heavy metal pollution in soil based on spectral data according to claim 4, characterized in that: If the detection result is a low confidence label, the following process is triggered: Based on the changes in the nonlinear mapping characteristic factor group of the current soil sample, the degree of model fitting distortion is judged. If it is higher than the fitting deviation threshold, the local fitting sub-model is called, and the local fitting sub-model is used to predict the heavy metal concentration in the current test area.

6. The method for detecting heavy metal pollution in soil based on spectral data according to claim 5, characterized in that: The process of determining the degree of model fitting distortion includes: Calculate the degree of deviation between the nonlinear mapping characteristic factor group of the current soil sample and the historical average characteristic factor group; If the degree of deviation is higher than a preset fitting deviation threshold, it is determined that there is fitting distortion, and the local sub-model is called to refit the path.

7. The method for detecting heavy metal pollution in soil based on spectral data according to claim 6, characterized in that: The model weight adjustment includes: When the reliability level label is a medium credibility label, weight adjustment is performed on multiple models based on the first credibility factor, the second credibility factor, and the third credibility factor; If the first credibility factor of any model is greater than the preset residual threshold, the weight of the current model is reduced; If the second credibility factor of any model is greater than the preset discreteness threshold, the weight of the current model is reduced; If the third credibility factor of any model is less than the similarity threshold, the weight of the current model is reduced; After the model weights are adjusted, the weight adjustment results are used to perform weighted summation on the prediction results of the multiple models to output the final predicted concentration result.

8. The method for detecting heavy metal pollution in soil based on spectral data according to claim 7, characterized in that: The process of constructing the hyperspectral feature sequence includes: Based on the collected spectral response curve, the main component characteristics of the spectrum, the band ratio characteristics and the slope change characteristics of the continuous spectrum band in the specified band of 450-2450nm are extracted; and a multidimensional combined feature vector is constructed as the hyperspectral feature sequence of the soil sample; The process of obtaining soil spectral reflectance images includes: Obtain spectral reflectance images of different soil surface layers in the test area; divide the spectral reflectance images into multiple sub-areas and extract the average spectral response curve of each sub-area; and extract the spectral response curves of all sub-areas in turn as input for subsequent processing.

9. A soil heavy metal pollution detection system based on spectral data, characterized in that: The method for implementing any one of claims 1 to 8 comprises: An acquisition module is configured to obtain a spectral reflectance image of the soil surface in the test area, extract a corresponding hyperspectral feature sequence, and construct a spectral response curve of the target soil; a calculation module configured to calculate a nonlinear mapping characteristic factor group and a credibility factor group of the soil sample based on the spectral response curve; a processing module configured to classify reliability level labels based on the credibility characterization values, and adjust a process for obtaining the predicted concentration according to the reliability level labels; The output module is configured to output the final predicted concentration result.

Citation Information

Patent Citations

  • A method and system for detecting heavy metals in soil

    CN113791040B

  • Method for estimating content of heavy metals in soil based on hyperspectral remote sensing technology

    CN114018833A

  • Soil heavy metal content inversion method fusing spectrum and spatial characteristics

    CN115236005A