Soil heavy metal pollution detection method and system based on spectral data

By constructing nonlinear mapping feature factor groups and confidence factor groups, combining the reliability characterization value and reliability level labels, the problem of insufficient accuracy in the estimation of heavy metal concentration in the existing technology in the medium and low concentration intervals is solved, and high-precision and robust detection in complex soil environments are achieved.

CN120369649AActive Publication Date: 2025-07-25WEIFANG XINBO PHYSICAL & CHEM TESTING CO LTD +1
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510863508.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-07-25
Estimated Expiration
2045-06-26

AI Technical Summary

Technical Problem

The soil heavy metal concentration estimation method based on linear relationships in the prior art has a low prediction accuracy in the medium and low concentration interval, ignoring the nonlinear spectral response of the complex soil component structure and heavy metal distribution state.

Method used

Using the soil heavy metal pollution detection method based on spectral data, the spectral reflectivity image of the soil surface is obtained, the high-spectral feature sequence is extracted, the spectral response curve is constructed, the nonlinear mapped feature factor group and the confidence factor group are calculated, the reliability characterization value is constructed, the reliability level label is divided, the model fitting distortion is judged, and the weight adjustment is performed, and the final predicted concentration result is output.

Benefits of technology

It improves the prediction accuracy of heavy metals in medium and low concentration intervals, enhances the robustness and adaptability of the detection system, can improve the stability and interpretability of the detection results in complex soil environments, and has self-correction capabilities, which are suitable for the variable and complex detection needs of the actual soil environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120369649A_ABST
    Figure CN120369649A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of soil detection, and discloses a soil heavy metal pollution detection method and system based on spectral data, and the detection method comprises the steps: obtaining a spectral reflectivity image of a soil surface layer of a to-be-detected region, and extracting a hyperspectral feature sequence to construct a spectral response curve of target soil; calculating a nonlinear mapping characteristic factor group and a credibility factor group based on the spectral response curve; constructing a credibility representation value according to the two, and dividing a reliability grade label; and judging whether model fitting distortion exists, whether a local fitting sub-model is called and whether model weight adjustment is carried out according to the label, and outputting a final prediction concentration result. The method disclosed by the invention has relatively high detection accuracy on relatively low-content soil heavy metals.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of soil detection, and in particular, to a method and system for detecting soil heavy metal pollution based on spectral data. Background Art

[0002] In recent years, with the development of spectral imaging technology, multispectral or hyperspectral sensing technology has been gradually applied to the identification and content estimation of heavy metal elements in soil. In the prior art, the collected soil spectra are mostly compared with the established spectral library, and spectral angle matching and linear regression models are used for concentration prediction.

[0003] Chinese Patent Publication No.: CN113791040B discloses a method and system for detecting soil heavy metals. The detection method combines near-infrared spectroscopy analysis and hyperspectral imaging technology. First, a photo and a spectral image of the soil are obtained, then spectral feature data is extracted from the spectral image, and the spectral angle is calculated with the spectral feature data in the spectral library. The similarity between the two is determined by the size of the spectral angle, so as to qualitatively analyze the types of soil heavy metals. Then, the concentration of soil heavy metals is quantitatively calculated by the calculation formula of the linear regression equation between the preset spectral feature data and the absolute concentration. The detection system includes a flying device, an analysis system, an optical camera, a multispectral sensor and a GPS positioning system carried on the flying device; the spectral library is set in the analysis system.

[0004] Although the above solution can be used for on-site and rapid analysis of the types and concentrations of heavy metals in soil, there are still the following problems: Based on the linear relationship for concentration estimation, the non-linear spectral response caused by the complex component structure of the soil and the occurrence state of heavy metals is ignored, resulting in low prediction accuracy in the medium and low concentration ranges. Summary of the Invention

[0005] In view of this, the present invention proposes a method and system for detecting soil heavy metal pollution based on spectral data to solve the problems existing in the prior art.

[0006] On the one hand, the method for detecting soil heavy metal pollution based on spectral data proposed by the present invention includes: Obtain the spectral reflectance image of the soil surface layer in the area to be measured, and extract the corresponding hyperspectral feature sequence to construct the spectral response curve of the target soil; Based on the spectral response curve, calculate the non-linear mapping feature factor group and the credibility factor group of the soil sample; Construct a credibility characterization value according to the credibility factor group and the non-linear mapping feature factor group; Divide the reliability level labels based on the credibility characterization value, and perform the following steps according to the reliability level labels: Based on the change of the non-linear mapping feature factor group of the soil sample, judge whether there is model fitting distortion to determine whether to call the local fitting sub-model; And, based on the credibility characterization value, judge whether to adjust the model weight; Output the final predicted concentration result; Wherein, the non-linear mapping feature factor group is calculated based on the fitting residuals of the multi-order non-linear function of the hyperspectral principal components, and the credibility factor group is calculated based on the residual amplitude and the spectral structure similarity.

[0007] Furthermore, the calculation process of the non-linear mapping feature factor group includes: Perform multi-order polynomial fitting on the hyperspectral feature sequence to obtain a predicted concentration sequence; Calculate the sum of squared residuals between the predicted concentration sequence and the original concentration label, and determine the first non-linear feature factor based on the ratio of the sum of squared residuals to the reference residual threshold; the original concentration label is the true heavy metal concentration value measured in the laboratory; Use the deviation degree between the spectral response curve and its kernel principal component projection curve as the second non-linear feature factor; Weightedly sum the first non-linear feature factor and the second non-linear feature factor to form the non-linear mapping feature factor group of the sample; After constructing the association relationship between all soil samples and the non-linear mapping feature factor group, store them.

[0008] Furthermore, the calculation process of the credibility factor group includes: Extract the residual between the spectral response curve of each soil sample and the model predicted concentration as the first credibility factor; Calculate the concentration deviation dispersion degree of the current soil sample under multiple model prediction paths as the second credibility factor; the multiple models include the global fitting model, the local fitting sub-model, the multi-fitting path model, the regularization fitting model and the kernel function regression model; Calculate the spectral structure cosine similarity between the current spectral response curve and the most similar sample in the spectral library as the third credibility factor; the spectral library is a database storing heavy metal types, corresponding contents and corresponding spectral response curves; Weightedly average the first credibility factor, the second credibility factor and the third credibility factor to construct the credibility factor group; And establish an index table between each soil sample and the credibility factor group.

[0009] Furthermore, the calculation process of the credibility characterization value includes: Analyze the credibility factor groups of different soil samples in the same area to be measured, and obtain the difference between the maximum credibility factor group and the minimum credibility factor group as the credibility fluctuation degree; Take the ratio of the credibility fluctuation degree to the preset credibility fluctuation threshold as the credibility characterization value.

[0010] Further, the process of dividing the reliability level labels includes: If the credibility characterization value is less than the preset low-credibility risk threshold, mark it as a high-credibility label; If the credibility characterization value is higher than the preset high-credibility risk threshold, mark it as a low-credibility label; If the credibility characterization value is between the low-credibility risk threshold and the high-credibility risk threshold, mark it as a medium-credibility label.

[0011] Further, if the detection result is a low-credibility label, trigger the following process: Based on the change of the non-linear mapping characteristic factor group of the current soil sample, judge the degree of model fitting distortion. If it is higher than the fitting deviation threshold, call the local fitting sub-model, which is used for predicting the heavy metal concentration in the current area to be measured.

[0012] Further, the process of judging the degree of model fitting distortion includes: Calculate the deviation degree between the non-linear mapping characteristic factor group of the current soil sample and the historical average characteristic factor group; If the deviation degree is higher than the preset fitting deviation threshold, it is determined that there is fitting distortion, and call the local sub-model to re-fit the path.

[0013] Further, the model weight adjustment includes: When the reliability level label is a medium-credibility label, adjust the weights of the multiple models based on the first credibility factor, the second credibility factor, and the third credibility factor; If the first credibility factor of any model is greater than the preset residual threshold, reduce the weight of the current model; If the second credibility factor of any model is greater than the preset dispersion threshold, reduce the weight of the current model; If the third credibility factor of any model is less than the similarity threshold, reduce the weight of the current model; After the model weight is adjusted, use the weight adjustment result to perform weighted summation on the prediction results of the multiple models, and output the final predicted concentration result.

[0014] Further, the process of constructing the hyperspectral feature sequence includes: Based on the collected spectral response curves, extract the spectral principal component features, band ratio features, and continuous spectral segment slope change features within the specified band of 450 - 2450 nm; construct a multi-dimensional combined feature vector as the hyperspectral feature sequence of the soil sample; The process of obtaining the soil spectral reflectance image includes: Obtain the spectral reflectance images of different soil surfaces within the area to be measured; divide the spectral reflectance image into multiple sub-regions and extract the average spectral response curve of each sub-region; sequentially extract the spectral response curves of all sub-regions as the input for subsequent processing.

[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: By introducing the construction of the hyperspectral feature sequence and the extraction of multi-dimensional combined features, the present invention improves the expression ability of spectral information for the occurrence state of heavy metals in soil, and can comprehensively reflect the spectral response characteristics of soil components at different concentration levels.

[0016] In the process of model construction, the present invention uses a non-linear mapping feature factor group to replace the traditional linear fitting method, and through the fusion of multi-order polynomial residuals and kernel principal component deviation degrees, improves the modeling ability in the non-linear response region, and is particularly suitable for the identification and fitting of medium and low concentration heavy metals.

[0017] By constructing the first, second, and third credibility factors and weighting them to form a credibility factor group, the stability and model adaptability of the prediction results can be comprehensively evaluated from multiple angles, not only paying attention to the predicted values, but also having the ability to judge the credibility of the prediction results themselves.

[0018] Introduce a division mechanism of credibility characterization values and reliability level labels in the detection process, so that the system can flexibly call local fitting sub-models or adjust the model weights according to different credibility levels, thereby enhancing the local adaptability while maintaining the generality of the overall model.

[0019] When the system determines that the detection result is medium credible or low credible, it can actively trigger the model path adjustment or re-fitting process according to the deviation between historical data and current features, so as to improve the robustness of the prediction results in scenarios with environmental disturbances or strong sample specificities.

[0020] Through the model weight dynamic adjustment mechanism, combined with the performance of different models in terms of residuals, dispersion, and similarity, the influence of misfitting models is effectively weakened, the interpretability and rationality of the overall prediction results are enhanced, and the risk of misjudgment under the dominance of a single model is avoided.

[0021] In addition, the linkage mechanism between the constructed spectral library and the model enables the system to have a certain self-correction ability when facing spectral feature heterogeneous samples, and can form a closed loop of "judgment - adjustment - re-prediction" during the detection process, improving the flexibility and intelligent level of the detection process.

[0022] On the other hand, the soil heavy metal pollution detection system based on spectral data proposed by the present invention includes: A collection module configured to obtain the spectral reflectance image of the soil surface layer in the area to be measured, and extract the corresponding hyperspectral feature sequence to construct the spectral response curve of the target soil; A calculation module configured to calculate the non-linear mapping feature factor group and the credibility factor group of the soil sample based on the spectral response curve; A processing module configured to divide the reliability level labels based on the credibility characterization value, and adjust the acquisition process of the predicted concentration according to the reliability level labels; An output module configured to output the final predicted concentration result.

[0023] It should be noted that the soil heavy metal pollution detection method and system based on spectral data of the present invention have the same beneficial effects, which will not be elaborated here. Description of the Drawings

[0024] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings: Figure 1 It is a flowchart of a soil heavy metal pollution detection method based on spectral data provided by an embodiment of the present invention.

[0025] Figure 2 It is a functional block diagram of a soil heavy metal pollution detection system based on spectral data provided by an embodiment of the present invention. Detailed Embodiments

[0026] Hereinafter, the exemplary embodiments disclosed in the present application will be described in more detail with reference to the drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be completely conveyed to those skilled in the art. It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the drawings and in combination with the embodiments.

[0027] Refer to Figure 1 As shown, an embodiment of the present invention provides a method for detecting soil heavy metal pollution based on spectral data, including: S1: Obtain the spectral reflectance image of the soil surface layer in the area to be measured, extract the corresponding hyperspectral feature sequence to construct the spectral response curve of the target soil; S2: Based on the spectral response curve, calculate the non-linear mapping feature factor group and the credibility factor group of the soil sample; S3: Construct a credibility characterization value according to the credibility factor group and the non-linear mapping feature factor group; S4: Divide the reliability level label based on the credibility characterization value, and perform the following steps according to the reliability level label: Based on the change of the non-linear mapping feature factor group of the soil sample, judge whether there is model fitting distortion to judge whether to call the local fitting sub-model; And, based on the credibility characterization value, judge whether to adjust the model weight; S5: Output the final predicted concentration result; Among them, the non-linear mapping feature factor group is calculated based on the multi-order non-linear function fitting residual of the hyperspectral principal component, and the credibility factor group is calculated based on the residual amplitude and the spectral structure similarity.

[0028] It can be understood that through the systematic extraction of the hyperspectral feature sequence in step S1, this method retains the important spectral band information of the soil reflectance in the 450-2450nm band, and constructs a more detailed spectral response curve through multi-dimensional features (such as principal components, ratio features, spectral band slopes, etc.), providing a higher information dimension and data resolution for subsequent modeling and error evaluation.

[0029] In step S2, by calculating the non-linear mapping feature factor group and the credibility factor group, a breakthrough in the traditional linear modeling method is achieved. The introduction of the non-linear mapping feature factor enables the model to identify the non-linear relationship between the soil heavy metal content and the spectral response, improving the prediction accuracy in the medium and low concentration ranges; while the introduction of the credibility factor realizes the self-diagnosis function, which can identify the stability and applicability of the current prediction result.

[0030] In step S3, constructing the credibility characterization value further aggregates multiple credibility factors into a single index, enabling the system to identify the overall reliability of sample prediction in a structured manner, avoiding misleading the overall judgment due to local fluctuations of a certain factor, and contributing to improving the robustness and responsiveness of the overall prediction system.

[0031] The reliability level label division and response strategy design in step S4 ensure that the system can not only output concentration values, but also has the ability to screen results and give risk warnings. If it is determined to be low credibility or medium credibility, it can actively judge whether there is model fitting distortion and trigger local fitting sub-models or model weight adjustment. This dynamic feedback and local adaptive mechanism enhance the generality and anti-interference ability under complex soil conditions.

[0032] Finally, before the result output in S5, the model path adjustment or prediction optimization is completed to ensure that the predicted concentration results output are at a high level in terms of reliability, stability and interpretability, and are applicable to the diverse and complex detection requirements in the actual soil environment. The overall process constitutes an intelligent detection system with a closed-loop of "feature construction - factor extraction - credibility diagnosis - model regulation - result output".

[0033] In some embodiments of the present application, the calculation process of the non-linear mapping feature factor group includes: Performing multi-order polynomial fitting on the hyperspectral feature sequence to obtain a predicted concentration sequence; Calculating the sum of squared residuals between the predicted concentration sequence and the original concentration label, and determining the first non-linear feature factor based on the ratio of the sum of squared residuals to the reference residual threshold; the original concentration label is the true heavy metal concentration value measured in the laboratory; Taking the deviation degree between the spectral response curve and its kernel principal component projection curve as the second non-linear feature factor; Weighted summing the first non-linear feature factor and the second non-linear feature factor to form the non-linear mapping feature factor group of the sample; After establishing the association relationship between all soil samples and the non-linear mapping feature factor group, they are stored.

[0034] It can be understood that applying multi-order polynomial regression analysis to the hyperspectral feature sequence, establishing a prediction model of soil spectral features and heavy metal concentration, and obtaining a predicted concentration sequence. Multi-order polynomial fitting can better capture the potential non-linear relationship between the spectrum and the concentration, and avoid complex phenomena that cannot be described by simple linear models.

[0035] By calculating the sum of squared residuals between the predicted concentration sequence and the true concentration (original concentration label) measured in the laboratory, the size of the fitting error is measured. Then, the ratio operation is performed on the sum of squared residuals and the preset reference residual threshold to obtain the first non-linear feature factor, which reflects the relative size of the model fitting error under the reference standard. The larger the value, the more significant the non-linear degree of the fitting.

[0036] Calculate the deviation degree between the calculated spectral response curve and its kernel principal component projection curve. The kernel principal component projection refers to the representation after mapping high-dimensional spectral data to a low-dimensional principal component space using kernel principal component analysis. The deviation degree reflects the difference between the original spectral data and its principal component projection. The larger the value, the greater the complexity or abnormality of the spectral structure of the sample, indicating that there may be non-linear mapping characteristics.

[0037] Perform weighted summation on the first non-linear characteristic factor and the second non-linear characteristic factor according to a preset weight to obtain the overall non-linear mapping characteristic factor group of the sample, which serves as an important basis for subsequent model judgment and processing.

[0038] Establish an association relationship between the non-linear mapping characteristic factor groups corresponding to all samples and the sample information and store it to facilitate subsequent quick query and model training update.

[0039] It should be noted that the deviation degree represents the degree of difference between two data sequences. Specifically, the original hyperspectral response curve is mapped to a low-dimensional space through kernel principal component analysis to obtain the principal component projection curve of the sample. Calculate the difference between the original spectral curve and this projection curve under a certain metric standard, such as Euclidean distance, cosine distance, or other similarity metrics. This difference value is the "deviation degree", which reflects the "deviation" size of the original spectral data compared to its low-dimensional principal component representation.

[0040] In the construction of the non-linear mapping characteristic factor group, first perform numerical normalization processing on the first non-linear characteristic factor and the second non-linear characteristic factor respectively to eliminate the dimensional difference and numerical range difference between the two. The normalization method can adopt linear normalization, mapping the numerical value of each factor to the interval of 0 to 1. Specifically, subtract the minimum value of each factor in all samples from each factor value, and then divide by the difference between the maximum value and the minimum value of the factor in all samples to ensure that all factor values after normalization are within the range of 0 to 1. After normalization, according to a large amount of historical sample data, calculate the correlation index between the two normalized factors and the actual detection accuracy, such as statistical quantities such as correlation coefficient or goodness of fit. Based on this correlation index, automatically determine the weight ratio of the two. The weight value reflects the contribution size of each factor to the accuracy of the detection result and satisfies that the sum of the weights is 1. Subsequently, use the calculated weights to perform weighted summation on the normalized first non-linear characteristic factor and the second non-linear characteristic factor to form the comprehensive value of the non-linear mapping characteristic factor group of the soil sample. This comprehensive value can more objectively reflect the non-linear characteristic performance of the sample, which helps to judge the fitting distortion and credibility analysis of the subsequent model.

[0041] In some embodiments of the present application, the calculation process of the credibility factor group includes: Extract the residual between the spectral response curve of each soil sample and the model predicted concentration as the first credibility factor; Calculate the concentration deviation dispersion of the current soil sample under multiple model prediction paths as the second credibility factor; the multiple models include a global fitting model, local fitting sub-models, a multi-fitting path model, a regularization fitting model, and a kernel function regression model; Calculate the spectral structure cosine similarity between the current spectral response curve and the most similar sample in the spectral library as the third credibility factor; the spectral library is a database storing heavy metal types, corresponding contents, and corresponding spectral response curves; Weightedly average the first credibility factor, the second credibility factor, and the third credibility factor to construct a credibility factor group; And establish an index table for each soil sample and the credibility factor group.

[0042] It should be noted that the credibility factor group aims to comprehensively reflect the consistency and reliability between the spectral data of soil samples and the model prediction results. The specific calculation process includes three aspects: The first credibility factor: Characterize the magnitude of the prediction error by extracting the residual (i.e., the difference) between the spectral response curve of the soil sample and the predicted concentration of the corresponding model. The smaller the residual, the more accurate the model prediction and the higher the credibility, and vice versa.

[0043] The second credibility factor (concentration deviation dispersion): For the same soil sample, use multiple model prediction paths (including a global fitting model, local fitting sub-models, a multi-fitting path model, a regularization fitting model, a kernel function regression model, etc.) to calculate the predicted values of its heavy metal concentration respectively. The concentration deviation dispersion is the deviation dispersion situation among these different predicted values, which is specifically obtained by calculating the statistical dispersion index of all predicted concentration values. The common method is to calculate the standard deviation or variance of the predicted concentration values. The smaller the standard deviation, the more consistent the model prediction results and the higher the credibility; the larger the standard deviation, the more dispersed the prediction results, the greater the model uncertainty, and the lower the credibility.

[0044] The third credibility factor: Calculate the spectral structure cosine similarity between the spectral response curve of the current soil sample and the most similar sample in the spectral library. The cosine similarity measures the cosine value of the angle between two spectral vectors. The value closer to 1 represents more similar spectral structures, indicating a high matching degree between the sample spectral characteristics and the data in the existing sample library, and more reliable model predictions.

[0045] After normalizing the above three factors respectively, weightedly average them according to the pre-determined weight ratio to construct a comprehensive credibility factor group value for comprehensively evaluating the credibility of sample predictions. The credibility factor group values corresponding to all samples are indexed with the samples for convenient subsequent query and analysis.

[0046] Method for obtaining concentration deviation dispersion: For a certain soil sample, multiple models are used to independently predict its heavy metal concentration, and a set of predicted concentration values is obtained. The standard deviation of the set of predicted concentration values is the concentration deviation dispersion of the sample.

[0047] In some embodiments of the present application, the calculation process of the credibility characterization value includes: Analyze the credibility factor groups of different soil samples in the same area to be measured, and obtain the difference between the maximum credibility factor group and the minimum credibility factor group as the credibility fluctuation degree; Take the ratio of the credibility fluctuation degree to the preset credibility fluctuation threshold as the credibility characterization value.

[0048] It can be understood that this method effectively reflects the stability and reliability of the prediction results by analyzing the credibility factor fluctuations of multiple soil samples in the same area. Using the difference between the maximum and minimum credibility factors as the fluctuation degree index can identify the degree of data consistency in the area, and then calculate the credibility characterization value through the ratio with the preset threshold, realizing the quantitative evaluation of the credibility of the soil heavy metal pollution detection results.

[0049] In some embodiments of the present application, the process of dividing the reliability level label includes: If the credibility characterization value is less than the preset low-risk threshold of credibility, mark it as a high-credibility label; If the credibility characterization value is higher than the preset high-risk threshold of credibility, mark it as a low-credibility label; If the credibility characterization value is between the low-risk threshold of credibility and the high-risk threshold of credibility, mark it as a medium-credibility label.

[0050] It should be noted that the methods for obtaining the preset low-risk threshold of credibility and the preset high-risk threshold of credibility are: collect the detection data of a large number of historical soil samples, and calculate the corresponding distribution of credibility characterization values. Through statistical analysis (quantile analysis or clustering analysis), determine the distribution interval of the credibility characterization value.

[0051] In this embodiment, the lower interval (the first 20%) in the distribution is used as the preset low-risk threshold of credibility, indicating a range of data with higher credibility; the higher interval (the last 20%) in the distribution is used as the high-risk threshold, indicating a range of data with lower credibility.

[0052] In some embodiments of the present application, if the detection result is a low-credibility label, the following process is triggered: Based on the change of the non-linear mapping feature factor group of the current soil sample, judge the degree of model fitting distortion. If it is higher than the fitting deviation threshold, call the local fitting sub-model, and the local fitting sub-model is used for predicting the heavy metal concentration in the current area to be measured.

[0053] It should be noted that when the detection result is marked with a low-confidence label, first, a detailed analysis is performed on the non-linear mapping feature factor group of the current soil sample to determine the degree of distortion of the model fitting. The specific steps include: comparing the non-linear mapping feature factor group of the current soil sample with the average non-linear mapping feature factor group in the same or similar regions in the historical database, and calculating the deviation degree between the two. This deviation degree is the difference in cosine similarity. Compare the calculated deviation degree with a preset fitting deviation threshold. The fitting deviation threshold is determined by historical data statistics and model performance tests, and is used to reflect the maximum tolerance of the normal model fitting range. If the deviation degree exceeds the fitting deviation threshold, it is determined that there is significant fitting distortion of the current model for this sample and it cannot accurately reflect the distribution of soil heavy metal concentrations. At this time, start the local fitting sub-model to re-fit specifically for the soil spectral characteristics and non-linear relationships in the current area to be measured.

[0054] It should be noted that the local fitting sub-model is a regional non-linear prediction model constructed based on support vector regression (SVR, Support Vector Regression), and is specifically used for supplementary prediction of spectral fitting distortion samples. The specific training method of this model is as follows: First, based on the historical spectral data of the area to be measured and the laboratory-measured heavy metal concentration labels, a local training sample set that meets the following conditions is selected: First, the Euclidean distance of the principal components of the sample spectral response curve and the current sample in the 450-2450 nm interval is less than the set similarity threshold; second, the weighted deviation degree between the non-linear mapping feature factor group and the current sample is less than the set fitting tolerance threshold. The samples that meet the above conditions are defined as the local sample set.

[0055] Then, using the hyperspectral feature sequence in this local sample set as the input and its corresponding laboratory concentration value as the output, construct an SVR model, and select the radial basis kernel function (RBF) as the kernel function form; determine the optimal penalty coefficient C and kernel width γ through grid search under cross-validation. This SVR model is the local fitting sub-model.

[0056] After the model training is completed, in the detection process, if the current sample is marked with a low-confidence label and the deviation degree between its non-linear mapping feature factor group and the historical average feature factor group is higher than the preset fitting deviation threshold, then automatically call this local SVR sub-model, input the spectral feature sequence of this sample, output the corrected predicted concentration result, and perform fusion weighted processing with the global model prediction result.

[0057] In some embodiments of the present application, the process of determining the degree of model fitting distortion includes: Calculating the deviation degree between the non-linear mapping feature factor group of the current soil sample and the historical average feature factor group; If the deviation degree is higher than the fitting deviation preset threshold, it is determined that there is fitting distortion, and the local sub-model is called to re-fit the path.

[0058] Specifically, the process of judging the degree of model fitting distortion includes the following steps: The first step is to obtain the historical average feature factor group. The system selects the measured samples in history that have similar soil types and compositions to the current area to be measured, screens out the samples with "highly credible labels" among them, and extracts two indicators in their non-linear mapping feature factor groups: the first non-linear feature factor (calculated from the ratio of the multi-order polynomial fitting residual to the threshold) and the second non-linear feature factor (calculated from the deviation degree between the spectral response curve and the kernel principal component projection curve). Calculate the mean values of these two factors respectively, and record them as the historical average feature factor group.

[0059] The second step is to extract the non-linear mapping feature factor group of the current soil sample to be measured.

[0060] The third step is to calculate the feature deviation degree between the current sample and the historical average sample. This deviation degree is obtained by using the weighted Euclidean distance calculation formula.

[0061] The fourth step is to compare the calculated deviation degree with the preset fitting deviation threshold. If the deviation degree is greater than the preset fitting deviation threshold, it is considered that there is a significant difference in non-linear feature mapping between the current sample and the historical credible sample, which may lead to prediction distortion of the global fitting model.

[0062] The fifth step is to, if it is judged that there is model fitting distortion, call the local fitting sub-model corresponding to the current area to replace the global model to predict the concentration of the current sample, so as to improve the adaptability and accuracy of the model.

[0063] In some embodiments of the present application, the model weight adjustment includes: When the reliability level label is a medium-credible label, the weights of multiple models are adjusted based on the first credibility factor, the second credibility factor, and the third credibility factor; If the first credibility factor of any model is greater than the preset residual threshold, the weight of the current model is reduced; If the second credibility factor of any model is greater than the preset dispersion threshold, the weight of the current model is reduced; If the third credibility factor of any model is less than the similarity threshold, the weight of the current model is reduced; After the model weights are adjusted, the weight adjustment results are used to perform weighted summation on the prediction results of multiple models, and the final predicted concentration result is output.

[0064] It should be noted that the purpose of the model weight adjustment process is to enhance the stability and accuracy of the combined model prediction by introducing a dynamic model scoring and weight allocation mechanism in the presence of certain uncertainties (i.e., when the detection result is marked as a "medium-confidence label"). The following is a step-by-step explanation of this process: When the confidence characterization value of a soil sample is between the "low-confidence risk threshold" and the "high-confidence risk threshold", the system marks the sample as a "medium-confidence label". The prediction of this sample has certain uncertainties, but it is not yet sufficient to trigger the local fitting sub-model. In this case, one cannot rely entirely on any single model, but rather needs to dynamically adjust the model weights based on the performance of each model on the current soil sample.

[0065] The first confidence factor (residual value): represents the error between the model-predicted concentration and the laboratory true concentration. If the residual value of a model on the current sample is too large, it indicates that the model has a poor fitting ability for this sample; the second confidence factor (multi-model prediction dispersion): represents the deviation degree between the predicted value of this model and the predicted values of other models. If the deviation is large, it means that the output of this model is not representative or lacks consistency with other models; the third confidence factor (spectral structure similarity): represents whether the sample on which the model is based is close to the similar samples in the spectral library. If the similarity is low, it indicates that the model has insufficient representativeness for the current sample.

[0066] If the first confidence factor is greater than the residual threshold (i.e., the prediction error is significantly too large), the model is judged to be inaccurate in fitting this sample, and its weight should be reduced; if the second confidence factor is greater than the dispersion threshold (i.e., the prediction result of this model differs greatly from other models), it indicates that the prediction of this model lacks robustness, and its weight should also be reduced; if the third confidence factor is less than the set similarity threshold (i.e., the sample has a large difference from the model training samples), then the reliability of the model inference is limited, and the weight should also be lowered.

[0067] These conditions are judged independently. Once a model triggers any one of the thresholds, it can be regarded as having a "poor confidence performance", and then its weight is adjusted. The adjustment amount of the weight is a preset value.

[0068] After the weights of all models are updated, the weighted average of the prediction results of multiple models will be calculated using the updated weights to obtain the final predicted concentration result. This method avoids the dominant role of a low-confidence model in the output and comprehensively utilizes the information of models with excellent performance to improve the overall prediction stability.

[0069] This mechanism does not simply "exclude" certain models, but carefully regulates the prediction contribution of each model through the quantitative calculation results of credibility factors. Especially when the credibility fluctuates but has not reached the level where the model needs to be rebuilt, this way of adjusting weights enables the system to still have good adaptability and prediction robustness. This mechanism provides a more realistic strategy for practical detection applications under the conditions of unstable current sample size and complex heterogeneous soil components.

[0070] In some embodiments of the present application, the construction process of the hyperspectral feature sequence includes: Based on the collected spectral response curves, extract the spectral principal component features, band ratio features, and continuous spectral segment slope change features within the specified wavelength band of 450 - 2450 nm; construct a multi-dimensional combined feature vector as the hyperspectral feature sequence of the soil sample; The process of obtaining the soil spectral reflectance image includes: Obtain the spectral reflectance images of different soil surfaces within the area to be measured; divide the spectral reflectance image into multiple sub-regions and extract the average spectral response curve of each sub-region; sequentially extract the spectral response curves of all sub-regions as the input for subsequent processing.

[0071] It should be noted that the construction process of the hyperspectral feature sequence includes: The hyperspectral feature sequence is a set of multi-dimensional feature vectors used to represent the spectral response characteristics of the soil, with advantages such as high information density and obvious feature differences. The specific construction process is as follows: Perform principal component analysis (PCA) on the original spectral response curves within the wavelength band of 450 - 2450 nm (i.e., the curve of the reflectance of each soil sub-region changing with the wavelength); select the principal component vectors with a cumulative contribution rate of more than 95%, and extract the corresponding principal component scores as the "spectral principal component features"; this part of the features is used to characterize the main change trend of the sample in the high-dimensional spectral space, which helps to reduce the dimension and remove redundancy.

[0072] Set several pairs of bands with diagnostic significance (such as 680 / 550 nm, 950 / 850 nm, etc.), and calculate the reflectance ratios between these band pairs; the band ratios can reveal the absorption differences between specific bands, spectral distortion, and the relative relationships of components such as moisture / organic matter; the obtained ratios form the "band ratio feature" vector.

[0073] In the entire spectral curve, divide several continuous band windows (such as every 50 nm as a segment), perform linear fitting on each segment, and extract the slope of the fitting line; at the same time, calculate the slope change rate between adjacent spectral segments to obtain the "slope difference feature"; this part of the features is used to identify the spectral change speed and local absorption edge information, which helps to identify the abnormal absorption behavior of trace elements in the soil.

[0074] The spectral principal component features, band ratio features, and continuous spectral band slope change features are spliced and combined to form a multi-dimensional combined feature vector of a unified dimension, that is, the hyperspectral feature sequence of the soil sample. The feature sequences of all sub-region samples are summarized together for subsequent feature factor calculation and model training.

[0075] In addition, the acquisition and processing process of the spectral reflectance image includes: The soil spectral reflectance image is the physical basis for obtaining hyperspectral features, and its acquisition and processing steps are as follows: Using a hyperspectral imaging device (such as a ground hyperspectral camera, an airborne imaging spectrometer, etc.), scan the area to be measured to obtain a spectral reflectance image data cube (a three-dimensional matrix of x, y, and λ), where the value range of λ is 450 - 2450 nm; each pixel records the reflectance of the soil at that position at different wavelengths.

[0076] The collected soil images are divided into several sub-regions according to the spatial dimension (such as 10×10 pixels or 50×50 pixels); the average value of the reflectance values of all pixels in each sub-region at each wavelength is taken to obtain the average spectral response curve of the region; the local noise and spatial stray interference are removed from the averaged curve, which reflects the representative soil spectral features of the region.

[0077] The spectral response curves (reflectance curves with wavelengths from 450 to 2450 nm) extracted from each sub-region are numbered in sequence; all sub-region curves are used as a set of samples and input into the subsequent hyperspectral feature extraction process to complete the construction of the hyperspectral feature sequence of the samples within the entire area to be measured.

[0078] Refer to Figure 2 As shown, the embodiment of the present invention provides a soil heavy metal pollution detection system based on spectral data, including: An acquisition module, configured to obtain the spectral reflectance image of the soil surface layer in the area to be measured and extract the corresponding hyperspectral feature sequence to construct the spectral response curve of the target soil; A calculation module, configured to calculate the non-linear mapping feature factor group and the credibility factor group of the soil sample based on the spectral response curve; A processing module, configured to divide the reliability level labels based on the credibility characterization value and adjust the acquisition process of the predicted concentration according to the reliability level labels; An output module, configured to output the final predicted concentration result.

[0079] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: it is still possible to modify the specific implementation manners of the present invention or make equivalent substitutions, and any modification or equivalent substitution that does not depart from the spirit and scope of the present invention should be covered within the protection scope of the claims of the present invention.

Claims

1. A method for detecting soil heavy metal pollution based on spectral data, characterized in that, Including: Obtain the spectral reflectance image of the soil surface layer in the area to be measured, extract the corresponding hyperspectral feature sequence, and construct the spectral response curve of the target soil; Based on the spectral response curve, calculate the non-linear mapping feature factor group and the credibility factor group of the soil sample; Construct a credibility characterization value according to the credibility factor group and the non-linear mapping feature factor group; Based on the credibility characterization value, divide the reliability level label, and perform the following steps according to the reliability level label: Based on the change of the non-linear mapping feature factor group of the soil sample, judge whether there is model fitting distortion to judge whether to call the local fitting sub-model; And, based on the credibility characterization value, judge whether to adjust the model weight; Output the final predicted concentration result; Wherein, the non-linear mapping feature factor group is calculated based on the multi-order non-linear function fitting residual of the hyperspectral principal component, and the credibility factor group is calculated based on the residual amplitude and the spectral structure similarity; 2. The method for detecting soil heavy metal pollution based on spectral data according to claim 1, wherein The calculation process of the non-linear mapping feature factor group includes: Perform multi-order polynomial fitting on the hyperspectral feature sequence to obtain a predicted concentration sequence; Calculate the sum of squared residuals between the predicted concentration sequence and the original concentration label, and determine the first non-linear feature factor based on the ratio of the sum of squared residuals to the reference residual threshold; the original concentration label is the true heavy metal concentration value measured in the laboratory; Take the deviation degree between the spectral response curve and its kernel principal component projection curve as the second non-linear feature factor; Weightedly sum the first non-linear feature factor and the second non-linear feature factor to form the non-linear mapping feature factor group of the sample; After establishing the association relationship between all soil samples and the non-linear mapping feature factor group, store them.

3. The method for detecting soil heavy metal pollution based on spectral data according to claim 2, wherein The calculation process of the credibility factor group includes: Extract the residual between the spectral response curve of each soil sample and the model predicted concentration as the first credibility factor; Calculate the concentration deviation dispersion degree of the current soil sample under multiple model prediction paths as the second credibility factor; the multiple models include a global fitting model, a local fitting sub-model, a multi-fitting path model, a regularization fitting model, and a kernel function regression model; Calculate the spectral structure cosine similarity between the current spectral response curve and the most similar sample in the spectral library as the third credibility factor; the spectral library is a database storing heavy metal types, corresponding contents, and corresponding spectral response curves; Weightedly average the first credibility factor, the second credibility factor, and the third credibility factor to construct the credibility factor group; And establish an index table between each soil sample and the credibility factor group.

4. The method for detecting soil heavy metal pollution based on spectral data according to claim 3, wherein, The calculation process of the credibility characterization value includes: Analyze the credibility factor groups of different soil samples in the same area to be measured, and obtain the difference between the maximum credibility factor group and the minimum credibility factor group as the credibility fluctuation degree; Take the ratio of the credibility fluctuation degree to the preset credibility fluctuation threshold as the credibility characterization value.

5. The method for detecting soil heavy metal pollution based on spectral data according to claim 4, characterized in that The process of dividing the reliability level label includes: If the credibility characterization value is less than the preset low credibility risk threshold, mark it as a high credibility label; If the credibility characterization value is higher than the preset high credibility risk threshold, mark it as a low credibility label; If the credibility characterization value is between the low credibility risk threshold and the high credibility risk threshold, it is marked as a medium credibility label.

6. The method for detecting soil heavy metal pollution based on spectral data according to claim 5, wherein If the detection result is a low credibility label, the following process is triggered: Based on the change of the non-linear mapping characteristic factor group of the current soil sample, judge the degree of model fitting distortion. If it is higher than the fitting deviation threshold, call the local fitting sub-model, which is used for predicting the heavy metal concentration in the current area to be measured.

7. The method for detecting soil heavy metal pollution based on spectral data according to claim 6, characterized in that, The process of judging the degree of model fitting distortion includes: Calculate the deviation degree between the non-linear mapping characteristic factor group of the current soil sample and the historical average characteristic factor group; If the deviation degree is higher than the preset fitting deviation threshold, it is determined that there is fitting distortion, and call the local sub-model to re-fit the path.

8. The method for detecting soil heavy metal pollution based on spectral data according to claim 7, wherein The model weight adjustment includes: When the reliability level label is a medium credibility label, adjust the weights of the multiple models based on the first credibility factor, the second credibility factor, and the third credibility factor; If the first credibility factor of any model is greater than the preset residual threshold, reduce the weight of the current model; If the second credibility factor of any model is greater than the preset dispersion threshold, reduce the weight of the current model; If the third credibility factor of any model is less than the similarity threshold, reduce the weight of the current model; After the model weight is adjusted, use the weight adjustment result to perform weighted summation on the prediction results of the multiple models, and output the final predicted concentration result.

9. The method for detecting soil heavy metal pollution based on spectral data according to claim 8, characterized in that, The construction process of the hyperspectral feature sequence includes: Based on the collected spectral response curve, extract the spectral principal component features, band ratio features, and continuous spectral segment slope change features within the specified band of 450 - 2450nm; construct a multi-dimensional combined feature vector as the hyperspectral feature sequence of the soil sample; The process of obtaining the soil spectral reflectance image includes: Obtain the spectral reflectance images of different soil surfaces in the area to be measured; divide the spectral reflectance image into multiple sub-regions and extract the average spectral response curve of each sub-region; sequentially extract the spectral response curves of all sub-regions as the input for subsequent processing.

10. A soil heavy metal pollution detection system based on spectral data, characterized in that, For implementing the method according to any one of claims 1 - 9, it includes: A collection module, configured to obtain the spectral reflectance image of the soil surface in the area to be measured, and extract the corresponding hyperspectral feature sequence to construct the spectral response curve of the target soil; A calculation module, configured to calculate the non-linear mapping characteristic factor group and the credibility factor group of the soil sample based on the spectral response curve; A processing module, configured to divide the reliability level label based on the credibility characterization value, and adjust the process of obtaining the predicted concentration according to the reliability level label; An output module, configured to output the final predicted concentration result.

Citation Information

Patent Citations

  • A method and system for detecting heavy metals in soil

    CN113791040B

  • Aviation hyperspectral image soil heavy metal concentration evaluation method based on gaussian process regression

    CN110174359A

  • Method for estimating content of heavy metals in soil based on hyperspectral remote sensing technology

    CN114018833A

  • Soil heavy metal content inversion method fusing spectrum and spatial characteristics

    CN115236005A

  • Method and system for detecting heavy metal pollution in soil sample

    CN119438539A