Soil heavy metal content inversion method, storage medium and electronic equipment
By constructing a low-dimensional feature space through covariance whitening-staining transformation and principal component analysis, and combining it with unsupervised clustering and support vector regression models, the problems of high sample acquisition cost and long data processing time in soil heavy metal pollution surveys were solved, and the heavy metal pollution status was quickly and efficiently clarified.
Patent Information
- Application Number
- CN202510939850.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-07-09
AI Technical Summary
Existing technologies in soil heavy metal pollution investigations have problems such as high sample acquisition costs, long data processing turnaround time, low model stability and poor timeliness, and cannot meet the needs of quickly and efficiently clarifying the heavy metal pollution status.
Covariance whitening-staining transformation and principal component analysis were used to construct a low-dimensional feature space. Unsupervised clustering and support vector regression models were combined to screen samples through Euclidean distance and construct an enhanced training set to achieve efficient prediction of soil heavy metal content.
It has achieved rapid and efficient clarification of heavy metal pollution status in a wide geographical space, with high prediction accuracy and fast acquisition process.
Smart Images

Figure CN120470559B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of hyperspectral remote sensing technology, and in particular to a soil heavy metal content inversion method, a storage medium and an electronic device. Background Art
[0002] With the development of urbanization, industrialization, and agricultural intensification, the problem of heavy metal pollution in soil has become increasingly prominent. The primary prerequisite for implementing soil heavy metal pollution prevention and control efforts is to understand the status of heavy metal pollution in the soil. Current methods for investigating heavy metal pollution in soil rely primarily on manual field sampling and testing and geostatistical interpolation. While manual field sampling and testing yields high accuracy, biases in the selection of discrete points during soil surveys often introduce significant uncertainty into the assessment of the scope and severity of heavy metal pollution. Furthermore, this method suffers from high cost, low efficiency, and poor timeliness. While geostatistical interpolation offers advantages such as low cost and large-scale statistical coverage, it relies heavily on the number of sample points, the uniformity of their spatial distribution, and model assumptions, resulting in poor accuracy and reliability in its predictions.
[0003] At the same time, emerging remote sensing detection technologies have the characteristics of being able to conduct rapid, dynamic and large-scale monitoring of target objects in a non-contact manner. Among them, hyperspectral remote sensing has gradually been widely used in soil heavy metal pollution monitoring research due to its advantages of multiple and continuous spectral bands. However, in the actual work of hyperspectral remote sensing detection of heavy metals in soil, there are often prominent problems such as high soil sample acquisition costs, long data processing turnaround time, low model stability and poor timeliness of soil heavy metal pollution investigations, which cannot meet the needs of quickly and efficiently clarifying the heavy metal pollution status in a wide geographical space. Therefore, it is necessary to provide a soil heavy metal content inversion method, storage medium and electronic equipment to meet the needs of quickly and efficiently clarifying the heavy metal pollution status in a wide geographical space. Summary of the Invention
[0004] The present invention aims to provide a soil heavy metal content inversion method, storage medium and electronic equipment. The specific technical solutions are as follows:
[0005] In a first aspect, the present invention provides a method for inverting soil heavy metal content, comprising:
[0006] Step S1: Use existing literature to obtain a heavy metal contaminated soil spectral library; collect soil samples in the target study area, measure the heavy metal content and spectral data of each soil sample, and obtain a local sample library. The data in the local sample library is the local sample set. D local; Eliminate abnormal samples in the spectral library; Eliminate spectral data in the local sample library that are not related to heavy metal content; Use covariance whitening-staining transformation to transform the spectral data in the spectral library to have the same mean and covariance as the local sample library to obtain a corrected spectral library sample set ;
[0007] Step S2: local sample set D local and the corrected spectral library sample set After standardization preprocessing, principal component analysis is used to analyze the local sample set. D local and the corrected spectral library sample set Perform feature space projection, retain the principal components with cumulative variance contribution rate ≥ 95%, and construct a low-dimensional feature space F reduced ;
[0008] Step S3: First, based on the silhouette coefficient optimization criterion, in the low-dimensional feature space F reduced Calculate the optimal number of clusters k optimal ; Secondly, combined with the optimal number of clusters k optimal , execute unsupervised clustering algorithm to obtain cluster division, and count local samples in low-dimensional feature space F reduced Cluster coverage within r ; Then, according to the cluster coverage r The range determines the Euclidean distance threshold T euclidean ;
[0009] Step S4: Select a test sample with unknown heavy metal content from the local samples S test , calculate the test sample S test The spectral library j The Euclidean distance between samples d j ; Screen all local samples that meet d j ≤ T euclidean The samples constitute the initial candidate set C initial ; For the preliminary candidate set C initial After the secondary adaptive selection criteria are used to divide the matching sample set, D selected ;
[0010] Step S5: Merge local sample sets D local and matching sample sets D selected , construct an enhanced training set D enhanced ;extract D enhanced The corresponding feature subset , and then use the support vector regression inversion model to test samples S test The heavy metal content value of the test sample is predicted S test The predicted value of heavy metal content is used to invert the soil heavy metal content; among them, B feature It is the characteristic band of the soil heavy metal spectrum response.
[0011] Optionally, in step S1, the method for eliminating abnormal samples in the spectral library includes: establishing a partial least squares model with the heavy metal content of the spectral library as the dependent variable and the spectral data of the local sample library as the independent variable, performing leave-one-out cross-validation, and calculating the absolute value of the residual between the true value and the predicted value of the heavy metal content of all samples in the spectral library; calculating the standard deviation of the absolute value of the residual of the heavy metal content of all samples in the spectral library s If the absolute value of the residual of the heavy metal content of a sample is greater than 3 s , it is considered that the prediction error of the heavy metal content of the sample deviates significantly from the overall trend, and it is an abnormal sample and is removed from the spectral library.
[0012] Optionally, in step S1, a corrected spectral library sample set is obtained. Methods include:
[0013] First, the covariance matrix of the local spectral data after orthogonal projection correction is decomposed using Equations 1) and 2) ;
[0014] Formula 1);
[0015] Formula 2);
[0016] In formula 1), Indicates the number of local samples; represents the orthogonal complement space; Represents the orthogonal complement space The transpose of
[0017] In formula 2), is the eigenvector matrix of the spectral covariance matrix of the local sample; for The transpose of is the eigenvalue diagonal matrix of the spectral covariance matrix of the local sample;
[0018] Secondly, use Equation 3) to construct the coloring operator ; Use Equation 4) to construct the whitening operator ;
[0019] Formula 3);
[0020] Formula 4);
[0021] In formula 3), express The square root diagonal matrix of ;
[0022] In formula 4), is the eigenvector matrix of the spectral covariance matrix of the spectral library; is the inverse diagonal matrix of the square roots of the spectral eigenvalues of the spectral library;
[0023] Finally, affine transformation is used to map the spectral library data into the local feature space to obtain the corrected spectral library sample set. ; The local feature space refers to the spatial structure or distribution shape defined by the statistical characteristics of the local samples; the statistical characteristics include the mean and covariance of the local samples; the affine transformation is expressed by formula 5);
[0024] Equation 5);
[0025] In formula 5), is the spectrum matrix of the spectrum library after whitening-staining transformation; is the spectrum matrix of the spectral library before whitening-staining transformation; is the mean vector of the spectrum in the spectral library; is the mean vector of the local sample library spectrum.
[0026] Optionally, in step S1, the orthogonal complement space is obtained The methods include:
[0027] First, an orthogonal residual space independent of the heavy metal content in the local sample library is obtained. , which is expressed by formula 6);
[0028] Equation 6);
[0029] In formula 6), X represents the spectrum matrix in the local sample library; Y Indicates the heavy metal content in the local sample library;W T for Y right X The regression weight vector W The transpose of
[0030] regression weight vector W It is expressed by formula 7);
[0031] Equation 7);
[0032] In formula 7), X T for X The transpose of X T X ) -1 for X T X The inverse matrix of
[0033] Secondly, for the orthogonal residual space Perform singular value decomposition to extract the principal component loading matrix P , the spectral matrix X Projection to orthogonal complement space , retaining the sensitive characteristics of heavy metal content;
[0034] Orthogonal complement space It is expressed by formula 8);
[0035] Equation 8);
[0036] In formula 8), P T for P The transpose of P T P ) -1 for P T P The inverse matrix of .
[0037] Optionally, in step S3, the optimal number of clusters k optimal It is expressed using formula 9);
[0038] Equation 9);
[0039] In formula 9), The local sample set involved in the calculation D local and the corrected spectral library sample set The total number of samples in ; The value range is 1 to Any natural number between Local sample set D local and the corrected spectral library sample set The silhouette coefficient of the sample in ; is the maximum number of candidate clusters; The value range is 2 to Any natural number between
[0040] The cluster coverage r It is expressed using formula 10);
[0041] Equation 10);
[0042] In formula 10), C j is the cluster set after clustering; cluster coverage r The value range includes r ≤0.25、0.25< r ≤0.5、0.5< r ≤0.75 and r >0.75;
[0043] The Euclidean distance threshold T euclidean Cluster coverage r The corresponding relationship is expressed by formula 11);
[0044] Equation 11).
[0045] Optionally, in step S4, the Euclidean distance d j It is expressed using formula 12);
[0046] Equation 12);
[0047] In formula 12), For test samples S test In low-dimensional feature space F reduced Middle The eigenvalue of the dimension; The spectral library j samples in the low-dimensional feature space F reduced Middle The eigenvalue of the dimension; The number of dimensions retained after dimensionality reduction; The value range is 1 to Any natural number between
[0048] The secondary adaptive selection criterion is expressed as Equation 13) and is used to determine the initial candidate set C initial The number of samples in N select ;
[0049] Equation 13);
[0050] In formula 13), N min Represents the preliminary candidate set C initial The minimum number of samples in ; Represents the preliminary candidate set C initial The maximum number of samples in ; Represents the preliminary candidate set C initial The number of samples in ;
[0051] when When , the closest Euclidean distance threshold is forced to be selected T euclidean Before N min samples;
[0052] when When the closest Euclidean distance threshold is retained T euclidean The first 20% of samples;
[0053] when When C initial All samples in .
[0054] Optionally, in step S5, the test sample S test The predicted value of heavy metal content is expressed by formula 14);
[0055] Equation 14);
[0056] In formula 14), For test samples S test Predicted values of heavy metal content; is the regression hyperplane parameter completed by training; SVR is the Gaussian radial basis kernel function; for SVR Mapping; B feature It is the characteristic band of the soil heavy metal spectrum response.
[0057] Optionally, the B feature The steps to obtain include:
[0058] First, background soil from heavy metal contaminated areas was selected as the reference sample. After pretreatment, a standard salt solution of a single target heavy metal was injected into each reference sample. N 0 standard soil samples with different heavy metal content gradient levels to obtain a standardized pollution sample set; among them, N 0>40;
[0059] Secondly, in a darkroom environment, a spectrometer is used to collect the spectral reflectance of each standard soil sample in each band within the range of 350-2500nm;
[0060] Finally, combined with the heavy metal content and full-band spectral reflectance of each standard soil sample, the variable projection importance algorithm is used to quantify the contribution of each band spectral reflectance to the prediction of heavy metal content, and then the soil heavy metal spectral response characteristic band is selected. B feature .
[0061] In a second aspect, the present invention provides a storage medium having computer program instructions stored thereon, which implement the soil heavy metal content inversion method when the computer program instructions are executed by a processor.
[0062] In a third aspect, the present invention provides an electronic device comprising: at least one processor, at least one memory, and computer program instructions stored in the memory, which implement the soil heavy metal content inversion method when the computer program instructions are executed by the processor.
[0063] The application of the technical solution of the present invention has at least the following beneficial effects:
[0064] The present invention provides a soil heavy metal content inversion method that can meet the needs of quickly and efficiently identifying heavy metal pollution conditions in a wide geographical space. Specifically, the present invention adopts step S1 to obtain a local sample set. D local and the corrected spectral library sample set Among them, the covariance whitening-staining transformation is helpful to improve the consistency of the spectral statistical characteristics of local samples and spectral library samples; step S2 is used to process D local and , construct a low-dimensional feature space F reduced ; Use step S3 in low-dimensional feature space F reducedCalculate the optimal number of clusters k optimal , combined with k optimal , execute unsupervised clustering algorithm to obtain cluster division, and count local samples in low-dimensional feature space F reduced Cluster coverage within r ; Then, according to the cluster coverage r The range determines the Euclidean distance threshold T euclidean ; Step S4 is used to select a test sample with unknown heavy metal content value from the local sample S test , calculate the test sample S test The spectral library j The Euclidean distance between samples d j ; Screen all local samples that meet d j ≤ T euclidean The samples constitute the initial candidate set C initial ; For the preliminary candidate set C initial After the secondary adaptive selection criteria are used to divide the matching sample set, D selected That is, the present invention combines step S3 and step S4 to achieve the matching sample set D selected Adaptive dynamic adjustment of sample quantity to ensure sufficient sample quantity; adopt step S5 to merge local sample sets D local and matching sample sets D selected , get the enhanced training set D enhanced ;extract D enhanced The corresponding feature subset , and then use the support vector regression inversion model to test samples S test The heavy metal content value of the test sample is predicted S test The predicted value of heavy metal content is used to invert the heavy metal content in soil, with high prediction accuracy and fast and efficient acquisition process. Therefore, the inversion method provided by the present invention is to S test The prediction results verified that this method can meet the needs of quickly and efficiently clarifying the heavy metal pollution status in a wide geographical space.
[0065] In addition to the above-described objects, features and advantages, the present invention has other objects, features and advantages. The present invention will be further described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are intended to explain the present invention and do not constitute an undue limitation of the present invention. In the accompanying drawings:
[0067] Figure 1 1 is a flow chart of a method for inverting soil heavy metal content in an embodiment;
[0068] Figure 2 This is a prediction result diagram of a soil heavy metal content inversion method in an embodiment. DETAILED DESCRIPTION
[0069] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention are within the scope of protection of the present invention.
[0070] Example:
[0071] See also Figure 1 , a soil heavy metal content inversion method, comprising:
[0072] Step S1: Obtain a heavy metal contaminated soil spectral library using existing literature (specifically: Ning Jing. Research on spectral feature identification and database construction of heavy metal contaminated soil [D]. Central South University, 2023. DOI: 10.27661 / d.cnki.gzhnu.2023.004518.); collect soil samples in the target study area, measure the heavy metal content and spectral data of each soil sample, and obtain a local sample library. The data in the local sample library is the local sample set. D local ; Eliminate abnormal samples in the spectral library; Eliminate spectral data in the local sample library that are not related to heavy metal content; Use covariance whitening-staining transformation to transform the spectral data in the spectral library to have the same mean and covariance as the local sample library to obtain a corrected spectral library sample set .
[0073] The method to obtain the local sample library is as follows:
[0074] Forty-two surface soil samples (i.e., 0–20 cm below the ground surface) were collected from the target study area using a completely randomized sampling route according to the FOREGS Geochemical Mapping Field Manual (Salminen, R., Tarvainen, T., Demetriades, A., Duris, M., Fordyce, FM, Gregorauskiene, V., Kahelin, H., Kivisilla, J., Klaver, G., Klein, H., Larson, JO, Lis, J., Locutura, J., Marsina, K., Mjartanova, H., Mouvet, C., O Connor, P., Odor, L., Ottonello, G., Paukola, T., Plant, JA, Reimann, C., Schermann, O., Siewers, U., Steenfelt, A., Van der Sluys, J., De Vivo, B., Williams, L., 1998. FOREGSGeochemical Mapping Field Manual. Espoo, Finland, Geological Survey of Finland. Geological Survey of Finland Guide, p. 38.), the geographic locations of soil samples were recorded by Global Positioning System (GPS) receivers;
[0075] The collected soil samples were air-dried, ground and sieved; the As content was determined by inductively coupled plasma optical emission spectrometry, with reference to the experimental section on page 76 of the existing literature (specifically: Zhao Qingling, Li Qingcai. Simultaneous determination of 54 components in soil samples by inductively coupled plasma optical emission spectrometry [J]. Rock and Mineral Testing, 2011, 30(01):75-78. DOI:10.15898 / j.cnki.11-2131 / td.2011.01.025.).
[0076] Step S2: local sample set D local and the corrected spectral library sample set After standardization preprocessing, principal component analysis (PCA) was used to analyze the local sample set. D local and the corrected spectral library sample set Perform feature space projection, retain the principal components with cumulative variance contribution rate ≥ 95%, and construct a low-dimensional feature space F reduced .
[0077] Step S3: First, based on the silhouette coefficient optimization criterion, in the low-dimensional feature space F reduced Calculate the optimal number of clusters k optimal ; Secondly, combined with the optimal number of clusters k optimal , execute unsupervised clustering algorithm (i.e. K-means) to obtain cluster division, and count local samples in low-dimensional feature space F reduced Cluster coverage within r ; Then, according to the cluster coverage r The range determines the Euclidean distance threshold T euclidean .
[0078] Step S4: Select a test sample with unknown heavy metal content from the local samples S test , calculate the test sample S test The spectral library j The Euclidean distance between samples d j ; Screen all local samples that meet d j ≤ T euclidean The samples constitute the initial candidate set C initial ; For the preliminary candidate set C initial After the secondary adaptive selection criteria are used to divide the matching sample set, D selected .
[0079] Step S5: Merge local sample sets D local and matching sample sets D selected , construct an enhanced training set D enhanced ;extract D enhanced The corresponding feature subset , and then use the support vector regression inversion model to test samples S test The heavy metal content value of the test sample is predicted S test The predicted value of heavy metal content is used to invert the soil heavy metal content; among them,B feature It is the characteristic band of the soil heavy metal spectrum response.
[0080] In step S1, the method for eliminating abnormal samples in the spectral library includes: establishing a partial least squares model (i.e., PLSR) with the heavy metal content of the spectral library as the dependent variable and the spectral data of the local sample library as the independent variable, performing leave-one-out cross-validation, and calculating the absolute value of the residual between the true value and the predicted value of the heavy metal content of all samples in the spectral library; calculating the standard deviation of the absolute value of the residual of the heavy metal content of all samples in the spectral library s If the absolute value of the residual of the heavy metal content of a sample is greater than 3 s , it is considered that the prediction error of the heavy metal content of the sample deviates significantly from the overall trend, and it is an abnormal sample and is removed from the spectral library.
[0081] In step S1, a corrected spectral library sample set is obtained. Methods include:
[0082] First, the covariance matrix of the local spectral data after orthogonal projection correction is decomposed using Equations 1) and 2) ;
[0083] Formula 1);
[0084] Formula 2);
[0085] In formula 1), Indicates the number of local samples; represents the orthogonal complement space; Represents the orthogonal complement space The transpose of
[0086] In formula 2), is the eigenvector matrix of the spectral covariance matrix of the local sample; for The transpose of is the eigenvalue diagonal matrix of the spectral covariance matrix of the local sample;
[0087] Secondly, use Equation 3) to construct the coloring operator ; Use Equation 4) to construct the whitening operator ;
[0088] Formula 3);
[0089] Formula 4);
[0090] In formula 3), express The square root diagonal matrix of ;
[0091] In formula 4), is the eigenvector matrix of the spectral covariance matrix of the spectral library; is the inverse diagonal matrix of the square roots of the spectral eigenvalues of the spectral library;
[0092] Finally, affine transformation is used to map the spectral library data into the local feature space to obtain the corrected spectral library sample set. ; The local feature space refers to the spatial structure or distribution shape defined by the statistical characteristics of the local samples; the statistical characteristics include the mean and covariance of the local samples; the affine transformation is expressed by formula 5);
[0093] Equation 5);
[0094] In formula 5), is the spectrum matrix of the spectrum library after whitening-staining transformation; is the spectrum matrix of the spectral library before whitening-staining transformation; is the mean vector of the spectrum in the spectral library; is the mean vector of the local sample library spectrum.
[0095] In step S1, the orthogonal complement space is obtained The methods include:
[0096] First, an orthogonal residual space independent of the heavy metal content in the local sample library is obtained. , which is expressed by formula 6);
[0097] Equation 6);
[0098] In formula 6), X represents the spectrum matrix in the local sample library; Y Indicates the heavy metal content in the local sample library; W T for Y right X The regression weight vector W The transpose of
[0099] regression weight vector W It is expressed by formula 7);
[0100] Equation 7);
[0101] In formula 7), X T for X The transpose of X T X) -1 for X T X The inverse matrix of
[0102] Secondly, for the orthogonal residual space Perform singular value decomposition to extract the principal component loading matrix P , the spectral matrix X Projection to orthogonal complement space , retaining the sensitive characteristics of heavy metal content;
[0103] Orthogonal complement space It is expressed by formula 8);
[0104] Equation 8);
[0105] In formula 8), P T for P The transpose of P T P ) -1 for P T P The inverse matrix of .
[0106] In step S3, the optimal number of clusters k optimal It is expressed using formula 9);
[0107] Equation 9);
[0108] In formula 9), The local sample set involved in the calculation D local and the corrected spectral library sample set The total number of samples in ; The value range is 1 to Any natural number between Local sample set D local and the corrected spectral library sample set The silhouette coefficient of the sample in ; is the maximum number of candidate clusters; The value range is 2 to Any natural number between
[0109] The cluster coverage r It is expressed using formula 10);
[0110] Equation 10);
[0111] In formula 10), C j is the cluster set after clustering; cluster coverage r The value range includes r ≤0.25、0.25< r ≤0.5、0.5< r ≤0.75 and r >0.75;
[0112] The Euclidean distance threshold T euclidean Cluster coverage r The corresponding relationship is expressed by formula 11);
[0113] Equation 11).
[0114] In step S4, the Euclidean distance d j It is expressed using formula 12);
[0115] Equation 12);
[0116] In formula 12), For test samples S test In low-dimensional feature space F reduced Middle The eigenvalue of the dimension; The spectral library j samples in the low-dimensional feature space F reduced Middle The eigenvalue of the dimension; The number of dimensions retained after dimensionality reduction; The value range is 1 to Any natural number between
[0117] The secondary adaptive selection criterion is expressed as Equation 13) and is used to determine the initial candidate set C initial The number of samples in N select ;
[0118] Equation 13);
[0119] In formula 13), N min Represents the preliminary candidate set C initial The minimum number of samples in this embodiment N min The value is 10; Represents the preliminary candidate set C initial The maximum number of samples in this embodiment The value is 200; Represents the preliminary candidate set C initial The number of samples in ;
[0120] when When , the closest Euclidean distance threshold is forced to be selected T euclidean Before N min samples;
[0121] when When the closest Euclidean distance threshold is retained T euclidean The first 20% of samples;
[0122] when When C initial All samples in .
[0123] In step S5, the test sample S test The predicted value of heavy metal content is expressed by formula 14);
[0124] Equation 14);
[0125] In formula 14), For test samples S test Predicted values of heavy metal content; is the regression hyperplane parameter completed by training; SVR is the Gaussian radial basis kernel function; for SVR Mapping; B feature It is the characteristic band of the soil heavy metal spectrum response.
[0126] described B feature The steps to obtain include:
[0127] First, background soil from heavy metal contaminated areas was selected as the reference sample. After pretreatment, a standard salt solution of a single target heavy metal (specifically As) was injected into each reference sample. N 0 standard soil samples with different heavy metal content gradient levels, and obtain a class of standardized pollution sample sets, see Table 1 for details; among them, N 0>40, in this embodiment N0 takes the value 41;
[0128] Secondly, in a darkroom environment, a PSR-3500 portable spectrometer was used to collect the spectral reflectance of each standard soil sample in the range of 350-2500nm (a total of 1024 bands); the spectrometer sampling bandwidth was 1.5nm (350-1000nm), 3.8nm (1000-1900nm), and 2.5nm (1900-2500nm); the absolute reflectance of the spectrometer must be obtained by whiteboard calibration before testing; a 1000W halogen lamp was used as the test light source in the spectral measurement, with a controlled 5° field of view and a 15° angle between the light source and the vertical direction; the light source distance was set to 30cm and the probe distance to 5cm. The reflectance spectrum of each standard soil sample was measured 10 times, and the arithmetic average of the spectral reflectance of each band was finally taken as the spectral reflectance of each band of the standard soil sample;
[0129] Finally, combined with the heavy metal content of each standard soil sample and the spectral reflectance of each band, the variable projection importance algorithm is used to quantify the contribution of the spectral reflectance of each band to the prediction of heavy metal content, and then the characteristic bands of soil heavy metal spectral response are selected. B feature .
[0130] Table 1 Class-standardized contaminated sample set
[0131]
[0132] The determination coefficient expressed by formula (15) R 2 The root mean square error (RMS) expressed by the indicator (Eq. 16) RMSE The relative analysis error expressed by the index and formula (17) RPD The prediction accuracy of Equation 14) is evaluated by using the following indicators. The prediction accuracy results are shown in Tables 2 and Figure 1 .
[0133] Equation 15);
[0134] Equation 16);
[0135] Equation 17);
[0136] In Equations 15) to 17), For test samples S test the number of The value range is 1 to Any natural number between For the Test samples Stest Detection value of heavy metal content; For the Test samples S test Predicted values of heavy metal content; For all The arithmetic mean of the sum.
[0137] Table 2 Prediction accuracy results of each indicator
[0138]
[0139] From Table 2 and Figure 1 Know, when R 2 Value and RPD The larger the value, the RMSE The smaller the value, the more Formula 14) S test The higher the prediction accuracy is. Specifically, R 2 The closer the value is to 1, the more effective Equation 14) is for the test sample. S test The higher the degree of fit of the predicted regression line, the closer it is to the ideal 1:1 line; RMSE The smaller the value, the more effective Equation 14) is for the test sample. S test The better the predictive ability of RPD When the value is <1.4, Equation 14) cannot be used for the test sample S test Make effective predictions; when 1.4< RPD When the value is <2, Equation 14) can be used to test the sample S test Make a rough forecast; RPD When the value is > 2, Equation 14) can be used to test the sample S test Has excellent predictive ability; in this embodiment RPD The value is close to 2, indicating that Equation 14) is effective for the test sample. S test Has good predictive ability. Figure 1 All test samples S test The predicted values are all at 95% (i.e. Figure 1 middle α =0.05), which is used to indicate the credible range of the predicted value. The uncertainty of the prediction is well controlled, and its robustness and generalization ability are strong.
[0140] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A soil heavy metal content inversion method, characterized in that: include: Step S1: Use existing literature to obtain a heavy metal contaminated soil spectral library; collect soil samples in the target study area, measure the heavy metal content and spectral data of each soil sample, and obtain a local sample library. The data in the local sample library is the local sample set. D local ; Eliminate abnormal samples in the spectral library; Eliminate spectral data in the local sample library that are not related to heavy metal content; Use covariance whitening-staining transformation to transform the spectral data in the spectral library to have the same mean and covariance as the local sample library to obtain a corrected spectral library sample set ; Step S2: local sample set D local and the corrected spectral library sample set After standardization preprocessing, principal component analysis is used to analyze the local sample set. D local and the corrected spectral library sample set Perform feature space projection, retain the principal components with cumulative variance contribution rate ≥ 95%, and construct a low-dimensional feature space F reduced ; Step S3: First, based on the silhouette coefficient optimization criterion, in the low-dimensional feature space F reduced Calculate the optimal number of clusters k optimal ; Secondly, combined with the optimal number of clusters k optimal , execute unsupervised clustering algorithm to obtain cluster division, and count local samples in low-dimensional feature space F reduced Cluster coverage within ρ ; Then, according to the cluster coverage ρ The range determines the Euclidean distance threshold T euclidean ; Step S4: Select a test sample with unknown heavy metal content from the local samples S test , calculate the test sample S test The spectral library j The Euclidean distance between samples d j ; Screen all local samples that meet d j ≤ T euclidean The samples constitute the initial candidate set C initial ; For the preliminary candidate set C initial After the secondary adaptive selection criteria are used to divide the matching sample set, D selected ; Step S5: Merge local sample sets D local and matching sample sets D selected , construct an enhanced training set D enhanced ;extract D enhanced The corresponding feature subset , and then use the support vector regression inversion model to test samples S test The heavy metal content value of the test sample is predicted S test The predicted value of heavy metal content is used to invert the soil heavy metal content; in, B feature It is the characteristic band of the soil heavy metal spectrum response.
2. The soil heavy metal content inversion method according to claim 1, characterized in that: In step S1, the method for eliminating abnormal samples in the spectral library includes: establishing a partial least squares model with the heavy metal content of the spectral library as the dependent variable and the spectral data of the local sample library as the independent variable, performing leave-one-out cross-validation, and calculating the absolute value of the residual between the true value and the predicted value of the heavy metal content of all samples in the spectral library; calculating the standard deviation of the absolute value of the residual of the heavy metal content of all samples in the spectral library σ If the absolute value of the residual of the heavy metal content of a sample is greater than 3 σ , it is considered that the prediction error of the heavy metal content of the sample deviates significantly from the overall trend, and it is an abnormal sample and is removed from the spectral library.
3. The soil heavy metal content inversion method according to claim 1, characterized in that: In step S1, a corrected spectral library sample set is obtained. Methods include: First, the covariance matrix of the local spectral data after orthogonal projection correction is decomposed using Equations 1) and 2) ; Formula 1); Formula 2); In formula 1), Indicates the number of local samples; represents the orthogonal complement space; Represents the orthogonal complement space The transpose of In formula 2), is the eigenvector matrix of the spectral covariance matrix of the local sample; for The transpose of is the eigenvalue diagonal matrix of the spectral covariance matrix of the local sample; Secondly, use Equation 3) to construct the coloring operator ; Use Equation 4) to construct the whitening operator ; Formula 3); Formula 4); In formula 3), express The square root diagonal matrix of ; In formula 4), is the eigenvector matrix of the spectral covariance matrix of the spectral library; is the inverse diagonal matrix of the square roots of the spectral eigenvalues of the spectral library; Finally, affine transformation is used to map the spectral library data into the local feature space to obtain the corrected spectral library sample set. ; The local feature space refers to the spatial structure or distribution shape defined by the statistical characteristics of the local samples; the statistical characteristics include the mean and covariance of the local samples; the affine transformation is expressed by formula 5); Equation 5); In formula 5), is the spectrum matrix of the spectrum library after whitening-staining transformation; is the spectrum matrix of the spectral library before whitening-staining transformation; is the mean vector of the spectrum in the spectral library; is the mean vector of the local sample library spectrum.
4. The soil heavy metal content inversion method according to claim 1, characterized in that: In step S1, the orthogonal complement space is obtained The methods include: First, an orthogonal residual space independent of the heavy metal content in the local sample library is obtained. , which is expressed by formula 6); Equation 6); In formula 6), X represents the spectrum matrix in the local sample library; Y Indicates the heavy metal content in the local sample library; W T for Y right X The regression weight vector W The transpose of regression weight vector W It is expressed by formula 7); Equation 7); In formula 7), X T for X The transpose of X T X ) -1 for X T X The inverse matrix of Secondly, for the orthogonal residual space Perform singular value decomposition to extract the principal component loading matrix P , the spectral matrix X Projection to orthogonal complement space , retaining the sensitive characteristics of heavy metal content; Orthogonal complement space It is expressed by formula 8); Equation 8); In formula 8), P T for P The transpose of P T P ) -1 for P T P The inverse matrix of .
5. The soil heavy metal content inversion method according to claim 1, characterized in that: In step S3, the optimal number of clusters k optimal It is expressed using formula 9); Equation 9); In formula 9), The local sample set involved in the calculation D local and the corrected spectral library sample set The total number of samples in ; The value range is 1 to Any natural number between Local sample set D local and the corrected spectral library sample set The silhouette coefficient of the sample in ; is the maximum number of candidate clusters; The value range is 2 to Any natural number between The cluster coverage ρ It is expressed using formula 10); Equation 10); In formula 10), C j is the cluster set after clustering; cluster coverage ρ The value range includes ρ ≤0.25、0.25< ρ ≤0.5、0.5< ρ ≤0.75 and ρ >0.75; The Euclidean distance threshold T euclidean Cluster coverage ρ The corresponding relationship is expressed by formula 11); Equation 11).
6. The soil heavy metal content inversion method according to claim 5, characterized in that: In step S4, the Euclidean distance d j It is expressed using formula 12); Equation 12); In formula 12), For test samples S test In low-dimensional feature space F reduced Middle The eigenvalue of the dimension; The spectral library j samples in the low-dimensional feature space F reduced Middle The eigenvalue of the dimension; The number of dimensions retained after dimensionality reduction; The value range is 1 to Any natural number between The secondary adaptive selection criterion is expressed as Equation 13) and is used to determine the initial candidate set C initial The number of samples in N select ; Equation 13); In formula 13), N min Represents the preliminary candidate set C initial The minimum number of samples in ; Represents the preliminary candidate set C initial The maximum number of samples in ; Represents the preliminary candidate set C initial The number of samples in ; when When , the closest Euclidean distance threshold is forced to be selected T euclidean Before N min samples; when When the closest Euclidean distance threshold is retained T euclidean The first 20% of samples; when When C initial All samples in .
7. The soil heavy metal content inversion method according to claim 6, characterized in that: In step S5, the test sample S test The predicted value of heavy metal content is expressed by formula 14); Equation 14); In formula 14), For test samples S test Predicted values of heavy metal content; is the regression hyperplane parameter completed by training; SVR is the Gaussian radial basis kernel function; for SVR Mapping; B feature It is the characteristic band of the soil heavy metal spectrum response.
8. The soil heavy metal content inversion method according to claim 7, characterized in that: described B feature The steps to obtain include: First, background soil from heavy metal contaminated areas was selected as the reference sample. After pretreatment, a standard salt solution of a single target heavy metal was injected into each reference sample. N 0 standard soil samples with different heavy metal content gradient levels to obtain a standardized pollution sample set; among them, N 0>40; Secondly, in a darkroom environment, a spectrometer is used to collect the spectral reflectance of each standard soil sample in each band within the range of 350-2500nm; Finally, combined with the heavy metal content and full-band spectral reflectance of each standard soil sample, the variable projection importance algorithm is used to quantify the contribution of each band spectral reflectance to the prediction of heavy metal content, and then the soil heavy metal spectral response characteristic band is selected. B feature .
9. A storage medium, characterized in that: Computer program instructions are stored thereon, and when the computer program instructions are executed by a processor, the soil heavy metal content inversion method according to any one of claims 1 to 8 is implemented.
10. An electronic device, characterized in that: include: At least one processor, at least one memory, and computer program instructions stored in the memory, which, when executed by the processor, implement the soil heavy metal content inversion method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Nuclear reactor digital twin key parameter autonomous optimization data inversion method
CN115618732A
Soil organic carbon spectral prediction method and apparatus based on spectrum-guided ensemble learning
WO2024259827A1