Source apportionment of soil heavy metals based on geostatistical analysis and apls-mlr

By combining the improved APLS-MLR method with geostatistical analysis, the limitation of APLS-MLR in eigenvalue decomposition in soil heavy metal source apportionment is solved, realizing accurate quantitative and visual analysis of pollution sources and supporting the prevention and control of heavy metal pollution in farmland soil.

CN116166923BActive Publication Date: 2026-04-24CHINA THREE GORGES UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA THREE GORGES UNIV
Filing Date
2022-12-11
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

The existing absolute principal component-multiple linear regression method APCS-MLR has limitations in eigenvalue decomposition and lacks visualization of results in soil heavy metal source apportionment, making it difficult to accurately analyze and determine specific pollution sources or pollution source types.

Method used

A partial least squares regression (APCS-MLR) method was improved by combining partial least squares regression with geostatistical analysis to establish an absolute partial least squares-multiple linear regression (APLS-MLR) method. The pollution sources were identified by Kriging interpolation and the contribution rate of the pollution sources was determined by combining multiple linear regression analysis.

Benefits of technology

It improves the accuracy and visualization of soil heavy metal source apportionment, and can accurately quantify the nature and contribution rate of pollution sources, supporting the prevention and control of heavy metal pollution in farmland soil.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116166923B_ABST
    Figure CN116166923B_ABST
Patent Text Reader

Abstract

The present application relates to a soil heavy metal source analysis method based on geostatistical analysis and APLS-MLR, comprising: sampling the soil of a research area, measuring the content of heavy metals in the soil sample, and pretreating; using the Kriging interpolation method to analyze the spatial content distribution characteristic map of the soil heavy metals in the research area; using the partial least squares method to analyze the heavy metal concentration data of the research area; establishing an absolute partial least squares-multiple linear regression method receptor model for soil heavy metal pollution source analysis; combining the spatial content distribution characteristic map of the soil heavy metals and the contribution rate of each pollution source to infer and determine the specific pollution source. The method of the present application can not only calculate and determine the number of pollution sources and the contribution rate of each pollution source, but also accurately determine the specific pollution source; the APLS-MLR method of the receptor model proposed in the present application solves the problem of the limitation of eigenvalue decomposition in principal component analysis in the APCS-MLR method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of soil heavy metal analysis, specifically involving a method for source apportionment of soil heavy metals based on geostatistical analysis and APLS-MLR. Background Technology

[0002] Heavy metal pollution in soil not only reduces soil activity and agricultural yields, but also enters the human body through the food chain, posing a threat to human health. Identifying the sources of farmland soil pollution is fundamental to the prevention, remediation, and treatment of heavy metal pollution in farmland soil. Therefore, developing quantitative analysis methods for the sources of heavy metal pollution in farmland soil is crucial and fundamental to solving the problem.

[0003] Currently, there are two levels of source apportionment for heavy metals in soil: one is qualitative identification of the main pollution source types, i.e., source identification; the other is not only qualitative analysis of pollution sources but also quantitative calculation of the contribution rate of different pollution sources, i.e., source apportionment. Researchers usually refer to both collectively as source apportionment. Currently, heavy metal source apportionment models are mainly divided into two categories: one is diffusion models that take pollution sources as the research object. Diffusion models start from the pollution source and assess the contribution of different source types to receptors based on the pollution source emission inventory and pollutant transport processes. However, they are affected by complex meteorological conditions and changes in chemical processes, resulting in large errors in model predictions and often unsatisfactory prediction results. The second category is receptor models that take polluted areas as the research object. Commonly used receptor models include Absolute Principal Component Multiple Linear Regression (APCS-MLR), Positive Matrix Factorization (PMF), Chemical Mass Balance (CMB), isotope labeling, and the Unmix model. These models have achieved good results in quantitative source apportionment.

[0004] Source apportionment research initially focused on the sources of particulate matter in the atmosphere and has gradually developed into a relatively complete technical system for atmospheric pollution source apportionment: pollution source inventory – air quality diffusion model – receptor model. Unlike air pollution, the hidden, cumulative, and geographically specific characteristics of soil pollution significantly limit research on soil pollutant source apportionment. Soil heavy metal pollution is particularly complex and uncertain.

[0005] The APCS-MLR receptor model is rarely used for source apportionment of heavy metals in soil. This model combines principal component analysis (PCA) with multiple linear regression (MLR). PCA can qualitatively analyze the pollution source corresponding to each principal component, and quantitatively determine the average contribution of the source to heavy metals and its contribution at each sampling point. However, the APCS-MLR receptor model has limitations in eigenvalue decomposition during principal component analysis, and the results lack visualization, cannot accurately predict the fingerprint spectrum of pollution sources, and have insufficient accuracy in interpreting the model results. It is also difficult to analyze and determine specific pollution sources or pollution source types.

[0006] Partial least squares regression is a novel multivariate statistical analysis method that utilizes information from dependent and independent variables. It combines multiple functions such as multiple linear regression, principal component analysis, and canonical correlation analysis, while integrating modeling and predictive data analysis methods with non-model-based data interpretation methods. This allows for better identification of pollution sources and improves the accuracy of source apportionment.

[0007] Therefore, this study proposes a method for analyzing the sources of heavy metals in soil that combines geostatistical analysis with receptor modeling. Partial least squares regression is used to improve the APCS-MLR receptor model method. Summary of the Invention

[0008] The technical problem of this invention is that the absolute principal component-multiple linear regression method (APCS-MLR) has better predictive performance in quantitative source apportionment than the diffusion model method, but it is rarely used for soil heavy metal source apportionment. At the same time, the APCS-MLR receptor model has certain limitations in the decomposition of eigenvalues ​​when performing principal component analysis, and the results lack intuitive visualization, making it difficult to analyze and determine specific pollution sources or pollution source types.

[0009] The purpose of this invention is to improve the absolute principal component-multiple linear regression method using partial least squares regression, proposing the Absolutely Partial Least Squares-Multiple Linear Regression (APLS-MLR) method for source apportionment of heavy metals in soil; and to combine geostatistical analysis of the spatial distribution characteristics of heavy metals with the APLS-MLR method to improve the accuracy and intuitiveness of source apportionment results, so as to facilitate further analysis and determination of the nature of pollution sources.

[0010] The technical solution of this invention is a method for apportioning soil heavy metal sources based on geostatistical analysis and APLS-MLR, comprising the following steps:

[0011] Step 1: Sampling of soil in the study area, measuring the content of heavy metals in the soil samples, and preprocessing the raw data obtained from the measurements;

[0012] Step 2: Use the Kriging interpolation method to analyze and obtain the spatial distribution characteristics of heavy metals in the soil of the study area, and identify the sources of soil pollution;

[0013] Step 3: Analyze the heavy metal concentration data of the study area using partial least squares method;

[0014] Step 4: Establish an absolute partial least squares-multiple linear regression receptor model for the analysis of heavy metal pollution sources in soil;

[0015] Step 4-1: Calculate the absolute partial least squares score for each soil sample;

[0016] The partial least squares score is obtained by multiplying the principal factor coefficient matrix obtained by partial least squares analysis with the standardized heavy metal content matrix, using soil heavy metal concentration as the independent variable.

[0017] The absolute partial least squares score of each soil sample is obtained by subtracting the partial least squares score of the 0 concentration sample from the partial least squares score of each soil sample.

[0018] Step 4-2: Using the absolute partial least squares score as the independent variable and the heavy metal concentration as the dependent variable, perform multiple linear regression analysis to obtain the regression coefficients and regression constants.

[0019] The obtained regression coefficients are used to convert the absolute partial least squares score of the sample into the concentration contribution of the pollution source corresponding to the principal component to each sample.

[0020] Step 4-3: Calculate the contribution rate of each principal component to the pollution source;

[0021] Step 5: Combining the spatial distribution characteristics of soil heavy metals from Step 2 with the number of pollution sources and the contribution rate of each pollution source analyzed by the absolute partial least squares-multiple linear regression receptor model, the specific pollution sources are inferred and determined.

[0022] Preferably, in step 1, EXCEL software is used to preprocess the original data, removing obviously erroneous attribute values, and missing values ​​in the original data are filled using the average value.

[0023] Preferably, in step 2, a histogram of the data is created using ArcGIS software, a logarithmic transformation is performed on the data that does not conform to the normal distribution, and the Kriging interpolation method is applied to draw a spatial distribution characteristic map of heavy metal content in the soil to analyze potential pollution sources.

[0024] Furthermore, step 3 specifically includes the following sub-steps:

[0025] Step 3-1: Annotate the raw data;

[0026]

[0027]

[0028] (1)

[0029] (2)

[0030] in For a standardized matrix of independent variables, For a standardized dependent variable matrix, The element values ​​of the matrix representing the independent variables. This represents the element values ​​of the dependent variable matrix. n Indicates the number of samples. p Indicates the number of independent variables. q Indicates the number of dependent variables; Represents the true value of the independent variable in the sample. Indicates the first j The mean of a sample of independent variables. Indicates the first j The variance of multiple samples of each independent variable. Represents the true value of the dependent variable in the sample. Indicates the first j The mean of multiple samples of a dependent variable. Indicates the first j Variance of multiple samples for each dependent variable;

[0031] Step 3-2: First round of principal component extraction;

[0032] Step 3-2-1: Extract the first principal component of the independent variable and analyze the matrix. Perform feature decomposition;

[0033] (3)

[0034] in This represents the first principal component of the independent variable. This represents the unit eigenvector corresponding to the largest eigenvalue;

[0035] Step 3-2-2: Extract the first principal component of the dependent variable and analyze the matrix. Perform feature decomposition;

[0036] (4)

[0037] in This represents the first principal component of the dependent variable. This represents the unit eigenvector corresponding to the largest eigenvalue;

[0038] Step 3-2-3: Calculate the residual matrix

[0039] (5)

[0040] (6)

[0041] in , These represent the residual matrices of the independent variable matrix and the dependent variable matrix, respectively. This represents the vector of regression coefficients of the independent variables when the principal components are extracted for the first time; This represents the vector of regression coefficients of the dependent variable when the principal components are extracted for the first time.

[0042] In formula (5)

[0043]

[0044] In formula (6)

[0045]

[0046] Step 3-3: A new round of principal component extraction;

[0047] make = , = Using the principal component extraction method in step 3-2, a new round of principal component extraction is performed on the residual matrix;

[0048] (7)

[0049] (8)

[0050] (9)

[0051] (10)

[0052] In equation (7), the subscript h Indicates the first h Secondary principal component extraction. Indicates the independent variable's first... h Principal components, This represents the unit eigenvector corresponding to the largest eigenvalue of the residual matrix of the independent variables when the principal components are extracted for the h-th time. , They represent the first h、h +1 extraction of principal components: residual matrix of independent variables.

[0053] Indicates the dependent variable's first... h Principal components, This represents the unit eigenvector corresponding to the largest eigenvalue of the residual matrix of the dependent variable. , They represent the first h、h The dependent variable residual matrix at +1 principal component extraction; Indicates the first h The regression coefficient vector of the independent variables when extracting principal components; Indicates the first h The regression coefficient vector of the dependent variable when extracting principal components.

[0054] In equation (9)

[0055]

[0056] In formula (10)

[0057]

[0058] Steps 3-4: Complete principal component extraction and determine the number of extracted principal components based on cross-validation.

[0059] Step 3-4-1: The dependent variable... Ingredients Cross-validity Defined as:

[0060] (11)

[0061] in Dependent variable The sum of squared prediction errors for The sum of squared errors;

[0062] (12)

[0063] (13)

[0064] in n Indicates the number of samples. for At sample points The actual value on for At sample points Fitted values ​​on; For the first Predicted values ​​for each sample point;

[0065] Step 3-4-2: Analyze the components of the dependent variable Y Cross-validity Defined as:

[0066] (14)

[0067] (15)

[0068] (16)

[0069] In the formula q Indicates the quantity of the dependent variable. This represents the sum of squared prediction errors for Y; This represents the sum of squared errors in Y;

[0070] Step 3-4-3: Determine the optimal number of principal components; based on the components Cross-validity The maximum value corresponding to h The value is used to determine the optimal number of principal components. r .

[0071] In step 4, the formula for calculating the concentration contribution of the pollution source corresponding to the principal component to each sample is as follows:

[0072]

[0073] In the formula Indicates the number of principal components. For the first i The concentration of certain heavy metals, For the constant term of multiple linear regression, These are the regression coefficients for multiple linear regression. Main component p The absolute partial least squares score;

[0074] Main component p for The content contribution of all samples The average value is the principal component. p The corresponding average absolute contribution of pollution sources;

[0075] principal component p The corresponding pollution source contribution rate is the ratio of its average absolute contribution to the total contribution of all sources.

[0076] Furthermore, in step 5, based on the spatial distribution characteristic map of heavy metals in the soil obtained in step 2, the principal component matrix obtained in step 3 using the partial least squares method, and the contribution rate of pollution source factors obtained in step 4, combined with the field investigation and verification of the study area, the specific pollution sources in the study area are inferred and determined.

[0077] Compared with the prior art, the beneficial effects of the present invention include:

[0078] 1) The method of this invention combines receptor model with geostatistical analysis, which can not only calculate and determine the number of pollution sources and the contribution rate of each pollution source, but also accurately identify specific pollution sources, which is conducive to carrying out the prevention and control of heavy metal pollution in farmland soil.

[0079] 2) This invention combines absolute partial least squares (APLS) with multiple linear regression to propose an APLS-MLR method for receptor models. This solves the problem of limitations in eigenvalue decomposition during principal component analysis in the APLS-MLR method. The APLS-MLR of this invention uses the absolute partial least squares score as the independent variable and the heavy metal concentration as the dependent variable to perform multiple linear regression analysis. The obtained regression coefficients are used to transform the absolute partial least squares score of the sample into the concentration contribution of the pollution source corresponding to the principal component to each sample, which improves the model regression effect and the accuracy of the calculated pollution source contribution rate is better. Attached Figure Description

[0080] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0081] Figure 1 This is a schematic flowchart of the soil heavy metal source analysis method according to an embodiment of the present invention.

[0082] Figure 2 This is a spatial distribution map of soil heavy metal As content in the study area calculated according to an embodiment of the present invention.

[0083] Figure 3 This is a spatial distribution map of the heavy metal Hg content in the soil of the study area calculated according to an embodiment of the present invention.

[0084] Figure 4 This is a spatial distribution map of the heavy metal Cr content in the soil of the study area calculated according to an embodiment of the present invention.

[0085] Figure 5 This is a spatial distribution map of the heavy metal Hg content in the soil of the study area calculated according to an embodiment of the present invention.

[0086] Figure 6 This is a spatial distribution map of the heavy metal Pb content in the soil of the study area calculated according to an embodiment of the present invention.

[0087] Figure 7This is a schematic diagram illustrating the contribution rate of different pollutants to heavy metal accumulation calculated in an embodiment of the present invention. Detailed Implementation

[0088] like Figure 1 As shown, the method for source apportionment of heavy metals in soil based on geostatistical analysis and APLS-MLR includes the following steps:

[0089] Step 1: Preprocess the raw data. After setting up sampling points in the study area, measure the content of heavy metals in the soil samples, and then preprocess the data using EXCEL software to remove attribute values ​​with obvious errors. Missing values ​​are replaced by the average value.

[0090] In the sample data, a small number of missing values ​​were found for Hg, Cd, and Pb, which were eventually replaced by their total metal average. A number of outliers were found in the sample. Considering all factors, since the sampling points were reasonable, the sample processing steps were rigorous, and the detection instruments were accurate, the measured heavy metal content was not erroneous. Therefore, the few outliers were retained.

[0091] Step 2: Analyze the spatial distribution characteristics of heavy metals in the soil of the study area using the Kriging interpolation method to identify soil pollution sources;

[0092] By analyzing the spatial variations in heavy metal concentrations across the study area, potential sources of heavy metal pollution in the soil can be identified.

[0093] Spatial distribution maps of all total metals were generated using ordinary kriging interpolation. The mean error (ME) of the ordinary kriging interpolation was close to 0, proving that the predicted values ​​were accurate. The root mean square standard error (RMSSE) values ​​were between 0.968 and 1.032, indicating that the standard error was accurate. The spatial distribution results of heavy metals obtained in the example are as follows: Figure 2-6 As shown.

[0094] Step 3: Partial least squares analysis was performed on the heavy metal concentration data of the study area, and the number of principal components was determined by cross-validation. The coefficient matrices of the obtained principal components and the five heavy metals are shown in Table 1.

[0095]

[0096] Step 4: Establish an APLS-MLR receptor model for soil heavy metal pollution source analysis; based on the obtained regression coefficients and regression constants, calculate the final contribution rate of each heavy metal. Experimental results are shown in Table 2.

[0097] Table 2. Data on the contribution of pollution sources to heavy metals

[0098]

[0099] Step 5: Based on the pollution sources identified by the spatial distribution characteristics of heavy metals through geostatistical analysis and the pollution sources and their contribution rates analyzed by the APLS-MLR receptor model, combined with field investigation and verification of the study area, determine the specific pollution sources and their contribution rates.

[0100] In principal component 1, As and Pb account for a large proportion. According to the characteristic map of total metal content interpolated by ordinary kriging space, the spatial distribution of As and Pb is quite similar. According to actual investigation, there are a large number of chemical enterprises in the southernmost and northernmost parts of the study area, and the area is traversed by a tributary of the Yangtze River from south to north. Therefore, it can be analyzed that the pollution source of source 1 is chemical pollution. From the spatial distribution map of Pb content, it can be seen that the high value area is located in the southern part of the study area, which is a transportation hub. Therefore, it is inferred that source 1 is a mixed source of traffic pollution and chemical irrigation water discharge.

[0101] In principal component 2, the heavy metal with a relatively large loading is Cr, which is composed of... Figure 4 It is known that there is almost no pollution of heavy metal Cr in the study area. Numerous studies have shown that the parent material of the soil is the main cause of Cr pollution. Therefore, it is inferred that source 2 is a "natural source".

[0102] In principal component 3, the heavy metals with relatively large loadings are Hg and Cd. Spatially, they are... Figure 2 It can be seen that the high-value Hg areas are concentrated, mainly in the eastern part of the study area, with a clear boundary from the low-value areas. The Hg element variation coefficient is 75%, which belongs to medium-high variation, indicating that the polluted area is significantly affected by human factors. The survey found that rivers and irrigation canals pass through the high-value areas, and there are enterprises discharging Hg wastewater around the rivers. Therefore, it is speculated that the accumulation of Hg in the soil of the high-value areas may be caused by long-term river irrigation with polluted water. Moreover, Hg and Cd elements are often used in the manufacture and use of pesticides. Therefore, it can be inferred that source 3 is an "agricultural source". The final contribution rate of heavy metal pollution sources in the paddy field soil of the study area is as follows: Figure 7 As shown.

[0103] This invention analyzes practical cases, and based on the spatial distribution characteristic map and principal component matrix of total metal and heavy metal content in soil obtained from the analysis, combined with field investigation of the study area, the specific pollution sources are qualitatively analyzed. Finally, based on the APLS-MLR receptor model, the source analysis of heavy metals in farmland soil is realized, and the contribution rate of each pollution source is obtained.

Claims

1. A method for source apportionment of heavy metals in soil based on geostatistical analysis and APLS-MLR, characterized in that, Includes the following steps: Step 1: Sampling of soil in the study area, measuring the content of heavy metals in the soil samples, and preprocessing the raw data obtained from the measurements; Step 2: Use the Kriging interpolation method to analyze and obtain the spatial distribution characteristics of heavy metals in the soil of the study area, and identify the sources of soil pollution; Step 3: Analyze the heavy metal concentration data of the study area using partial least squares method; Step 4: Establish an absolute partial least squares-multiple linear regression receptor model for the analysis of heavy metal pollution sources in soil; Step 4-1: Calculate the absolute partial least squares score for each soil sample; The partial least squares score is obtained by multiplying the principal factor coefficient matrix obtained by partial least squares analysis with the standardized heavy metal content matrix, using soil heavy metal concentration as the independent variable. The partial least squares score of each soil sample is subtracted from the partial least squares score of the zero-concentration sample to obtain the absolute partial least squares score of each sample. Step 4-2: Using the absolute partial least squares score as the independent variable and the heavy metal concentration as the dependent variable, perform multiple linear regression analysis to obtain the regression coefficients and regression constants. Using the obtained regression coefficients, the absolute partial least squares score of the sample is converted into the concentration contribution of the pollution source corresponding to the principal component to each sample. Step 4-3: Calculate the contribution rate of each principal component to the pollution source; Step 5: Combining the spatial distribution characteristics of soil heavy metals from Step 2 with the number of pollution sources and the contribution rate of each pollution source analyzed by the absolute partial least squares-multiple linear regression receptor model, the specific pollution sources are inferred and determined.

2. The method for analyzing the sources of heavy metals in soil according to claim 1, characterized in that, In step 1, the original data is preprocessed using EXCEL software to remove obviously erroneous attribute values, and missing values ​​in the original data are filled using the average value.

3. The method for analyzing the sources of heavy metals in soil according to claim 2, characterized in that, In step 2, a histogram of the data is created using ArcGIS software. Logarithmic transformation is performed on the data that does not conform to the normal distribution. Kriging interpolation method is applied to draw a spatial distribution characteristic map of heavy metal content in the soil and potential pollution sources are analyzed.

4. The method for analyzing the sources of heavy metals in soil according to claim 3, characterized in that, Step 3 specifically includes the following sub-steps: Step 3-1: Standardize the raw data; ; ; ;(1) ;(2) in For a standardized matrix of independent variables, For a standardized dependent variable matrix, The element values ​​of the matrix representing the independent variables. This represents the element values ​​of the dependent variable matrix. n Indicates the number of samples. p Indicates the number of independent variables. q Indicates the number of dependent variables; Represents the true value of the independent variable in the sample. Indicates the first j The mean of a sample of independent variables. Indicates the first j The variance of multiple samples of each independent variable. This represents the true value of the dependent variable in the sample. Indicates the first j The mean of multiple samples of a dependent variable. Indicates the first j Variance of multiple samples for each dependent variable; Step 3-2: First round of principal component extraction; Step 3-2-1: Extract the first principal component of the independent variable and analyze the matrix. Perform feature decomposition; ;(3) in This represents the first principal component of the independent variable. This represents the unit eigenvector corresponding to the largest eigenvalue; Step 3-2-2: Extract the first principal component of the dependent variable and analyze the matrix. Perform feature decomposition; ;(4) in This represents the first principal component of the dependent variable. This represents the unit eigenvector corresponding to the largest eigenvalue; Step 3-2-3: Calculate the residual matrix ;(5) ;(6) in , These represent the residual matrices of the independent variable matrix and the dependent variable matrix, respectively. This represents the vector of regression coefficients of the independent variables when the principal components are extracted for the first time; This represents the vector of regression coefficients of the dependent variable when the principal components are extracted for the first time; In formula (5) ; In formula (6) ; Step 3-3: A new round of principal component extraction; make = , = Using the principal component extraction method in step 3-2, a new round of principal component extraction is performed on the residual matrix; ;(7) ; (8) ; (9) ; (10) In equation (7), the subscript h Indicates the first h Secondary principal component extraction. Indicates the independent variable's first... h Principal components, Indicates the first h The unit eigenvector corresponding to the largest eigenvalue of the residual matrix of the independent variables when extracting principal components for the second time. , They represent the first h、h +1 principal component extraction residual matrix; Indicates the dependent variable's first... h Principal components, This represents the unit eigenvector corresponding to the largest eigenvalue of the residual matrix of the dependent variable. , They represent the first h、h The dependent variable residual matrix at +1 principal component extraction; Indicates the first h The regression coefficient vector of the independent variables when extracting principal components; Indicates the first h The vector of regression coefficients of the dependent variable when extracting principal components; In equation (9) ; In formula (10) ; Steps 3-4: Complete principal component extraction and determine the number of extracted principal components based on cross-validation. Step 3-4-1: The dependent variable... Ingredients Cross-validity Defined as: ;(11) in Dependent variable The sum of squared prediction errors for The sum of squared errors; ;(12) ; (13) in n Indicates the number of samples. for At sample points The actual value on for At sample points Fitted values ​​on; For the first Predicted values ​​for each sample point; Step 3-4-2: Analyze the components of the dependent variable Y Cross-validity Defined as: ;(14) ;(15) ; (16) In the formula q Indicates the quantity of the dependent variable. This represents the sum of squared prediction errors for the dependent variable Y; This represents the sum of squared errors of the dependent variable Y; Step 3-4-3: Determine the optimal number of principal components; based on the components Cross-validity The maximum value corresponding to h The value is used to determine the optimal number of principal components. r .

5. The method for analyzing the sources of heavy metals in soil according to claim 4, characterized in that, In step 5, based on the spatial distribution characteristic map of heavy metals in the soil obtained in step 2, the principal component matrix obtained in step 3 using the partial least squares method, and the contribution rate of pollution source factors obtained in step 4, combined with the field investigation and verification of the study area, the specific pollution sources in the study area are inferred and determined.

Citation Information

Patent Citations

  • Cr element soil moisture content correction method based on an XRF detection technology

    CN113567652A

  • Soil heavy metal spectral feature extraction and optimization method based on wavelength frequency selection

    CN114354666A