A method for soil heavy metal source analysis based on linear discriminant analysis
By using linear discriminant analysis and multivariate linear regression receptor models in the analysis of soil heavy metal pollution sources, the soil heavy metal pollution sources and their contribution rates are identified and analyzed, and the problems of low efficiency and poor results in the existing technology are solved, and more accurate and efficient pollution source analysis is achieved.
Patent Information
- Application Number
- CN202210906863.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-29
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-07-29
AI Technical Summary
Existing methods for analyzing soil heavy metal pollution sources, such as the APCS-MLR receptor model, have problems such as the limitations of eigenvalue decomposition and the inability to utilize category prior knowledge, resulting in low efficiency and poor results.
A multivariate linear regression (MLR) receptor model based on linear discriminant analysis (LDA) was used to identify the source of heavy metal pollution in soil and its contribution rate through LDA dimensionality reduction and combined with geostatistical methods.
The rapid and accurate analysis of the various sources of heavy metal pollution in farmland soil and their contribution rate has been achieved, and the limitations of unsupervised learning cannot use category prior knowledge, reducing uncertainty.
Smart Images

Figure CN115329272B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of soil heavy metal analysis, and relates to a soil heavy metal source analysis method based on linear discriminant analysis. Background Art
[0002] With the rapid development of industrialization in recent years, heavy metal pollution in soil has become increasingly serious. Heavy metals that enter the soil, due to their concealment, difficulty in degradation, and enrichment, will not only affect the normal growth of crops, but also enter the human body through the food chain, causing harm to human health. The prevention and control of heavy metal pollution in farmland soil has become a major strategic need in my country. Identification of the source of farmland soil pollution is the basis for the prevention, control and restoration of heavy metal pollution in farmland soil. Therefore, the development of quantitative analysis methods for the source of heavy metal pollution in farmland soil has become the key and basis for solving the problem of heavy metal pollution in soil.
[0003] Source analysis of heavy metal pollution in soil generally refers to qualitative and quantitative analysis of pollution sources. Currently, the most commonly used source analysis models include isotope ratio analysis, positive definite matrix factor analysis model, UNMIX model, and absolute factor analysis / multiple linear regression analysis (APCS-MLR) model.
[0004] In the existing technology, the paper with the document number 2022,38(03):212-219 discloses a method for heavy metal source analysis in paddy field soil based on the APCS-MLR receptor model. The innovation of the paper is to identify the source through principal component analysis, and linearly regress the main pollution factors obtained with the concentration of soil pollution elements. The regression coefficient is used to calculate the contribution rate of the pollution factor to the pollution element. However, it was found in the study that the APCS-MLR receptor model has certain limitations in the decomposition of eigenvalues when performing principal component analysis. Moreover, if the user has a certain prior knowledge of the observed object and has mastered some characteristics of the data, but cannot intervene in the processing process through parameterization and other methods, the expected effect may not be achieved and the efficiency is not high.
[0005] The present invention can use the prior knowledge and experience of categories in the dimensionality reduction process through LDA, overcoming the limitation that unsupervised learning such as PCA cannot use prior knowledge of categories. The ALDS-MLR receptor model first uses linear discriminant analysis for source identification, then linearly regresses the main pollution factors obtained with the concentration of soil pollution elements, and then combines geostatistical methods to explore the sources of heavy metals in the soil and their contribution rate, in order to provide a theoretical basis for the scientific prevention and control and remediation of local soil heavy metal pollution. Summary of the invention
[0006] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a method for quantitative analysis of the sources of heavy metal pollution in farmland soil, which can solve the limitation problem that unsupervised learning such as PCA cannot use category prior knowledge and reduce uncertainty.
[0007] In order to solve the above technical problems, the technical solution adopted by the present invention is:
[0008] A method for soil heavy metal source analysis based on linear discriminant analysis includes the following steps:
[0009] Step 1: Set up monitoring points in the study area, collect soil samples, process soil samples, and measure the content of heavy metals in soil samples;
[0010] Step 2: Conduct descriptive statistical analysis on soil heavy metals in the studied area;
[0011] Step 3: Use geostatistical methods to analyze the spatial distribution characteristics of heavy metals in the soil of the study area and identify the sources of soil pollution;
[0012] Step 4: Establish the ALDS-MLR receptor model for soil heavy metal pollution source analysis, and use the ALDS-MLR receptor model for soil heavy metal pollution source analysis to analyze the soil heavy metal pollution sources and their contribution rates;
[0013] Step 5: Based on the pollution sources identified by the spatial distribution characteristics of heavy metals through geostatistical analysis and the pollution sources and their contribution rates analyzed by the ALDS-MLR receptor model, clear pollution sources and their contribution rates are obtained;
[0014] In step 2, the heavy metal content is checked by a histogram to see if it conforms to a normal distribution;
[0015] In step 3, the heavy metal content data that does not conform to the normal distribution are logarithmically transformed, and the ordinary Kriging interpolation method is used to draw the spatial characteristic distribution map of soil heavy metals to analyze the pollution source;
[0016] In step 4, the following sub-steps were used to establish the ALDS-MLR receptor model for soil heavy metal pollution source apportionment:
[0017] Step 4-1: Perform correlation analysis on the collected heavy metal concentration data to obtain the correlation coefficients between the heavy metals;
[0018] Step 4-2: Use linear discriminant analysis (LDA) to reduce the dimension, maximize the distance between classes, minimize the distance within classes, optimize the objective function, and obtain the optimal projection matrix ω;
[0019] Step 4-3: Obtain absolute linear discriminant scores and establish a multivariate linear regression model;
[0020] In step 4-2, the following sub-steps are included:
[0021] Step 4-2-1: Calculate the intra-class scatter matrix
[0022]
[0023]
[0024] Among them C i is the i-th sample, μ i is the mean of the i-th sample, and x is the true value of the sample.
[0025] Step 4-2-2: Calculate the hashing of each class mean point relative to the sample center to obtain the inter-scatter matrix S b
[0026]
[0027] Where C is the number of sample categories, μ i is the mean of the ith sample, and μ is the mean of all samples.
[0028]
[0029] Step 4-2-3: Optimize the objective function to maximize the distance between classes and minimize the variance within classes, and maximize the objective function to achieve the maximum;
[0030]
[0031] Where ω is the projection matrix, C i is the i-th sample, μ i is the mean of the i-th sample, and x is the true value of the sample. With S b Substituting the above objective function into the equation, it can be simplified to:
[0032]
[0033] Where S b , They are the inter-class dispersion and intra-class dispersion, respectively.
[0034] Step 4-2-4: Calculate the matrix Perform eigendecomposition on the matrix;
[0035] Step 4-2-5: Calculate the eigenvalues obtained in step 4-2-4. The largest d eigenvalues and the corresponding eigenvalue vectors (ω1, ω2, ω3, …ω d)Get the projection matrix ω;
[0036] Step 4-2-6: For each sample feature x in the sample set i , converted into a new sample z i =ω T x i ;
[0037] Step 4-2-7 obtains the output sample set. D`={(z1,y1),(z2,y2),Λ,(z i ,y i ), Λ, (z n ,y n )}, where z i is the i-th new sample after transformation, y i is the category to which the i-th sample belongs, and n represents the number of samples;
[0038] In step 4-3, the following sub-steps are included:
[0039] Step 4-3-1: Calculate the absolute linear discriminant score;
[0040] Step 4-3-2: Use the absolute linear discriminant score as the independent variable and the metal concentration as the dependent variable to perform regression analysis to obtain the regression coefficient and regression constant;
[0041] Step 4-3-3: Take the absolute linear discriminant score as the independent variable and the metal concentration as the dependent variable for regression analysis to obtain the regression coefficient and regression constant term; subtract the main factor score of the sample with zero concentration from the main factor score obtained by linear discriminant analysis to obtain the ALDS of each sample; take ALDS as the independent variable and the heavy metal element content as the dependent variable for multiple linear regression. The obtained regression coefficient can convert ALDS into the concentration contribution of the pollution source corresponding to the main factor to each sample. The formula is:
[0042] Where: Z i0 is the sample with zero concentration of heavy metal element i, mg·kg -1 ; is the average value of the content of heavy metal element i, mg·kg -1 ; δ i is the standard deviation of the content of heavy metal element i, mg·kg -1 . b io is the constant term of multiple linear regression, b pi is the regression coefficient of multiple linear regression, ALDSp is the absolute linear discriminant score of factor p, bpi×ALDSp is the absolute linear discriminant score of factor p for c iThe average value of bpi×ALDSp of all samples is the average absolute contribution of the pollution source corresponding to factor p. The contribution rate of the pollution source corresponding to factor p is the ratio of its average absolute contribution to the contribution of all sources.
[0043] In step 5, the pollution sources identified by the spatial distribution characteristics of heavy metals by geostatistical analysis and the pollution sources and their contribution rates analyzed by the ALDS-MLR receptor model are used to obtain clear pollution sources and their contribution rates.
[0044] Compared with the prior art, the present invention has the following technical effects:
[0045] The present invention proposes an algorithm based on linear discriminant analysis and multiple linear regression, which is an improved receptor model for soil source analysis. It can quickly and accurately analyze the pollution sources of heavy metals in farmland soil and their contribution rates, in order to provide a theoretical basis for the scientific prevention, control and restoration of local soil heavy metal pollution.
[0046] LDA dimension reduction has not been used in previous studies on the source analysis of heavy metals in soil. LDA can use prior knowledge and experience of categories in the process of dimension reduction, overcoming the limitation that unsupervised learning such as PCA cannot use prior knowledge of categories. The ALDS-MLR receptor model first uses linear discriminant analysis for source identification, and then linearly regresses the main pollution factors obtained with the concentration of soil pollution elements. The regression coefficient is used to calculate the contribution rate of pollution factors to pollution elements.
[0047] The ALDS-MLR model is a receptor model that combines two statistical methods, linear discriminant analysis and multiple linear regression. The analysis results are more reliable and accurate, overcoming the limitation of unsupervised learning such as PCA that cannot use category prior knowledge. The content of heavy metals in the soil was determined, the pollution level of heavy metals was analyzed, and a mixed method, including correlation analysis, linear discriminant analysis, absolute linear discriminant analysis / multiple linear regression analysis (ALDS-MLR), was used to combine geostatistical methods to explore the sources of heavy metals in the soil and their contribution rate, thereby providing a theoretical basis for scientific prevention and control and restoration of heavy metal pollution in farmland soil. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 A flow chart of a method for analyzing the sources of heavy metals in soil provided by an embodiment of the present invention;
[0049] Figure 2 This is the spatial distribution map of soil heavy metal As content in Zhijiang and Dangyang cities according to the embodiment of the present invention;
[0050] Figure 3 This is a spatial distribution map of soil heavy metal Hg content in Zhijiang and Dangyang cities according to an embodiment of the present invention;
[0051] Figure 4 This is the spatial distribution map of soil heavy metal Cr content in Zhijiang and Dangyang cities according to the embodiment of the present invention;
[0052] Figure 5 This is a spatial distribution map of soil heavy metal Hg content in Zhijiang and Dangyang cities according to an embodiment of the present invention;
[0053] Figure 6 This is a spatial distribution diagram of soil heavy metal Pb content in Zhijiang and Dangyang cities according to an embodiment of the present invention;
[0054] Figure 7 A schematic diagram of the contribution of different factors to heavy metal accumulation provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0055] A method for soil heavy metal source analysis based on linear discriminant analysis includes the following steps:
[0056] Step 1: Set up monitoring points in the study area, collect soil samples, process soil samples, and measure the content of heavy metals in soil samples;
[0057] Step 2: Conduct descriptive statistical analysis on soil heavy metals in the studied area;
[0058] Step 3: Use geostatistical methods to analyze the spatial distribution characteristics of heavy metals in the soil of the study area and identify the sources of soil pollution;
[0059] Step 4: Use the ALDS-MLR receptor model to analyze the sources of soil heavy metal pollution and their contribution rates;
[0060] Step 5: Based on the pollution sources identified by the spatial distribution characteristics of heavy metals through geostatistical analysis and the pollution sources and their contribution rates analyzed by the ALDS-MLR receptor model, clear pollution sources and their contribution rates are obtained;
[0061] In step 1, collect soil samples from the cultivated layer, and collect 0-20cm samples for planting general crops. In order to ensure the representativeness of the samples, adopt the scheme of collecting mixed samples, set up 3-7 sampling areas for each soil unit, and a single sampling area can be a naturally divided field. It can also be composed of multiple fields, and its range is preferably about 200m×200m;
[0062] In step 2, the descriptive statistical analysis of soil heavy metals includes: maximum value, minimum value, mean value, standard deviation, coefficient of variation CV, and the histogram is used to test whether the heavy metal content conforms to the normal distribution;
[0063] In step 3, the heavy metal content data that does not conform to the normal distribution in step 2 is log-log transformed, and the ordinary Kriging interpolation method is used to draw the spatial characteristic distribution map of soil heavy metals to analyze the pollution source;
[0064] In step 4, when establishing the ALDS-MLR receptor model, the following sub-steps are included:
[0065] Step 4.1: Perform correlation analysis on the collected heavy metal concentration data to obtain the correlation coefficients between the heavy metals. The closer the correlation is to 1, the stronger the correlation between the variables is, and the closer the correlation is to 0, the weaker the correlation between the variables is. The significance level P value is generally less than 0.05 to be statistically significant. Generally, P<0.05 is significant and P<0.01 is very significant.
[0066] Step 4.2: Use linear discriminant analysis (LDA) to reduce the dimension, maximize the distance between classes, minimize the distance within classes, optimize the objective function, and obtain the optimal projection matrix ω;
[0067] In step 4.2, the following sub-steps are also included:
[0068] Step 4.2.1: Calculate the intra-class scatter matrix
[0069]
[0070]
[0071] Among them C i is the i-th sample, μ i is the mean of the i-th sample, and x is the true value of the sample.
[0072] Step 4.2.2: Calculate the hashing of each class mean point relative to the sample center and obtain the inter-scatter matrix S b
[0073]
[0074] Where C is the number of sample categories, μ i is the mean of the ith sample, and μ is the mean of all samples.
[0075]
[0076] Step 4.2.3: Optimize the objective function to maximize the inter-class distance and minimize the intra-class variance, and maximize the objective function to achieve its maximum.
[0077]
[0078] Where ω is the projection matrix, Ci is the i-th sample, μ i is the mean of the ith sample, and x is the true value of the sample. With S b Substituting the above objective function into the equation, it can be simplified to:
[0079]
[0080] Where S b , They are the inter-class dispersion and intra-class dispersion, respectively.
[0081] Step 4.2.4: Calculate the matrix Perform eigendecomposition on the matrix;
[0082] Step 4.2.5: Calculate the eigenvalues obtained in step 4.2.4. The largest d eigenvalues and the corresponding eigenvalue vectors (v1, ω2, ω3, …ω d ) to get the projection matrix ω.
[0083] Step 4.2.6: For each sample feature x in the sample set i , converted into a new sample z i =ω T x i .
[0084] Step 4.2.7 obtains the output sample set. D`={(z1,y1),(z2,y2),Λ,(z i ,y i ), Λ, (z n ,y n )}, where z i is the i-th new sample after transformation, y i is the category to which the i-th sample belongs, and n represents the number of samples;
[0085] Step 4.3: Obtain the absolute linear discriminant score and establish a multivariate linear regression model;
[0086] In step 4.3, the following sub-steps are included:
[0087] Step 4.3.1: Calculate the absolute linear discriminant score;
[0088] Step 4.3.2: Use the absolute linear discriminant score as the independent variable and the metal concentration as the dependent variable to perform regression analysis to obtain the regression coefficient and regression constant;
[0089] Step 4.3.3: Take the absolute linear discriminant score as the independent variable and the metal concentration as the dependent variable for regression analysis to obtain the regression coefficient and regression constant term; subtract the main factor score of the sample with zero concentration from the main factor score obtained by linear discriminant analysis to obtain the ALDS of each sample; take ALDS as the independent variable and the heavy metal element content as the dependent variable for multiple linear regression. The obtained regression coefficient can convert ALDS into the concentration contribution of the pollution source corresponding to the main factor to each sample. The formula is:
[0090] Where: Z i0 is the sample with zero concentration of heavy metal element i, mg·kg -1 ; is the average value of the content of heavy metal element i, mg·kg -1 ; δ i is the standard deviation of the content of heavy metal element i, mg·kg -1 . b io is the constant term of multiple linear regression, b pi is the regression coefficient of multiple linear regression, ALDSp is the absolute linear discriminant score of factor p, bpi×ALDSp is the absolute linear discriminant score of factor p for c i The average value of bpi×ALDSp of all samples is the average absolute contribution of the pollution source corresponding to factor p. The contribution rate of the pollution source corresponding to factor p is the ratio of its average absolute contribution to the contribution of all sources.
[0091] In step 5, the pollution sources identified by the spatial distribution characteristics of heavy metals by geostatistical analysis and the pollution sources and their contribution rates analyzed by the ALDS-MLR receptor model are used to obtain clear pollution sources and their contribution rates.
[0092] Among them, ALDS is absolute linear discriminant analysis; MLR is multivariate linear regression.
[0093] The technical solution of the present invention is described below by an embodiment:
[0094] like Figure 1 As shown, a soil heavy metal source analysis method based on linear discriminant analysis includes the following steps:
[0095] Step 1: Collect surface soil samples, collect cultivated layer soil samples, and collect 0-20cm samples for planting crops. Take a mixed sample collection plan to ensure the representativeness of the sample.
[0096] Step 2: Conduct descriptive statistical analysis on soil heavy metals in the studied area;
[0097] The descriptive statistics of heavy metal concentrations in soil samples are shown in the following table: It can be seen that the content of different heavy metals in the soil varies greatly, and the average values are AS (11.98 mg kg -1 )、Hg(0.06mg·kg -1 ), Cr(74.86mg·kg -1 )、Cd(0.17mg·kg -1 )、Pb(28.81mg·kg -1 ). Among them, the average values of Cd, As, and Pb are 1.54, 1.14, and 1.12 times the background values of Hubei Province, and the average values of Cr and Hg are close to the background values. The average values of heavy metal content did not exceed the agricultural land soil pollution risk screening values of the corresponding heavy metals specified in the "Soil Environmental Quality Agricultural Land Soil Pollution Risk Control Standard (Trial)" (GB 15618-2018). The coefficient of variation of heavy metals in farmland soil in Yichang City is Hg>Cd>As>Pb>Cr, among which Hg has the largest coefficient of variation, followed by Cd, and the spatial heterogeneity is strong, which may be caused by human activities.
[0098] Table 1 Descriptive statistics of heavy metal concentrations in soil samples (n=721)
[0099]
[0100] Step 3: Use geostatistical methods to analyze the spatial distribution characteristics of heavy metals in the soil of the study area and identify the sources of soil pollution;
[0101] The spatial variation of heavy metal concentration in the study area was analyzed to identify the potential sources of soil heavy metal pollution. The spatial distribution map of each total metal was generated by ordinary kriging interpolation. The average error (ME) of ordinary kriging interpolation was close to 0, which proved that the predicted value was accurate. The root mean square standard error (RMSSE) value was between 0.968 and 1.032, indicating that the standard error was accurate. The spatial distribution results of heavy metals are shown in Figure 2. Figure 2 — Figure 6 As shown. The spatial distribution characteristics of Pb and As are similar, showing obvious point source pollution. The high Pb values are mainly distributed in the south, and the most obvious is in the southeast. The spatial distribution of Hg is uniform and similar to the background value of soil Hg in Hubei Province. The Hg pollution in this area is good, and there are some high-value areas mainly in the east. The spatial distribution characteristics of Cr and Cd are similar. The heavy metal Cd pollution in the study area is the most serious, showing non-point source pollution. The high-value area is from north to south, and the content is highest in the south. The high-value areas of Cr are mainly distributed in the southeast, and show a trend from northeast to southwest.
[0102] Step 4: Use the ALDS-MLR receptor model for soil heavy metal pollution source analysis to analyze the soil heavy metal pollution sources and their contribution rates.
[0103] Step 4.1: By analyzing the correlation coefficients between each heavy metal, the larger the correlation coefficient, the stronger the relationship between the heavy metals, and the more likely they are to have similar pollution sources; the results of the correlation analysis are shown in the following table:
[0104] Table 2 Correlation analysis of heavy metals in soil of Yichang City
[0105]
[0106] Note: ** indicates that the correlation is significant at the 0.01 level (two-tailed); * indicates that the correlation is significant at the 0.05 level.
[0107] Step 4.2: Linear discriminant analysis LDA reduces dimension, maximizes inter-class distance, minimizes intra-class distance, optimizes objective function, and obtains the best projection matrix ω. The projection coefficient matrix is shown in the following table:
[0108] Table 3 Soil heavy metal projection coefficient matrix
[0109]
[0110] Step 4.3: Obtain the absolute linear discriminant score and suggest a multiple linear regression model; according to the obtained regression coefficient and regression constant term, calculate the final contribution rate of each heavy metal. The experimental results are shown in the following table:
[0111] Table 4 Contribution rate of each heavy metal pollution source
[0112]
[0113] Get the mapping of each heavy metal pollution source, such as Figure 7 As shown:
[0114] Step 5: Based on the pollution sources identified by the spatial distribution characteristics of heavy metals through geostatistical analysis and the pollution sources and their contribution rates analyzed by the ALDS-MLR receptor model, clear pollution sources and their contribution rates are obtained;
[0115] Source 1 has a greater contribution rate to AS, Pb and Cr. From the correlation, we can see that there is a significant correlation between AS, Pb and Cr. Figure 2 , Figure 4 and Figure 6 It can be seen that the spatial distribution of the three heavy metals has both similarities and differences. Since Pb is a car exhaust emission and the southern part of the study area is a transportation hub, the spatial distribution map of Pb ( Figure 6) distribution also shows an increasing trend in the south, so it is inferred that source 1 is a traffic pollution source. The change trends of Cr and As are highly consistent, and the region is traversed by the Yangtze River tributaries, and there are a large number of chemical activities in the high-value area, and the problem of "chemicals surrounding the river" is serious. Therefore, source 1 is a mixed source of traffic sources and river irrigation water.
[0116] Source 2 has a large load of heavy metals including Cd, Figure 5 It can be seen that the heavy metal Cd pollution in the study area is the most serious, showing non-point source pollution. The high-value area is from north to south, and the content is highest in the south. According to the investigation, there are a large number of chemical plants near the southern tributaries, and Cd is widely used in various chemical industries, so it is inferred that source 2 is an "industrial source".
[0117] Source 3 has a large load of heavy metals including Hg. From the perspective of spatial distribution, Figure 3 It can be seen that the Hg high-value area is concentrated, mainly distributed in the eastern part of the study area, with a clear boundary with the low-value area. The coefficient of variation of the Hg element is 47%, which is a medium-to-high variation, indicating that the polluted area is greatly affected by human factors. The survey found that there are rivers and irrigation canals passing through the high-value area, and there are enterprises discharging Hg wastewater around the river. Therefore, it is speculated that the accumulation of Hg in the soil of the high-value area may be caused by long-term river pollution irrigation. Studies have shown that Hg and As are important components of pesticides. Repeated application of pesticides containing Hg or inorganic As was widely used in agriculture before it was banned, but due to the difficulty of heavy metal degradation, it is still accumulated in the soil. Therefore, source 3 is an "agricultural source".
[0118] Through correlation analysis, linear discriminant analysis and geostatistical analysis, the pollution sources of these five heavy metals in Yichang farmland soil can be roughly divided into three main sources: mixed sources of traffic sources and river irrigation water, industrial sources and agricultural sources. According to the quantitative source analysis of the ALDS-MLR receptor model, mixed sources have a large contribution rate to As, Cr, and Pb, which are 71.08%, 97.44%, and 73.07%, respectively. Industrial sources have a large contribution rate to As, Cd, and Pb, which are 12.20%, 11.80%, and 9.19%, respectively. Agricultural sources have a large contribution rate to Hg and Cd, which are 68.17% and 33.69%, respectively.
Claims
1. A method for soil heavy metal source analysis based on linear discriminant analysis, characterized in that: The following steps are involved: Step 1: Set up monitoring points in the study area, collect soil samples, process soil samples, and measure the content of heavy metals in soil samples; Step 2: Conduct descriptive statistical analysis on soil heavy metals in the studied area; Step 3: Use geostatistical methods to analyze the spatial distribution characteristics of heavy metals in the soil of the study area and identify the sources of soil pollution; Step 4: Establish the ALDS-MLR receptor model for soil heavy metal pollution source analysis, and use the ALDS-MLR receptor model for soil heavy metal pollution source analysis to analyze the soil heavy metal pollution sources and their contribution rates; Step 5: Based on the pollution sources identified by the spatial distribution characteristics of heavy metals through geostatistical analysis and the pollution sources and their contribution rates analyzed by the ALDS-MLR receptor model, clear pollution sources and their contribution rates are obtained; In step 4, the following sub-steps were used to establish the ALDS-MLR receptor model for soil heavy metal pollution source apportionment: Step 4-1: Perform correlation analysis on the collected heavy metal concentration data to obtain the correlation coefficients between the heavy metals; Step 4-2: Use linear discriminant analysis (LDA) to reduce the dimension, maximize the distance between classes, minimize the distance within classes, optimize the objective function, and obtain the optimal projection matrix ω; Step 4-3: Obtain absolute linear discriminant scores and establish a multivariate linear regression model; In step 4-2, the following sub-steps are included: Step 4-2-1: Calculate the intra-class scatter matrix S w ; Where C is the number of sample categories, C i is the class to which the i-th sample belongs, μ i is the mean of the i-th sample, and x is the true value of the sample; Step 4-2-2: Calculate the hashing of each class mean point relative to the sample center to obtain the inter-class scatter matrix S b ; Where C is the number of sample categories, μ i is the mean of the ith sample, μ is the mean of all samples, N is the number of samples, N i is the i-th sample; Step 4-2-3: Optimize the objective function to maximize the distance between classes and minimize the variance within classes, and maximize the objective function to achieve the maximum: Where ω is the projection matrix, C is the number of sample categories, C i is the i-th sample, μ i is the mean of the i-th sample, μ is the mean of all samples, and x is the true value of the sample; replace S in step 4-2-1 with S in step 4-2-2 w With S b Substituting the above objective function into the equation, it can be simplified to: Where S b , S w are the inter-class scatter matrix and the intra-class scatter matrix respectively, and ω is the projection matrix; Step 4-2-4: Calculate the matrix Perform eigendecomposition on the matrix; Step 4-2-5: Calculate the eigenvalues obtained in step 4-2-4. The largest d eigenvalues and the corresponding eigenvalue vectors (ω1, ω2, ω3, …ω d ) to get the projection matrix ω, where ω d is the projection vector represented by the dth eigenvalue; Step 4-2-6: For each sample feature x in the sample set i , converted into a new sample z i =ω T x i ; Step 4-2-7: Get the output sample set; D`={(z1,y1),(z2,y2),…,(z i ,y i ),…,(z N ,y N )}, where z i is the i-th new sample after transformation, y i is the category to which the i-th sample belongs, and N represents the number of samples; In step 4-3, the following sub-steps are included: Step 4-3-1: Calculate the absolute linear discriminant score; The absolute linear discriminant score was obtained by subtracting the main factor score of the 0 concentration sample from the main factor score obtained by linear discriminant analysis to obtain the ALDS of each sample; Step 4-3-2: Take the absolute linear discriminant score as the independent variable and the metal concentration as the dependent variable for regression analysis to obtain the regression coefficient and regression constant term; the obtained regression coefficient can be used to convert ALDS into the concentration contribution of the pollution source corresponding to the main factor to each sample, and the formula is: Where p is the number of factors, c i is the concentration of heavy metal i, b io is the constant term of multiple linear regression, b pi is the regression coefficient of multiple linear regression, ALDS p is the absolute linear discriminant score of factor p, b pi ×ALDS p is the factor p for c i The content contribution of all samples is pi ×ALDS p The average value is the average absolute contribution of the pollution source corresponding to factor p, where the contribution rate of the pollution source corresponding to factor p is the ratio of its average absolute contribution to the contribution of all sources.
2. The method according to claim 1, characterized in that In step 2, a histogram is used to check whether the heavy metal content conforms to a normal distribution.
3. The method according to claim 1, characterized in that In step 3, the heavy metal content data that do not conform to the normal distribution are logarithmically transformed, and the ordinary Kriging interpolation method is used to draw the spatial characteristic distribution map of soil heavy metals to analyze the pollution source.
4. The method according to claim 1, characterized in that: In step 5, ALDS is absolute linear discriminant analysis and MLR is multivariate linear regression.
Citation Information
Patent Citations
Surface soil heavy metal pollution source quantitative identification method based on enrichment factor value calculation
CN109900682A
Urban surface soil heavy metal pollution analysis and evaluation method
CN111046572A