Multi-model-based cross-regional farmland soil heavy metal source analysis and prediction combined method
By combining multiple models, including APCS-MLR, UNMIX, input flux, and geographic detector models, and optimizing the hyperparameters of random forest, the problem of inaccurate results in the analysis and prediction of soil heavy metal pollution sources was solved, and accurate quantification and precise prediction were achieved across regions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA THREE GORGES UNIV
- Filing Date
- 2025-12-29
- Publication Date
- 2026-04-28
AI Technical Summary
Existing technologies for the analysis and prediction of heavy metal pollution sources in soil suffer from problems such as inaccurate results, inability to accurately locate pollution sources, and neglect of regional differences. In particular, the calculation results of a single model cannot verify its correctness and applicability.
A multi-model approach was adopted, combining the absolute principal component score-multiple linear regression model (APCS-MLR) to verify the number of factors in the UNMIX model, using the input flux model to calculate the pollution sources in the region, combining the geographic detector model to explore potential pollution sources, and optimizing the hyperparameters of the random forest model to improve prediction accuracy.
It achieves accurate quantitative analysis and precise prediction of cross-regional soil heavy metal pollution sources, can identify regional pollution sources, improves the reliability and prediction accuracy of model results, and avoids the smoothness error of geostatistical interpolation.
Smart Images

Figure CN121935522A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of soil pollution detection technology, specifically involving a combined method for cross-regional farmland soil heavy metal source apportionment and prediction based on multiple models. Background Technology
[0002] Soil plays a crucial role in isolating pollution, but it is also susceptible to contaminants, especially heavy metals. Excessive heavy metals in farmland soil not only alter its physicochemical properties, threaten biodiversity, and reduce food yields, but also accumulate in the human body through the food chain, threatening human health. While changes in parent material, topography, and inherent soil properties are sufficient to influence heavy metal pollution, human activities such as mining, industrial emissions, transportation, and the application of pesticides and fertilizers exacerbate this process. Therefore, establishing a method for source apportionment and prediction of heavy metals in farmland soil to comprehensively and accurately locate pollution sources is crucial for pollution control. Common methods for heavy metal source apportionment, such as positive definite matrix factor analysis (PMF), stable isotope analysis, UNMIX method, principal component analysis (PCA), and correlation analysis, are widely used. Common heavy metal prediction methods, such as Kriging interpolation and inverse distance weighted interpolation, are frequently employed in geostatistical interpolation.
[0003] Given the characteristics of soil heavy metal pollution—its insidious nature, persistence, and difficulty in biodegradation—and the complex and ever-changing nature of soil pollution problems, coupled with the inherent limitations of single receptor models, current work on source apportionment and prediction of soil heavy metal pollution inevitably reveals the following shortcomings: Qualitative source apportionment methods such as principal component analysis and correlation analysis can only classify various heavy metals and determine whether they originate from the same pollution source. These methods are sensitive to the distribution and linearity of the data; the presence of nonlinear relationships, outliers, or a large number of outliers can affect the results. However, soil pollution is complex and multifaceted, influenced by multiple pollution sources. Therefore, applying qualitative source apportionment methods to heavy metal source apportionment may not yield complete results.
[0004] While positive definite matrix factor analysis (PMF) and UNMIX models can quantitatively analyze the contribution of each pollution source to each heavy metal, the PMF model struggles to determine a reliable number of factors if the dataset does not conform to a normal distribution. UNMIX can automatically calculate a reliable number of pollution factors, but this result lacks specific real-world pollution comparisons, making it difficult to accurately identify pollution sources. The stable isotope method, due to its high cost, can only be applied to the analysis of small-scale heavy metal sources.
[0005] Input flux model (IFX) can determine the main sources of heavy metal pollution by calculating the input flux of each pollution source to each heavy metal. However, calculating the input flux only for the entire study area can lead to an excessively high input flux of a certain pollution source in a certain zone, which will affect the heavy metal source tracing of pollution sources in other zones. Furthermore, the calculation results of a single model cannot be compared and verified, and the correctness of the model results cannot be confirmed.
[0006] In geostatistical interpolation methods, Kriging interpolation yields poor prediction results when the dataset has a poor normal distribution. While inverse distance weighted interpolation does not require the data to be normally distributed, its prediction accuracy is generally lower than that of Kriging interpolation, making it only suitable for preliminary predictions of heavy metal distribution. Furthermore, to maintain image smoothness, both methods will not mark points causing heavy metal hotspots as hotspots in the predicted image.
[0007] Whether it's principal component analysis or correlation analysis in multivariate statistical methods, or PMF models or UNMIX models in receptor models, or input flux models, the results obtained by existing studies through these models are usually the pollution sources of a certain heavy metal in an entire study area. However, the heavy metal pollution sources in certain sub-regions of the study area are easily overlooked due to the overall calculation, making it impossible to accurately identify the pollution sources and their geographical locations.
[0008] Therefore, it is necessary to design a combined method based on multiple models for cross-regional farmland soil heavy metal source apportionment and prediction to solve the above problems. Summary of the Invention
[0009] This invention provides a combined method for cross-regional heavy metal source apportionment and prediction in farmland soil based on multiple models. This method utilizes the PCA model, the computational basis of the Absolute Principal Component Score-Multiple Linear Regression (APCS-MLR), to help the UNMIX model verify the accuracy of the number of factors calculated by the model. Simultaneously, it compares the quantitative results of pollution sources from APCS-MLR with those from UNMIX to verify the correctness of the model results. The IFX model can calculate the input flux of pollution sources in each study area, incorporating field investigation data to support the source tracing of the receptor model results, enabling cross-regional pollution source identification and correlation analysis, and tracing overlooked heavy metal pollution sources in each area. The Geographic Detector (GDM) model can explore the impact of potential pollution sources on heavy metal distribution, providing data support for establishing pollution sources when actual survey data is insufficient. Furthermore, in predicting heavy metal distribution, this invention uses random forest as the basic machine learning model, employs geostatistical interpolation to optimize the dataset, giving it spatial properties, and utilizes an improved sparrow search algorithm to determine the hyperparameters of the random forest prediction model, thereby improving the accuracy of the prediction.
[0010] To achieve the above-mentioned technical effects, the technical solution adopted by the present invention is as follows: The combined approach of multiple models for cross-regional source apportionment and prediction of heavy metals in farmland soil includes the following steps: S1, Raw dataset collection and preprocessing; S2, Principal component analysis of soil heavy metal data was performed using the PCA model; S3. The APCS-MLR model and the UNMIX model were used to perform quantitative source apportionment of soil heavy metal concentrations. S4. Combining the input flux model results of each region, the contribution rate of pollutants resolved by the receptor model, the actual survey situation, and the GDM model results, a specific analysis of pollution sources is conducted. S5, based on the environmental dataset in S4, uses geostatistical interpolation, receptor model and geographic detector to optimize the dataset, and optimizes the random forest model RF based on the improved sparrow optimization algorithm, so as to predict the spatial distribution map of heavy metals in the soil.
[0011] Preferably, the specific method of step S1 is as follows: S101. Divide the study area into multiple zones where samples can be collected. Collect soil samples from each zone and send them to the laboratory to test the heavy metal content. Based on previous experience in soil heavy metal source apportionment studies, identify potential pollution sources and pathways in the soil. Collect samples from each pollution source in each zone and send them to the laboratory to test the heavy metal content. Collect environmental data that may indirectly affect changes in soil heavy metal content in the current year from the internet as the basic dataset for heavy metal prediction. S102, use Excel software to normalize the heavy metal concentration data detected in the soil samples to obtain the normalized dataset corresponding to the concentration dataset. Use Excel software to calculate the summary table of heavy metals in potential pollution sources, including the maximum, minimum and median values.
[0012] Preferably, the specific method of step S2 is as follows: S201. The detected soil heavy metal content was imported into SPSS software for factor dimensionality reduction. The KMO test and Bartlett's test of sphericity were used to determine whether the data were suitable for principal component analysis. For the KMO test, the KMO value should be greater than 0.5, and for the Bartlett's test of sphericity, the p value should be less than 0.05 or 0.01. In step S202, after completing the suitability tests for factor analysis, including the KMO and Bartlett's test of sphericity, and selecting datasets that meet the criteria, principal component analysis (PCA) was used to reduce the dimensionality of the soil heavy metal data. The factor structure was optimized using the variance-maximum orthogonal rotation technique, ultimately generating a visualization of eigenvalues, variance contribution rates, and cumulative contribution rates. A PCA result table was also output to clarify the number of effective principal components extracted and their explanatory power. Datasets that met the criteria for factor analysis were selected, and PCA was performed. After variance-maximum orthogonal rotation, a visualization of soil heavy metal eigenvalues and cumulative variance contribution rates was generated, resulting in a PCA result table and the corresponding number of principal components.
[0013] Preferably, the specific method of step S3 is as follows: S301. Import the preprocessed normalized data from S1 into EPA Unmix 6.0 software. According to the software user manual, the normalized data should ideally include the total concentration normalized data of heavy metals. In the lower right corner of the main window, you can track the status of the species as designated as "Total," "Tracer," or "Norm." Usually, the normalized species and the total species are the same, so the source composition obtained is the mass fraction. S302, in the built-in functions of Unmix software, select the function of automatically generating initial factors. If the original dataset cannot automatically initialize the number of factors, the number of principal components parsed by the PCA model is used as the minimum number of factors. Manually set the number of factors required for the model to run, and then set the model to perform calculations. S303 sets the criteria for determining the optimal pollutant. S304, the number of factors that meet the results of S303 after the statistical software is run; S305. Using the built-in functions of EPA Unmix 6.0 software, we obtained the contribution of each factor to various metals and the fitting results of the actual and predicted values of each heavy metal after the model was run. In EXCEL software, we used the source component spectrum obtained by running EPA Unmix 6.0 software to calculate the contribution ratio of each pollutant factor to the overall soil heavy metal concentration. S306. Based on the number of factors resolved by the UNMIX model, select an appropriate number of principal components. In SPSS software, calculate the principal component scores obtained after PCA analysis. Add data "0" to the last row of the soil heavy metal dataset. Then, perform PCA analysis using the soil heavy metal dataset with 0 to obtain the principal component scores of the row with data "0". In EXCEL, subtract the principal component scores of the row with data "0" from the principal component scores of each soil heavy metal sample to obtain the absolute principal component scores of each soil heavy metal sample. S307: Add the absolute principal component scores to the soil heavy metal dataset in SPSS, and use the software's built-in linear regression function to calculate the goodness of fit between heavy metal concentration and absolute principal component scores. If the goodness of fit is too low, it is necessary to appropriately add principal components with eigenvalues less than but close to 1 from the PCA analysis and recalculate S306-S307. S308, calculate the average value of the absolute principal components as the basis for calculating the contribution of pollutants to heavy metals; S309, the APCS-MLR model can calculate known pollution sources containing principal components, as well as unknown pollution sources not included in the principal components by PCA analysis; S310, Calculate the contribution of pollution sources according to the contribution calculation formula of known pollution sources and unknown pollution sources; S311. Compare the final results of the APCS-MLR and UNMIX models. If there are differences between the results of the two models, select the best model result by combining the summary table of heavy metal concentrations in potential pollution sources. Preferably, the criteria for determining the optimal pollutant in S302 are as follows: a. Of all the sampled data points input into the model, at least 80% of the data have residual analysis results between -3 and +3 and conform to a normal distribution; b. The fitting coefficient R² is greater than or equal to 0.8; c. Minimum signal-to-noise ratio greater than or equal to 2.
[0014] Preferably, the specific method of step S4 is as follows: S401. First, calculate the input flux of heavy metals in the pollution source samples collected throughout the study area to obtain the proportion of input flux of each heavy metal for each pollution source. Then, verify the results with the APCS-MLR and UNMIX model analysis and the actual survey results to make a preliminary judgment on the real pollution sources. S402. When the number of factors resolved by the UNMIX model exceeds the number of "known sources" in the APCS-MLR model, it means that the results resolved by the UNMIX model include the factors that the APCS-MLR model lost due to the defects of PCA analysis—the unknown sources. This means that the results resolved by the UNMIX model are reasonable. S403. If the input flux results differ significantly from the receptor model results, it indicates that a single pollution source cannot represent the pollution source of a certain heavy metal in the entire study area. There must be other pollution sources that affect the distribution of the heavy metal. If this occurs, it is necessary to calculate the input flux of heavy metals in pollution sources in each survey area and calculate the comparison coefficient (COP) between the input flux of heavy metals in different pollution sources in each survey area and the input flux of heavy metals in different pollution sources in the entire study area. This will determine that a certain heavy metal in the entire study area has different main pollution sources in different zones, thus achieving cross-regional pollution source identification. This result can make up for the shortcomings of the receptor model. S404 If the data from the field survey is insufficient, and the computational data of the input flux is insufficient to support the source analysis of all types of heavy metals in the study, then the factor detector model in the GDM model needs to be used. It can study the factors and potential mechanisms that affect the spatial heterogeneity of heavy metals, explore the contribution of a certain influencing factor to the distribution of heavy metal pollution, and use the datasets collected in S1 that may have an indirect impact on the changes in soil heavy metal content to analyze unknown pollution sources that still cannot be analyzed after input flux.
[0015] Preferably, the key computational metrics for the GDM model are the q-value and the p-value: The q-value quantifies the explanatory power of various influencing factors on the target variable by comparing the ratio of in-stratum variance to population variance. The p-value verifies statistical reliability. In practical applications, a significance test is often constructed using Monte Carlo simulation. When p < 0.05, the factor can be considered to have statistical significance in explaining spatial differentiation.
[0016] Preferably, the specific method of step S5 is as follows: S501. Since the environmental dataset in S404 is data that indirectly affects the distribution of heavy metals in the soil, such as the distance between sampling points and mining areas, highways, factories, etc., this data often ignores the spatial heterogeneity of heavy metal distribution. Therefore, when optimizing the dataset, two measures can be adopted: First, use geostatistical interpolation to obtain a rough spatial distribution of heavy metals; second, use the results of the receptor model as spatial weight data, perform geostatistical interpolation, and add the results of both to the environmental dataset to give it spatial heterogeneity. S502, the environmental dataset is collected based on pollution pathways that may affect the distribution of heavy metals. This relies heavily on previous research experience. Due to time constraints during data collection, the real-time performance of the data may not be very strong. Therefore, when using it, it is necessary to first use the GDM factor detector model to explore environmental data that contributes to each heavy metal element. When predicting the distribution of heavy metals, environmental data that does not contribute significantly should be deleted to reduce data interference. S503, when predicting heavy metal distribution using the Random Forest model, it is necessary to determine its hyperparameters. Manually determining hyperparameters is often not optimal. This paper proposes a hyperparameter optimization method based on the Sparrow Search Algorithm (SSA). This method automatically searches within a given upper and lower limit range of hyperparameters using the Sparrow Search Algorithm, aiming to minimize the loss function (such as mean squared error, mean absolute error, etc.) to find the hyperparameter combination that optimizes model performance. S504, the general process for setting the optimization algorithm; S505, optimized sparrow search algorithm.
[0017] a. The initial population generated by chaotic mapping has better ergodicity and randomness, and can be more evenly distributed in the search space, avoiding getting trapped in local optima. Compared with simple random initialization, chaotic mapping can better explore the search space and reduce the risk of the algorithm converging to a local optimum too early. b. An improved producer (discoverer) update strategy: By introducing a dynamic exponential decay factor, the discoverer's position update strategy emphasizes exploration in early iterations and development in later iterations. When the warning value is large, a normally distributed random perturbation is added to enhance the exploration capability. c. An improved follower update strategy adaptively adjusts the step size based on the individual's ranking in the population. For individuals with poor fitness (i>self.pop / 2), an exponential decay approach is used to move closer to the current worst individual. For individuals with good fitness, an update strategy based on the best individual and a random matrix is used. S506. After finding the optimal hyperparameters of the random forest model using the optimized sparrow search algorithm, the spatial distribution map of heavy metals in the soil is accurately predicted based on the environmental dataset with spatial properties. The spatial distribution map predicted by machine learning and the geostatistical interpolation prediction map are compared, and the prediction map with higher accuracy is used for heavy metal source analysis. Preferably, the general process of setting the optimization algorithm includes: a. Define the range of hyperparameters: Based on experience or preliminary experiments, determine the reasonable range of values for each hyperparameter; b. Initialize the sparrow population: Each individual sparrow represents a set of hyperparameter combinations; c. Constructing a random forest model: For each individual sparrow, construct a random forest model using its corresponding hyperparameter combination and train it on the training set; d. Evaluate model performance: Evaluate the model's performance on the validation set and calculate the loss function value; e. Update sparrow locations: According to the rules of the sparrow optimization algorithm, update the locations of the sparrow population, i.e., adjust the hyperparameter combination; f. Iterative optimization: Repeat steps 3-5 until the preset maximum number of iterations is reached or the convergence condition is met; g. Output optimal hyperparameters: Select the combination of hyperparameters that minimizes the loss function from the final sparrow population as the optimal solution.
[0018] Preferably, the optimized sparrow search algorithm includes: a. The initial population generated by chaotic mapping has better ergodicity and randomness, and can be more evenly distributed in the search space, avoiding getting trapped in local optima. Compared with simple random initialization, chaotic mapping can better explore the search space and reduce the risk of the algorithm converging to a local optimum too early. b. An improved producer (discoverer) update strategy: By introducing a dynamic exponential decay factor, the discoverer's position update strategy emphasizes exploration in early iterations and development in later iterations. When the warning value is large, a normally distributed random perturbation is added to enhance the exploration capability. c. An improved follower update strategy adaptively adjusts the step size based on the individual's ranking in the population. For individuals with poor fitness (i>self.pop / 2), an exponential decay approach is used to move closer to the current worst individual. For individuals with good fitness, an update strategy based on the best individual and a random matrix is adopted.
[0019] The beneficial effects of this invention are as follows: 1. This invention employs a combination of PCA principal component analysis model from multivariate statistical analysis methods and APCS-MLR and UNMIX models from receptor models. The results of each model are cross-validated to make the analysis results more reasonable. Furthermore, when determining the optimal analysis results, the content of each heavy metal in the pollution source is taken into account to make the results more consistent with the actual survey results. This enables rapid and accurate quantitative source analysis of heavy metal pollution sources in soil, accurately calculates the contribution rate of each pollution source, quantifies it, and makes the quantitative results more reliable.
[0020] 2. This invention uses an input flux model to calculate the input flux of heavy metals in pollution sources in the overall study area, quantifies the results as contribution rates, and compares them with the analysis results of the receptor model to determine the pollution source of each heavy metal. For heavy metals with different analysis results from the two models, the input flux of pollution sources to heavy metals in each partition of the study area is calculated. The input flux of the partition is compared with the comprehensive input of the overall study area, and pollution sources exceeding the threshold are identified as the pollution sources of heavy metals in that partition, resulting in more accurate identification results.
[0021] 3. This invention combines the analytical results of the geographic detector model and the receptor model. The absolute principal component score or factor contribution matrix obtained from the receptor model is used as the dependent variable. The geographic detector determines the degree of influence of each environmental variable on the dependent variable. The environmental variables with large contributions are selected as pollution pathways to indirectly identify pollution sources. This enables automated identification of pollution sources when there is insufficient field survey data. The identification speed is higher and the reliability is stronger.
[0022] 4. The present invention uses geostatistical interpolation to make the dataset spatially heterogeneous, which conforms to the second law of geography. It also uses the analytical results of the receptor model to give the dataset spatial weights of potential pollution sources. The information contained in the dataset is more relevant to the heavy metals in the collected soil.
[0023] 5. This invention uses a chaotic mapping algorithm to optimize the population size in the sparrow search optimization algorithm, introduces a dynamic exponential decay factor and adaptively adjusts the step size, which enhances the exploration ability of producers in the algorithm, making the optimization algorithm faster and less prone to getting trapped in local optima.
[0024] 6. This invention combines an improved environmental dataset and an improved sparrow optimization algorithm in a random forest model to achieve accurate prediction of the spatial distribution of heavy metals in soil.
[0025] 7. This invention uses a precisely predicted spatial distribution map of heavy metals in soil to optimize the source apportionment results, so that the results no longer misidentify hotspots of heavy metal pollution due to the smoothness of geostatistical interpolation. Attached Figure Description
[0026] Figure 1 This is a schematic diagram showing the comprehensive fitting coefficients and minimum signal-to-noise ratio results of the UNMIX model in an embodiment of the present invention.
[0027] Figure 2 This is a schematic diagram showing the fitting coefficient results of each heavy metal in the UNMIX model in an embodiment of the present invention.
[0028] Figure 3 This is a schematic diagram showing the contribution rate of each pollutant to heavy metals in the UNMIX model in this embodiment of the invention.
[0029] Figure 4 This is a schematic diagram showing the contribution rate of each pollutant to heavy metals in the APCS-MLR model in this embodiment of the invention.
[0030] Figure 5 This is a schematic diagram of the input flux results of heavy metals in each pollution source in the input flux model of this invention embodiment.
[0031] Figure 6This is a schematic diagram of the data collection zones and sampling point results in the study area of this invention.
[0032] Figure 7 This is a schematic diagram of the geostatistical interpolation results of eight heavy metal elements in an embodiment of the present invention.
[0033] Figure 8 This is a schematic diagram illustrating the degree of explanation of each environmental variable of the GDM model for the factors of the UNMIX model in an embodiment of the present invention.
[0034] Figure 9 This is a schematic diagram of the geostatistical interpolation reclassification results of Cd and As in an embodiment of the present invention.
[0035] Figure 10 This is a schematic diagram of the geostatistical interpolation results of the UNMIX model factors in an embodiment of the present invention.
[0036] Figure 11 This is a schematic diagram of the spatial distribution of Cd elements based on the optimized best dataset and the best model prediction in an embodiment of the present invention. Detailed Implementation
[0037] Example 1: The combined approach of multiple models for cross-regional source apportionment and prediction of heavy metals in farmland soil includes the following steps: S1, Raw dataset collection and preprocessing; S2, Principal component analysis of soil heavy metal data was performed using the PCA model; S3. The APCS-MLR model and the UNMIX model were used to perform quantitative source apportionment of soil heavy metal concentrations. S4. Combining the input flux model results of each region, the contribution rate of pollutants resolved by the receptor model, the actual survey situation, and the GDM model results, a specific analysis of pollution sources is conducted. S5, based on the environmental dataset in S4, uses geostatistical interpolation, receptor model and geographic detector to optimize the dataset, and optimizes the random forest model RF based on the improved sparrow optimization algorithm, so as to predict the spatial distribution map of heavy metals in the soil.
[0038] Preferably, the specific method of step S1 is as follows: S101. Divide the study area into multiple zones where samples can be collected. Collect soil samples from each zone and send them to the laboratory to test the heavy metal content. Based on previous experience in soil heavy metal source apportionment studies, identify potential pollution sources and pathways in the soil. Collect samples from each pollution source in each zone and send them to the laboratory to test the heavy metal content. Collect environmental data that may indirectly affect changes in soil heavy metal content in the current year from the internet as the basic dataset for heavy metal prediction. S102, use Excel software to normalize the heavy metal concentration data detected in the soil samples to obtain the normalized dataset corresponding to the concentration dataset. Use Excel software to calculate the summary table of heavy metals in potential pollution sources, including the maximum, minimum and median values.
[0039] Preferably, the specific method of step S2 is as follows: S201. The detected soil heavy metal content was imported into SPSS software for factor dimensionality reduction. The KMO test and Bartlett's test of sphericity were used to determine whether the data were suitable for principal component analysis. For the KMO test, the KMO value should be greater than 0.5, and for the Bartlett's test of sphericity, the p value should be less than 0.05 or 0.01. In step S202, after completing the suitability tests for factor analysis, including the KMO and Bartlett's test of sphericity, and selecting datasets that meet the criteria, principal component analysis (PCA) was used to reduce the dimensionality of the soil heavy metal data. The factor structure was optimized using the variance-maximum orthogonal rotation technique, ultimately generating a visualization of eigenvalues, variance contribution rates, and cumulative contribution rates. A PCA result table was also output to clarify the number of effective principal components extracted and their explanatory power. Datasets that met the criteria for factor analysis were selected, and PCA was performed. After variance-maximum orthogonal rotation, a visualization of soil heavy metal eigenvalues and cumulative variance contribution rates was generated, resulting in a PCA result table and the corresponding number of principal components.
[0040] Preferably, the specific method of step S3 is as follows: S301. Import the preprocessed normalized data from S1 into EPA Unmix 6.0 software. According to the software user manual, the normalized data should ideally include the total concentration normalized data of heavy metals. In the lower right corner of the main window, you can track the status of the species as designated as "Total," "Tracer," or "Norm." Usually, the normalized species and the total species are the same, so the source composition obtained is the mass fraction. S302, in the built-in functions of Unmix software, select the function of automatically generating initial factors. If the original dataset cannot automatically initialize the number of factors, the number of principal components parsed by the PCA model is used as the minimum number of factors. Manually set the number of factors required for the model to run, and then set the model to perform calculations. S303 sets the criteria for determining the optimal pollutant. S304, the number of factors that meet the results of S303 after the statistical software is run; S305. Using the built-in functions of EPA Unmix 6.0 software, we obtained the contribution of each factor to various metals and the fitting results of the actual and predicted values of each heavy metal after the model was run. In EXCEL software, we used the source component spectrum obtained by running EPA Unmix 6.0 software to calculate the contribution ratio of each pollutant factor to the overall soil heavy metal concentration. S306. Based on the number of factors resolved by the UNMIX model, select an appropriate number of principal components. In SPSS software, calculate the principal component scores obtained after PCA analysis. Add data "0" to the last row of the soil heavy metal dataset. Then, perform PCA analysis using the soil heavy metal dataset with 0 to obtain the principal component scores of the row with data "0". In EXCEL, subtract the principal component scores of the row with data "0" from the principal component scores of each soil heavy metal sample to obtain the absolute principal component scores of each soil heavy metal sample. S307: Add the absolute principal component scores to the soil heavy metal dataset in SPSS, and use the software's built-in linear regression function to calculate the goodness of fit between heavy metal concentration and absolute principal component scores. If the goodness of fit is too low, it is necessary to appropriately add principal components with eigenvalues less than but close to 1 from the PCA analysis and recalculate S306-S307. S308, calculate the average value of the absolute principal components as the basis for calculating the contribution of pollutants to heavy metals; S309, the APCS-MLR model can calculate known pollution sources containing principal components, as well as unknown pollution sources not included in the principal components by PCA analysis; S310, Calculate the contribution of pollution sources according to the contribution calculation formula of known pollution sources and unknown pollution sources; S311. Compare the final results of the APCS-MLR and UNMIX models. If there are differences between the results of the two models, select the best model result by combining the summary table of heavy metal concentrations in potential pollution sources. Preferably, the criteria for determining the optimal pollutant in S302 are as follows: a. Of all the sampled data points input into the model, at least 80% of the data have residual analysis results between -3 and +3 and conform to a normal distribution; b. The fitting coefficient R² is greater than or equal to 0.8; c. Minimum signal-to-noise ratio greater than or equal to 2.
[0041] Preferably, the specific method of step S4 is as follows: S401. First, calculate the input flux of heavy metals in the pollution source samples collected throughout the study area to obtain the proportion of input flux of each heavy metal for each pollution source. Then, verify the results with the APCS-MLR and UNMIX model analysis and the actual survey results to make a preliminary judgment on the real pollution sources. S402. When the number of factors resolved by the UNMIX model exceeds the number of "known sources" in the APCS-MLR model, it means that the results resolved by the UNMIX model include the factors that the APCS-MLR model lost due to the defects of PCA analysis—the unknown sources. This means that the results resolved by the UNMIX model are reasonable. S403. If the input flux results differ significantly from the receptor model results, it indicates that a single pollution source cannot represent the pollution source of a certain heavy metal in the entire study area. There must be other pollution sources that affect the distribution of the heavy metal. If this occurs, it is necessary to calculate the input flux of heavy metals in pollution sources in each survey area and calculate the comparison coefficient (COP) between the input flux of heavy metals in different pollution sources in each survey area and the input flux of heavy metals in different pollution sources in the entire study area. This will determine that a certain heavy metal in the entire study area has different main pollution sources in different zones, thus achieving cross-regional pollution source identification. This result can make up for the shortcomings of the receptor model. S404 If the data from the field survey is insufficient, and the computational data of the input flux is insufficient to support the source analysis of all types of heavy metals in the study, then the factor detector model in the GDM model needs to be used. It can study the factors and potential mechanisms that affect the spatial heterogeneity of heavy metals, explore the contribution of a certain influencing factor to the distribution of heavy metal pollution, and use the datasets collected in S1 that may have an indirect impact on the changes in soil heavy metal content to analyze unknown pollution sources that still cannot be analyzed after input flux.
[0042] Preferably, the key computational metrics for the GDM model are the q-value and the p-value: The q-value quantifies the explanatory power of various influencing factors on the target variable by comparing the ratio of in-stratum variance to population variance. The p-value verifies statistical reliability. In practical applications, a significance test is often constructed using Monte Carlo simulation. When p < 0.05, the factor can be considered to have statistical significance in explaining spatial differentiation.
[0043] Preferably, the specific method of step S5 is as follows: S501. Since the environmental dataset in S404 is data that indirectly affects the distribution of heavy metals in the soil, such as the distance between sampling points and mining areas, highways, factories, etc., this data often ignores the spatial heterogeneity of heavy metal distribution. Therefore, when optimizing the dataset, two measures can be adopted: First, use geostatistical interpolation to obtain a rough spatial distribution of heavy metals; second, use the results of the receptor model as spatial weight data, perform geostatistical interpolation, and add the results of both to the environmental dataset to give it spatial heterogeneity. S502, the environmental dataset is collected based on pollution pathways that may affect the distribution of heavy metals. This relies heavily on previous research experience. Due to time constraints during data collection, the real-time performance of the data may not be very strong. Therefore, when using it, it is necessary to first use the GDM factor detector model to explore environmental data that contributes to each heavy metal element. When predicting the distribution of heavy metals, environmental data that does not contribute significantly should be deleted to reduce data interference. S503, when predicting heavy metal distribution using the Random Forest model, it is necessary to determine its hyperparameters. Manually determining hyperparameters is often not optimal. This paper proposes a hyperparameter optimization method based on the Sparrow Search Algorithm (SSA). This method automatically searches within a given upper and lower limit range of hyperparameters using the Sparrow Search Algorithm, aiming to minimize the loss function (such as mean squared error, mean absolute error, etc.) to find the hyperparameter combination that optimizes model performance. S504, the general process for setting the optimization algorithm; S505, optimized sparrow search algorithm.
[0044] a. The initial population generated by chaotic mapping has better ergodicity and randomness, and can be more evenly distributed in the search space, avoiding getting trapped in local optima. Compared with simple random initialization, chaotic mapping can better explore the search space and reduce the risk of the algorithm converging to a local optimum too early. b. An improved producer (discoverer) update strategy: By introducing a dynamic exponential decay factor, the discoverer's position update strategy emphasizes exploration in early iterations and development in later iterations. When the warning value is large, a normally distributed random perturbation is added to enhance the exploration capability. c. An improved follower update strategy adaptively adjusts the step size based on the individual's ranking in the population. For individuals with poor fitness (i>self.pop / 2), an exponential decay approach is used to move closer to the current worst individual. For individuals with good fitness, an update strategy based on the best individual and a random matrix is used. S506. After finding the optimal hyperparameters of the random forest model using the optimized sparrow search algorithm, the spatial distribution map of heavy metals in the soil is accurately predicted based on the environmental dataset with spatial properties. The spatial distribution map predicted by machine learning and the geostatistical interpolation prediction map are compared, and the prediction map with higher accuracy is used for heavy metal source analysis. Preferably, the general process of setting the optimization algorithm includes: a. Define the range of hyperparameters: Based on experience or preliminary experiments, determine the reasonable range of values for each hyperparameter; b. Initialize the sparrow population: Each individual sparrow represents a set of hyperparameter combinations; c. Constructing a random forest model: For each individual sparrow, construct a random forest model using its corresponding hyperparameter combination and train it on the training set; d. Evaluate model performance: Evaluate the model's performance on the validation set and calculate the loss function value; e. Update sparrow locations: According to the rules of the sparrow optimization algorithm, update the locations of the sparrow population, i.e., adjust the hyperparameter combination; f. Iterative optimization: Repeat steps 3-5 until the preset maximum number of iterations is reached or the convergence condition is met; g. Output optimal hyperparameters: Select the combination of hyperparameters that minimizes the loss function from the final sparrow population as the optimal solution.
[0045] Preferably, the optimized sparrow search algorithm includes: a. The initial population generated by chaotic mapping has better ergodicity and randomness, and can be more evenly distributed in the search space, avoiding getting trapped in local optima. Compared with simple random initialization, chaotic mapping can better explore the search space and reduce the risk of the algorithm converging to a local optimum too early. b. An improved producer (discoverer) update strategy: By introducing a dynamic exponential decay factor, the discoverer's position update strategy emphasizes exploration in early iterations and development in later iterations. When the warning value is large, a normally distributed random perturbation is added to enhance the exploration capability. c. An improved follower update strategy adaptively adjusts the step size based on the individual's ranking in the population. For individuals with poor fitness (i>self.pop / 2), an exponential decay approach is used to move closer to the current worst individual. For individuals with good fitness, an update strategy based on the best individual and a random matrix is adopted.
[0046] Example 2: A combined approach based on multiple models for cross-regional source apportionment and prediction of heavy metals in farmland soil includes the following steps: S1: Raw dataset collection and preprocessing: S101, taking Yichang City as the study area, divides it into 15 regions. Soil samples are collected from each region and sent to the laboratory to test the content of heavy metals (Cd, Hg, As, Pb, Cr, Cu, Zn, Ni). Based on previous experience in soil heavy metal source apportionment studies, potential pollution sources and pollution pathways in the soil are determined. Samples of each pollution source are collected from each region and sent to the laboratory to test the content of heavy metals. Data that may have an indirect impact on changes in soil heavy metal content in the current year are collected online as the basic dataset for heavy metal prediction. Based on previous research and comprehensive analysis of relevant soil environmental data, the types, degrees, boundaries, and spatiotemporal evolution characteristics of heavy metal pollution in the project area were clarified. Potential pollution sources in the pilot area, including key enterprises, solid waste dumps, livestock farming, residential areas, irrigation water, and agricultural inputs, were identified. Considering the region's hydrology, topography, land use, economic development, and social environment, the main pollution sources were preliminarily determined to be atmospheric dustfall, livestock and poultry manure, agricultural inputs, and irrigation water. The concentration of heavy metals in arable soil is affected indirectly or directly by a variety of factors, leading to changes in heavy metal concentration. Based on historical research, elevation, topography, and geomorphology are the main factors influencing the heterogeneity of heavy metals in soil. Rainfall and temperature further affect the migration of heavy metals in soil by influencing soil moisture. The accumulation of heavy metals in soil is closely related to agricultural activities. Population density reflects the intensity of human production and living activities to a certain extent. The more frequent human activities, the more domestic waste and sewage irrigation will lead to the accumulation of heavy metals. In addition, motor vehicles, agricultural machinery and equipment will emit exhaust gases containing heavy metals. The deposition of exhaust gases will cause heavy metal pollution in the soil. Soil pH value can directly or indirectly affect the activity, mobility and geochemical behavior of heavy metals in the soil. Heavy metals in the soil, such as Cr, Ni and Cd, are inherited from the weathering of the parent material through biogeochemical cycles. Therefore, when collecting environmental datasets, the following data can be collected: distance from the sampling point to the road, distance from the waterway, distance from the mining area, slope, distance from the factory, elevation, soil use type, soil type, population density, annual average temperature, annual average precipitation, fertilizer use, pesticide use, normalized vegetation index, total phosphorus, total nitrogen, and total potassium. S102, use Excel software to normalize the heavy metal concentration data detected in the soil samples to obtain the normalized dataset corresponding to the concentration dataset, and use Excel software to calculate the summary table (including maximum value, minimum value and median) of each heavy metal in the pollution sources (atmospheric dust, livestock and poultry manure, agricultural inputs and irrigation water). S2: Principal component analysis of soil heavy metal data using PCA model: S201. The detected heavy metal content in the soil was imported into SPSS software for factor reduction. The KMO test and Bartlett's test of sphericity were used to determine whether the data were suitable for principal component analysis. For the KMO test, the KMO value should be greater than 0.5, and for the Bartlett's test of sphericity, the p-value should be less than 0.05 or 0.01. Currently, with 3 principal components, the KMO value is 0.665, and the p-value is less than 0.01. As shown in Table 1, after the maximum variance orthogonal rotation, the cumulative variance of the three principal components of S202 is 62.108%. The total variance interpretation plots for the three principal component analyses are shown in Table 1 below: Table 1: Explanation of variance for the three principal components;
[0047] Since the PCA analysis results are used as the basis for the calculation of the APCS-MLR model, the cumulative variance of the three principal components, which is 62.108%, is relatively insufficient to support the subsequent calculations of the model. Therefore, principal components with eigenvalues close to 1 need to be included in the principal component analysis as well. The results are shown in Table 2. Table 2: Explanation of variance for the five principal components;
[0048] In Table 2, the eigenvalues of principal components 4 and 5 both exceed 0.8. The maximum number of principal components does not exceed the total number of heavy metals. Therefore, it is necessary to further determine the selection of the number of principal components by referring to the results of the UNMIX model. S3. Quantitative source apportionment of soil heavy metal concentrations was performed using the APCS-MLR and UNMIX models: S301. Import the normalized data into the EPA Unmix 6.0 software and specify the normalized total concentration as "Total" and "Norm". After running S302 with UNMIX software, the optimal number of factors was determined to be 5. The model's minimum fit coefficient R² reached 0.89, exceeding the system requirement of 0.8, indicating that the model can explain 89% of the species variance. Simultaneously, the minimum signal-to-noise ratio also reached 2.37, exceeding the system's minimum requirement of 2, further validating the reliability of the analytical results. The results are as follows... Figure 1 As shown.
[0049] S303, the number of factors that meet the S303 result after running the statistical software - the number of pollution factors is 5; S304, the software also provides the goodness of fit between the actual and predicted values for each heavy metal, such as Figure 2 As shown.
[0050] Figure 2Among them, the goodness of fit for Pb was low at 0.41, but the goodness of fit for the other 7 heavy metal elements was higher than 0.6, reaching as high as 1.0. Therefore, the UNMIX model is not very reliable for tracing the source of Pb pollution factors and needs to be combined with the results analysis of the APCS-MLR model. The contribution of factors analyzed by the UNMIX model to heavy metals is as follows: Figure 3 As shown.
[0051] Cd, Hg, and As originate from factors 2, 1, and 3, respectively; Pb, Cu, and Zn have similar origins, with the main contribution coming from factor 5; Cr and Ni have similar origins, with the main contribution coming from factor 4. S305, the factor number analysis result of the UNMIX model is 5, so when selecting principal components, 4 principal components are selected because the final calculation result of APCS-MLR includes 4 known principal component factors and 1 unknown factor. The unknown factor represents the 30% variance that the principal component analysis failed to accumulate. Then, in SPSS software, the principal component scores obtained after PCA analysis are calculated. Data "0" is added to the last row of the soil heavy metal dataset, and then PCA analysis is performed using the soil heavy metal dataset with 0 to obtain the principal component score of the row with data "0". In EXCEL, the principal component score of each soil heavy metal sample is subtracted from the principal component score of the row with data "0" to obtain the absolute principal component score of each soil heavy metal sample. S306: The absolute principal component scores were added to the soil heavy metal dataset in SPSS. The software's built-in linear regression function was used to calculate the goodness of fit between heavy metal concentration and absolute principal component scores. If the goodness of fit was too low, principal components with eigenvalues less than but close to 1 from the PCA analysis needed to be added appropriately, and the calculations for S305-S306 were repeated. When the number of principal components was 3, the goodness of fit between the 3 absolute principal component scores and the 8 heavy metal elements was shown in Table 3. The goodness of fit (R^2) for heavy metals Cd, As and Pb was less than 0.6, and the highest goodness of fit for the 8 heavy metals was Hg - 0.761. Table 3: Results of regression analysis based on 3 principal components and 8 heavy metals;
[0052] The goodness of fit of the four absolute principal component scores and the linear regression of heavy metals is shown in Table 4. The goodness of fit of all eight heavy metals is higher than 0.6, and the best goodness of fit is 0.922. Although the overall goodness of fit of the linear regression of heavy metal elements will be higher when the number of principal components is 5, considering that the optimal factor calculated by the UNMIX model is 5, the number of basic factors for the operation of APCS-MLR is selected as 4 (4 known sources and 1 unknown source), which corresponds to the UNMIX model. Table 4: Regression analysis results based on 4 principal components and 8 heavy metals;
[0053] S307, the average value of the absolute principal components is used as the basis for calculating the contribution of pollutants to heavy metals. The average values of the four absolute principal components are calculated to be 1.525839923, 1.726180077, 0.821980033, and 2.589669902, respectively. The S308 APCS-MLR model can calculate known pollution sources containing principal components, as well as unknown pollution sources not included in the principal components by PCA analysis. S309, calculate the contribution of pollution sources based on the contribution calculation formulas for known and unknown pollution sources. The calculation formula for known pollution sources is as follows: ; represents the linear regression coefficient of the m-th absolute principal component on the i-th heavy metal. This represents the average absolute principal component score of the m-th component. The intercept term represents the basic content level of heavy metal elements in the absence of other influencing factors (i.e., all independent variables are 0).
[0054] The formula for calculating unknown pollution sources is as follows: ; S310. Compare the final results of the APCS-MLR and UNMIX models. If there are differences between the two models, select the best model result by combining the summary table of heavy metal concentrations in potential pollution sources. Visualize the results of the APCS-MLR model, such as... Figure 4 As shown: Compared to the UNMIX model, the APCS-MLR model tends to group the sources of As, Cu, and Zn together. Cd is not considered a unique pollution source in its results either. However, considering the low goodness of fit (0.602 and 0.638) for Cd and As in the APCS-MLR model, its analysis of Cd and As is unreliable. The UNMIX model, on the other hand, shows a goodness of fit of 0.41 for Pb. Therefore, it is necessary to use third-party data summaries, specifically a summary table of heavy metal elements detected in the pollution source, as shown in Table 5. Table 5: Summary of Heavy Metal Element Content in Pollution Sources;
[0055] Among the pollution sources, the contents of As and Pb are very different from those of Cu and Zn, and the contents of Cd are also different from those of Cr and Ni. After comparison and verification, the results of the UNMIX model are more consistent with the actual situation. However, the Pb element in the model results is classified into the same category as Cu and Zn. This is because heavy metal pollution comes from multiple pollution sources. The pollution sources of Pb, Cu, and Zn partially overlap, but other pollution sources occupy the dominant position of Pb and Cu and Zn pollution, respectively. S4: Integrating input flux models, receptor models (contribution rate of pollutants), actual survey results, and GDM model results from each zone, conduct in-depth analysis of pollution sources: S401: Calculate the heavy metal input flux of pollution source samples in the study area, obtain the proportion of input flux of different heavy metals for each pollution source, verify with APCS-MLR, UNMIX models, and actual survey results, preliminarily determine the actual pollution sources, and visualize the input flux results of each pollution source, such as... Figure 5 As shown.
[0056] Figure 5 The input flux results for known pollution sources in the overall study area are provided. Except for As, atmospheric dustfall is included as a pollution source for all other heavy metals. Livestock manure also accounts for a significant proportion of the pollution sources for Hg, Cu, Zn, and Ni, and this cannot be ignored. S402: The UNMIX model resolves more factors than the APCS-MLR model (known sources), indicating that its results include some unknown sources lost by the APCS-MLR due to PCA defects. The resolution results are reasonable. S403: The input flux model and the receptor model show significant differences in the allocation of heavy metal sources, indicating that a single pollution source cannot represent the source of a certain heavy metal pollution in the study area. It is necessary to calculate the heavy metal input flux of pollution sources in each survey area to identify the main pollution sources of a certain heavy metal in different zones, thus compensating for the shortcomings of the receptor model and the overall regional input flux calculation. Theoretically, this step requires calculating the input flux of heavy metals from each pollution source in the 15 zones of the study area, as shown in the diagram. Figure 6 As shown.
[0057] Taking Pb (which has poor resolution in the UNMIX model), Cd and As (which have poor resolution in the APCS-MLR model), and Hg (which has a good fit of 0.8 or higher in both models) as examples, geostatistical interpolation was used to plot the spatial distribution of the eight heavy metal elements.
[0058] Figure 7The spatial distribution of As (As) differs significantly from that of Cu (Cu) and Zn (Zn). Areas exceeding As pollution standards are concentrated in the southern part of Yiling District, Yidu City, and the eastern part of Longzhouping in Changyang County. In these areas, we calculated the Coefficient of Performance (COP) ratio of heavy metal input flux from each pollution source in the corresponding sub-region to the total heavy metal input flux from all pollution sources in the overall study area. Only the COP of irrigation water in Yidu City was 6.381, significantly higher than the overall heavy metal input flux from irrigation water in the study area. This indicates that... Figure 5 The source of pollution of As (irrigation water) in the middle is Yidu City, but the COP of irrigation water in other areas is much lower than 1, and only the input flux COP of livestock and poultry manure is close to 1, indicating that other sources of pollution of As (livestock and poultry manure) are from Yiling and Changyang areas. Unlike the analysis of As, Cd levels in nearly half of the study area exceeded risk control limits, particularly in Wufeng Fujiayan, Changyang, Yidu, and Dianjun areas. Figure 7 Among them, pollution is the most serious. After calculating their comparison coefficients, the COP of atmospheric deposition in Dianjun area and Yidu City is greater than 1, at 1.274 and 3.727 respectively. The COP of livestock and poultry manure in Fujiayan Town of Wufeng County and Longzhouping Town of Changyang County is greater than 2, at 4.102 and 2.601 respectively, indicating that Cd pollution is not as severe as expected. Figure 5 The pollution sources shown are atmospheric dustfall, but different zones have different main pollution sources. The source can be determined based on the COP of the pollution sources in different zones. Finally, there is the source analysis of Pb and Hg. Both of these elements, whether from the detection concentration of heavy metals in the soil or from their pollution distribution map, indicate that the pollution originated from the Xingshan County area. However, due to insufficient funds, only irrigation water samples were collected from this area, and only the Cu element irrigation water input flux COP reached 0.958, which was insufficient to analyze the source of Pb and Hg. S404: Due to insufficient field survey data, the input flux calculation data could not support the analysis of all heavy metal sources. Using the GDM model factor detector, based on the S1 dataset, we explored the spatial heterogeneity factors and mechanisms of heavy metals, analyzed unknown pollution sources, and the factor detector in the geographic detector could determine the degree of influence of each explanatory variable on the dependent variable. The results are as follows: Figure 8 As shown.
[0059] Figure 8Among the environmental variables, all p-values were less than 0.05. Among them, population density (POP) had the highest explanatory power for Hg, with a q value of 0.395, while the distance between the sampling point and the road (LOAD) had the highest explanatory power for Pb. These two variables with the highest explanatory power can be used to trace the sources of Pb and Hg. For example, population density increases the amount of various types of garbage (Hg pollution), while the traffic of cars leads to wear and tear on brakes, tires, and lubricating oil (Pb pollution). S5: Based on the S404 environmental dataset, the data is optimized through geostatistical interpolation, receptor model, and geographic detector. The improved Sparrow Optimization Algorithm (SSA) is then used to optimize the Random Forest (RF) model to predict the spatial distribution map of heavy metals in the soil. The specific steps are as follows: S501: The environmental dataset contains data indirectly affecting the distribution of heavy metals in soil. Spatial heterogeneity is ignored. Kriging interpolation, inverse distance weighted interpolation, and empirical Bayesian interpolation are used to obtain a rough spatial distribution of heavy metals. Then, the spatial distribution of the eight heavy metals, including their concentrations, is divided into nine levels for each interpolation method. The average concentration range of each level is then taken and rounded to integers to improve data continuity. Taking Cd and As, the most polluted metals in the study area, as examples, the original and processed spatial distribution maps of their inverse distance weighted interpolation are shown below. Figure 9 As shown.
[0060] Figure 9 The values of heavy metal elements are expanded values to facilitate concentration stratification in ArcGIS software. However, the applicability of this method is limited. If the prediction accuracy of geostatistical interpolation is high based on the collected soil heavy metal dataset, the spatially stratified dataset processing can represent most of the spatial characteristics of heavy metal elements in the dataset. However, if the prediction accuracy of geostatistical interpolation is not high enough, the spatial characteristics of heavy metals in the spatially stratified dataset will be greatly weakened. At this point, the results of the Optimal Acceptor Model (UNMIX) are needed. The UNMIX model, after factor analysis of the soil heavy metal dataset, yields a factor contribution matrix that shows the weighted proportion of each pollutant's contribution to each sample. In the UNMIX model results, factors 2 and 3 contribute significantly to Cd and As. Inverse distance weighted interpolation of these two factors yields the following results: Figure 10 As shown.
[0061] Figure 10 The weighted contribution distributions of factors 2 and 3 to heavy metals Cd and As are similar to the spatial distributions of Cd and As. By integrating environmental data, spatial hierarchy data, and spatial weight data related to heavy metals into one dataset, the spatial distribution of heavy metals in soil can be accurately predicted. S502: Environmental data collection relies on past experience and its real-time performance is questionable. In S404, there are 17 environmental factors used to analyze the dependent variable, but few of them meet the minimum requirements of the GDM model (q<0.05). One reason is that the actual pollution sources differ from past research experience, and another is that the real-time performance of the data is insufficient, which limits the model results. Therefore, it is necessary to use the GDM factor detector model to screen environmental data that contribute to heavy metal elements and have high real-time performance, and delete data with weak contributions to reduce interference. S503: Set the upper and lower limits of the hyperparameters of the random forest. The parameters include the number of decision trees in the random forest, the maximum depth of the decision trees, the minimum number of samples required for node splits, the minimum number of samples required for leaf nodes, and the number of features considered when searching for the best split. Based on SSA, the hyperparameters of RF are optimized to minimize the loss function. The optimal combination of hyperparameters is automatically searched, and the root mean square error (RMSE), mean absolute error (MAE), and goodness of fit (R^2) generated during random forest prediction are used as evaluation indicators of the loss function. S504: Setting the SSA optimization algorithm process: a. Determine the reasonable range of values for hyperparameters; b. Initialize the sparrow population, with each individual representing a combination of hyperparameters; c. Construct and train an RF model using various hyperparameter combinations; d. Evaluate the model performance on the validation set and calculate the loss function value; e. Update the sparrow's position and adjust the hyperparameter combination; f. Iterate and optimize until the conditions are met; g. Output the optimal combination of hyperparameters that minimizes the loss function; S505: Optimized SSA Algorithm a. Use chaotic mapping to generate the initial population to enhance ergodicity and randomness, and avoid local optima; b. Improve the discoverer update strategy by introducing a dynamic exponential decay factor and adding random perturbations to enhance exploration when the warning value is large; c. Improve the follower update strategy and adaptively adjust the step size based on individual rankings; Taking the Cd element as an example, the results of the original SSA and the improved SSA algorithms are compared, and the results are shown in Table 6: Table 6: Cd element prediction results based on multiple algorithms;
[0062] SSA-RF: Original SSA optimized Random Forest (RF); ISSA-RF: Improved SSA algorithm optimized RF; GDM-ISSA-RF: After removing interference variables using GDM, the RF is optimized using an improved SSA algorithm; GSR-ISSA-RF: Adds spatially layered data to GDM-ISSA-RF; GSW-ISSA-RF: Based on GSR-ISSA-RF, spatial weight data is added; As can be seen, after optimizing the dataset using spatial data and the GDM model, and searching for the optimal hyperparameters of the random forest using the improved SSA algorithm, good results can be achieved in predicting the spatial distribution of Cd elements. S506: Using the optimized SSA, the optimal hyperparameters for RF were found, and the heavy metal content was predicted using the optimal hyperparameters for the eight heavy metal elements. The results are shown in Table 7. Table 7: Optimal prediction results for 8 heavy metal elements;
[0063] Based on an environmental dataset containing spatial properties, the spatial distribution map of heavy metals in soil was accurately predicted. The prediction maps were compared with those obtained through machine learning and geostatistical interpolation, and the one with higher accuracy was used for heavy metal source apportionment. Taking Cd, the most heavily polluted element, as an example, its predicted spatial distribution map is shown below. Figure 11 As shown: Comparing the Cd element prediction map obtained by statistical interpolation with the factor 2 interpolation map of the UNMIX model, the spatial distribution of Cd elements predicted by the GSW-ISSA-RF model is more refined. (Comparison with the Kriging interpolation map of Cd elements is also provided.) Figure 11 The pollution in Sandouping Town and its surrounding areas in Yiling District, eastern Wufeng County, and southern Yuan'an County was highlighted. By combining the COP coefficient of the pollution source input flux in these areas, the pollution sources of Cd can be further explored. The above embodiments are merely preferred technical solutions of the present invention and should not be considered as limitations on the present invention. The embodiments and features described in these embodiments can be arbitrarily combined without conflict. The scope of protection of the present invention should be limited to the technical solutions described in the claims, including equivalent substitutions of the technical features described in the claims. That is, equivalent substitutions and improvements within this scope are also within the scope of protection of the present invention.
Claims
1. A combined method for source apportionment and prediction of heavy metals in cross-regional farmland soil based on multiple models, characterized by: Includes the following steps: S1, Raw dataset collection and preprocessing; S2, Principal component analysis of soil heavy metal data was performed using the PCA model; S3. The APCS-MLR model and the UNMIX model were used to perform quantitative source apportionment of soil heavy metal concentrations. S4. Combining the input flux model results of each region, the contribution rate of pollutants resolved by the receptor model, the actual survey situation, and the GDM model results, a specific analysis of pollution sources is conducted. S5, based on the environmental dataset in S4, uses geostatistical interpolation, receptor model and geographic detector to optimize the dataset, and optimizes the random forest model RF based on the improved sparrow optimization algorithm, so as to predict the spatial distribution map of heavy metals in soil.
2. The combined method for source apportionment and prediction of heavy metals in cross-regional farmland soil based on multiple models as described in claim 1, characterized in that: The specific method for step S1 is as follows: S101. Divide the study area into multiple zones where samples can be collected. Collect soil samples from each zone and send them to the laboratory to test the heavy metal content. Based on previous experience in soil heavy metal source apportionment studies, identify potential pollution sources and pathways in the soil. Collect samples from each pollution source in each zone and send them to the laboratory to test the heavy metal content. Collect environmental data that may indirectly affect changes in soil heavy metal content in the current year from the internet as the basic dataset for heavy metal prediction. S102, use Excel software to normalize the heavy metal concentration data detected in the soil samples to obtain the normalized dataset corresponding to the concentration dataset. Use Excel software to calculate the summary table of heavy metals in potential pollution sources, including the maximum, minimum and median values.
3. The combined method for source apportionment and prediction of heavy metals in cross-regional farmland soil based on multiple models as described in claim 1, characterized in that: The specific method for step S2 is as follows: S201. The detected soil heavy metal content was imported into SPSS software for factor dimensionality reduction. The KMO test and Bartlett's test of sphericity were used to determine whether the data was suitable for principal component analysis. S202. After completing the applicability test of factor analysis and selecting the dataset that meets the conditions, principal component analysis is used to reduce the dimensionality of soil heavy metal data. The factor structure is optimized by the variance maximum orthogonal rotation technique. Finally, a visualization result chart containing eigenvalues, variance contribution rate and cumulative contribution rate is generated. At the same time, a principal component analysis result table is output to clarify the number of effective principal components extracted and their explanatory power. Select datasets that meet the criteria for factor analysis based on the test results, perform principal component analysis, and generate plots of soil heavy metal eigenvalues and cumulative variance contribution rates after maximum variance orthogonal rotation. Obtain the principal component analysis result table and the corresponding number of principal components.
4. The combined method for source apportionment and prediction of heavy metals in cross-regional farmland soil based on multiple models as described in claim 1, characterized in that: The specific method for step S3 is as follows: S301. Import the preprocessed normalized data from S1 into EPA Unmix 6.0 software. According to the software user manual, the normalized data should ideally include the total concentration normalized data of heavy metals. In the lower right corner of the main window, you can track the status of species as specified as "Total", "Tracer", or "Norm". S302, in the built-in functions of Unmix software, select the function of automatically generating initial factors. If the original dataset cannot automatically initialize the number of factors, the number of principal components parsed by the PCA model is used as the minimum number of factors. Manually set the number of factors required for the model to run, and then set the model to perform calculations. S303 sets the criteria for determining the optimal pollutant. S304, the number of factors that meet the results of S303 after the statistical software is run; S305. Using the built-in functions of EPA Unmix 6.0 software, we obtained the contribution of each factor to various metals and the fitting results of the actual and predicted values of each heavy metal after the model was run. In EXCEL software, we used the source component spectrum obtained by running EPA Unmix 6.0 software to calculate the contribution ratio of each pollutant factor to the overall soil heavy metal concentration. S306. Based on the number of factors resolved by the UNMIX model, select an appropriate number of principal components. In SPSS software, calculate the principal component scores obtained after PCA analysis. Add data "0" to the last row of the soil heavy metal dataset. Then, perform PCA analysis using the soil heavy metal dataset with 0 to obtain the principal component scores of the row with data "0". In EXCEL, subtract the principal component scores of the row with data "0" from the principal component scores of each soil heavy metal sample to obtain the absolute principal component scores of each soil heavy metal sample. S307: Add the absolute principal component scores to the soil heavy metal dataset in SPSS, and use the software's built-in linear regression function to calculate the goodness of fit between heavy metal concentration and absolute principal component scores. If the goodness of fit is too low, it is necessary to appropriately add principal components with eigenvalues less than but close to 1 from the PCA analysis and recalculate S306-S307. S308, calculate the average value of the absolute principal components as the basis for calculating the contribution of pollutants to heavy metals; S309, the APCS-MLR model can calculate known pollution sources containing principal components, as well as unknown pollution sources not included in the principal components by PCA analysis; S310, Calculate the contribution of pollution sources according to the contribution calculation formula of known pollution sources and unknown pollution sources; S311. Compare the final results of the APCS-MLR and UNMIX models. If there are differences between the results of the two models, select the best model result by combining the summary table of heavy metal concentrations in potential pollution sources.
5. The combined method for source apportionment and prediction of heavy metals in cross-regional farmland soil based on multiple models as described in claim 4, characterized in that: The criteria for determining the optimal pollutant in S302 are as follows: a. Of all the sampled data points input into the model, at least 80% of the data have residual analysis results between -3 and +3 and conform to a normal distribution; b. The fitting coefficient R² is greater than or equal to 0.8; c. Minimum signal-to-noise ratio greater than or equal to 2.
6. The combined method for source apportionment and prediction of heavy metals in cross-regional farmland soil based on multiple models as described in claim 1, characterized in that: The specific method for step S4 is as follows: S401. First, calculate the input flux of heavy metals in the pollution source samples collected throughout the study area to obtain the proportion of input flux of each heavy metal for each pollution source. Then, verify the results with the APCS-MLR and UNMIX model analysis and the actual survey results to make a preliminary judgment on the real pollution sources. S402. When the number of factors resolved by the UNMIX model exceeds the number of "known sources" in the APCS-MLR model, it means that the results resolved by the UNMIX model include the factors that the APCS-MLR model lost due to the defects of PCA analysis, i.e., unknown sources, and the results resolved by the UNMIX model are reasonable. S403. If the input flux results differ significantly from the receptor model results, it indicates that for the entire study area, a single pollution source cannot represent the pollution source of a certain heavy metal in the entire study area. There must be other pollution sources that affect the distribution of the heavy metal. If this occurs, calculate the input flux of heavy metals in pollution sources in each survey area, and calculate the comparison coefficient COP between the input flux of heavy metals in different pollution sources in each survey area and the input flux of heavy metals in different pollution sources in the entire study area. This will determine that a certain heavy metal in the entire study area has different main pollution sources in different zones, thus achieving cross-regional pollution source identification. S404 If the data from the field survey is insufficient, and the computational data of the input flux is insufficient to support the source analysis of all types of heavy metals in the study, then the factor detector model in the GDM model needs to be used. It can study the factors and potential mechanisms that affect the spatial heterogeneity of heavy metals, explore the contribution of a certain influencing factor to the distribution of heavy metal pollution, and use the datasets collected in S1 that may have an indirect impact on the changes in soil heavy metal content to analyze unknown pollution sources that still cannot be analyzed after input flux.
7. The combined method for source apportionment and prediction of heavy metals in cross-regional farmland soil based on multiple models as described in claim 6, characterized in that: The key computational metrics for the GDM model are the q-value and the p-value: The q-value quantifies the explanatory power of various influencing factors on the target variable by comparing the ratio of in-stratum variance to population variance. The p-value verifies statistical reliability. In practical applications, a significance test is often constructed using Monte Carlo simulation. When p < 0.05, the factor can be considered to have statistical significance in explaining spatial differentiation.
8. The combined method for source apportionment and prediction of heavy metals in cross-regional farmland soil based on multiple models according to claim 7, characterized in that: The specific method for step S5 is as follows: S501 employs two measures when optimizing the dataset: first, it uses geostatistical interpolation to obtain a rough spatial distribution of heavy metals; second, it uses the results of the receptor model as spatial weight data for geostatistical interpolation, and adds the results of both to the environmental dataset to give it spatial heterogeneity. When using S502, the factor detector model of GDM is first used to explore the environmental data that contributes to each heavy metal element. When predicting the distribution of heavy metals, environmental data that does not contribute significantly are deleted to reduce data interference. S503, a hyperparameter optimization method for random forests based on the Sparrow Search Algorithm (SSA), optimizes hyperparameters by automatically searching within a given upper and lower limit range of hyperparameters using the Sparrow Search Algorithm, aiming to minimize the loss function and find the hyperparameter combination that optimizes model performance. S504, the general process for setting the optimization algorithm; S505, optimized sparrow search algorithm; S506: After finding the optimal hyperparameters of the random forest model using the optimized sparrow search algorithm, the spatial distribution map of heavy metals in the soil is accurately predicted based on the environmental dataset with spatial properties. The spatial distribution map predicted by machine learning and the geostatistical interpolation prediction map are compared, and the prediction map with higher accuracy is used for heavy metal source analysis.
9. The combined method for source apportionment and prediction of heavy metals in cross-regional farmland soil based on multiple models as described in claim 8, characterized in that: The general process of setting an optimization algorithm includes: a. Define the range of hyperparameters: Based on experience or preliminary experiments, determine the reasonable range of values for each hyperparameter; b. Initialize the sparrow population: Each individual sparrow represents a set of hyperparameter combinations; c. Constructing a random forest model: For each individual sparrow, construct a random forest model using its corresponding hyperparameter combination and train it on the training set; d. Evaluate model performance: Evaluate the model's performance on the validation set and calculate the loss function value; e. Update sparrow locations: According to the rules of the sparrow optimization algorithm, update the locations of the sparrow population, i.e., adjust the hyperparameter combination; f. Iterative optimization: Repeat step ce until the preset maximum number of iterations is reached or the convergence condition is met; g. Output optimal hyperparameters: Select the combination of hyperparameters that minimizes the loss function from the final sparrow population as the optimal solution.
10. The combined method for source apportionment and prediction of heavy metals in cross-regional farmland soil based on multiple models according to claim 9, characterized in that: Optimizing the sparrow search algorithm includes: a. Generate the initial population through chaotic mapping; b. Improved producer update strategy: By introducing a dynamic exponential decay factor, the position update strategy of the discoverer emphasizes exploration in the early iterations and development in the later iterations. When the warning value is large, a normally distributed random perturbation is added. c. Improved follower update strategy: The step size is adaptively adjusted according to the individual's ranking in the population. For individuals with poor fitness (i > self.pop / 2), an exponential decay method is used to move closer to the current worst individual. For individuals with good fitness, an update strategy based on the best individual and a random matrix is used.