Machine learning-based soil heavy metal pollution source identification and risk assessment method and system

By using machine learning-based methods, a soil heavy metal pollution source identification and risk assessment system was constructed. This system solves the accuracy problem of traditional methods in separating pollution sources and assessing risks, and realizes quantitative identification and risk assessment of soil heavy metal pollution sources, providing a scientific basis for decision-making.

CN122134133APending Publication Date: 2026-06-02山东省地质调查院(山东省自然资源厅矿产勘查技术指导中心)

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
山东省地质调查院(山东省自然资源厅矿产勘查技术指导中心)
Filing Date
2026-03-06
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately separate the contributions of different pollution sources to heavy metal elements in soil, especially in cases of complex mixed pollution or highly overlapping pollution source characteristics. The identification results are unstable and lack quantitative characterization of the natural background, affecting the accuracy of pollution risk assessment.

Method used

Using a machine learning-based approach, a heavy metal element content input matrix is ​​constructed, and k-means clustering analysis, multi-parent decomposition, and principal component analysis are performed to decompose it into a source spectrum matrix and a source contribution matrix. By combining conservative elements and time series data, the background contribution value of each pollution source is extracted, and a dynamic risk classification threshold is introduced to achieve adaptive adjustment of pollution risk.

Benefits of technology

It enables quantitative identification and risk assessment of soil heavy metal pollution sources, distinguishes between natural background and anthropogenic pollution, improves the accuracy of pollution source identification and the scientific nature of risk assessment, and provides a visualized decision-making basis for continuous spatial distribution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122134133A_ABST
    Figure CN122134133A_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for identifying and assessing soil heavy metal pollution sources based on machine learning. The method constructs a feature dataset by collecting data on heavy metal element content and spatial environmental attributes of soil in a target area; it decomposes the input matrix using multivariate analysis and k-means clustering to obtain a source spectrum matrix and a source contribution matrix, thus identifying pollution sources and analyzing their contributions; further, it establishes a background contribution model by combining conservative elements and time series information to extract the background contribution value of each pollution source within the target time period; it calculates single-factor and comprehensive risk indices based on pollution source reconstruction and background correction, and generates soil heavy metal pollution risk level information by combining dynamic risk grading thresholds; finally, it generates a soil heavy metal pollution risk distribution map of the target area through spatial interpolation and zoning processing, simultaneously achieving quantitative analysis of pollution sources, background value correction, risk index calculation, and spatial visualization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of environmental monitoring and soil pollution prevention and control, and in particular to a method and system for identifying and assessing the risk of heavy metal pollution sources in soil based on machine learning. Background Technology

[0002] Heavy metal pollution in soil is a globally significant environmental issue, primarily stemming from various anthropogenic activities such as industrial emissions, agricultural fertilization, waste landfilling, and transportation. Heavy metals in soil are difficult to degrade and readily migrate and accumulate, potentially harming ecosystems and human health through the food chain. Therefore, the identification, quantitative analysis, and risk assessment of heavy metal pollution in soil are of paramount importance.

[0003] Currently, methods for identifying and assessing the risk of heavy metal pollution sources in soil mainly include chemical analysis, statistical models, and empirical judgment. While these methods can analyze the distribution patterns and potential sources of pollutants to some extent, they still suffer from the following technical problems: Traditional methods struggle to accurately separate the contributions of different pollution sources to various heavy metal elements in the soil when dealing with multiple pollution sources, especially in cases of complex mixed pollution or highly overlapping pollution source characteristics, leading to unstable identification results. Existing methods often lack quantitative characterization of the natural background contribution of pollution sources, making it difficult to distinguish between natural background and anthropogenic pollution, thus affecting the accurate assessment of pollution risk. Soil heavy metal pollution exhibits significant spatial heterogeneity and temporal dynamic changes. Traditional statistical methods mostly rely on static averages, ignoring the temporal information and spatial distribution characteristics of sampling points, resulting in limited applicability of pollution risk assessment results across different regions or time periods.

[0004] To address the above issues, a technical solution is needed that can simultaneously achieve pollution source identification, background contribution prediction, dynamic risk index calculation, and spatial risk zoning. Summary of the Invention

[0005] The purpose of this invention is to provide a method and system for identifying and assessing soil heavy metal pollution sources based on machine learning, in order to solve the aforementioned technical problems. This method achieves quantitative identification and pollution risk assessment of soil heavy metal pollution sources, distinguishes between natural background and anthropogenic pollution, introduces a dynamic risk grading threshold, and allows the pollution risk level to adaptively adjust with changes in soil physicochemical properties and pollution source background characteristics. Based on spatial interpolation and risk zoning, a continuous pollution risk distribution map is generated, enabling visualization of soil heavy metal pollution risk and providing a scientific and intuitive basis for pollution control and environmental management.

[0006] The objective of this invention can be achieved through the following technical solutions: A machine learning-based method for identifying and assessing the risk of heavy metal pollution sources in soil, comprising: S1. Obtain soil heavy metal element content data and spatial and environmental attribute data corresponding to the sampling points within the target area; preprocess the data to construct a feature dataset for pollution source analysis. S2. Based on the feature dataset, construct a heavy metal element content input matrix, and decompose the input matrix using multivariate analysis to obtain a source spectrum matrix characterizing the characteristics of different pollution sources and a source contribution matrix characterizing the degree of pollution source contribution at each sampling point, thereby realizing the identification and quantitative analysis of soil heavy metal pollution sources. S3. Based on the source contribution matrix, extract the contribution data of each pollution source in the target time period, establish a background contribution model by combining conservative element information, and calculate the background contribution value of each pollution source in the target time period. S4. Based on the source spectrum matrix, source contribution matrix and background contribution value, the pollution source reconstruction and background correction are performed on the heavy metal element content at each sampling point, and the pollution risk index of each heavy metal element at the sampling point is calculated. S5. Based on the pollution risk index and combined with the dynamic risk classification threshold, classify and determine the risk of heavy metal pollution in the soil at the sampling points, and generate corresponding risk level information. S6. Based on the risk level information and the spatial location information of the sampling points, generate the spatial distribution results of soil heavy metal pollution risk to characterize the spatial differentiation characteristics of soil heavy metal pollution risk in the target area.

[0007] Furthermore, sampling points were set up in the target area to collect surface soil samples and determine the content of heavy metal elements in the soil samples. Record the geographic coordinates of the sampling points and obtain the terrain parameters and land use type information of the sampling points; The raw data includes parameters such as heavy metal element content, geographic coordinates, topography, and land use type. Missing values ​​were handled in the original data; noise was identified and removed from the original data; and the heavy metal element content parameters were standardized. Construct feature variables including geographic coordinate parameters, terrain parameters, and land use parameters.

[0008] Furthermore, an input matrix is ​​constructed based on the heavy metal element content parameters: in, This represents the content parameter value of the j-th type of heavy metal element at the i-th sampling point. , ; Let the number of pollution source categories be K; Perform k-means clustering analysis on the input matrix X to divide the sampling points into K clusters. Calculate the mean of each cluster for each heavy metal element to form the initial source spectrum matrix elements: in, Let k represent the set of samples in the k-th cluster. The number of samples in the sample set is represented by the initial source spectrum matrix. The source spectrum matrix is ​​constructed using it as the initial input for multi-parent decomposition: in, This represents the characteristic combination of the k-th type of pollution source with the j-th type of heavy metal element. ; Standardize the input matrix X: in, Let be the mean value of the j-th type of heavy metal element. Standard deviation Calculate the standardized matrix for the standardized data. covariance matrix ; The covariance matrix is ​​analyzed through eigenvalue decomposition. The data is decomposed, and eigenvalues ​​and eigenvectors are output. The eigenvectors represent the principal component directions of the data, and the eigenvalues ​​represent the variance contribution of each principal component. Based on the magnitude of the eigenvalues, the first K principal components are selected to construct the principal component matrix V; using the principal component matrix V and the standardized data matrix... Calculate the contribution matrix: ; in, This represents the contribution of the k-th type of pollution source at the i-th sampling point; Calculate the error function based on the current values ​​of the contribution matrix C and the source spectrum matrix S. , The Frobenius norm is used to adjust the source spectrum matrix S using gradient descent to minimize this error; when the error function If the number of iterations falls below the set threshold or reaches the preset maximum number, the iteration stops.

[0009] Furthermore, the contribution data of all pollution sources at different time periods are extracted through the source contribution matrix C; Determine the target time period Based on the pollution source contribution data within the target time period, the contribution sequence for the corresponding time period is extracted from the source contribution matrix. ; By filtering the extracted time-period data, high-frequency fluctuations are removed, and a smooth background contribution sequence is output. Regression analysis was used to fit the smoothed background contribution sequence with the corresponding conservative element to establish a background contribution sequence model, which reflects the changing trend of pollution sources in the target period. Based on the background contribution sequence model, the background contribution value of each pollution source in the target time period is output.

[0010] Furthermore, the method for extracting contribution data of all pollution sources at different time periods through the source contribution matrix includes: grouping the sampling points according to a preset time interval based on the sampling time information of the sampling points, constructing a correspondence between sampling points and time periods, and determining a target time period within the time periods. The target time period is a continuous time interval. Or a time set consisting of multiple discrete sampling times; within each time period, select the set of sampling points belonging to that time period. Extract the corresponding row vector from the source contribution matrix C. This forms a set of contribution data for the k-th type of pollution source during that time period; Statistical calculations are performed on the contribution data set of the same pollution source within the same time period to obtain the time period contribution value of the pollution source within the time period. The statistical calculations include at least one of the following: mean, median, or weighted average. For the target time period Select the set of sampling points belonging to the target time period. Then, the corresponding source contribution amount is extracted from the source contribution matrix to construct the contribution amount sequence of the k-th type of pollution source in the target time period. in, This represents the contribution of the i-th sampling point to the k-th type of pollution source during the target time period.

[0011] Furthermore, the smoothing process employs a moving average method: in, This represents the background contribution value of the k-th type of pollution source during the target time period, where n is the number of sampling points during the target time period; a conservative element content sequence corresponding to the target time period is selected. , The smoothed background contribution sequence is fitted to the conservative element using regression analysis to establish a background contribution sequence model. The regression relationship of the background contribution sequence model is expressed as follows: in, and For regression coefficients, For residual terms; Based on the background contribution sequence model, calculate the background contribution value of the k-th pollution source during the target time period: in, and The regression coefficient represents the background contribution of the k-th pollution source during the target time period. This represents the statistical mean of the content of conservative elements within the target time period.

[0012] Furthermore, for any sampling point i and any heavy metal element j, the pollution source reconstruction value of the heavy metal element at that sampling point is calculated based on the source contribution matrix C and the source spectrum matrix S: in, This represents the content of the j-th type of heavy metal element at the i-th sampling point due to the contribution of the pollution source; based on the reconstructed pollution source value and the background contribution value of each pollution source during the target time period, the net pollution contribution of the j-th type of heavy metal element at the sampling point is calculated: Based on the net pollution contribution For any sampling point i, the j-th type of heavy metal element... = , Let be the soil environmental quality standard limit value corresponding to the j-th heavy metal element; when When ≤0, it indicates that the heavy metal element is at the natural background level; when A value greater than 0 indicates that the heavy metal element poses a pollution risk to the soil environment, and the higher the value, the higher the pollution risk.

[0013] Introducing a comprehensive risk index : m represents the number of heavy metal elements involved in the risk assessment. Let be the weighting coefficient of the j-th heavy metal element, used to reflect the relative risk contribution of different heavy metal elements to the soil environment and human health, and satisfying the following: Based on the single-factor risk index and comprehensive risk index The pollution risk of sampling points is classified into risk levels according to risk grading thresholds, including low risk, medium risk and high risk.

[0014] Furthermore, the risk classification threshold is a dynamic threshold, which includes a first risk threshold. Second risk threshold First risk threshold Second risk threshold The method for determining it is as follows: Based on the benchmark risk thresholds for corresponding heavy metal elements in the soil environmental quality standards , Based on this, a threshold adjustment factor is introduced to correct the baseline risk threshold, resulting in a dynamic risk grading threshold, wherein: in, The standardized threshold adjustment factor includes soil physicochemical property parameters and pollution source background contribution characteristic parameters. The soil physicochemical property parameters include at least one of soil pH value and organic matter content, and the pollution source background contribution characteristic parameters include at least the background contribution values ​​of each pollution source. Let be the weighting coefficients of the corresponding threshold adjustment factor, and satisfy . ; Based on the dynamic risk grading threshold, the first risk threshold Second risk threshold The risk index is classified and determined to generate soil heavy metal pollution risk level information that is adaptively adjusted according to changes in soil environmental conditions and pollution source background characteristics.

[0015] Furthermore, the spatial location information of each sampling point and the corresponding soil heavy metal pollution risk level information are obtained. The risk level information is the risk level result calculated based on the source contribution matrix, background contribution value, risk index and dynamic risk classification threshold. The spatial location information of the sampling points is spatially associated with the corresponding risk level information to construct a spatial mapping relationship between sampling points and risk levels. Based on the spatial mapping relationship, a spatial interpolation method is used to convert the risk level information of discrete sampling points into continuous spatial distribution data. The spatial interpolation or spatial expansion method includes inverse distance weighting, kriging interpolation, or a combination thereof. Based on continuous spatial distribution data, the target area is divided into risk level zones, which are at least low-risk, medium-risk and high-risk areas. The risk level zoning results are overlaid with the spatial base map in the geographic information system to generate a soil heavy metal pollution risk zoning map of the target area. The risk zoning map is used to characterize the spatial distribution characteristics of soil heavy metal pollution risk in different areas.

[0016] Furthermore, a machine learning-based system for identifying and assessing the risk of heavy metal pollution sources in soil includes: The data acquisition and preprocessing module is used to acquire soil heavy metal element content data of sampling points in the target area, as well as the geographic coordinates, terrain parameters and land use type parameters corresponding to the sampling points. The module performs missing value processing, noise identification and removal and standardization on the data to construct a feature dataset for pollution source analysis. The pollution source analysis module is used to construct a heavy metal element content input matrix based on the feature dataset, decompose the input matrix using a multivariate analysis method, and output a source spectrum matrix that characterizes the combination of heavy metal element characteristics of different pollution sources and a source contribution matrix that characterizes the degree of pollution source contribution at each sampling point. The background contribution modeling module is used to extract the contribution data of each pollution source in the target time period based on the source contribution matrix, perform time period statistics and smoothing on the contribution data, and establish a background contribution model in combination with conservative element information to calculate the background contribution value of each pollution source in the target time period. The risk index calculation module is used to reconstruct the pollution source and correct the background of the heavy metal element content at each sampling point based on the source spectrum matrix, source contribution matrix and background contribution value, and to calculate the single-factor risk index and comprehensive risk index of each heavy metal element at each sampling point. The dynamic risk classification module is used to classify and determine the risk of heavy metal pollution in the soil at each sampling point based on the risk index and the dynamic risk classification threshold constructed based on the soil environmental quality standard, and generate corresponding pollution risk level information. The risk zoning generation module is used to spatially correlate the pollution risk level information with the spatial location information of the sampling points, generate a continuous spatial distribution result of soil heavy metal pollution risk in the target area through spatial interpolation or spatial expansion methods, and output a soil heavy metal pollution risk zoning map.

[0017] Compared with the prior art, the present invention has the following technical effects: By constructing an input matrix and performing k-means clustering analysis, multi-gene decomposition, and principal component analysis, the heavy metal element content at sampling points can be decomposed into a source spectrum matrix and a source contribution matrix, enabling quantitative analysis of different pollution sources and their contributions to soil heavy metal pollution. Combining conservative elements and time-series data, smoothing and regression modeling of pollution source contributions effectively extracts the background contribution values ​​of each pollution source during the target time period, thereby distinguishing between natural background and anthropogenic pollution impacts and improving the accuracy of pollution source identification.

[0018] By introducing a dynamic risk classification threshold, and integrating soil physicochemical properties, background contribution characteristics of pollution sources, and risk index, the risk level can be adaptively adjusted according to changes in soil environmental conditions and pollution source characteristics, thereby improving the scientific rigor and practicality of risk assessment.

[0019] Based on the spatial information and risk level of sampling points, spatial interpolation and risk zoning are used to realize the continuous spatial distribution mapping and visualization of pollution risks, providing an intuitive basis for decision-making in pollution control and environmental management. Attached Figure Description

[0020] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation

[0021] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. These embodiments are based on the technical solution of the present invention and provide detailed implementation methods and specific operating procedures. However, the scope of protection of the present invention is not limited to the following embodiments.

[0022] like Figure 1 The method for identifying and assessing soil heavy metal pollution sources based on machine learning, as shown, includes: S1. Obtain soil heavy metal element content data and corresponding spatial and environmental attribute data for sampling points within the target area. Preprocess the data to construct a feature dataset for pollution source apportionment. This dataset completes data collection and preliminary processing, ensuring a complete and reliable data foundation for subsequent analysis. Spatial attributes (such as coordinates, topography, and land use type) and environmental parameters can be used to distinguish pollution sources and influencing factors. Preprocessing (missing value handling, noise removal, and standardization) ensures data accuracy and comparability. S2. Based on the feature dataset, an input matrix of heavy metal element content is constructed. The input matrix is ​​decomposed using multivariate analysis to obtain a source spectrum matrix characterizing the characteristics of different pollution sources and a source contribution matrix characterizing the contribution of each sampling point to the pollution source, thus realizing the identification and quantitative analysis of soil heavy metal pollution sources. The original data is transformed into matrix form, and multivariate analysis is used to separate the characteristic patterns of different pollution sources. This can not only identify the pollution source type, but also quantify the contribution of each sampling point to the influence of each pollution source, providing a mathematical basis for the quantitative analysis of pollution sources.

[0023] S3. Based on the source contribution matrix, extract the contribution data of each pollution source in the target time period, and establish a background contribution model by combining conservative element information to calculate the background contribution value of each pollution source in the target time period. Through time series extraction and conservative element modeling, distinguish between natural background and anthropogenic pollution, calculate the background contribution value of each pollution source in the target time period, provide a benchmark for pollution risk assessment, and avoid overestimating the pollution level.

[0024] S4. Based on the source spectrum matrix, source contribution matrix, and background contribution value, the heavy metal element content at each sampling point is reconstructed and corrected for pollution sources, and the pollution risk index of each heavy metal element at the sampling point is calculated. Through pollution source reconstruction and background correction, the observed values ​​are decomposed into pollution source contribution and background values, thereby calculating the single-factor risk index and comprehensive risk index, quantifying the pollution level and risk magnitude, and providing data basis for risk level classification.

[0025] S5. Based on the pollution risk index and combined with the dynamic risk classification threshold, the soil heavy metal pollution risk of the sampling points is classified and determined, and the corresponding risk level information is generated. The quantitative risk indicators are converted into risk levels, and the thresholds are automatically adjusted according to environmental conditions and pollution source background to achieve dynamic and adaptive risk classification, which facilitates management and decision-making.

[0026] S6. Based on risk level information and spatial location information of sampling points, generate spatial distribution results of soil heavy metal pollution risk to characterize the spatial differentiation characteristics of soil heavy metal pollution risk in the target area; generate continuous distribution maps of the risk levels of discrete sampling points through spatial interpolation and visualization to intuitively display the spatial distribution of soil heavy metal pollution and high-risk areas, providing decision support for environmental supervision, pollution control and land management.

[0027] Specifically, sampling points are set up in the target area to collect surface soil samples and determine the content of heavy metal elements in the soil samples; Record the geographic coordinates of the sampling points and obtain the terrain parameters and land use type information of the sampling points; This process generates raw data including heavy metal element content parameters, geographic coordinate parameters, topographic parameters, and land use type parameters; it provides spatial attribute information for subsequent pollution source apportionment and spatial risk analysis. Geographic coordinates are used for spatial positioning and interpolation analysis; topographic parameters affect pollution migration and deposition; land use type affects pollutant sources and exposure risks.

[0028] Missing values ​​are handled in the raw data; noise is identified and removed from the raw data; heavy metal element content parameters are standardized; and data loss issues that may occur during sampling or testing are addressed to ensure data integrity and the reliability of subsequent analysis.

[0029] Feature variables, including geographic coordinate parameters, topographic parameters, and land use parameters, are constructed. This forms the input feature set for the pollution source identification model, enabling subsequent matrix factorization and multivariate analysis to comprehensively consider environmental factors and spatial distribution characteristics.

[0030] Specifically, an input matrix is ​​constructed based on heavy metal element content parameters: in, This represents the content parameter value of the j-th type of heavy metal element at the i-th sampling point. , It provides a standardized input data structure for cluster analysis and principal component analysis, enabling the processing of content information of different elements within the same framework. Let the number of pollution source categories be K; Perform k-means clustering analysis on the input matrix X to divide the sampling points into K clusters. Calculate the mean of each cluster for each heavy metal element to form the initial source spectrum matrix elements: in, Let k represent the set of samples in the k-th cluster. Indicates the number of samples in the initial source spectrum matrix. The source spectrum matrix is ​​constructed as the initial input for multi-source decomposition: k-means clustering is used to group sampling points according to heavy metal characteristics, thereby roughly distinguishing possible pollution source types. The initial source spectrum matrix is ​​formed by the mean of each cluster, providing reasonable initial values ​​for multi-source decomposition, improving the iteration convergence speed and avoiding local optima. The initial source spectrum matrix reflects the typical characteristics of each type of pollution source with different heavy metal elements.

[0031] in, This represents the characteristic combination of the k-th type of pollution source with the j-th type of heavy metal element. ; Standardize the input matrix X: in, Let be the mean value of the j-th type of heavy metal element. Standard deviation Calculate the standardized matrix for the standardized data. covariance matrix ; The covariance matrix is ​​analyzed through eigenvalue decomposition. The data is decomposed and outputs eigenvalues ​​and eigenvectors. The eigenvectors represent the principal component directions of the data, and the eigenvalues ​​represent the variance contribution of each principal component. This process eliminates the differences in the dimensions and numerical ranges of different heavy metal elements, ensuring that the results of clustering and principal component analysis are not affected by the scale of a single element.

[0032] Based on the magnitude of the eigenvalues, the first K principal components are selected to construct the principal component matrix V; using the principal component matrix V and the standardized data matrix... Calculate the contribution matrix: ; in, This represents the contribution of the k-th type of pollution source at the i-th sampling point; it eliminates the differences in the dimensions and numerical ranges of different heavy metal elements, ensuring that the results of clustering and principal component analysis are not affected by the scale of a single element. Principal component selection avoids interference from noise and redundant information, improving the accuracy and robustness of pollution source identification.

[0033] Calculate the error function based on the current values ​​of the contribution matrix C and the source spectrum matrix S. , The Frobenius norm is used to adjust the source spectrum matrix S using gradient descent to minimize this error; when the error function The iteration stops when the number of iterations falls below a set threshold or reaches a preset maximum. This yields more accurate pollution source characteristics and contributions. Through iterative optimization, the source spectrum matrix can accurately reflect the distribution characteristics of different pollution sources in terms of heavy metal elements.

[0034] Specifically, the contribution data of all pollution sources at different time periods are extracted through the source contribution matrix C; by grouping the matrix according to time or sampling point attributes, the contribution data set of each pollution source in each time period is obtained.

[0035] Determine the target time period Based on the pollution source contribution data within the target time period, the contribution sequence for the corresponding time period is extracted from the source contribution matrix. Clearly define the time frame for analysis or the sampling period of interest.

[0036] By filtering the extracted time period data to remove high-frequency fluctuations, a smooth background contribution sequence is output; the contribution of the sampling points within that time period is selected from the source contribution matrix to form the target time period sequence.

[0037] Regression analysis was used to fit the smoothed background contribution sequence with the corresponding conservative elements to establish a background contribution sequence model, reflecting the changing trend of pollution sources during the target period. Conservative elements—elements stable under natural conditions and not easily affected by human pollution—were used as reference factors to model the background contribution of pollution sources. Regression analysis correlated the smoothed background contribution with the changes in conservative elements, yielding a predictive model for the background contribution of each pollution source over time.

[0038] Based on the background contribution sequence model, the background contribution value of each pollution source in the target time period is output. The model prediction results are then converted into specific numerical values ​​for subsequent pollution source reconstruction, net pollution calculation, and risk index calculation.

[0039] Specifically, the method for extracting contribution data of all pollution sources at different time periods through the source contribution matrix includes: grouping the sampling points according to a preset time interval based on the sampling time information of the sampling points, constructing the correspondence between sampling points and time periods, and determining the target time period within the time period. The target time period is a continuous time interval. Or a time set consisting of multiple discrete sampling times; within each time period, select the set of sampling points belonging to that time period. Extract the corresponding row vectors from the source contribution matrix C. This process generates a data set representing the contribution of pollution source type k within a given time period. The sampled data is organized by time to form a time series structure, mapping sampling points to time periods and providing temporal data support for subsequent background contribution analysis. The contribution of each sampling point in the matrix is ​​associated with its corresponding time period, forming a time-divided data set of pollution source contributions.

[0040] Statistical calculations are performed on the contribution data set of the same pollution source within the same time period to obtain the time-period contribution value of the pollution source within that time period. The statistical calculations include calculating at least one of the mean, median, or weighted average; the influence of fluctuations in single-point data is eliminated to obtain a representative contribution value within the time period. Stable pollution source contribution indicators are provided to facilitate the construction of background contribution sequences.

[0041] Target time period Select the set of sampling points belonging to the target time period. Then, the corresponding source contribution amounts are extracted from the source contribution matrix to construct the contribution sequence of the k-th type of pollution source during the target time period. The pollution source contribution data for the target time period are precisely selected and a time series is formed. This represents the contribution of the i-th sampling point to the k-th type of pollution source within the target time period. The contributions within the target time period are organized into a sequence.

[0042] Specifically, the smoothing process uses a moving average method: in, This represents the background contribution value of the k-th pollution source during the target time period, where n is the number of sampling points during the target time period; conservative element content sequences corresponding to the target time period are selected. It eliminates the impact of single-point high-frequency fluctuations and obtains representative background contribution values ​​of pollution sources within the target time period; it provides smooth time series data.

[0043] The smoothed background contribution sequence is fitted to the conservative element using regression analysis to establish a background contribution sequence model. The regression relationship of the background contribution sequence model is expressed as follows: in, and For regression coefficients, The residual term is used to link the background contribution with the conservative element through regression fitting, reflecting the changing trend of the pollution source during the target period. Based on the background contribution sequence model, calculate the background contribution value of the k-th pollution source during the target time period: in, and The regression coefficient represents the background contribution of the k-th pollution source during the target time period. This represents the statistical mean of the content of conservative elements within the target time period. It distinguishes the natural background contribution of pollution sources from the actual sampling data, providing a reliable reference for risk assessment.

[0044] Specifically, for any sampling point i and any heavy metal element j, the pollution source reconstruction value of the heavy metal element at that sampling point is calculated based on the source contribution matrix C and the source spectrum matrix S: in, This represents the content of the j-th type of heavy metal element at the i-th sampling point, contributed by the pollution source. The source contribution matrix C and the source spectrum matrix S are combined to reconstruct the source composition of each heavy metal element at each sampling point. Based on the reconstructed pollution source values ​​and the background contribution values ​​of each pollution source within the target time period, the net pollution contribution of the j-th type of heavy metal element at the sampling point is calculated. The reconstructed values ​​are compared with the background contribution values, and the natural background portion is removed to obtain the net pollution amount caused by anthropogenic or exogenous pollution. Based on net pollution contribution For any sampling point i, the j-th type of heavy metal element... = , Let be the soil environmental quality standard limit value corresponding to the j-th heavy metal element; when When ≤0, it indicates that the heavy metal element is at the natural background level; when A value greater than 0 indicates that the heavy metal element poses a pollution risk to the soil environment, and the higher the value, the higher the pollution risk. The net pollution level is compared with environmental standards to quantify the degree of pollution risk.

[0045] Introducing a comprehensive risk index : m represents the number of heavy metal elements involved in the risk assessment. Let be the weighting coefficient of the j-th heavy metal element, used to reflect the relative risk contribution of different heavy metal elements to the soil environment and human health, and satisfying the following: Based on single-factor risk index and comprehensive risk index The pollution risk of sampling sites is classified into low, medium, and high risk levels according to risk grading thresholds. This converts the quantitative risk index into understandable level information, facilitating pollution control decision-making.

[0046] Specifically, the risk grading threshold is a dynamic threshold, which includes the first risk threshold. Second risk threshold Dynamic thresholds allow for more flexible risk assessment, reflecting actual soil physicochemical properties and differences in pollution source background, thus improving the accuracy of risk assessment. The first risk threshold... Second risk threshold The method for determining it is as follows: Based on the benchmark risk thresholds for corresponding heavy metal elements in the soil environmental quality standards , Based on this, a threshold adjustment factor is introduced to correct the baseline risk threshold, resulting in a dynamic risk grading threshold, where: in, The standardized threshold adjustment factor includes soil physicochemical property parameters and pollution source background contribution characteristic parameters. The soil physicochemical property parameters include at least one of soil pH value and organic matter content, and the pollution source background contribution characteristic parameters include at least the background contribution values ​​of each pollution source. Let be the weighting coefficients of the corresponding threshold adjustment factor, and satisfy . Based on the baseline risk threshold, an adaptive threshold is output by adjusting it according to multiple factors. Based on the dynamic risk classification threshold, the first risk threshold Second risk threshold The system categorizes and classifies risk indices, generating soil heavy metal pollution risk level information that adaptively adjusts to changes in soil environmental conditions and pollution source background characteristics. Based on a baseline risk threshold, it is modified according to multiple factors to output an adaptive threshold. Soil physicochemical properties influence heavy metal migration, bioavailability, and environmental risk, making the threshold more consistent with actual environmental conditions and improving the scientific rigor of the assessment.

[0047] Specifically, spatial location information of each sampling point and corresponding soil heavy metal pollution risk level information are obtained. The risk level information is the result of risk level calculation based on source contribution matrix, background contribution value, risk index and dynamic risk classification threshold; spatial risk attributes of sampling points are established.

[0048] Spatially correlate the spatial location information of sampling points with their corresponding risk level information to construct a spatial mapping relationship between sampling points and risk levels; binding the risk information of discrete points with their spatial location facilitates spatial interpolation and visualization.

[0049] Based on spatial mapping relationships, spatial interpolation methods are used to convert the risk level information of discrete sampling points into continuous spatial distribution data. Spatial interpolation or spatial extension methods include inverse distance weighting, kriging interpolation, or combinations thereof. Binding the risk information of discrete points to their spatial location facilitates spatial interpolation and visualization. This provides clear regional risk level classifications, supporting land management, pollution control, and planning decisions.

[0050] Based on continuous spatial distribution data, the target area is divided into risk level zones, which are at least low-risk, medium-risk and high-risk areas. By overlaying the risk level zoning results with the spatial base map in the geographic information system, a soil heavy metal pollution risk zoning map of the target area is generated. This risk zoning map is used to characterize the spatial distribution characteristics of soil heavy metal pollution risk in different areas. A visualization map is then created to intuitively display the spatial distribution characteristics of soil heavy metal pollution risk in different areas.

[0051] Specifically, a machine learning-based system for identifying and assessing the risk of heavy metal pollution sources in soil includes: The data acquisition and preprocessing module is used to acquire soil heavy metal element content data of sampling points in the target area, as well as the geographic coordinates, terrain parameters and land use type parameters corresponding to the sampling points. It performs missing value processing, noise identification and removal and standardization processing on the data to construct a feature dataset for pollution source analysis. The pollution source analysis module is used to construct an input matrix of heavy metal element content based on the feature dataset, decompose the input matrix using multivariate analysis, and output a source spectrum matrix that characterizes the combination of heavy metal element characteristics of different pollution sources and a source contribution matrix that characterizes the degree of pollution source contribution at each sampling point. The background contribution modeling module is used to extract the contribution data of each pollution source in the target time period based on the source contribution matrix, perform time period statistics and smoothing on the contribution data, and build a background contribution model by combining conservative element information to calculate the background contribution value of each pollution source in the target time period. The risk index calculation module is used to reconstruct the pollution source and correct the background of the heavy metal element content at each sampling point based on the source spectrum matrix, source contribution matrix and background contribution value, and to calculate the single-factor risk index and comprehensive risk index of each heavy metal element at each sampling point. The dynamic risk classification module is used to classify and determine the risk of heavy metal pollution in soil at each sampling point based on the risk index and the dynamic risk classification threshold constructed based on the soil environmental quality standard, and generate corresponding pollution risk level information. The risk zoning generation module is used to spatially correlate pollution risk level information with the spatial location information of sampling points, generate continuous spatial distribution results of soil heavy metal pollution risk in the target area through spatial interpolation or spatial expansion methods, and output a soil heavy metal pollution risk zoning map.

[0052] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.

[0053] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for identifying and assessing the risk of heavy metal pollution sources in soil based on machine learning, characterized in that, include: S1. Obtain soil heavy metal element content data and spatial and environmental attribute data corresponding to the sampling points within the target area; preprocess the data to construct a feature dataset for pollution source analysis. S2. Construct a heavy metal element content input matrix based on the heavy metal element content parameters, and decompose the input matrix using multivariate analysis to obtain a source spectrum matrix characterizing the characteristics of different pollution sources and a source contribution matrix characterizing the degree of pollution source contribution at each sampling point. S3. Extract the contribution data of each pollution source in the target time period, combine it with the conservative element information to establish a background contribution model, and calculate the background contribution value of each pollution source in the target time period. S4. Based on the source spectrum matrix, source contribution matrix and background contribution value, the pollution source reconstruction and background correction are performed on the heavy metal element content at each sampling point, and the pollution risk index of each heavy metal element at the sampling point is calculated. S5. Based on the pollution risk index and combined with the dynamic risk classification threshold, classify and determine the risk of heavy metal pollution in the soil at the sampling points, and generate corresponding risk level information. S6. Based on the risk level information and the spatial location information of the sampling points, generate the spatial distribution results of soil heavy metal pollution risk to characterize the spatial differentiation characteristics of soil heavy metal pollution risk in the target area.

2. The method according to claim 1, characterized in that, Sampling points were set up in the target area to collect surface soil samples and determine the content of heavy metal elements in the soil samples. Record the geographic coordinates of the sampling points and obtain the terrain parameters and land use type information of the sampling points; The raw data includes parameters such as heavy metal element content, geographic coordinates, topography, and land use type. Missing values ​​were handled in the original data; noise was identified and removed from the original data; and the heavy metal element content parameters were standardized. Construct feature variables including geographic coordinate parameters, terrain parameters, and land use parameters.

3. The method according to claim 2, characterized in that, Construct an input matrix based on the heavy metal element content parameters: in, This represents the content parameter value of the j-th type of heavy metal element at the i-th sampling point. , ; Let the number of pollution source categories be K; Perform k-means clustering analysis on the input matrix X to divide the sampling points into K clusters. Calculate the mean of each cluster for each heavy metal element to form the initial source spectrum matrix elements: in, Let k represent the set of samples in the k-th cluster. The number of samples in the sample set is represented by the initial source spectrum matrix. The source spectrum matrix is ​​constructed using it as the initial input for multi-parent decomposition: in, This represents the characteristic combination of the k-th type of pollution source with the j-th type of heavy metal element. ; Standardize the input matrix X: in, Let be the mean value of the j-th type of heavy metal element. Standard deviation, Calculate the standardized matrix for the standardized data. covariance matrix ; The covariance matrix is ​​analyzed through eigenvalue decomposition. The data is decomposed, and eigenvalues ​​and eigenvectors are output. The eigenvectors represent the principal component directions of the data, and the eigenvalues ​​represent the variance contribution of each principal component. Based on the magnitude of the eigenvalues, the first K principal components are selected to construct the principal component matrix V; using the principal component matrix V and the standardized data matrix... Calculate the contribution matrix: ; in, This represents the contribution of the k-th type of pollution source at the i-th sampling point; Calculate the error function based on the current values ​​of the contribution matrix C and the source spectrum matrix S. , The Frobenius norm is used to adjust the source spectrum matrix S using gradient descent to minimize this error; when the error function If the number of iterations falls below the set threshold or reaches the preset maximum number, the iteration stops.

4. The method according to claim 3, characterized in that, The contribution data of all pollution sources at different time periods are extracted using the source contribution matrix C. Determine the target time period Based on the pollution source contribution data within the target time period, the contribution sequence for the corresponding time period is extracted from the source contribution matrix. ; By filtering the extracted time-period data, high-frequency fluctuations are removed, and a smooth background contribution sequence is output. Regression analysis was used to fit the smoothed background contribution sequence with the corresponding conservative element to establish a background contribution sequence model, which reflects the changing trend of pollution sources in the target period. Based on the background contribution sequence model, the background contribution value of each pollution source in the target time period is output.

5. The method according to claim 4, characterized in that, The method for extracting contribution data of all pollution sources at different time periods through the source contribution matrix includes: grouping the sampling points according to a preset time interval based on the sampling time information of the sampling points, constructing a correspondence between the sampling points and time periods, and determining the target time period within the time period. The target time period is a continuous time interval. Or a time set consisting of multiple discrete sampling times; within each time period, select the set of sampling points belonging to that time period. Extract the corresponding row vector from the source contribution matrix C. This forms a set of contribution data for the k-th type of pollution source during that time period. Statistical calculations are performed on the contribution data set of the same pollution source within the same time period to obtain the time period contribution value of the pollution source within the time period. The statistical calculations include at least one of the following: mean, median, or weighted average. For the target time period Select the set of sampling points belonging to the target time period. Then, the corresponding source contribution amount is extracted from the source contribution matrix to construct the contribution amount sequence of the k-th type of pollution source in the target time period. in, This represents the contribution of the i-th sampling point to the k-th type of pollution source during the target time period.

6. The method according to claim 4, characterized in that, The smoothing process uses a moving average method: in, This represents the background contribution value of the k-th type of pollution source during the target time period, where n is the number of sampling points during the target time period; a conservative element content sequence corresponding to the target time period is selected. , The smoothed background contribution sequence is fitted to the conservative element using regression analysis to establish a background contribution sequence model. The regression relationship of the background contribution sequence model is expressed as follows: in, and For regression coefficients, For residual terms; Based on the background contribution sequence model, calculate the background contribution value of the k-th pollution source during the target time period: in, and The regression coefficient represents the background contribution of the k-th pollution source during the target time period. This represents the statistical mean of the content of conservative elements within the target time period.

7. The method according to claim 5, characterized in that, For any sampling point i and any heavy metal element j, calculate the pollution source reconstruction value of the heavy metal element at that sampling point based on the source contribution matrix C and the source spectrum matrix S: in, This represents the content of the j-th type of heavy metal element at the i-th sampling point due to the contribution of the pollution source; based on the reconstructed pollution source value and the background contribution value of each pollution source during the target time period, the net pollution contribution of the j-th type of heavy metal element at the sampling point is calculated: Based on the net pollution contribution For any sampling point i, the j-th type of heavy metal element... = , The soil environmental quality standard limit for the j-th heavy metal element; Introducing a comprehensive risk index : m represents the number of heavy metal elements involved in the risk assessment. Let be the weighting coefficient of the j-th heavy metal element, used to reflect the relative risk contribution of different heavy metal elements to the soil environment and human health, and satisfying the following: Based on the single-factor risk index and comprehensive risk index The pollution risk of sampling points is classified into risk levels according to the risk classification threshold.

8. The method according to claim 7, characterized in that, The risk classification threshold is a dynamic threshold, which includes a first risk threshold. Second risk threshold First risk threshold Second risk threshold The method for determining it is as follows: Based on the benchmark risk thresholds for corresponding heavy metal elements in the soil environmental quality standards , Based on this, a threshold adjustment factor is introduced to correct the baseline risk threshold, resulting in a dynamic risk grading threshold, wherein: in, The standardized threshold adjustment factor includes soil physicochemical property parameters and pollution source background contribution characteristic parameters. The soil physicochemical property parameters include at least one of soil pH value and organic matter content, and the pollution source background contribution characteristic parameters include at least the background contribution values ​​of each pollution source. Let be the weight coefficients of the corresponding threshold adjustment factor, and satisfy . ; Based on the dynamic risk grading threshold, the first risk threshold Second risk threshold The risk index is classified and determined to generate soil heavy metal pollution risk level information that is adaptively adjusted according to changes in soil environmental conditions and pollution source background characteristics.

9. The method according to claim 8, characterized in that, The spatial location information of each sampling point and the corresponding soil heavy metal pollution risk level information are obtained. The risk level information is the risk level result calculated based on the source contribution matrix, background contribution value, risk index and dynamic risk classification threshold. The spatial location information of the sampling points is spatially associated with the corresponding risk level information to construct a spatial mapping relationship between sampling points and risk levels. Based on the aforementioned spatial mapping relationship, a spatial interpolation method is used to convert the risk level information of discrete sampling points into continuous spatial distribution data; Based on continuous spatial distribution data, the target area is divided into risk level zones, which are at least low-risk, medium-risk and high-risk areas. The risk level zoning results are overlaid with the spatial base map in the geographic information system to generate a soil heavy metal pollution risk zoning map of the target area. The risk zoning map is used to characterize the spatial distribution characteristics of soil heavy metal pollution risk in different areas.

10. A machine learning-based system for identifying and assessing the risk of heavy metal pollution sources in soil, characterized in that, The system employs any one of claims 1-9, and the system comprises: The data acquisition and preprocessing module is used to acquire soil heavy metal element content data of sampling points in the target area, as well as the geographic coordinates, terrain parameters and land use type parameters corresponding to the sampling points. The module performs missing value processing, noise identification and removal and standardization on the data to construct a feature dataset for pollution source analysis. The pollution source analysis module is used to construct a heavy metal element content input matrix based on the feature dataset, decompose the input matrix using a multivariate analysis method, and output a source spectrum matrix that characterizes the combination of heavy metal element characteristics of different pollution sources and a source contribution matrix that characterizes the degree of pollution source contribution at each sampling point. The background contribution modeling module is used to extract the contribution data of each pollution source in the target time period based on the source contribution matrix, perform time period statistics and smoothing on the contribution data, and establish a background contribution model in combination with conservative element information to calculate the background contribution value of each pollution source in the target time period. The risk index calculation module is used to reconstruct the pollution source and correct the background of the heavy metal element content at each sampling point based on the source spectrum matrix, source contribution matrix and background contribution value, and to calculate the single-factor risk index and comprehensive risk index of each heavy metal element at each sampling point. The dynamic risk classification module is used to classify and determine the risk of heavy metal pollution in the soil at each sampling point based on the risk index and the dynamic risk classification threshold constructed based on the soil environmental quality standard, and generate corresponding pollution risk level information. The risk zoning generation module is used to spatially correlate the pollution risk level information with the spatial location information of the sampling points, generate a continuous spatial distribution result of soil heavy metal pollution risk in the target area through spatial interpolation or spatial expansion methods, and output a soil heavy metal pollution risk zoning map.