Soil heavy metal pollutant tracing and concentration prediction method, device and medium
By combining principal component analysis and geographic detectors, the problem of inaccurate identification of heavy metal pollution sources in existing technologies has been solved, realizing automated, refined and rapid identification of pollution sources. This breaks through the limitations of existing methods and provides accurate pollution source location and concentration prediction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-04
- Publication Date
- 2026-03-03
AI Technical Summary
Existing methods for tracing heavy metal pollution sources are difficult to accurately identify the specific locations and contributions of multiple potential industrial pollution sources, and quantitative analysis is challenging. Stable isotope methods are costly, and principal component and cluster analysis are not effective for qualitative analysis.
By combining principal component analysis and a geographic detector, the initial range of pollution sources is determined through principal component analysis, and a more precise pollution source is screened out using a geographic detector. The pollution source type is determined by the factor loading matrix, the spatial distribution is determined by the factor score matrix, and the pollution source name and location are automatically identified by combining geographic information.
It achieves automated identification of pollution sources, with detailed and accurate results, reduced human intervention, fast identification speed, and strong quantitative analysis capabilities.
Smart Images

Figure CN114548598B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of soil pollution remediation, and in particular to a method, equipment and medium for tracing the source and predicting the concentration of heavy metal pollutants in soil. Background Technology
[0002] With the increasing severity of heavy metal pollution in soil, source tracing is crucial for soil remediation. Current research commonly employs source apportionment methods for heavy metal pollutants, including stable element isotope analysis, principal component analysis, cluster analysis, and positive definite matrix factorization, all of which offer some explanation for soil pollution sources.
[0003] However, the stable isotope method is mainly applicable to a few heavy metal elements with stable isotopes. Furthermore, due to the need to compare various forms of heavy metals, it is labor-intensive and costly, and is currently often used for source identification of heavy metals in small-area, fixed-point soil samples. Principal component analysis and cluster analysis are more qualitative analyses and difficult to use for quantitative analysis. The quantitative source obtained from the positive definite factor matrix is only a rough estimate of industrial, agricultural, or parent material sources. When multiple potential industrial pollution sources exist in a region, it is impossible to accurately determine which industrial source the pollution originates from, making it difficult to accurately identify the contribution and location of different pollution sources. Summary of the Invention
[0004] This invention provides a method, equipment, and medium for tracing the source and predicting the concentration of heavy metal pollutants in soil. It uses a combination of principal component analysis and geographic detectors to accurately identify the name and location of pollution sources.
[0005] In a first aspect, embodiments of the present invention provide a method for tracing the source of heavy metal pollutants in soil, comprising:
[0006] Principal component analysis was performed on the matrix of heavy metal concentrations in soil from multiple sampling points within a region.
[0007] Based on the principal component analysis results, the names and geographical locations of multiple pollution sources within the region were determined;
[0008] A geographic detector is used to select the final pollution source from the multiple pollution sources.
[0009] Secondly, embodiments of the present invention provide a method for predicting soil heavy metal concentrations, including:
[0010] Obtain the geographical location to be predicted;
[0011] The predicted geographical location is input into the soil heavy metal concentration prediction model as described in claim 5 or 6 to obtain the heavy metal concentration of the predicted geographical location.
[0012] Thirdly, embodiments of the present invention also provide an electronic device, the electronic device comprising:
[0013] One or more processors;
[0014] Memory, used to store one or more programs.
[0015] When the one or more programs are executed by the one or more processors, the one or more processors implement the soil heavy metal pollutant tracing or soil heavy metal concentration prediction method described in the above embodiments.
[0016] Fourthly, embodiments of the present invention also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the soil heavy metal pollutant tracing or soil heavy metal concentration prediction method described in the above embodiments.
[0017] The technical effects of the embodiments of the present invention are as follows:
[0018] 1. This invention combines principal component analysis results with a geographic detector. Principal component analysis determines the initial selection range of pollution sources, and the geographic detector determines the degree of influence of each pollution source on the principal component score, thereby screening out more accurate pollution sources. This achieves automated determination of pollution sources and reduces the intervention of human experience in the selection of pollution sources.
[0019] 2. The embodiments of the present invention use factor loading matrix to determine the type of pollution source and factor score matrix to determine the spatial distribution of pollution source. Then, based on the type and spatial distribution of pollution source, the specific pollution source name and geographical location are determined. This overcomes the limitation of positive definite factor matrix model, which can only identify pollution source type (such as industrial source, traffic source, geological background, etc.), and the identification results are more refined.
[0020] 3. This embodiment of the invention locates the pollution source contributing the most to the pollution source near the sub-region with the highest principal component factor score based on the spatial distribution of the principal component factor score; then, combined with the geographical information of the area to be identified, it automatically identifies the name and location of the pollution source contributing the most to the pollution source. No human intervention is required, achieving automated identification of pollution sources with accurate results and fast identification speed.
[0021] 4. This embodiment of the invention utilizes the basic principle of factor detectors in geographic detectors, taking the spatial distribution of the principal component factor scores as the dependent variable and the spatial distribution of the distance from the sampling point to each pollution source as the independent variable. It selects the pollution sources with the highest influence on the dependent variable, further determining the precise range of pollution sources. It fully utilizes the variable distribution obtained from principal component analysis, eliminating the need to introduce other variables, making it simple, fast, and highly accurate. Attached Figure Description
[0022] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0023] Figure 1 This is a flowchart of a method for tracing the source of heavy metal pollutants in soil provided by an embodiment of the present invention;
[0024] Figure 2 This is a schematic diagram of the spatial distribution of factor scores of two principal component factors provided in an embodiment of the present invention;
[0025] Figure 3 This is a comparison chart of the results of the prediction model and the positive definite factor matrix model provided in the embodiments of the invention;
[0026] Figure 4 This is a flowchart of a method for predicting soil heavy metal concentration provided in an embodiment of the present invention;
[0027] Figure 5 This is a comparison chart between the predicted concentration provided by this application, the predicted concentration of PMF, and the actual concentration, as provided in the embodiments of this invention.
[0028] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0029] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0030] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0031] In the description of this invention, it should also be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0032] Figure 1 This is a flowchart of a method for tracing the source of heavy metal pollutants in soil, provided by an embodiment of the present invention. It is applicable to situations where the heavy metal pollution source within a region is determined based on the heavy metal concentration at some sampling points. This embodiment is executed by electronic equipment. (Combined with...) Figure 1 The method provided in this embodiment specifically includes:
[0033] S110. Perform principal component analysis on the matrix composed of soil heavy metal concentrations from multiple sampling points within a region.
[0034] The region refers to a pre-selected geographical area. In this embodiment, the heavy metal pollution sources in this region will be traced to determine their names and geographical locations. The multiple sampling points are some geographical locations within the geographical area. In this embodiment, the heavy metal concentrations in the soil at these geographical locations are sampled to form an initial matrix for principal component analysis, in order to identify potential pollution sources based on the principal component analysis results.
[0035] Specifically, firstly, an initial matrix for principal component analysis is generated based on the soil heavy metal concentrations from multiple sampling points. Each element in the initial matrix corresponds to a two-dimensional spatial coordinate of a geographical location, and each element value corresponds to the heavy metal concentration at that location. This embodiment is based on the assumption that the total concentration of pollutants equals the sum of the contributions from all pollution sources.
[0036] For example, suppose the pre-selected geographical area is A square region was divided into nine grid points with a step size of 500m for both length and width. Four grid points with spatial coordinates [0,0], [500,500], [500,1000], and [1000,500] were selected as sampling points. The chromium (Cr) concentrations at these four sampling points were collected as 47.86 mg / kg, 53.78 mg / kg, 68.41 mg / kg, and 46.34 mg / kg, respectively. The following initial matrix was then constructed: .
[0037] In one specific implementation, the Tongling Yiduo Industrial Park in Anhui Province is pre-selected as the geographical area. Located in a prefecture-level city in central Anhui Province, the park possesses relatively abundant mineral resources, with a long history of mining pyrite and copper. Established in the early 21st century, after years of development, the park has initially formed a leading industry focused on copper processing and utilization, machinery manufacturing, and electronic components. The planned total area of the park is 14.2 km². 2 The park is adjacent to several railways and highways, and its western boundary is a river that flows from south to northwest, with its upper reaches passing through mining areas and smelters. Mining areas are also distributed in the southwest of the park. The prevailing wind direction in the park is from southwest to northeast. After determining the pre-selected area, soil heavy metal concentrations were collected from multiple sampling points within the park. Specific concentration values are not detailed here; Table 1 only presents a statistical overview of the soil heavy metal concentrations in the park.
[0038] Table 1
[0039]
[0040] As shown in Table 1, the park contains 8 major heavy metals, so the initial matrix is a three-dimensional matrix with a third dimension of 8.
[0041] After obtaining the initial matrix, principal component analysis (PCA) is performed on it. The PCA results are used to identify potential pollution sources within the region. Specifically, PCA essentially transforms at least one pollution source affecting heavy metal concentration into at least one uncorrelated principal component factor. Each principal component factor characterizes the complex of at least one pollution source. Optionally, for source identification purposes, the rotation method used in PCA is varimax rotation, and factors with eigenvalues greater than 1 are extracted as principal component factors after rotation.
[0042] Optionally, the principal component analysis results include: the factor loading matrix and factor score matrix of the principal component factors.
[0043] The factor loading matrix is a matrix composed of the coefficients of the equations relating each principal component factor as the independent variable and the heavy metal concentration at each sampling point as the dependent variable. It reflects the contribution rate of each principal component factor to the heavy metal concentration at each sampling point and is used to determine the type of potential pollution source.
[0044] The factor score matrix is a matrix composed of the coefficients of the equations relating heavy metal concentration at each sampling point to each principal component factor. It reflects the contribution rate of heavy metal concentration at each sampling point to each principal component factor and is used to determine the spatial distribution of potential pollution sources.
[0045] S120. Based on the principal component analysis results, determine the names and geographical locations of multiple pollution sources within the region.
[0046] Optionally, based on the principal component analysis results, the names and geographical locations of multiple pollution sources within the region are determined, specifically including the following steps:
[0047] Step 1: Determine the type of pollution source in the region based on the factor loading matrix.
[0048] Specifically, firstly, the rotated component matrix is obtained based on the factor loading matrix. The rotated component matrix is a matrix composed of the coefficients of the equations relating each principal component factor as the independent variable and the concentration of each heavy metal in the region as the dependent variable, reflecting the contribution rate of each principal component factor to the concentration of each heavy metal. In the specific implementation of the aforementioned industrial park, after performing principal component analysis on the initial matrix, the rotated component matrix shown in Table 2 is obtained based on the factor loading matrix.
[0049] Table 2
[0050]
[0051] As shown in Table 2, two principal component factors, Factor 1 and Factor 2, were extracted after Varimax rotation. Factor 1 and Factor 2 are uncorrelated, and each factor represents a combination of at least one pollution source. Each row of the rotated component matrix corresponds to a heavy metal, and each column corresponds to a principal component factor. The value of each element is the contribution of the principal component factor in its column to the heavy metal concentration in its row. For example, the element in the first row and first column indicates that Factor 1 contributes 0.82 to the Cr concentration.
[0052] After obtaining the rotated component matrix, the type of potential pollution source (e.g., anthropogenic or natural source) is determined based on the contribution of principal component factors to heavy metal concentrations. Optionally, if the contribution rate of a principal component factor to the concentration of each of multiple metals (e.g., three) is greater than a set threshold (e.g., 0.8), then the pollution source (one or more) corresponding to the principal component factor is considered to be an emission source of multiple heavy metals. Combining source apportionment expertise and existing research, the type of pollution source can be determined, such as an anthropogenic pollution source.
[0053] In the specific implementation of the aforementioned industrial park, both principal component factors were able to explain the sources of heavy metals in the soil. Specifically, Factor 1 contributed 0.82, 0.97, 0.98, 0.88, 0.98, 0.93, and 0.96 to the seven heavy metals chromium, copper, zinc, mercury, arsenic, lead, and cadmium, respectively, explaining the vast majority of their sources. Factor 2 contributed 0.95 to nickel, explaining its source. The above analysis indicates that the heavy metal elements in the soil of this industrial park mainly originate from Factor 1, and the pollution source corresponding to Factor 1 is anthropogenic pollution.
[0054] Step 2: Determine the spatial distribution of principal component factor scores based on the factor score matrix.
[0055] The element values in the factor score matrix are called principal component factor scores, reflecting the contribution rate of the heavy metal concentration at each sampling point to each principal component factor. Based on the principal component factor scores corresponding to each sampling point and the spatial coordinates of the sampling point, the spatial distribution of the principal component factor scores can be determined.
[0056] Step 3: Based on the type of pollution source and the spatial distribution of the principal component factor scores, determine the names and geographical locations of multiple pollution sources within the region.
[0057] Optionally, firstly, based on the spatial distribution of the principal component factor scores, determine the sub-region with the highest principal component factor scores within the region.
[0058] The sub-region with the highest factor score is used to determine the name and location of the pollution source, and the area near the sub-region is where the potential pollution source is most likely to appear. Figure 2 This is a schematic diagram of the spatial distribution of factor scores for two principal component factors provided in an embodiment of the present invention, corresponding to the specific implementation shown in Table 2. The left figure shows the spatial distribution of the score for factor 1, and the right figure shows the spatial distribution of the score for factor 2. Figure 2 As shown, the areas with higher factor 1 scores are mainly concentrated on the west side of the park, so the west side is taken as the sub-region with the highest score.
[0059] Then, based on the type of pollution source and the sub-region, the names and geographical locations of multiple pollution sources are determined.
[0060] Specifically, this embodiment provides two strategies for identifying pollution sources. One strategy is to identify pollution sources starting from the sub-region with the highest factor score. The vicinity of the sub-region is where potential pollution sources are most likely to occur. By combining the geographical information of the area under study (e.g., a park map), anthropogenic pollution sources near the sub-region are identified. In the specific implementation of the industrial park mentioned above, an anthropogenic pollution source near the west side of the park is a river flowing upstream through the mining area; therefore, this river is considered a pollution source.
[0061] Another approach is to identify other pollution sources solely based on the type of pollution source, without considering sub-regions. In the specific implementation of the aforementioned industrial park, the park is primarily used for copper processing, and the mines outside the park are also mainly used for copper production. Other human-caused pollution sources exist, including atmospheric deposition (from smelting in the mining area), transportation (such as main road transportation), and other human activities (residential areas).
[0062] By combining the results from the two strategies, five pollution sources were identified: metal recycling plants, rivers, main roads, mining areas, and communities.
[0063] S130. Using a geographic detector, select the final pollution source from the multiple pollution sources.
[0064] To ensure greater accuracy of the identified pollution sources, this embodiment employs a geographic detector to assess the reliability of each source. Specifically, it utilizes the factor detector within the geographic detector. After inputting the independent and dependent variables into the factor detector, the output q-value for each independent variable characterizes its influence on the dependent variable, thus achieving the purpose of independent variable selection.
[0065] Optionally, the spatial distribution of the principal component factor scores is used as the dependent variable, and the spatial distribution of the distance from the sampling point to each pollution source is used as the independent variable. A geographic detector is used to evaluate the degree of influence of each pollution source on the dependent variable. Based on the degree of influence of each pollution source on the dependent variable, the final pollution source is selected from the multiple pollution sources.
[0066] In the specific implementation plan of the aforementioned industrial park, five specific pollution sources were identified: metal recycling plants, rivers, main roads, mining areas, and communities. This embodiment will utilize geographic detection methods to identify the most reliable pollution source from these five sources as the final pollution source.
[0067] Specifically, the spatial distribution of the shortest distance from the sampling point to each pollution source—X1 (metal recycling plant), X2 (river), X3 (main road), X4 (mining area), and X5 (community)—is used as five independent variables. The spatial distribution of the principal component factor scores is used as dependent variables Y1 (factor 1) and Y2 (factor 2). The independent and dependent variables are input into a factor detector to obtain the q-value of Xi against Yj, where i = 1, 2, 3, 4, 5, and j = 1, 2. The calculation principle is as follows:
[0068] (5)
[0069] Where, q x These are index values used to measure the spatial correlation between Xi and Yj, where h = 1, 2, 3, ..., L, L represents the level (sub-region or subclass) of factor Xi, and N and N h These represent the number of samples in the entire region and in each stratum h, respectively. The symbol σ 2 and σ h 2 These represent the variances of Yj in the entire region and in each h layer, respectively.
[0070] All variables in the formula are automatically determined by the geographic detector based on the input data, requiring no human intervention. The final calculated q value is a value between 0 and 1, representing the degree to which Yj is influenced by Xi. The larger the q value, the stronger the spatial correlation between Xi and Yj, thus enabling the matching of principal component factors and pollution sources through factor detection.
[0071] According to calculations by the geospatial detector, the q values for soil principal component factor 1 corresponding to metal recycling and processing plants (X1), rivers (X2), and communities (X5) are 0.52, 0.51, and 0.46, respectively, which have some explanatory power for the spatial distribution of principal component factor 1. This indicates that, apart from rivers, the spatial distribution of soil principal component factor 1 is related to other pollution sources. X1 (metal recycling and processing plants) and X4 (mining areas) have relatively low q values corresponding to soil principal component factor 2, with a maximum value of 0.23, which also have some explanatory power for the spatial distribution of principal component factor 2. Furthermore, X3 (main road) has q values less than 0.05 for both principal component factor 1 and principal component factor 2, indicating a low explanatory power for both factors. Therefore, excluding main roads from the pollution source range, the range of soil heavy metal pollution sources is further narrowed down to: metal recycling and processing plants (X1), rivers (X2), mining areas (X4), and communities (X5).
[0072] The technical effects of this embodiment are as follows:
[0073] 1. This invention combines principal component analysis results with a geographic detector. Principal component analysis determines the initial selection range of pollution sources, and the geographic detector determines the degree of influence of each pollution source on the principal component score, thereby screening out more accurate pollution sources. This achieves automated determination of pollution sources and reduces the intervention of human experience in the selection of pollution sources.
[0074] 2. The embodiments of the present invention use factor loading matrix to determine the type of pollution source and factor score matrix to determine the spatial distribution of pollution source. Then, based on the type and spatial distribution of pollution source, the specific pollution source name and geographical location are determined. This overcomes the limitation of positive definite factor matrix model, which can only identify pollution source type (such as industrial source, traffic source, geological background, etc.), and the identification results are more refined.
[0075] 3. This embodiment of the invention locates the pollution source contributing the most to the pollution source near the sub-region with the highest principal component factor score based on the spatial distribution of the principal component factor score; then, combined with the geographical information of the area to be identified, it automatically identifies the name and location of the pollution source contributing the most to the pollution source. No human intervention is required, achieving automated identification of pollution sources with accurate results and fast identification speed.
[0076] 4. This embodiment of the invention utilizes the basic principle of factor detectors in geographic detectors, taking the spatial distribution of the principal component factor scores as the dependent variable and the spatial distribution of the distance from the sampling point to each pollution source as the independent variable. It selects the pollution sources with the highest influence on the dependent variable, further determining the precise range of pollution sources. It fully utilizes the variable distribution obtained from principal component analysis, eliminating the need to introduce other variables, making it simple, fast, and highly accurate.
[0077] Based on the above and following embodiments, this embodiment establishes a metal concentration prediction model according to the finally determined pollution source. Optionally, after selecting the final pollution source from the plurality of pollution sources using a geographic detector, the method further includes: establishing a heavy metal concentration prediction model using a linear regression method based on the final pollution source, wherein the prediction model uses the distance between any geographical location in the region and the final pollution source as the independent variable, and the soil heavy metal concentration at the geographical location as the dependent variable.
[0078] This embodiment establishes a regression equation based on the final pollution source to quantitatively identify the impact of the pollution source on heavy metal concentration. The regression equation uses the distance from any geographical location to the final pollution source as the independent variable and the soil heavy metal concentration at that location as the dependent variable. Specifically, considering that the propagation characteristics of these heavy metals from the final pollution source to the soil decrease with increasing distance, the distance from the sampling point to the final pollution source is used to quantify the contribution of the pollution source. Optionally, the following regression equation is constructed:
[0079]
[0080] Where c represents the heavy metal concentration at any geographical location, and D represents the distance of that geographical location from the final pollution source; A and b is the constant to be fitted.
[0081] After constructing the regression equation, the undetermined parameters in the regression equation are fitted using the heavy metal concentration at the sampling point and the distance from the final pollution source, resulting in the final regression model. The fitted regression model is then tested for goodness of fit and significance. A model that meets the requirements for goodness of fit and significance can be used to predict the heavy metal concentration at any location.
[0082] Optionally, there are multiple final pollution sources; based on the final pollution sources, a heavy metal concentration prediction model is established using a linear regression method, including: using a multiple linear regression method to determine the optimal combination of independent variables for the metal concentration prediction model based on multiple final pollution sources, wherein the optimal combination of independent variables includes: the distance between any geographical location in the region and at least one final pollution source; and establishing the prediction model with the optimal combination of independent variables as independent variables and the soil heavy metal concentration at the geographical location as the dependent variable.
[0083] When there are multiple sources of pollution, the prediction model uses a multiple linear regression equation, with the independent variable being the distance between a geographical location and at least one source of pollution. Table 3 shows the regression model for heavy metal concentrations in the soil of the aforementioned industrial park.
[0084] Table 3
[0085]
[0086]
[0087] Among them, D mine D community and D river Let represent a geographical location and the locations of a metal mine, river, and processing plant, respectively. It can be seen that the dependent variable of the equation is the concentration of various heavy metals, and the independent variable factor is D. mine D community and D river In the process of fitting the equation, in addition to metal processing plants, rivers, and mines, the regression equation originally considered main roads as the final source of pollution. However, after calculation, this variable had a very small impact on the concentration of various heavy metals and was excluded from the combination of independent variables in the regression model.
[0088] Table 3 also shows that rivers flowing through the mines are the most likely sources of pollution for copper, zinc, mercury, arsenic, lead, and cadmium. The mining area is the most significant source of chromium, while the community provides some explanation for the sources of nickel and mercury. R² is used to measure the explanatory power of variables through regression equations. All eight models in Table 3 have R² values higher than 0.5, with the regression equation for Pb showing the highest good fit at 0.89.
[0089] Compared with the positive definite factor matrix (PMF) model, the prediction model provided in this embodiment can quantitatively identify pollution sources. Figure 3 This is a comparison chart of the results of the prediction model and the positive definite factor matrix model provided in the embodiments of the invention. Combined with... Figure 3For pollution control in a specific area, if only the type of pollution source is known but its location is uncertain, effective control may be difficult to implement. This is because some areas may have several pollution sources, all belonging to the same category. For example, in the specific implementation of the aforementioned industrial park, the pollution source type of both the mine and the factory is industrial activity. The source tracing results obtained from the PMF model are insufficient to further pinpoint the pollution source. However, the source tracing results and prediction model of this application can be used in conjunction with distance to quantify whether a possible pollution source (with known location) is a major source factor, thus providing a powerful tool for pollution control in complex environments.
[0090] Figure 4 This is a flowchart of a method for predicting soil heavy metal concentration provided by an embodiment of the present invention. It is applicable to predicting the heavy metal concentration at any location within a region based on an established prediction model. In this embodiment, the method is executed by an electronic device. (Combined with...) Figure 4 The method provided in this embodiment specifically includes:
[0091] S210, Obtain the geographical location to be predicted.
[0092] S220. Input the geographical location to be predicted into the soil heavy metal concentration prediction model described in any of the above embodiments to obtain the heavy metal concentration of the geographical location to be predicted.
[0093] The predictive models constructed in any of the above embodiments rely on the distance from the final pollution source to a given geographical location to quantify the source contribution. Therefore, theoretically, the heavy metal concentration at that location should exhibit a negative correlation with the distance from that location to the main pollution source.
[0094] This embodiment used a model to predict heavy metal concentrations at multiple geographical locations. The prediction results showed that the heavy metal content in the soil decreased significantly with increasing distance from the river. For example, after the lead content rapidly decreased from 300 mg / kg (387 m from the river) to 38 mg / kg (1900 m from the river), it remained at approximately 40 mg / kg. Furthermore, copper, zinc, arsenic, and cadmium showed similar trends. This validates the effectiveness of the prediction model.
[0095] Figure 5 This is a comparison chart between the predicted concentration provided by this application, the predicted concentration of PMF, and the actual concentration, as provided in the embodiments of this invention. It can be seen that the concentration predicted by this application is closer to the actual concentration.
[0096] Optionally, after obtaining the heavy metal concentration of the geographical location to be predicted, the method further includes: after obtaining the heavy metal concentration of multiple geographical locations to be predicted, if the changing trend of the distance between two geographical locations to be predicted and a final pollution source is the same as the changing trend of the heavy metal concentration of the two geographical locations to be predicted, it is considered that there is an interaction effect of multiple pollution sources between the two geographical locations to be predicted.
[0097] This embodiment analyzes the impact of multiple pollution sources using prediction results from multiple geographical locations. In general, the variation of heavy metal concentration with distance, combined with the interaction detection function of source factors, verifies and analyzes the characteristics of the identified pollution sources. Specifically, with the pollution source as the center and a 500-meter gradient, the variation trends of most heavy metal concentrations are consistent with the pollution sources identified by the prediction model; heavy metal concentrations gradually decrease with distance from rivers, roads, and communities. However, some heavy metals exhibit oscillating trends, indicating the existence of interactive effects from multiple pollution sources. Through similar pollution source characteristic analysis, the impact of pollution sources on surrounding areas can be better analyzed.
[0098] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention, such as... Figure 6 As shown, the device includes a processor 50, a memory 51, an input device 52, and an output device 53; the number of processors 50 in the device can be one or more. Figure 6 Taking a processor 50 as an example; the processor 50, memory 51, input device 52, and output device 53 in the device can be connected via a bus or other means. Figure 6 Taking the example of a connection between China and Israel via a bus.
[0099] The memory 51, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the method, equipment, and medium for tracing the source and predicting the concentration of heavy metal pollutants in soil according to an embodiment of the present invention. The processor 50 executes various functional applications and data processing of the device by running the software programs, instructions, and modules stored in the memory 51, thereby realizing the aforementioned method, equipment, and medium for tracing the source and predicting the concentration of heavy metal pollutants in soil.
[0100] The memory 51 may primarily include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a given function; the data storage area may store data created based on terminal usage. Furthermore, the memory 51 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory, or other non-volatile solid-state storage device. In some instances, the memory 51 may further include memory remotely located relative to the processor 50, which can be connected to the device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0101] Input device 52 can be used to receive input digital or character information, and to generate key signal inputs related to user settings and function control of the device. Output device 53 may include display devices such as a display screen.
[0102] This invention also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements a method, device, and medium for tracing the source and predicting the concentration of heavy metal pollutants in soil according to any embodiment.
[0103] The computer storage medium of this invention can be any combination of one or more computer-readable media. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0104] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.
[0105] Program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0106] Computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages—such as Java, Smalltalk, and C++—as well as conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0107] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the technical solutions of the embodiments of the present invention.
Claims
1. A method for tracing soil heavy metal pollutants, characterized in that, The method comprises the following steps: performing principal component analysis on a matrix composed of soil heavy metal concentrations of a plurality of sampling points in a region, wherein the principal component analysis result comprises a factor loading matrix and a factor score matrix of a principal component factor representing a pollution source; determining the names and geographical positions of a plurality of pollution sources in the region according to the principal component analysis result; specifically, determining the types of pollution sources in the region according to the factor loading matrix; determining the spatial distribution of the principal component factor scores according to the factor score matrix; determining a sub-region with the highest principal component factor score in the region according to the spatial distribution of the principal component factor scores; and determining the names and geographical positions of a plurality of pollution sources according to the types of pollution sources and the sub-region; specifically, including two pollution source determination strategies, one of which is to determine pollution sources from the sub-region with the highest factor score, and the other of which is to determine other pollution sources from the types of pollution sources without considering the sub-region, and the determination results of the two strategies are combined to obtain a plurality of pollution sources; using the spatial distribution of the principal component factor scores as the dependent variable and the spatial distribution of the distances from the sampling points to each pollution source as the independent variable, using a geographic detector to evaluate the influence degree of each pollution source on the dependent variable; and selecting the final pollution source from the plurality of pollution sources according to the influence degree of each pollution source on the dependent variable. The optimal independent variable combination of the metal concentration prediction model is determined by using a multiple linear regression method according to multiple final pollution sources, wherein the optimal independent variable combination comprises a distance between any geographical position in the region and at least one final pollution source; the prediction model is established by taking the optimal independent variable combination as an independent variable and taking the soil heavy metal concentration of the geographical position as a dependent variable; specifically, the following regression equation is constructed: wherein c represents the heavy metal concentration at any geographical position, D represents the distance between the geographical position and the final pollution source; A and b are constants to be fitted; after the form of the regression equation is constructed, the undetermined parameters in the regression equation are fitted by using the metal concentration at the sampling point and the distance between the sampling point and the final pollution source, so that the final regression model is obtained.
2. A soil heavy metal concentration prediction method, comprising: obtaining a geographical position to be predicted; inputting the geographical position to be predicted into the soil heavy metal concentration prediction model of claim 1 to obtain the heavy metal concentration of the geographical position to be predicted.
3. The concentration prediction method according to claim 2, characterized by, After obtaining the heavy metal concentration of the geographical position to be predicted, the method further comprises: After obtaining the heavy metal concentrations of a plurality of geographical positions to be predicted, if the change trend of the distances between two geographical positions to be predicted and a final pollution source is the same as the change trend of the heavy metal concentrations of the two geographical positions to be predicted, it is considered that there is an interaction effect between the two geographical positions to be predicted.
4. An electronic device, comprising: One or more processors; a memory for storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the soil heavy metal pollutant tracing method of claim 1 and the soil heavy metal concentration prediction method of claim 2 or 3. The program is executed by the processor to implement the soil heavy metal pollutant tracing method of claim 1 and the soil heavy metal concentration prediction method of claim 2 or 3.
5. A computer-readable storage medium having stored thereon a computer program, characterized in that,
Citation Information
Patent Citations
Target soil property content prediction method based on soil transfer function
CN111508569A