Method and device for analyzing source-sink relationship between historical mine and surrounding farmland soil
By using a geographically weighted regression model and sample analysis, the problem of quantitative and accurate source apportionment of heavy metal pollution in farmland soil around mines was solved, and quantitative analysis of mine pollution was achieved.
Patent Information
- Application Number
- CN202411417454.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-11
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2044-10-11
AI Technical Summary
Existing technologies are insufficient for quantitative and accurate source apportionment of heavy metal pollution in farmland soils surrounding mines, especially when pollution sources around mines are complex and have a high geological background.
A geographically weighted regression model was adopted. By collecting relevant data on mine production and historical pollution information, samples were collected and analyzed to form a database of results. The correlation coefficient between the heavy metal content of surrounding farmland soil and mine solid waste was calculated by using a local regression analysis model to weight geographical location. The results were verified by combining geographical, meteorological and historical mining activities.
It enables quantitative and accurate source apportionment of heavy metals in soils surrounding mines, solving the problems of data discontinuity and low accuracy caused by spatial heterogeneity, and providing quantitative analysis capabilities for mine pollution.
Smart Images

Figure CN119337333B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of heavy metal pollution causation analysis, and more specifically, to a method and apparatus for analyzing the source-sink relationship between soil from historical legacy mines and surrounding farmland. Background Technology
[0002] In soil, the distribution of heavy metals is uneven, concealed, cumulative, regional, and complex. In recent years, in order to more accurately describe the sources of heavy metals in soil and their source-sink relationships with surrounding pollution, analytical methods have been developed, including emission inventory methods, chemical mass balance models (element / isotope ratios), and increasingly multivariate statistical models (factor analysis, cluster analysis, positive matrix factorization (PMF), UNMIX models, etc.), advanced statistical algorithms (conditional inference trees, random forests, etc.), and spatial analysis methods (geographic detectors, etc.).
[0003] However, traditional multivariate statistical methods, such as factor analysis (FA), principal component analysis (PCA), and cluster analysis (CA), while effective in source identification, typically provide only qualitative results and are insufficient for quantitative source apportionment. This means that while these methods can identify potential pollution sources, they cannot accurately assess the specific contribution of each source to soil heavy metal pollution. Secondly, chemical mass balance (CMB) models and isotope ratio rules require frequent monitoring of source and receptor samples in the study area, creating emission inventories, and continuously updating the emission source composition spectrum. While these methods can quantitatively evaluate the contribution of each pollution source, they may introduce errors during the processing.
[0004] In particular, the sources of pollution in farmland surrounding mines are complex, often compounded by high geological background levels. Furthermore, due to the influence of various factors such as mining scale, mining methods, geological conditions, and climatic conditions, the degree of pollution in farmland surrounding mines often varies significantly. Therefore, how to quantitatively and accurately analyze the source-sink relationship of heavy metals in the soil of mines and surrounding farmland is an urgent problem to be solved. Summary of the Invention
[0005] In view of this, the embodiments of this application aim to provide a method and apparatus for analyzing the source-sink relationship between historical legacy mines and surrounding farmland soil, so as to solve the problem that it is difficult to quantitatively and accurately apportion heavy metals in soils around mines.
[0006] Firstly, this specification provides a method for analyzing the source-sink relationship between soils from historically abandoned mines and surrounding farmland, including:
[0007] Collect heavy metal pollution data, including mining production-related data and historical pollution information;
[0008] Samples were collected and analyzed to obtain sample analysis data. The samples included mine solid waste samples, acidic wastewater samples, irrigation water samples, bottom sediment samples, and farmland soil samples.
[0009] Heavy metal pollution data and sample analysis data are collected and summarized to form a results database;
[0010] Extract sample input data from the results database;
[0011] The sample input data was input into a geographic weighted regression model to obtain the correlation coefficient between the heavy metal content of the surrounding farmland soil and the heavy metal content of the mine solid waste sample;
[0012] Among them, the geographically weighted regression model is constructed by weighting geographical location on the basis of the local regression analysis model. The independent variable of the local regression analysis model is the heavy metal content in the mine solid waste sample, and the dependent variable is the heavy metal content in the farmland soil sample. Geographical location represents the geographical relationship between the farmland soil sample and the mine solid waste sample. The rationality and accuracy of the correlation coefficient are determined based on the heavy metal content in the acidic wastewater sample, irrigation water sample and sediment sample.
[0013] According to the first aspect, in one possible implementation, the parameters of the local regression analysis model are determined by ordinary least squares, and the bandwidth of the geographically weighted regression model is determined by the Akaike Information Criterion.
[0014] According to the first aspect, in one possible implementation, after collecting heavy metal pollution data, the analytical method further includes: verifying and supplementing the heavy metal pollution data.
[0015] According to the first aspect, in one possible implementation, after inputting the sample input data into a geographically weighted regression model to obtain the correlation coefficient between the heavy metal content of the surrounding farmland soil and the mine, the analytical method further includes:
[0016] Based on geographical conditions, meteorological conditions, and historical mining activities, the correlation coefficient was tested.
[0017] Based on the evaluation system of mine pollution source-sink relationship, the pollution of mines is analyzed and evaluated. The evaluation system of mine pollution source-sink relationship is constructed based on geographical conditions, meteorological conditions, mineral characteristics, mining technology and mining years. The correlation coefficient is significantly correlated with mineral characteristics and mining years.
[0018] According to the first aspect, in one possible implementation, before collecting heavy metal pollution data, the analysis method further includes: determining the priority of mines based on geographic information spatial overlay analysis and literature review, wherein the geographic information spatial overlay analysis includes spatial analysis and layer overlay, and the priority represents the degree of attention paid to the mines by the academic community or public opinion.
[0019] According to the first aspect, in one possible implementation, before collecting samples for analysis, the analysis method further includes: determining the sample placement locations based on mine pollution diffusion conditions, pollution correlations, and empirical analysis methods.
[0020] According to the first aspect, in one possible implementation, collecting samples for analysis includes:
[0021] Determine the single-factor analysis indicators and testing methods;
[0022] The sample was analyzed based on single-factor analysis indicators and testing methods.
[0023] According to the first aspect, in one possible implementation, the conditions for the diffusion of mine pollution include meteorological conditions, topographic features, the physical and chemical properties of pollutants, and industrial and agricultural activities.
[0024] According to the first aspect, in one possible implementation, collecting and summarizing heavy metal pollution data and sample analysis data includes: collecting and summarizing heavy metal pollution data and sample analysis data using a data management system.
[0025] Secondly, this specification provides an analytical apparatus for the source-sink relationship between soil from historically abandoned mines and surrounding farmland, comprising:
[0026] The pollution data collection unit is used to collect heavy metal pollution data, which includes mining production-related data and historical pollution information.
[0027] The results database creation unit is used to obtain sample analysis data, collect and summarize heavy metal pollution data and sample analysis data to form the results database;
[0028] The sample extraction unit is used to extract sample input data from the results database;
[0029] The correlation coefficient calculation unit is used to input sample input data into a geographic weighted regression model to obtain the correlation coefficient between the heavy metal content of surrounding farmland soil and the heavy metal content of mine solid waste samples.
[0030] Among them, the geographically weighted regression model is constructed by weighting geographical location on the basis of the local regression analysis model. The independent variable of the local regression analysis model is the heavy metal content in the mine solid waste sample, and the dependent variable is the heavy metal content in the farmland soil sample. Geographical location represents the geographical relationship between the farmland soil sample and the mine solid waste sample. The rationality and accuracy of the correlation coefficient are determined based on the heavy metal content in the acidic wastewater sample, irrigation water sample and sediment sample.
[0031] According to the second aspect, in one possible implementation, the parameters of the local regression analysis model are determined by ordinary least squares, and the bandwidth of the geographically weighted regression model is determined by the Akaike Information Criterion.
[0032] According to the second aspect, in one possible implementation, the pollution data collection unit further includes a verification and supplementation module for verifying and supplementing heavy metal pollution data.
[0033] According to the second aspect, in one possible implementation, the correlation coefficient calculation unit further includes:
[0034] The verification module is used to verify correlation coefficients based on geographical conditions, meteorological conditions, and historical mining activities.
[0035] The analysis and evaluation module is used to analyze and evaluate mine pollution based on the mine pollution source-sink relationship evaluation system. The mine pollution source-sink relationship evaluation system is constructed based on geographical conditions, meteorological conditions, mineral characteristics, mining technology and mining years. The correlation coefficient is significantly correlated with mineral characteristics and mining years.
[0036] According to the second aspect, in one possible implementation, the pollution data collection unit further includes a priority determination module for determining the priority of the mine based on geographic information spatial overlay analysis and research literature, wherein the geographic information spatial overlay analysis includes spatial analysis and layer overlay, and the priority characterizes the degree of attention paid to the mine by the academic community or public opinion.
[0037] According to the second aspect, in one possible implementation, the results database creation unit further includes a site layout module for determining the layout sites of samples based on mine pollution diffusion conditions, pollution correlations, and empirical analysis methods.
[0038] According to the second aspect, in one possible implementation, the conditions for the diffusion of mine pollution include meteorological conditions, topographic features, the physical and chemical properties of pollutants, and industrial and agricultural activities.
[0039] According to the second aspect, in one possible implementation, the results database creation unit uses the data management system to obtain sample analysis data and collects and summarizes heavy metal pollution data and sample analysis data.
[0040] Thirdly, this specification provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the steps in the analysis method described above.
[0041] Fourthly, this specification provides a computer-readable storage medium storing a computer program that performs the steps of the analysis method described above.
[0042] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0043] Unlike existing technologies, this invention provides a method for analyzing the source-sink relationship between historical legacy mines and surrounding farmland soil. This method involves collecting heavy metal pollution data, including mine production-related information and historical pollution data; collecting and analyzing samples, including mine solid waste samples, acidic wastewater samples, irrigation water samples, sediment samples, and farmland soil samples; collecting and summarizing the heavy metal pollution data and sample analysis data to form a results database; extracting sample input data from the results database; and inputting the sample input data into a geographically weighted regression model to obtain the correlation coefficient between the heavy metal content of surrounding farmland soil and the heavy metal content of mine solid waste samples. The geographically weighted regression model is constructed by weighting geographical location on a local regression analysis model. The independent variable of the local regression analysis model is the heavy metal content in the mine solid waste sample, and the dependent variable is the heavy metal content in the farmland soil sample. Geographical location represents the geographical relationship between the farmland soil sample and the mine solid waste sample. The rationality and accuracy of the correlation coefficient are determined based on the heavy metal content in the acidic wastewater sample, irrigation water sample, and sediment sample. This scheme uses a geographic weighted regression model obtained by weighting geographic location on the basis of local regression analysis model to calculate the correlation coefficient. Since this model can reflect the spatial heterogeneity of the sample, it solves the problems of data discontinuity and low accuracy caused by spatial heterogeneity on the basis of quantitative analysis, and realizes quantitative and accurate source apportionment of heavy metals in the soil around the mine. Attached Figure Description
[0044] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0045] Figure 1 A flowchart illustrating an analysis method for the source-sink relationship between soil in historical legacy mines and surrounding farmland, provided as an embodiment of this application;
[0046] Figure 2 A schematic diagram of an analysis device for the source-sink relationship of soil in historical abandoned mines and surrounding farmland provided for an embodiment of this application;
[0047] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0048] Unless otherwise defined, the technical or scientific terms used in the embodiments of this specification shall have the ordinary meaning understood by one of ordinary skill in the art to which this specification pertains. The terms "first," "second," and similar terms used in the embodiments of this specification do not indicate any order, quantity, or importance, but are merely used to avoid confusion of constituent elements.
[0049] Unless the context otherwise requires, throughout this specification, "a plurality of" means "at least two," and "including" is interpreted as open-ended or encompassing, that is, "including, but not limited to." In the description of this specification, terms such as "one embodiment," "some embodiments," "exemplary embodiment," "example," "specific example," or "some examples" are intended to indicate that a particular feature, structure, material, or characteristic associated with that embodiment or example is included in at least one embodiment or example of this specification. The illustrative representations of the above terms do not necessarily refer to the same embodiment or example.
[0050] The technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this specification.
[0051] like Figure 1 The diagram shown is a flowchart illustrating an analysis method for the source-sink relationship of soil in historical abandoned mines and surrounding farmland provided by an embodiment of the present invention. The analysis method may specifically include the following steps.
[0052] S110: Collect heavy metal pollution data, including mining production-related data and historical pollution information.
[0053] S120: Collect samples for analysis to obtain sample analysis data, including samples of mine solid waste, acidic wastewater, irrigation water, sediment, and farmland soil;
[0054] S130: Collect and summarize heavy metal pollution data and sample analysis data to form a results database;
[0055] S140: Extract sample input data from the results database;
[0056] S150: Input the sample input data into the geographic weighted regression model to obtain the correlation coefficient between the heavy metal content of the surrounding farmland soil and the heavy metal content of the mine solid waste sample;
[0057] Among them, the geographically weighted regression model is constructed by weighting geographical location on the basis of the local regression analysis model. The independent variable of the local regression analysis model is the heavy metal content in the mine solid waste sample, and the dependent variable is the heavy metal content in the farmland soil sample. Geographical location represents the geographical relationship between the farmland soil sample and the mine solid waste sample. The rationality and accuracy of the correlation coefficient are determined based on the heavy metal content in the acidic wastewater sample, irrigation water sample and sediment sample.
[0058] The aforementioned methods cover mines containing both non-ferrous and ferrous metals, as well as typical non-metallic minerals such as pyrite, coal shale, and phosphate rock. Mine production-related data can include mineral type, mining history, mining methods, beneficiation processes, on-site beneficiation and smelting conditions, solid waste properties and storage conditions, distance to surrounding farmland, pollution prevention measures, irrigation methods and sources for surrounding farmland, etc. Historical pollution information can include heavy metal pollution data, documents, and journal reports related to the mine being analyzed. When collecting samples, solid waste samples can be collected from historical mines, and acidic wastewater samples, irrigation water samples, sediment samples, and farmland soil samples can be collected from the surrounding area for analysis to obtain comprehensive data. During the mining, beneficiation, and smelting processes of metal mines, acidic wastewater with low pH and containing heavy metals is generated. This wastewater can pollute the surrounding environment. Irrigation water and sediment samples are used to analyze water pollution around the mine, while farmland soil samples are used to analyze soil pollution around the mine. Information technology tools such as data management systems can be used to collect and summarize heavy metal pollution data and sample analysis data to form a results database. The extracted sample input data can be sample analysis data from mine solid waste samples and agricultural land soil samples. Sample analysis data from acidic wastewater samples, irrigation water samples, and sediment samples can be used as controls to determine the rationality and accuracy of the data, but not directly in the calculation. Finally, a geographically weighted regression model is used to calculate the correlation coefficient between the heavy metal content of surrounding agricultural land soil and the mine. Geographical location can characterize the distance and orientation of the agricultural land soil sample and the mine. Geographical weighting of the local regression analysis model can be applied based on the geographical location of the samples, weighting the regression coefficients and predicted values (intercepts) of the corresponding regression equations. The correlation coefficient is a measure of the degree of linear correlation between variables, generally represented by the letter r. The correlation coefficient is a value between -1 and 1; the larger the absolute value, the stronger the correlation. Furthermore, the obtained correlation coefficients can be added to the results database to form an updated results database for subsequent analysis and use.
[0059] This embodiment calculates the correlation coefficient by using a geographic weighted regression model obtained by weighting geographic location on the basis of a local regression analysis model. Since this model can reflect the spatial heterogeneity of the sample, it solves the problems of data discontinuity and low accuracy caused by spatial heterogeneity on the basis of quantitative analysis, and realizes quantitative and accurate source apportionment of heavy metals in the soil around the mine.
[0060] In at least one embodiment of this application, in order to utilize mature regression analysis and determine a reasonable analysis range, this application determines the local regression analysis model based on ordinary least squares (OLS) and the bandwidth of the geographically weighted regression model based on the Akaike information criterion. OLS is a mature regression analysis method that ensures the reliability of the regression analysis. In the regression model, bandwidth is an important parameter that directly affects the estimation accuracy and precision of the model. The choice of bandwidth involves balancing the smoothness and fit of the data; too small a bandwidth may lead to overfitting, while too large a bandwidth may lead to underfitting. After determining the bandwidth, the location of the mine to be analyzed can be used as the regression point to determine the farthest sample point. Methods for determining the bandwidth of the geographically weighted regression model include cross-validation (CV), the Akaike information criterion (or minimum information criterion, AIC), and the generalized cross-validation criterion (GVC). Simulation experiments and experience show that the CV and GVC criteria generally tend to determine a slightly smaller bandwidth, and a smaller bandwidth reduces the deviation of the estimated value of the regression function. AIC requires fewer model parameters, balancing simplicity and accuracy; therefore, this solution uses AIC to determine the bandwidth.
[0061] For example, the heavy metal content in mine solid waste samples and heavy metal content in farmland soil samples were used to construct regression equations using ordinary least squares, as follows.
[0062] y i = β0 + β1x i +ε i (1)
[0063] In the formula, y i It is the value of the dependent variable corresponding to the i-th sampling position, x i β is the value of the independent variable at the i-th sampling position, β0 is the y-intercept, β1 is the regression coefficient estimated for the independent variable at the i-th position, and ε i This is the error term. i There can be N, where N is a positive integer.
[0064] Geographically weighted regression models provide local coefficient estimates by adding geographic location to the function, as shown below.
[0065] y i = β0 (u i ,v i ) + β1 (u i ,v i )x i +εi (2)
[0066] In the formula, u i and v i The coordinates of the i-th sampling position are β0(u i v i ) is the intercept at the i-th sampling position, β1(u i v i ) represents the regression coefficient of the independent variable at location i. Therefore, by geographically weighting the intercept and regression coefficients of the regression equation in the local regression analysis model, the regression equation of the geographically weighted regression model can be constructed.
[0067] The geographically weighted regression model uses a weighted regression function to estimate the regression coefficients. Based on equation (2), it can be transformed into equation (3) below.
[0068] β (u i ,v i ) = (X T W (u i ,v i ) X) -1 X T W (u i ,v i Y (3)
[0069] In the formula, matrix X corresponds to the independent variable x. i Matrix Y corresponds to the dependent variable y i W(ui, vi) is the weight matrix chosen to ensure that closer observations have a greater impact on the results.
[0070] Based on the above formula, and combined with the formula for calculating the Pearson correlation coefficient, we can obtain the formula for calculating the correlation coefficient (r) (4).
[0071]
[0072] In the formula, R 2 i R represents the coefficient of determination. 2 This coefficient represents the extent to which the heavy metal content in solid waste explains the heavy metal content in soil or sediment in a geographically weighted regression model, β1(u i v i ) is related to β(u) in equation (3) i v i The same regression coefficient. The correlation coefficient r can be calculated. i To determine whether the dependent and independent variables are strongly or weakly correlated, a higher correlation coefficient value indicates that the heavy metal content in soil or sediment is more affected by the heavy metal content in solid waste.
[0073] In at least one embodiment of this application, to ensure the accuracy and comprehensiveness of heavy metal pollution data, the analysis method further includes verifying and supplementing the heavy metal pollution data. Verification and supplementation methods include on-site investigation, personnel interviews, and comparisons of water resource extraction and utilization. By verifying and supplementing the heavy metal pollution data, the accuracy and comprehensiveness of the heavy metal pollution data can be guaranteed.
[0074] In at least one embodiment of this application, in order to test the correlation coefficient of heavy metal pollution and obtain reliable analytical results, the analytical method further includes: testing the correlation coefficient based on geographical conditions, meteorological conditions and historical mining activities; and analyzing and evaluating the pollution of the mine based on the mine pollution source-sink relationship evaluation system, wherein the mine pollution source-sink relationship evaluation system is constructed based on geographical conditions, meteorological conditions, mineral characteristics, mining technology and mining years, and the correlation coefficient is significantly correlated with mineral characteristics and mining years (i.e., a positive correlation after significance testing).
[0075] The calculation of correlation coefficients can sometimes deviate significantly due to factors such as the randomness of sampling points and the sample size. In such cases, the correlation coefficient can be tested by considering the geographical conditions, meteorological conditions, and historical mining activities of the sampling points, allowing for appropriate adjustments, such as changing sampling points or adjusting the sample size. After obtaining accurate correlation coefficients, the pollution of the mine can be analyzed and evaluated based on the constructed mine pollution source-sink relationship evaluation system. Specific evaluation indicators can be set as needed, such as pollution severity, pollution range, urgency of remediation, difficulty of remediation, and rate of return on remediation. Different mineral types and different mining years are reflected in different source-sink relationships in the correlation coefficients. For example, the source-sink relationship between lead-zinc mines and surrounding farmland soil is stronger than that between sand and gravel mines. This is reflected in the fact that the correlation coefficient calculated for lead-zinc mines will be higher than that for sand and gravel mines in terms of heavy metal source-sink relationships. The same principle applies to mining years; mines with longer mining years will have higher calculated correlation coefficients than those with shorter mining years.
[0076] In at least one embodiment of this application, in order to distinguish the severity and urgency of pollution sources, the analysis method further includes: determining the priority of mines based on geographic information spatial overlay analysis and research literature, wherein the geographic information spatial overlay analysis includes spatial analysis and layer overlay, and the priority represents the degree of attention paid to mines by the academic community or public opinion.
[0077] Spatial analysis is a technique in Geographic Information Systems (GIS) that uses geographic data to analyze spatial relationships, patterns, and trends. Spatial analysis can reveal hidden patterns in data, predict future developments, or assess certain characteristics of geospatial data. Layer overlay is a commonly used technique in GIS that involves stacking multiple geographic data layers according to certain rules to generate new information. Each layer can contain different types of geographic data, such as topography, land use, and transportation networks. Spatial analysis and layer overlay can be combined with information on water sources and farmland for analysis. For example, mines surrounding water source protection areas and farmland can be prioritized for inclusion in surveys. High-priority mines can be configured with higher sampling densities, thus more accurately and comprehensively reflecting the source-sink relationships of heavy metals between the mine and the surrounding soil.
[0078] In at least one embodiment of this application, to obtain reasonable sampling locations, the analytical method further includes: determining the sample placement locations based on mine pollution diffusion conditions, pollution correlations, and empirical analysis. Mine pollution diffusion conditions may include meteorological conditions, topographic features, the physical and chemical properties of pollutants, and industrial and agricultural activities. For example, a region with high precipitation has better pollution diffusion conditions. Pollution correlations may include geographical location and industrial and agricultural activities. For example, if the soil in area A has a high heavy metal content and there is a nearby mine B, although mine B is not included in the source-sink relationship, it can be included in the analysis because of its geographical proximity to area A. Empirical analysis is a scientific analytical method that relies primarily on empirical knowledge to analyze and understand things. This method is more intuitive and closer to reality when compared with theoretical analysis methods. The basic idea of empirical analysis is to conduct in-depth exploration and analysis of the research object through the experience, wisdom, and collective wisdom of the analysts, thereby drawing reasonable conclusions or selecting the best solution.
[0079] In at least one embodiment of this application, in order to analyze the collected samples, the analysis method further includes: determining single-element analysis indicators and testing methods; and analyzing the samples according to the single-element analysis indicators and testing methods. Single-element analysis indicators refer to evaluation standards for solid waste, and the testing indicators include more than eight heavy metals, which can be expanded if necessary. Single-element analysis indicators can be determined based on task requirements, the mining conditions of the area to be analyzed, and the heavy metal pollution status of the area to be analyzed. Testing methods and the optimal method can be determined by comprehensively considering factors such as detection cost, sample content, and detection limit.
[0080] In at least one embodiment of this application, information technology can be used to collect and summarize heavy metal pollution data and sample analysis data. For example, a data management system can be used to collect and summarize heavy metal pollution data and sample analysis data. Furthermore, software and hardware products related to fixed and mobile devices can be developed for different scenarios to conduct source tracing analysis and management of heavy metal pollution.
[0081] In at least one embodiment of this application, the principles and single-element analysis indicators for determining the deployment points can be based on national and industry standards. Relevant national and industry standards include HJ / T 20-1998, HJ 91.1-2019, DZ / T0295-2016, and DZ / T 0289-2015, etc.
[0082] This application also provides an analytical device for analyzing the source-sink relationship between soils from historically abandoned mines and surrounding farmland, corresponding to the above-mentioned analytical method.
[0083] like Figure 2 As shown, an analysis device 200 for analyzing the source-sink relationship of soil in historical legacy mines and surrounding farmland is provided as an embodiment of the present invention. The analysis device includes:
[0084] Pollution data collection unit 210 is used to collect heavy metal pollution data, which includes mining production-related data and historical pollution information.
[0085] The results database creation unit 220 is used to obtain sample analysis data, collect and summarize heavy metal pollution data and sample analysis data to form the results database;
[0086] The sample extraction unit 230 is used to extract sample input data from the results database;
[0087] The correlation coefficient calculation unit 240 is used to input the sample input data into the geographic weighted regression model to obtain the correlation coefficient between the heavy metal content of the surrounding farmland soil and the heavy metal content of the mine solid waste sample.
[0088] Among them, the geographically weighted regression model is constructed by weighting geographical location on the basis of the local regression analysis model. The independent variable of the local regression analysis model is the heavy metal content in the mine solid waste sample, and the dependent variable is the heavy metal content in the farmland soil sample. Geographical location represents the geographical relationship between the farmland soil sample and the mine solid waste sample. The rationality and accuracy of the correlation coefficient are determined based on the heavy metal content in the acidic wastewater sample, irrigation water sample and sediment sample.
[0089] The aforementioned apparatus includes mines containing non-ferrous and ferrous metal ores, as well as typical non-metallic minerals such as pyrite, coal shale, and phosphate rock. Mine production-related data can include mineral type, mining history, mining methods, beneficiation processes, on-site beneficiation and smelting conditions, solid waste properties and storage conditions, distance to surrounding farmland, pollution prevention measures, irrigation methods and sources of irrigation water for surrounding farmland, etc. Historical pollution information can include heavy metal pollution data, documents, and journal reports related to the mine to be analyzed. When collecting samples, solid waste samples can be collected from historical mine sites, and acidic wastewater samples, irrigation water samples, sediment samples, and farmland soil samples can be collected from the surrounding area for analysis to obtain comprehensive data. During the mining, beneficiation, and smelting processes of metal mines, acidic wastewater with low pH and containing heavy metals is generated. This wastewater can pollute the surrounding environment. Irrigation water and sediment samples are used to analyze water pollution around the mine, while farmland soil samples are used to analyze soil pollution around the mine. Information technology tools such as data management systems can be used to collect and summarize heavy metal pollution data and sample analysis data to form a results database. The extracted sample input data can be sample analysis data from mine solid waste samples and agricultural land soil samples. Sample analysis data from acidic wastewater samples, irrigation water samples, and sediment samples can be used as controls to determine the rationality and accuracy of the data, but not directly in the calculation. Finally, a geographically weighted regression model is used to calculate the correlation coefficient between the heavy metal content of surrounding agricultural land soil and the mine. Geographical location can characterize the distance and orientation of the agricultural land soil sample and the mine. Geographical weighting of the local regression analysis model can be applied based on the geographical location of the samples, weighting the regression coefficients and predicted values (intercepts) of the corresponding regression equations. The correlation coefficient is a measure of the degree of linear correlation between variables, generally represented by the letter r. The correlation coefficient is a value between -1 and 1; the larger the absolute value, the stronger the correlation. Furthermore, the obtained correlation coefficients can be added to the results database to form an updated results database for subsequent analysis and use.
[0090] This embodiment calculates the correlation coefficient by using a geographic weighted regression model obtained by weighting geographic location on the basis of a local regression analysis model. Since this model can reflect the spatial heterogeneity of the sample, it solves the problems of data discontinuity and low accuracy caused by spatial heterogeneity on the basis of quantitative analysis, and realizes quantitative and accurate source apportionment of heavy metals in the soil around the mine.
[0091] In at least one embodiment of this application, in order to utilize mature regression analysis and determine a reasonable analysis range, this application determines the local regression analysis model based on ordinary least squares (OLS) and the bandwidth of the geographically weighted regression model based on the Akaike information criterion. OLS is a mature regression analysis method that ensures the reliability of the regression analysis. In the regression model, bandwidth is an important parameter that directly affects the estimation accuracy and precision of the model. The choice of bandwidth involves balancing the smoothness and fit of the data; too small a bandwidth may lead to overfitting, while too large a bandwidth may lead to underfitting. After determining the bandwidth, the location of the mine to be analyzed can be used as the regression point to determine the farthest sample point. Methods for determining the bandwidth of the geographically weighted regression model include cross-validation, the Akaike information criterion (or minimum information criterion, AIC), and the generalized cross-validation criterion (GCV). Simulation experiments and experience show that the CV and GCV criteria generally tend to determine a slightly smaller bandwidth, and a smaller bandwidth reduces the deviation of the estimated value of the regression function. AIC requires fewer model parameters, balancing simplicity and accuracy; therefore, this solution uses AIC to determine the bandwidth.
[0092] In at least one embodiment of this application, to ensure the accuracy and comprehensiveness of heavy metal pollution data, the pollution data collection unit further includes a verification and supplementation module for verifying and supplementing the heavy metal pollution data. By verifying and supplementing the heavy metal pollution data, the accuracy and comprehensiveness of the heavy metal pollution data can be guaranteed.
[0093] In at least one embodiment of this application, in order to verify the correlation coefficient of heavy metal pollution and obtain reliable analytical results, the correlation coefficient calculation unit further includes: a verification module for verifying the correlation coefficient based on geographical conditions, meteorological conditions, and historical mining activities; and an analysis and evaluation module for analyzing and evaluating mine pollution based on a mine pollution source-sink relationship evaluation system, wherein the mine pollution source-sink relationship evaluation system is constructed based on geographical conditions, meteorological conditions, mineral characteristics, mining technology, and mining years, and the correlation coefficient is significantly correlated with mineral characteristics and mining years.
[0094] The calculation of correlation coefficients can sometimes deviate significantly due to factors such as the randomness of sampling points and the sample size. In such cases, the correlation coefficient can be tested by considering the geographical conditions, meteorological conditions, and historical mining activities of the sampling points, allowing for appropriate adjustments, such as changing sampling points or adjusting the sample size. After obtaining accurate correlation coefficients, the pollution of the mine can be analyzed and evaluated based on the constructed mine pollution source-sink relationship evaluation system. Specific evaluation indicators can be set as needed, such as pollution severity, pollution range, urgency of remediation, difficulty of remediation, and rate of return on remediation. Different mineral types and different mining years are reflected in different source-sink relationships in the correlation coefficients. For example, the source-sink relationship between lead-zinc mines and surrounding farmland soil is stronger than that between sand and gravel mines. This is reflected in the fact that the correlation coefficient calculated for lead-zinc mines will be higher than that for sand and gravel mines in terms of heavy metal source-sink relationships. The same principle applies to mining years; mines with longer mining years will have higher calculated correlation coefficients than those with shorter mining years.
[0095] In at least one embodiment of this application, in order to distinguish the severity and urgency of pollution sources, the pollution data collection unit further includes a priority determination module, which is used to determine the priority of the mine based on geographic information spatial overlay analysis and research literature. The geographic information spatial overlay analysis includes spatial analysis and layer overlay, and the priority represents the degree of attention paid to the mine by the academic community or public opinion.
[0096] Spatial analysis is a technique in Geographic Information Systems (GIS) that uses geographic data to analyze spatial relationships, patterns, and trends. Spatial analysis can reveal hidden patterns in data, predict future developments, or assess certain characteristics of geospatial data. Layer overlay is a commonly used technique in GIS that involves stacking multiple geographic data layers according to certain rules to generate new information. Each layer can contain different types of geographic data, such as topography, land use, and transportation networks. Spatial analysis and layer overlay can be combined with information on water sources and farmland for analysis. For example, mines surrounding water source protection areas and farmland can be prioritized for inclusion in surveys. High-priority mines can be configured with higher sampling densities, thus more accurately and comprehensively reflecting the source-sink relationships of heavy metals between the mine and the surrounding soil.
[0097] In at least one embodiment of this application, to obtain reasonable sampling locations, the results database creation unit further includes a sampling location layout module, used to determine the sampling locations based on mine pollution diffusion conditions, pollution correlations, and empirical analysis methods. Mine pollution diffusion conditions may include meteorological conditions, topographic features, the physical and chemical properties of pollutants, and industrial and agricultural activities. For example, a region with high precipitation has better pollution diffusion conditions. Pollution correlations may include geographical location, industrial and agricultural activities, etc. For example, if the soil in area A has a high heavy metal content and there is a nearby mine B, although mine B is not included in the source-sink relationship, it can be included in the analysis due to its geographical proximity. Empirical analysis is a scientific analysis method that mainly relies on empirical knowledge to analyze and understand things. This method is more intuitive and closer to reality when compared with theoretical analysis methods. The basic idea of empirical analysis is to conduct in-depth exploration and analysis of the research object through the experience, wisdom, and collective wisdom of the analysts, thereby drawing reasonable conclusions or selecting the best solution.
[0098] In at least one embodiment of this application, information technology can be used to collect and summarize heavy metal pollution data and sample analysis data. For example, the results database creation unit uses a data management system to obtain sample analysis data and collects and summarizes the heavy metal pollution data and sample analysis data. Furthermore, software and hardware products related to fixed and mobile equipment can be developed for different scenarios to conduct source tracing analysis and management of heavy metal pollution.
[0099] like Figure 3 The figure shows a schematic diagram of an electronic device according to an embodiment of the present application. As shown, the electronic device 30 includes one or more processors 31 and a memory 32.
[0100] The processor 31 may be a central processing unit (CPU) or other form of processing unit with data processing and / or instruction execution capabilities, and may control other components in the electronic device 30 to perform desired functions.
[0101] The memory 32 may include one or more computer program products. The memory may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 31 may execute the program instructions to implement the analysis methods and / or other desired functions described in the foregoing embodiments of this application. Various contents such as input signals, signal components, and noise components may also be stored in the computer-readable storage medium.
[0102] In one example, the electronic device 30 may also include an input device 33 and an output device 34, which are interconnected via a bus system and / or other forms of connection mechanism (not shown).
[0103] Of course, for the sake of simplicity, Figure 3 Only some of the components of the electronic device 30 relevant to this application are shown in this illustration; components such as buses, input / output interfaces, etc., are omitted. In addition, the electronic device 30 may include any other suitable components depending on the specific application.
[0104] Based on embodiments of this application, this application also provides a computer-readable storage medium, wherein computer instructions are used to cause a computer to perform the steps in the analysis methods of the foregoing embodiments.
[0105] Based on the embodiments of this application, this application also provides a processor for running a program, wherein the program executes the steps of the analysis methods in the foregoing embodiments.
[0106] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing custom logic functions or processes, and the scope of preferred embodiments of the invention includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as will be understood by those skilled in the art to which embodiments of the invention pertain.
[0107] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-including system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.
[0108] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the described embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system.
[0109] The features described above in the disclosed embodiments can be substituted or combined with each other, enabling those skilled in the art to implement or use this application. The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the invention. Any person skilled in the art can make many possible variations and modifications to the technical solutions of the present invention using the methods and techniques disclosed above, or modify them into equivalent embodiments with equivalent changes, without departing from the scope of the present invention. Therefore, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the present invention's technical solutions still fall within the protection scope of the present invention.
Claims
1. A method for analyzing the source-sink relationship between soil from historically abandoned mines and surrounding farmland, characterized in that, include: Collect heavy metal pollution data, which includes mining production-related data and historical pollution information. Samples were collected and analyzed to obtain sample analysis data, including mine solid waste samples, acidic wastewater samples, irrigation water samples, sediment samples, and farmland soil samples. The heavy metal pollution data and the sample analysis data are collected and summarized to form a results database; Extract sample input data from the results database; The sample input data is input into a geographic weighted regression model to obtain the correlation coefficient between the heavy metal content of the surrounding farmland soil and the heavy metal content of the mine solid waste sample; After obtaining the correlation coefficient between the heavy metal content of the surrounding farmland soil and the solid waste sample from the mine, the correlation coefficient was tested based on geographical conditions, meteorological conditions, and historical mining activities. Based on the mine pollution source-sink relationship evaluation system, the pollution of the mine is analyzed and evaluated. The mine pollution source-sink relationship evaluation system is constructed based on geographical conditions, meteorological conditions, mineral characteristics, mining technology and mining years. The correlation coefficient is significantly correlated with mineral characteristics and mining years. The geographically weighted regression model is constructed by weighting geographical locations on the basis of the local regression analysis model. The independent variable of the local regression analysis model is the heavy metal content in the mine solid waste sample, and the dependent variable is the heavy metal content in the farmland soil sample. The geographical location represents the geographical relationship between the farmland soil sample and the mine solid waste sample. The rationality and accuracy of the correlation coefficient are determined based on the heavy metal content in the acidic wastewater sample, the irrigation water sample, and the sediment sample. The parameters of the local regression analysis model are determined using ordinary least squares, and the bandwidth of the geographically weighted regression model is determined using the Akaike Information Criterion. The process of obtaining the correlation coefficient between the heavy metal content of the surrounding farmland soil and the heavy metal content of the mine solid waste sample includes the following steps: The heavy metal content in mine solid waste samples and agricultural land soil samples were used to construct regression equations using the ordinary least squares method, as shown in formula (1): (1) In the formula, It is the value of the dependent variable corresponding to the i-th sampling position, x i β is the value of the independent variable at the i-th sampling position, β0 is the y-intercept, β1 is the regression coefficient estimated for the independent variable at the i-th position, and ε i It is the error term, x i There are N, where N is a positive integer; Geographically weighted regression models provide local coefficient estimates by adding geographic location to the function, as shown below: (2) In the formula, and These are the coordinates of the i-th sampling position. It is the intercept at the i-th sampling position. It is the regression coefficient of the independent variable at position i; The geographically weighted regression model uses a weighted regression function to estimate the regression coefficients. Based on equation (2), it can be transformed into the following equation (3): (3) In the formula, matrix X corresponds to the independent variable x. i Matrix Y corresponds to the dependent variable , It is the weight matrix that is chosen to ensure that closer observations have a greater impact on the results; Based on the above formulas, and combined with the formula for calculating the Pearson correlation coefficient, we can obtain the formula (4) for calculating the correlation coefficient (r): (4) In the formula, Coefficient of determination This coefficient represents the extent to which the heavy metal content in solid waste explains the heavy metal content in soil or sediment in a geographically weighted regression model. It is in equation (3) The same regression coefficient can be used to calculate the correlation coefficient. To determine whether the dependent and independent variables are strongly or weakly correlated, a higher correlation coefficient value indicates that the heavy metal content in soil or sediment is more affected by the heavy metal content in solid waste.
2. The analytical method according to claim 1, characterized in that, Following the collection of heavy metal pollution data, the process also includes: verifying and supplementing the heavy metal pollution data.
3. The analytical method according to claim 1, characterized in that, Before collecting heavy metal pollution data, the process also includes: determining the priority of the mine based on geographic information spatial overlay analysis and literature review, wherein the geographic information spatial overlay analysis includes spatial analysis and layer overlay, and the priority represents the degree of attention paid to the mine by the academic community or public opinion.
4. The analytical method according to claim 1, characterized in that, Before analyzing the collected samples, the method further includes: determining the location of the sample placement points based on mine pollution diffusion conditions, pollution correlation, and empirical analysis.
5. The analytical method according to claim 4, characterized in that, The analysis of the collected samples includes: determining single-factor analysis indicators and testing methods; The sample is analyzed based on the single-factor analysis indicators and the testing method.
6. The analytical method according to claim 4, characterized in that, The conditions for the diffusion of mine pollution include meteorological conditions, topographic features, the physical and chemical properties of pollutants, and industrial and agricultural activities.
7. The analytical method according to claim 1, characterized in that, The step of collecting and summarizing the heavy metal pollution data and the sample analysis data includes: using a data management system to collect and summarize the heavy metal pollution data and the sample analysis data.
8. An analytical apparatus for analyzing the source-sink relationship of soil in historical legacy mines and surrounding farmland, applied to the analytical method described in any one of claims 1 to 7, characterized in that, include: A pollution data collection unit is used to collect heavy metal pollution data, wherein the heavy metal pollution data includes mining production-related data and historical pollution status information. The results database creation unit is used to obtain sample analysis data, collect and summarize the heavy metal pollution data and the sample analysis data to form the results database; A sample extraction unit is used to extract sample input data from the results database; The correlation coefficient calculation unit is used to input the sample input data into the geographic weighted regression model to obtain the correlation coefficient between the heavy metal content of the surrounding farmland soil and the heavy metal content of the mine solid waste sample; The geographically weighted regression model is constructed by weighting geographical locations on the basis of the local regression analysis model. The independent variable of the local regression analysis model is the heavy metal content in the mine solid waste sample, and the dependent variable is the heavy metal content in the farmland soil sample. The geographical location represents the geographical relationship between the farmland soil sample and the mine solid waste sample. The rationality and accuracy of the correlation coefficient are determined based on the heavy metal content in the acidic wastewater sample, irrigation water sample, and sediment sample.