A method for spatial prediction of ammonium nitrogen pollution in soil of rare earth mining area
By using a hierarchical Bayesian three-dimensional surface mean model, combined with spatial hierarchical heterogeneity and Bayesian theory, the problem of traditional methods failing to handle spatial hierarchical heterogeneity and sample sparsity in the prediction of ammonium nitrogen pollution in rare earth mining areas is solved, and more accurate and robust prediction of pollutant spatial distribution is achieved.
Patent Information
- Application Number
- CN202411464259.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-21
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2044-10-21
AI Technical Summary
Traditional methods fail to adequately consider spatial stratification heterogeneity and the influence of samples at different levels when dealing with ammonium nitrogen pollution in rare earth mining areas, resulting in inaccurate predictions of pollutant spatial distribution. In particular, it is difficult to capture the spatial structure of pollutants in complex terrain and multi-layered structures when borehole data is sparse.
A hierarchical Bayesian three-dimensional surface mean model is adopted, combining spatial hierarchical heterogeneity theory and Bayesian theory. The spatial autocorrelation and heterogeneity of pollutants are identified by Moran's index and geographic detector. The parameters are estimated by using mixed covariance matrix and Markov chain Monte Carlo sampling method, and a robust pollutant spatial distribution prediction model is constructed.
It improves the accuracy and robustness of pollutant spatial distribution prediction, accurately captures pollutant distribution in complex environments and multi-level structures, enhances the model's adaptability and predictive ability, and can still provide stable prediction results, especially in cases of sparse samples.
Smart Images

Figure CN119442859B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the field of ammonium nitrogen pollution space prediction, and particularly relates to a hierarchical Bayesian model, a mixed covariance matrix and environmental protection in a mining area. BACKGROUND
[0002] The in-situ leaching method of ion-adsorbed rare earth ore is prone to cause serious water and soil pollution problems in the mining area, among which the ammonium nitrogen pollution is particularly prominent. The ammonium nitrogen pollution in the rare earth mining area shows significant spatial autocorrelation and heterogeneity. The spatial autocorrelation is mainly affected by environmental factors such as rainfall, terrain slope and mining activities, forming a three-dimensional spatial autocorrelation pattern of pollution. The spatial heterogeneity is closely related to the existing form of ammonium nitrogen and the leaching process of ammonium nitrogen in the soil.
[0003] With the closure of the mine, the source of ammonium nitrogen pollution is cut off, but the ammonium nitrogen remaining in the soil during the leaching process shows different degrees of aggregation in the soil at different depths due to the leaching effect, eventually leading to spatial heterogeneity in the vertical direction. This complex spatial distribution of pollutants, combined with the complex topography of the mining area, makes the traditional prediction method perform poorly in dealing with the spatial distribution of pollutants in the mining area, and there are still some key problems: first, the spatial stratified heterogeneity is not fully considered, many existing methods do not consider the spatial stratified heterogeneity, and the heterogeneity of the pollutant distribution between different geological layers or depths is not reasonably modeled; second, the influence of potential samples in different layers is ignored, some methods do not fully consider the influence of potential samples between different layers in modeling, especially in the presence of strong heterogeneity and weak autocorrelation, it is difficult to accurately reflect the correlation between pollutants in different layers; third, the spatial structure is not accurately captured due to the sparsity of drilling data. Even if the spatial stratified heterogeneity and the influence of samples in different layers are considered in the modeling scheme, in the case of sparse drilling data, it is often difficult to accurately capture the spatial structure of the pollutants. Since the spatial structure is usually expressed by a semi-variogram function or a covariance function, the parameters of these functions need sufficient sample support to ensure stable estimation, and in the stratified heterogeneity modeling, the reduction of the number of local layer samples will lead to inaccurate capture of the local layer spatial structure, thereby affecting the prediction performance of the overall model.
[0004] These problems have a direct impact on the accuracy and robustness of model prediction, especially in the spatial distribution prediction task of pollutants in the mining area, there are factors such as sample sparsity, strong heterogeneity and complex topography, and the defects of traditional methods are particularly prominent. Therefore, for the mining area with obvious stratified characteristics, it is urgent to develop a robust modeling method to improve the accuracy of the spatial distribution prediction of pollutants, to meet the needs of environmental protection and sustainable development of the mining area. SUMMARY
[0005] To this end, the present application provides a method for spatial prediction of ammonium nitrogen pollution in rare earth mining area soil, based on spatial stratified heterogeneity theory and Bayesian theory, combined with spatial interpolation technology, an innovative stratified Bayesian three-dimensional curved surface mean model is proposed to solve the problem of insufficient precision in identifying and processing spatial heterogeneity of pollutants.
[0006] The specific implementation steps of the model are:
[0007] Step 1: Based on the characteristics of rare earth mining process and ammonia nitrogen pollution, combined with mountain terrain and leaching path, the sampling area is divided and the drilling position is determined, and the sample collection and laboratory analysis are completed.
[0008] Step 2: Identify the spatial pattern of ammonium nitrogen pollutants. The spatial autocorrelation and spatial stratified heterogeneity of ammonium nitrogen pollutants are identified by Moran's index and geographic detector, and the interlayer samples and intralayer samples are divided, and the interlayer samples are the global samples removed from the mean.
[0009] Step 3: According to the spatial pattern of the ammonium nitrogen pollutants, the spatial modeling is carried out. The specific implementation steps are: first, the global spatial structure information extraction mechanism is designed; second, the mixed covariance matrix is designed; third, the Markov chain Monte Carlo sampling method is used to calculate the posterior probability distribution of the covariance function parameters of the intralayer spatial process and the interlayer spatial process.
[0010] Step 4: According to the posterior probability distribution of the covariance function parameters of the intralayer spatial process and the posterior probability distribution of the covariance function parameters of the interlayer spatial process, the model prediction based on Bayesian theory is carried out, and the uncertainty estimation of the spatial prediction of ammonium nitrogen pollution is obtained.
[0011] Preferably, the Moran's index is denoted as I, and its mathematical expression is:
[0012]
[0013] Where n is the number of samples, x i and x j represent the soil ammonium nitrogen concentration at the i and j positions, respectively, represents the mean of the variable in the study area, W ij is the spatial weight matrix between x i and x j , the value range of I is -1~1, I=-1 represents high negative correlation, I=1 represents high positive correlation, and I=0 represents no correlation.
[0014] Preferably, the geographic detector is denoted as q, and its mathematical expression is:
[0015]
[0016] wherein σ 2 represents the total variance of the observation samples in the whole study area, N is the number of observation samples in the whole study area, N h represents the number of observation samples of each stratum, L represents the number of divided spatial layers, represents the variance of the observation samples of different strata h, q ranges from 0 to 1, the closer q is to 1, the stronger the heterogeneity of the regional unit is, and when q = 0, it represents that the regional unit does not exist heterogeneity.
[0017] Preferably, the specific design steps of the global spatial structure information extraction mechanism in step 3 are: using the global sample subset obtained by the random sampling method with replacement to perform multiple parameter fitting to obtain the prior information of the global spatial relationship in the hierarchical framework, and the prior information is the prior probability distribution of the estimated variogram parameters of the intralayer spatial process and the interlayer spatial process; and more robust parameter estimation is ensured.
[0018] The prior probability distribution of the variogram parameters is set to f prior (Θ):
[0019] f prior (Θ) = {sill, range, nugget}
[0020] Wherein, Θ represents the parameter group {sill, range, nugget}, sill represents the sill value of the variogram, range represents the range of the variogram, and nugget represents the nugget value of the variogram.
[0021] Preferably, the specific implementation method of the mixed covariance matrix is: superimposing the corresponding covariances of the different spatial processes of the sampling points and the points to be estimated to realize the consideration of the influence of different spatial processes at the position to be estimated. The spatial autocorrelation and the spatial hierarchical heterogeneity can be effectively combined in a unified framework to obtain the intralayer spatial process and the interlayer spatial process. The intralayer spatial process refers to the spatial autocorrelation of the pollutants within the same geological layer or depth layer, and the degree of similarity between the sample points is reflected by calculating the covariance between the sample points in the same layer. The greater the covariance, the stronger the spatial autocorrelation of the pollutant distribution between the sample points; the interlayer spatial process refers to the spatial hierarchical heterogeneity of the difference of the pollutant distribution between different geological layers or depth layers, and the covariance between different layers is small or even negative, which means that there is a significant difference or reverse trend of the pollutants between different strata.
[0022] Preferably, the expression of the mixed covariance matrix calculation is as follows:
[0023]
[0024] wherein, denotes the covariance between the point s i and the sampling points s j , denotes the spatial distance between the point s i and the sampling points s j , in and Θ between represent the prior probability distribution of the covariance function parameters of the intra-layer spatial process and the inter-layer spatial process, respectively, in f (·) and f (·) represent the covariance functions of the intra-layer spatial process and the inter-layer spatial process, respectively, between L i and L j represent the layers where the point s i and the sampling points s j are located.
[0025] Preferably, the Markov Chain Monte Carlo sampling method (MCMC) is expressed as:
[0026] f posterior (Θ|Y(s))∝f(Y(s)∣Θ)&f prior (Θ)
[0027] wherein, f posterior (Θ|Y(s)) represents the conditional probability distribution of the parameter Θ given the observed sample, f(Y(s)∣Θ) represents the conditional distribution probability of the pollutant concentration given the parameter Θ, and f prior (Θ) is the prior probability distribution of the variogram parameters.
[0028] Compared with the prior art, the present application has the beneficial effects that:
[0029] 1. The present application introduces spatial hierarchical heterogeneity modeling, which can accurately identify the pollution heterogeneity between different depth layers by layering the spatial distribution of pollutants according to different depths and strata, break through the limitation of traditional spatial interpolation methods which only rely on single-level spatial structure modeling, and can deal with the distribution of pollutants with significant hierarchical heterogeneity, solve the problem that traditional interpolation methods cannot handle strong heterogeneity in the vertical direction, improve the ability to capture complex pollution distribution in rare earth mine areas, and be suitable for pollution prediction in complex environments and multi-level structures.
[0030] 2, The application designs a strategy for obtaining prior information based on a repeated sampling mechanism. By borrowing global spatial information from other areas, additional support is provided for local parameter estimation, thereby improving the robustness of parameter estimation. The repeated sampling method can effectively increase the sample size of parameter estimation, thereby ensuring that robust pollutant spatial distribution prediction results can still be obtained in the case of sparse samples. Not only does this improve the reliability of parameter estimation in sparse data scenarios, but it also enhances the adaptability and predictive ability of the model, effectively addressing the shortcomings of traditional methods that rely on large sample data.
[0031] 3, The application can effectively handle complex spatial structure problems and provide a robust tool for pollutant spatial distribution modeling by constructing a mixed covariance matrix to capture the spatial relationship within and between layers and address the impact of spatial heterogeneity on model prediction performance. BRIEF DESCRIPTION OF DRAWINGS
[0032] Figure 1 A hierarchical Bayesian three-dimensional curved surface mean model implementation flowchart is provided for the application.
[0033] Figure 2 An implementation diagram for the global spatial structure information extraction mechanism is provided.
[0034] Figure 3 A specific implementation diagram for the mixed covariance matrix is provided.
[0035] Figure 4 An intuitive diagram of the model and other models for dividing the spatial boundary of ammonium nitrogen pollution is provided.
[0036] Figure 5 A visualization effect diagram of the model prediction uncertainty is provided. DETAILED DESCRIPTION
[0037] To make the purpose, technical scheme and advantages of the application clearer and more apparent, the application will be further described below with reference to the embodiments. Obviously, the described embodiments are part of the embodiments of the application, not all. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the application.
[0038] The application provides a method for predicting the spatial distribution of ammonium nitrogen pollution in rare earth mining areas. The method mainly uses spatial hierarchical heterogeneity theory and Bayesian theory, combined with spatial interpolation technology, to propose an innovative hierarchical Bayesian three-dimensional curved surface mean model (HBM-MSN). The implementation steps of the model are as shown in Figure 1 The specific implementation steps are as follows:
[0039] Step 1: Based on the characteristics of rare earth mining process and ammonia nitrogen pollution, combined with mountain terrain and leaching path, the sampling area is divided and the drilling position is determined, and the sample collection and laboratory analysis are completed.
[0040] Step 2: Identify the spatial pattern of ammonium nitrogen pollutants. The spatial autocorrelation and spatial stratified heterogeneity of ammonium nitrogen pollutants are identified by the Moran index and the geographic detector, respectively, and the interlayer samples and intralayer samples are divided, and the interlayer samples are the global samples after removing the mean.
[0041] Step 3: Spatial modeling according to the spatial pattern of ammonium nitrogen pollutants. The specific implementation steps are: the first step is to design the global spatial structure information extraction mechanism; the second step is to design the mixed covariance matrix; the third step is to calculate the posterior probability distribution of the covariance function parameters of the intralayer spatial process and the interlayer spatial process by Markov Chain Monte Carlo sampling method (MCMC), which are p(Θ I ) and p(Θ B ) respectively. Figure 1
[0042] Step 4: Bayesian-based model prediction according to the posterior distribution of the intralayer covariance parameter probability distribution and the interlayer covariance parameter probability distribution, to obtain the uncertainty estimation of ammonium nitrogen pollution spatial prediction, the specific effect is shown in Figure 5 , it can be seen that the top soil of the rare earth mining area soil used for prediction is slightly polluted by ammonium nitrogen, and the risk level is low, and the high uncertainty area is mainly concentrated in the bottom layer in the southwest.
[0043] Preferably, the Moran index is denoted as I, and its mathematical expression is:
[0044]
[0045] where n is the number of samples, x i and x j represent the soil ammonium nitrogen concentration at the i-th and j-th positions, respectively, represents the mean of the variable in the study area, W ij is the spatial weight matrix between x i and x j , the value range of I is -1-1, I=-1 represents high negative correlation, I=1 represents high positive correlation, and I=0 represents no correlation.
[0046] Preferably, the geographic detector is denoted as q, and its mathematical expression is:
[0047]
[0048] where σ 2 Total variance of the whole study area, N is the number of observation samples of the whole study area, N h L represents the number of layers divided in each stratum, L represents the number of layers divided in each stratum,
[0049] Preferably, the global spatial structure information extraction mechanism is designed, and the specific design steps are as follows: a plurality of parameter fittings are performed by using a global sample subset obtained by a random sampling method with replacement to obtain prior information of global spatial relationship in a hierarchical framework, the global sample is the intra-layer sample and the inter-layer sample, the prior information is a prior probability distribution of estimated variogram parameters of the intra-layer spatial process and the inter-layer spatial process, and the prior probability distribution is respectively Figure 1 Θ B and Θ I ; more robust parameter estimation is ensured; an implementation diagram of the global spatial structure information extraction mechanism is shown in Figure 2 The subset sample obtained by random sampling fitting obtains a variogram p(Θ G After a plurality of repeated sampling fitting processes, a large number of parameter potential values are obtained;
[0050] The prior probability distribution of the variogram parameter is set as f prior (Θ):
[0051] f prior (Θ) = {sill, range, nugget}
[0052] Wherein, Θ represents a parameter group {sill, range, nugget}, sill represents a sill value of the variogram, range represents a range of the variogram, and nugget represents a nugget value of the variogram.
[0053] Preferably, the mixed covariance matrix is designed, and the specific implementation method is as follows: the covariances corresponding to different spatial processes of the to-be-estimated point and the sampling point are superimposed to realize the consideration of the influence of different spatial processes at the to-be-estimated position. The specific implementation diagram is shown in Figure 3As shown. The hybrid covariance matrix design effectively combines spatial autocorrelation and spatial stratification heterogeneity within a unified framework to obtain intra-layer and inter-layer spatial processes. The intra-layer spatial process refers to the spatial autocorrelation of pollutants within the same geological layer or depth layer. By calculating the covariance between sample points within the same layer, the degree of similarity between these points is reflected. The larger the covariance, the stronger the autocorrelation of pollutant distribution between sample points. The inter-layer spatial process refers to the spatial stratification heterogeneity of pollutant distribution differences between different geological layers and depth layers. By calculating the covariance between sample points at different levels, a smaller or even negative inter-layer covariance means that there are significant differences or reverse trends in pollutant distribution between different strata.
[0054] Preferably, the expression for calculating the mixed covariance matrix is as follows:
[0055]
[0056] in, Indicates the point s to be estimated i and sampling point s j Covariance between Indicates the point s to be estimated i and sampling point s j Spatial distance between them, Θ in and Θ between f represents the prior probability distribution of the covariance function parameters of the intra-layer and inter-layer spatial processes, respectively. in (·) and f between (·) represent the covariance functions of intra-layer and inter-layer spatial processes, respectively. i and L j They represent the points to be estimated, s and s, respectively. i and sampling point s j The floor it is located on.
[0057] Preferably, the Markov chain Monte Carlo sampling method (MCMC) is expressed as follows:
[0058] f posterior (Θ∣Y(s))∝f(Y(s)∣Θ)&f prior (Θ)
[0059] Among them, f posterior (Θ|Y(s)) represents the conditional probability distribution of parameter Θ given the observed samples, and f(Y(s)|Θ) represents the conditional probability distribution of ammonium nitrogen pollutant concentration given parameter Θ. prior (Θ) represents the prior probability distribution of the parameters of the mutation function.
[0060] According to two indexes of mean absolute error and root mean square error, the performance of the HBM-MSN of the application and three-dimensional ordinary Kriging model (3D-OK) and three-dimensional layered mean surface model (3D-MSN) are compared, and a direct view diagram is shown in Figure 4 As shown in the figure, the relationship between the ammonium nitrogen concentration and the gray scale under the action of different models can be seen, and the visual pollution threshold is that the ammonium nitrogen concentration is greater than 7 mg / kg. According to the X and Y axes in the figure, the specific horizontal and spatial position of the ammonium nitrogen pollutant can be seen. (c) in the figure is the HBM-MSN described in the application, and the figure shows that it well captures the three-dimensional spatial distribution boundary of the pollutant, especially the detail change of the high concentration area is more obvious, and accurately shows the accumulation characteristics of the pollutant, and clearly demarcates the boundary of different pollution areas. (a) and (b) in the figure are 3D-OK and 3D-MSN respectively, and the figure shows that the spatial boundary division effect of ammonium nitrogen of the two models also shows the spatial distribution of the pollutant, but they are rough in detail processing, especially in the description of the distribution boundary of the pollutant, there is a certain ambiguity, and it is difficult to accurately capture the local concentration change. The HBM-MSN shows the gradual change process of different depths and regions in the process of showing the gradual change of the pollutant concentration, and compared with the 3D-OK and the 3D-MSN, the HBM-MSN is not fine enough in the performance of local concentration change, especially the 3D-OK, which is prone to produce too smooth prediction results, and cannot accurately show the local heterogeneity. As can be directly seen from the figure, the HBM-MSN can identify the change rule of the pollutant at different levels, especially in the vertical direction, the heterogeneity of the ammonium nitrogen pollutant is more clearly shown. This is because the HBM-MSN combines the spatial relationship within and between layers through the mixed covariance matrix, so as to more accurately reflect the distribution difference of the pollutant in different strata.
[0061] The above is only the preferred embodiment of the application, and is not used to limit the application. For those skilled in the art, the application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and scope of the application shall be included in the protection scope of the application.
Claims
1. A method for spatial prediction of ammonium nitrogen pollution in rare earth mining areas, characterized in that: The method includes: Step 1: Based on the characteristics of rare earth mining technology and ammonia nitrogen pollution, combined with the mountain topography and leaching path, the sampling area was divided and the drilling location was determined to complete the sample collection and laboratory analysis. Step 2: Identify the spatial patterns of ammonium nitrogen pollutants. Use Moran's index and a geographic detector to identify the spatial autocorrelation and spatial stratification heterogeneity of ammonium nitrogen pollutants, and divide them into inter-stratum samples and intra-stratum samples. The inter-stratum samples are global samples with the mean removed. Step 3: Spatial modeling is performed based on the spatial pattern of the ammonium nitrogen pollutants. The steps are as follows: First, design a global spatial structure information extraction mechanism; second, design a mixed covariance matrix; third, use the Markov chain Monte Carlo sampling method to calculate the posterior probability distribution of the covariance function parameters of the intra-layer spatial process and the inter-layer spatial process. Step 4: Based on the posterior probability distribution of the covariance function parameters of the intra-layer spatial process and the posterior probability distribution of the covariance function parameters of the inter-layer spatial process, perform model prediction based on Bayesian theory to obtain an uncertainty estimate of the spatial prediction of ammonium nitrogen pollution. The steps of the global spatial structure information extraction mechanism design in step 3 are as follows: using a subset of global samples obtained by random sampling with replacement to perform multiple parameter fittings to obtain prior information of global spatial relationships in the hierarchical framework. The prior information is the prior probability distribution of the estimated variogram parameters of intra-layer spatial processes and inter-layer spatial processes. Let the prior probability distribution of the parameters of the mutation function be f. prior (Θ): f prior (Θ)={sill,range,nugget} Where Θ represents the parameter group {sill, range, nugget}, still represents the sill value of the mutation function, range represents the range of the mutation function, and nugget represents the nugget value of the mutation function. The method for designing the hybrid covariance matrix is as follows: superimpose the covariances of the point to be estimated and the sampling point in different spatial processes; The expression for calculating the mixed covariance matrix is as follows: in, Indicates the point s to be estimated i and sampling point s j Covariance between Indicates the point s to be estimated i and sampling point s j Spatial distance between them, Θ in and Θ between f represents the prior probability distribution of the covariance function parameters of the intra-layer and inter-layer spatial processes, respectively. in (·) and f between (·) represent the covariance functions of intra-layer and inter-layer spatial processes, respectively. i and L j They represent the points to be estimated, s and s, respectively. i and sampling point s j The floor it is located on.
2. The method for spatial prediction of ammonium nitrogen pollution in rare earth mining areas as described in claim 1, characterized in that: The Moran index is denoted as I, and its mathematical expression is: Where n is the number of samples, x i and x j These represent the soil ammonium nitrogen concentrations at positions i and j, respectively. w represents the mean of the variable within the study area. ij It is x i and x j The spatial weight matrix between them, where I ranges from -1 to 1, I = -1 indicates a high negative correlation, I = 1 indicates a high positive correlation, and I = 0 indicates no correlation. The geographic detector is denoted as q, and its mathematical expression is: Where, σ 2 N represents the total variance of the observed samples in the entire study area, and N is the number of observed samples in the entire study area. h This represents the number of observation samples for each stratum, and L represents the number of spatial layers. q represents the variance of the observed samples in different strata h. The value of q ranges from 0 to 1. The closer q is to 1, the stronger the heterogeneity of the regional unit. When q = 0, it means that there is no heterogeneity in the regional unit.
3. The method for spatial prediction of ammonium nitrogen pollution in rare earth mining areas as described in claim 1, characterized in that: The Markov chain Monte Carlo sampling method is expressed as: f posterior (Θ∣Y(s))∝f(Y(s)∣Θ)&f prior (Θ) Among them, f posterior (Θ|Y(s)) represents the conditional probability distribution of parameter Θ given the observed sample, and f(Y(s)|Θ) represents the conditional probability distribution of pollutant concentration given parameter Θ. prior (Θ) represents the prior probability distribution of the parameters of the mutation function.
Citation Information
Patent Citations
Method for establishing heavy metal content space model of coal mining area
CN107545103A
Screening method and screening tree of uranium mine area soil radioactive contamination spatial interpolation method
CN115795240A