Soil pollution accumulation prediction method based on multi-factor dynamic driving

By using a method based on geographic detectors and backpropagation neural networks, combined with carbon neutrality policies, and dynamically adjusting the weights of driving factors, the problem of factor changes not being considered in the prediction of soil pollutant accumulation is solved, and more accurate predictions of future pollutant accumulation are achieved.

CN121834288APending Publication Date: 2026-04-10北京市科学技术研究院资源环境研究所(北京市土地修复工程技术研究中心)
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-14
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing technologies fail to effectively consider the dynamic changes of various driving factors in predicting soil pollutant accumulation, resulting in an inability to accurately predict future pollutant accumulation amounts and risks.

Method used

A geospatial detector-based method was used to identify the main driving factors and their synergistic enhancement factor combinations. Combined with BP neural network model training, the weight coefficients and thresholds were determined through iterative learning. Combined with carbon neutrality policy scenario correction parameters, the factor weights were dynamically adjusted using sliding window analysis to establish a multi-factor dynamic-driven soil pollution accumulation prediction model.

Benefits of technology

It breaks through the limitations of focusing only on major influencing factors, dynamically reflects changes in driving factors, and improves the accuracy and timeliness of soil pollutant accumulation prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121834288A_ABST
    Figure CN121834288A_ABST
Patent Text Reader

Abstract

The invention provides a soil pollution accumulation prediction method based on multi-factor dynamic driving, and the method comprises the steps: analyzing the driving capability of each potential driving factor for the accumulation characteristics of soil pollutants based on a geographic detector method, and recognizing a main driving factor and a factor combination with the cooperative enhancement driving capability; inputting the identified main driving factors into a BP neural network model for training, determining a weight coefficient and a threshold value between an input layer and an output layer through iterative learning, and establishing a pollutant accumulation prediction model; in combination with historical driving factor data and carbon neutralization policy scene parameters, correcting parameters in the prediction model and generating driving factor data of a year to be predicted; and performing time sequence processing on the driving factor data by using a sliding window analysis method, dynamically adjusting the weight coefficient and contribution rate of each factor, and identifying future main driving factors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of soil pollution investigation and remediation technology, specifically relating to a method for predicting the accumulation of soil pollution based on multi-factor dynamic driving. Background Technology

[0002] Current research on soil pollutant accumulation typically relies on long-term sampling and monitoring methods. However, this approach requires substantial human and material resources, making it difficult to ensure continuous and stable monitoring. Research based on predictive models, such as mass balance models, accumulation rate models, and regression models, is a current focus. These methods primarily establish the relationship between various driving factors and pollutant accumulation, building predictive models based on identifying the main driving factors of pollutant spatial distribution and their driving capacity.

[0003] There are many potential influencing factors on soil pollutant accumulation, such as atmospheric deposition and transportation. Under my country's carbon neutrality policy, the pursuit of green development and green transformation may lead to industrial transformation or even restructuring. Energy conservation, emission reduction, and energy saving involve industries such as coal, petrochemicals, chemicals, steel, power, photovoltaics, and new energy vehicles, affecting pollutant emissions during industrial production and transportation. This causes changes in the driving force and contribution rate of various factors to pollutant accumulation in soil. However, current research mainly builds predictive models based on the current main driving factors and their driving force, without considering the changes in the driving force and contribution rate of each factor over time. Therefore, it cannot provide support for accurately predicting future soil pollutant accumulation and accumulation risks. Summary of the Invention

[0004] The present invention aims to at least partially solve one of the technical problems in the related art.

[0005] Therefore, the first objective of this invention is to propose a method for predicting soil pollution accumulation based on multi-factor dynamic driving.

[0006] The second objective of this invention is to propose a soil pollution accumulation prediction device based on multi-factor dynamic driving.

[0007] To achieve the above objectives, a first aspect of the present invention proposes a method for predicting soil pollution accumulation based on multi-factor dynamic driving, comprising: S1, based on the geographic detector method, analyze the driving ability of each potential driving factor on the accumulation characteristics of soil pollutants, and identify the main driving factors and factor combinations with synergistic driving ability. S2, the main driving factors identified are input into the BP neural network model for training. The weight coefficients and thresholds between the input layer and the output layer are determined through iterative learning to establish a pollutant accumulation prediction model. S3 combines historical driving factor data with carbon neutrality policy scenario parameters to correct the parameters in the prediction model and generate driving factor data for the year to be predicted. S4 utilizes the sliding window analysis method to perform time-series processing on the driving factor data, dynamically adjusts the weight coefficients and contribution rates of each factor, and identifies the main driving factors in the future.

[0008] In one embodiment of the present invention, S1 includes: S11, by calculating the variance explained by the driving factors and pollutant accumulation characteristics. Quantitative driving capability; S12 uses the information entropy method to calculate the mutual information between driving factors and pollutant accumulation characteristics, and identifies factor combinations with significant synergistic enhancement effects.

[0009] In one embodiment of the present invention, S2 includes: S21 uses a sigmoid function as the activation function for the hidden layer nodes and a linear activation function for the output layer. S22 optimizes the weight parameters using the Adam optimizer and defines the loss function.

[0010] In one embodiment of the present invention, S3 further includes: S31, calculate the average annual emission reduction rate of each driving factor based on the linear regression model; S32 corrects the contribution rate parameter of the driving factor using an exponential decay model.

[0011] In one embodiment of the present invention, S4 includes: S41, determine the sliding window length based on the autocorrelation coefficient of the time series; S42 optimizes the sliding window length using the information entropy minimization criterion.

[0012] To achieve the above objectives, a second aspect of the present invention provides a soil pollution accumulation prediction device based on multi-factor dynamic driving, comprising: The driving factor analysis module is used to analyze the driving ability of each potential driving factor on the accumulation characteristics of soil pollutants based on the geographic detector method, and to identify the main driving factors and combinations of factors with synergistic driving ability. The model training module is used to input the identified main driving factors into the BP neural network model for training. Through iterative learning, the weight coefficients and thresholds between the input layer and the output layer are determined to establish a pollutant accumulation prediction model. The parameter calibration and data generation module is used to combine historical driving factor data with carbon neutrality policy scenario parameters to calibrate the parameters in the prediction model and generate driving factor data for the year to be predicted. The time series processing and weight adjustment module is used to perform time series processing on the driving factor data using the sliding window analysis method, dynamically adjust the weight coefficients and contribution rates of each factor, and identify the main driving factors in the future.

[0013] This invention discloses a method and apparatus for predicting soil pollution accumulation based on multi-factor dynamic driving. It establishes an accumulation prediction model by integrating multiple driving factors of soil pollutant accumulation and their synergistic effects, thus overcoming the limitation of focusing only on the main influencing factors. It also focuses on the dynamic changes in the driving capacity and contribution rate of each factor in pollutant accumulation, thus overcoming the limitation of treating the driving capacity and contribution rate of factors as fixed values.

[0014] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0015] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 This is a flowchart of a soil pollution accumulation prediction method based on multi-factor dynamic driving according to an embodiment of the present invention; Figure 2 This is a schematic diagram illustrating the trend of the contribution rate of the spatial distribution of polycyclic aromatic hydrocarbons (PAHs), a major factor (GDP), over time according to an embodiment of the present invention. Figure 3 This is a schematic diagram showing the changing trend of four PAHs concentrations in soil over time according to an embodiment of the present invention. Figure 4 This is a structural diagram of a soil pollution accumulation prediction device based on multi-factor dynamic driving according to an embodiment of the present invention. Detailed Implementation

[0016] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0017] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0018] The following description, with reference to the accompanying drawings, describes a method and apparatus for predicting soil pollution accumulation based on multi-factor dynamic driving according to an embodiment of the present invention.

[0019] Example 1 Figure 1 This is a flowchart of a soil pollution accumulation prediction method based on multi-factor dynamic driving according to an embodiment of the present invention, such as... Figure 1 As shown, it includes: S1, based on the geographic detector method, analyzes the driving ability of each potential driving factor on the accumulation characteristics of soil pollutants, and identifies the main driving factors and combinations of factors with synergistic driving ability.

[0020] Specifically, this step utilizes the Geographical Detector method to quantitatively analyze the driving force of various potential driving factors on soil pollutant accumulation characteristics, thereby identifying the main driving factors and their synergistic enhancing effects. The Geographical Detector method is a statistical model based on spatial heterogeneity analysis. Its core principle is to calculate the explanatory power of factors on the spatial differentiation of dependent variables (such as the concentration of polycyclic aromatic hydrocarbons (PAHs) in soil). (Value), to assess its driving ability. The value ranges from [0, 1]. The larger the value, the stronger the explanatory power of the factor for the spatial distribution of pollutant accumulation.

[0021] In some implementations, data on various potential drivers of soil pollutant accumulation are first collected, including but not limited to industrial emission intensity, traffic density, land use type, climate conditions, and population density. These factors are then registered with spatial distribution data of soil pollutant concentrations to form a unified spatial dataset. Subsequently, a geographic detector model is used to independently detect each factor and calculate its... The value is used to quantify its driving force on pollutant accumulation.

[0022] Furthermore, the synergistic effects among multiple factors are analyzed using an interaction detector. If the interaction between two factors... The value is significantly higher than that of each individual The sum of the values ​​indicates a synergistic enhancing effect between the two. This method supports the identification of nonlinear and non-equivalent combinations of factors, thereby revealing the complex driving mechanisms of pollutant accumulation.

[0023] Specifically, geographic detector models typically employ 95% confidence intervals for significance testing to ensure that identified driving factors are statistically significant. Furthermore, spatial stratification is performed using either the Jenks method or the equidistant method to improve the model's spatial resolution.

[0024] Furthermore, S1 includes: S11 quantifies the driving force by calculating the variance explained by the driving factors and the characteristics of pollutant accumulation.

[0025] The technical implementation of this step is based on the principle of analysis of variance in statistics, specifically by comparing driving factors. Conditional variance of pollutant accumulation characteristics under certain conditions Compared to the total variance without considering this factor This is used to calculate its explanatory power for pollutant accumulation. The formula is defined as:

[0026] in, The overall variance represents the characteristics of pollutant accumulation (such as polycyclic aromatic hydrocarbon concentration), reflecting the randomness when not affected by any driving factors; Indicating in driving factors The conditional variance of pollutant accumulation characteristics under the influence of factors such as GDP, traffic density, and industrial emissions reflects the degree to which these factors constrain pollutant distribution. The value range is [0, 1]. The larger the value, the stronger the driving factor. The stronger the ability to explain the accumulation of pollutants, the more significant its driving force.

[0027] In its implementation, this step first spatializes the multi-source environmental data to construct a spatial correspondence between driving factors and pollutant accumulation characteristics. Then, a geographic detector algorithm is used to detect each driving factor and calculate its corresponding... The algorithm divides spatial data into several sub-regions using a sliding window or partitioned statistics method, and calculates the mean and variance of pollutant accumulation characteristics within each sub-region, thereby assessing the explanatory power of the driving factors for spatial heterogeneity.

[0028] S12 uses the information entropy method to calculate the mutual information between driving factors and pollutant accumulation characteristics, and identifies factor combinations with significant synergistic enhancement effects.

[0029] In some implementations, this step is based on the principle of mutual information (MI) in information theory, by calculating driving factors. Characteristics of pollutant accumulation Mutual information values ​​between Assess their correlation. Mutual information measures the statistical dependence between two random variables; the higher the value, the stronger the correlation. right The higher the information contribution, the better. In this invention, driving factors include, but are not limited to, industrial emissions, traffic density, land use type, and climate conditions, while pollutant accumulation characteristics are mainly indicated by the concentration distribution of polycyclic aromatic hydrocarbons (PAHs) in soil. This is achieved through joint probability distribution. With marginal distribution , The logarithm of the ratio, formula It can effectively capture nonlinear relationships, and is especially suitable for identifying the synergistic effects of multiple factors.

[0030] Specifically, the calculation of mutual information depends on the estimation of probability distributions. This invention employs a method based on histogram or kernel density estimation to... , and Discretization is typically performed by dividing pollutant concentrations into several intervals (e.g., 5-10), and calculating the frequency distribution within each interval based on historical monitoring data. The threshold for mutual information value is set as follows. This is to screen out factors that have a significant impact on pollutant accumulation. Furthermore, to identify synergistic enhancement effects, it is also necessary to calculate the joint mutual information of multiple factors. And compare its difference with the sum of single-factor mutual information, if This indicates and There is a synergistic enhancement effect.

[0031] S2, the identified main driving factors are input into the BP neural network model for training. The weight coefficients and thresholds between the input layer and the output layer are determined through iterative learning to establish a pollutant accumulation prediction model.

[0032] In some implementations, a backpropagation (BP) neural network model consists of an input layer, hidden layers, and an output layer. The input layer contains... Each hidden layer contains several nodes, corresponding to the identified main driving factors, such as industrial emission intensity, traffic density, and land use type. The output layer consists of one node, representing the predicted concentration of pollutants (such as polycyclic aromatic hydrocarbons, PAHs) in the soil. The number of nodes in the hidden layer can be adjusted according to the complexity of the input features, and is usually determined using empirical formulas or cross-validation methods. For example, the number of nodes in the hidden layer can be set to... or 1.5 times that. Each node is connected via learnable weights, where Indicates the input layer node. This represents a hidden layer node. Activation functions typically employ the Sigmoid or ReLU functions to enhance the model's non-linear fitting capability.

[0033] Specifically, the mean squared error (MSE) is used as the loss function during model training. ,in For the sample size, This represents the actual pollutant concentration. These are the model's predicted values. Learning rate. The value is set to 0.01, and the maximum number of iterations is set to 25,000 to ensure that the model converges to the optimal solution. During training, the weights and thresholds are continuously adjusted through the backpropagation algorithm, so that the prediction error gradually decreases.

[0034] Furthermore, S2 includes: S21 uses a sigmoid function as the activation function for the hidden layer nodes, and a linear activation function for the output layer.

[0035] In some implementations, sigmoid functions Used for hidden layer nodes, it is essentially a continuous, differentiable nonlinear function that maps any real input to the interval [0, 1], thereby enhancing the model's ability to fit complex nonlinear relationships. This function is widely used in neural networks to introduce nonlinear characteristics, enabling the network to learn and express the nonlinear response relationship between driving factors and pollutant concentration. The number of hidden layer nodes can be adjusted according to the number of input factors. Adjustments are made to the model complexity, and the optimal number of layers and nodes are usually determined using empirical formulas or cross-validation methods.

[0036] The output layer uses a linear activation function, meaning the output value is directly equal to a linear combination of the weighted inputs, in the form of: ,in For connection weights, For output of the hidden layer, This is the bias term. Linear activation functions are suitable for regression tasks, as they can directly output continuous values, i.e., predicted concentrations of soil pollutants, avoiding the limitations imposed by nonlinear functions on the output range.

[0037] Specifically, the model is set to a maximum of 25,000 iterations to ensure sufficient convergence during training. Weight coefficients and thresholds are optimized using backpropagation, and the loss function is typically expressed as mean squared error (MSE). ,in For the sample size, For the true value, This is the predicted value. The training process continuously adjusts the network parameters using gradient descent to minimize the prediction error.

[0038] S22 optimizes the weight parameters using the Adam optimizer and defines the loss function.

[0039] In some implementations, the loss function is defined as ,in The model represents the first The predicted output for each sample, This represents the actual observed value of the sample. This represents the total number of samples. This loss function measures the deviation between the model output and the true value, and is the core basis for gradient calculation during the optimization process.

[0040] The Adam optimizer is an adaptive learning rate optimization method based on first- and second-order moment estimation. Its core idea is to dynamically adjust the learning rate of each parameter by calculating the first moment (mean) and second moment (uncentered variance) of the gradient. Specifically, Adam updates the momentum and RMSProp terms of the parameters in each iteration and normalizes the gradient based on these statistics, thereby achieving differentiated learning rate adjustments across different dimensions. Its update formula is:

[0041] in, For the first The parameter values ​​for the next iteration. The initial learning rate (usually set to) ), and These are the first-order and second-order moment estimates after bias correction, respectively. For smoothing terms (usually taken) (), used to prevent division by zero errors.

[0042] S3 combines historical driving factor data with carbon neutrality policy scenario parameters to correct the parameters in the forecasting model and generate driving factor data for the year to be forecasted.

[0043] Specifically, in some implementations, firstly, driving factor data from the past 5 to 10 years (i.e., a time window length of 5 to 10 years) are selected as model input. These driving factors include, but are not limited to, industrial emission intensity, transportation density, energy structure ratio, and land use type. Data collection must comply with the classification and quantification standards for pollutant source items in the "Technical Specification for Soil Environmental Monitoring" (HJ 656-2013). Subsequently, this historical data is input into a BP neural network model, the model structure of which includes an input layer containing... 1 node (corresponding to) The main driving factors are (1, 2, 3, 4, 5, 6, 7, 8, 9, 1 Nodes in the Output Layer, which is 1 Node in the Output Layer (corresponding to the soil pollutant concentration), and the number of hidden layer nodes is adaptively adjusted according to the input and output dimensions and data complexity.

[0044] During parameter calibration, parameters from carbon neutrality policy scenarios, such as the national plan's target for reducing carbon dioxide emissions per unit of GDP and Beijing's "Blue Sky Protection Campaign" requirements for VOCs and PM2.5 reduction, are incorporated to calculate the pollutant reduction rate. The reduction rate can be expressed as the rate of change in emissions per unit time. ,in This represents the change in emissions. The time interval is defined as follows. By coupling the emission reduction rate with historical driving factor data, the weight coefficients and thresholds in the BP neural network are dynamically adjusted to reflect the impact of policy implementation on the future trends of driving factors.

[0045] Furthermore, S3 includes: S31, calculate the average annual emission reduction rate of each driving factor based on a linear regression model.

[0046] In some implementations, this linear regression model uses a time variable As an independent variable, the average annual emission reduction rate of the driving factors As the dependent variable, the historical data are fitted using the least squares method to solve for the regression coefficients. and .in, Indicates the initial emission reduction rate. This represents the rate of change of emission reduction speed over time, i.e., the average annual rate of change. In practice, historical data for driving factors typically cover a time span of 5 to 10 years, with data frequency ranging from annual to quarterly, depending on the availability and characteristics of the factors. By constructing a time series dataset, the observed values ​​of driving factors are paired with their corresponding years, and the data is input into a regression algorithm for model training.

[0047] Specifically, the fitting accuracy of a regression model can be measured by the coefficient of determination R. 2 Statistical indicators such as mean squared error are used for evaluation. In this invention, it is recommended that R... 2 A value of no less than 0.85 is required to ensure that the model has a high explanatory power in fitting the emission reduction trends of the driving factors. Furthermore, the regression coefficients... The sign and magnitude of the emission reduction factor directly reflect the direction and intensity of the emission reduction. For example, if This indicates that the emissions of this factor are showing a downward trend year by year, which is in line with the carbon neutrality policy orientation.

[0048] S32 corrects the contribution rate parameter of the driving factor using an exponential decay model.

[0049] In some implementations, this step first acquires actual observational data of each driving factor over the past 5 to 10 years as input samples for the BP neural network. Through the training process of the neural network, the model can learn the nonlinear mapping relationship between each factor and pollutant concentration and output the factor weights under the current scenario. Subsequently, based on scenario settings such as carbon neutrality policies and national and local (e.g., Beijing) environmental protection goals, the emission reduction rate of pollutants is calculated, and the changing trends of each driving factor at future time points are derived accordingly. On this basis, an exponential decay model is used to correct the contribution rate of the factors, where... This indicates the contribution rate of the factor at the current moment. For predicting the time step (in years). This is the attenuation coefficient under the policy scenario, and its value range is usually within... It is used to characterize the degree to which the influence of factors on the accumulation of pollutants weakens over time.

[0050] Specifically, The value of is derived from the factor driving ability calculated by the geographic detector method, while The determination of this depends on the setting of policy scenarios and the fitting analysis of historical emission reduction data. For example, under a carbon neutrality scenario, if a certain factor is industrial emission intensity, then... Nonlinear regression estimation can be performed using historical emission decline curves to reflect the long-term decline trend driven by policies.

[0051] S4 can be generated by inputting unified latent space features into a diffusion model for low-step denoising, or directly input into the task decoder to output perceptual prediction results.

[0052] In some implementations, the length of the sliding window can be set according to the temporal variation characteristics of the driving factors, typically 5 to 10 years, consistent with the time span of historical data. For example, if the driving factor data covers the past 10 years, the window length can be set to 10 years, with each sliding step being 1 year, thereby generating multiple time window samples. The driving factor data within each window serves as the input sample, fed into a pre-trained BP neural network model, and the model outputs the predicted concentration values ​​of soil pollutants for the corresponding time period.

[0053] Furthermore, by calculating the changes in the weight coefficients of the driving factors within each window, their dynamic driving capabilities over different time periods can be identified. The updating of the weight coefficients relies on the model's backpropagation algorithm, whose error function is defined as... ,in Indicates the first Input data for each window, This represents the prediction error. The model updates the weight parameters iteratively by minimizing this error function, with a maximum learning iteration count set to 25,000 to ensure model convergence and stability.

[0054] Furthermore, S4 includes: S41, determine the sliding window length based on the autocorrelation coefficient of the time series.

[0055] Specifically, in the steps of this invention, the sliding window length The determination depends on the autocorrelation coefficient of the time series. ,in Indicates lag The autocorrelation coefficient, Time series The mean of the pollutant concentration. The core of this step lies in identifying the time-dependent structure of pollutants by analyzing the autocorrelation characteristics of the pollutant concentration time series, thereby providing a reasonable window length setting for subsequent sliding window analysis.

[0056] In some implementations, the pollutant concentration time series is first preprocessed, including missing value imputation, outlier removal, and standardization, to ensure data continuity and comparability. Subsequently, different lag orders are calculated. Autocorrelation coefficient And draw an autocorrelation graph (ACF). In some implementations, it can be set... The range of values ​​is This is to cover the short- and medium-term time dependence of pollutant accumulation processes. Through observation... Follow The decay trend is used to determine the lag order at which the autocorrelation coefficient first significantly falls below the confidence interval (typically at the 95% confidence level). And use this as the length of the sliding window. Reference basis, namely .

[0057] Furthermore, the autocorrelation coefficient The calculation must meet the requirement of statistical significance, and is usually judged using Bartlett's confidence interval, the formula of which is: ,in This represents the length of the time series. In practical applications, if... exist When the value first falls below the confidence interval, the sliding window length... It can be set to 6 to ensure that the model can capture key time-dependent features in the pollutant accumulation process.

[0058] S42 optimizes the sliding window length using the information entropy minimization criterion.

[0059] Adopting the information entropy minimization criterion Optimize sliding window length ,in Indicates the first Each driving factor in the window length The information entropy is calculated. The core objective of this step is to improve the information expression efficiency of driving factors in future time series predictions by dynamically adjusting the sliding window length, thereby more accurately identifying their dynamic driving capacity and contribution rate to soil pollutant accumulation.

[0060] In some implementations, sliding window analysis is a commonly used time series modeling method. Its basic principle is to divide continuous time series data into multiple segments of length 1. The subsequences are used as input samples for the model. In this invention, the sliding window length... The choice of window length directly affects the model's ability to capture the changing trends of driving factors. If the window is too short, it may fail to reflect the long-term trend of the factors; if the window is too long, it may introduce too much noise and reduce the model's sensitivity. Therefore, using information entropy as an evaluation metric to optimize the window length is key to improving the model's prediction accuracy.

[0061] Information entropy The calculation is based on the driving factor within the window length. The smaller the value of the probability distribution under the given window length, the lower the uncertainty of the factor and the more concentrated the information expression. This is achieved by minimizing the sum of the information entropy of all driving factors. By finding the optimal window length, the model can more effectively extract the dynamic features of driving factors in time series modeling.

[0062] Specifically, in this invention, the time series data length of the driving factor is the past Year, window length The search scope is usually set to This ensures a balance between time resolution and trend representation in the model. The optimization process can employ grid search or gradient descent, combined with cross-validation to ensure the model's generalization ability.

[0063] The present invention provides a method for predicting soil pollution accumulation based on multi-factor dynamic driving, which can dynamically reflect the changing trends of driving factors under the background of carbon neutrality policy, and improve the accuracy and timeliness of soil pollutant accumulation prediction.

[0064] Example 2 The following describes in detail an embodiment of the present invention, a method for predicting soil pollution accumulation based on multi-factor dynamic driving, with reference to the accompanying drawings.

[0065] The specific analysis steps are as follows: Based on the geographic detector method, we quantitatively analyze the driving ability of each potential driving factor on the accumulation characteristics of soil pollutants, and identify the main driving factors and multiple factors with synergistic driving abilities. The identified n main driving factors are input into the input layer of a backpropagation (BP) neural network, with the output being the soil pollutant concentration. Therefore, the input layer has n nodes, and the output layer has one node. The activation function at each node represents the functional relationship between the output and input. Each pair of connected nodes has a network parameter representing the weight of the signal passing through that node. These parameters, combined with the driving power of the quantitatively calculated main factors, are used to determine the weight values ​​corresponding to the minimum error through repeated training. The maximum number of learning iterations is set to 25,000. Finally, the weight coefficients and thresholds are determined, training stops, and the model is complete.

[0066] Based on data of each driving factor over the past 5-10 years, a regression model is established using a backpropagation neural network. Combining carbon neutrality and environmental protection goals set by national and Beijing environmental policies and measures under different scenarios, the pollutant emission reduction rate is calculated, and the corresponding parameters in the regression model are corrected. On this basis, data of each human driving factor in the year to be predicted are obtained. These driving factor data are input into a trained BP neural network to obtain the weight coefficients between these factors; based on the sliding window analysis method, i.e., a series of driving factor data samples obtained in units of a certain length sliding window, the future development data of each driving factor is obtained, and the dynamic driving capacity and contribution rate of each factor are analyzed; the main driving factors for the future accumulation of soil PAHs are identified. Figure 2 and 3 As shown, Figure 2 The contribution rate of polycyclic aromatic hydrocarbons (PAHs) to the spatial distribution of GDP as a major factor changes over time. Figure 3 The trend of the concentration of four PAHs in the soil over time is shown.

[0067] By inputting future data of each factor into the established cumulative prediction model, the cumulative trend and cumulative amount of soil pollutants under different future scenarios can be obtained.

[0068] Example 3 To achieve the above embodiments, such as Figure 4 As shown, this embodiment also provides a soil pollution accumulation prediction device 10 based on multi-factor dynamic driving. The device 10 includes a driving factor analysis module 100, a model training module 200, a parameter correction and data generation module 300, and a time series processing and weight adjustment module 400.

[0069] The driving factor analysis module 100 is used to analyze the driving ability of each potential driving factor on the accumulation characteristics of soil pollutants based on the geographic detector method, and to identify the main driving factors and factor combinations with synergistic driving ability. The model training module 200 is used to input the identified main driving factors into the BP neural network model for training, and to determine the weight coefficients and thresholds between the input layer and the output layer through iterative learning to establish a pollutant accumulation prediction model. The parameter correction and data generation module 300 is used to combine historical driving factor data with carbon neutrality policy scenario parameters to correct the parameters in the prediction model and generate driving factor data for the year to be predicted. The time series processing and weight adjustment module 400 is used to perform time series processing on the driving factor data using the sliding window analysis method, dynamically adjust the weight coefficients and contribution rates of each factor, and identify the main driving factors in the future.

[0070] Furthermore, the aforementioned driving factor analysis module 100 is also used for: The driving force is quantified by calculating the variance explained by the driving factors and the characteristics of pollutant accumulation. The mutual information between driving factors and pollutant accumulation characteristics is calculated using the information entropy method to identify factor combinations with significant synergistic enhancement effects.

[0071] Furthermore, the aforementioned model training module 200 is also used for: A sigmoid function is used as the activation function for the hidden layer nodes, and a linear activation function is used for the output layer. The weight parameters are optimized using the Adam optimizer, and a loss function is defined at the same time.

[0072] Furthermore, the parameter correction and data generation module 300 is also used for: The average annual emission reduction rate of each driving factor was calculated based on a linear regression model. The contribution rate parameter of the driving factor is corrected by an exponential decay model.

[0073] Furthermore, the aforementioned timing processing and weight adjustment module 400 is also used for: The sliding window length is determined based on the autocorrelation coefficient of the time series. The sliding window length is optimized using the information entropy minimization criterion.

[0074] This invention provides a soil pollution accumulation prediction device based on multi-factor dynamic driving, which can dynamically reflect the changing trends of driving factors under the background of carbon neutrality policy, thereby improving the accuracy and timeliness of soil pollutant accumulation prediction.

[0075] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0076] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.

Claims

1. A method for predicting accumulation of soil pollution based on multi-factor dynamic driving, characterized in that, The method comprises the following steps: S1, based on the geographical detector method, analyze the driving ability of each potential driving factor on the accumulation characteristics of soil pollutants, identify the main driving factors and the combination of factors with synergistic enhancement driving ability; S2, input the identified main driving factors into the BP neural network model for training, determine the weight coefficient and threshold value between the input layer and the output layer through iterative learning, and establish a pollutant accumulation prediction model; S3, combine historical driving factor data and carbon neutralization policy scenario parameters to correct the parameters in the prediction model and generate driving factor data for the year to be predicted; S4, use the sliding window analysis method to process the driving factor data in time sequence, dynamically adjust the weight coefficient and contribution rate of each factor, and identify the main driving factors in the future.

2. The method of claim 1, wherein, The S1 comprises: S11, explaining variance between the driving factors and the pollutant accumulation characteristics by calculation quantifying the driving capacity; S12, calculate the mutual information between the driving factor and the accumulation characteristics of the pollutants by using the information entropy method, and identify the factor combination with significant synergistic enhancement effect.

3. The method of claim 1, wherein, The S2 comprises: S21, use S-shaped function as the activation function of the hidden layer node, and use linear activation function for the output layer; S22, optimize the weight parameters through the Adam optimizer, and define the loss function.

4. The method of claim 1, wherein, The S3 further comprises: S31, calculate the annual emission reduction speed of each driving factor based on the linear regression model; S32, correct the contribution rate parameters of the driving factors through the exponential decay model.

5. The method of claim 1, wherein, The S4 comprises: S41, determine the sliding window length according to the autocorrelation coefficient of the time sequence; S42, optimize the sliding window length by using the information entropy minimization criterion.

6. A soil pollution accumulation prediction device based on multi-factor dynamic driving, characterized in that, The method comprises: a driving factor analysis module, configured to analyze the driving ability of each potential driving factor on the accumulation characteristics of soil pollutants based on the geographical detector method, and identify the main driving factors and the combination of factors with synergistic enhancement driving ability; a model training module, configured to input the identified main driving factors into the BP neural network model for training, determine the weight coefficient and threshold value between the input layer and the output layer through iterative learning, and establish a pollutant accumulation prediction model; a parameter correction and data generation module, configured to combine historical driving factor data and carbon neutralization policy scenario parameters to correct the parameters in the prediction model and generate driving factor data for the year to be predicted; a time sequence processing and weight adjustment module, configured to use the sliding window analysis method to process the driving factor data in time sequence, dynamically adjust the weight coefficient and contribution rate of each factor, and identify the main driving factors in the future.

7. The apparatus of claim 6, wherein, The driving factor analysis module is further configured to: quantify the driving ability by calculating the variance explanation rate between the driving factor and the accumulation characteristics of the pollutants; calculate the mutual information between the driving factor and the accumulation characteristics of the pollutants by using the information entropy method, and identify the factor combination with significant synergistic enhancement effect.

8. The apparatus of claim 6, wherein, The model training module is further configured to: use S-shaped function as the activation function of the hidden layer node, and use linear activation function for the output layer; optimize the weight parameters through the Adam optimizer, and define the loss function.

9. The apparatus of claim 6, wherein, The parameter correction and data generation module is further configured to: calculate the annual emission reduction speed of each driving factor based on the linear regression model; correct the contribution rate parameters of the driving factors through the exponential decay model.

10. The apparatus of claim 6, wherein, The time sequence processing and weight adjustment module is further configured to: The length of the sliding window is determined according to the autocorrelation coefficient of the time series; The length of the sliding window is optimized by using the information entropy minimization criterion.