Methods and models for inversion of key parameters in numerical model of biodegradation of groundwater pollutants based on neural network algorithm
Through the neural network algorithm-based method, the challenges of parameter dynamics, complexity and high-dimensionality in the biodegradation of groundwater pollutants are solved, and more efficient and accurate simulation is achieved, reducing computational costs.
Patent Information
- Application Number
- CN202411478765.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-22
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-10-22
AI Technical Summary
The prior art faces challenges such as parameter dynamics, complexity and high dimensionality when dealing with groundwater pollutants, resulting in increased computational costs and inefficient simulations.
Using a method based on neural network algorithm, a multi-layer neural network architecture under the framework of physical information neural network is built by preprocessing and dimensionality reduction, and the relevant reaction equations of the natural attenuation process of pollutants are embedded, biochemical reaction mechanisms and data-driven models are coupled, and the neural network model is trained using Adam's optimization algorithm to invert key parameters.
It improves the simulation efficiency and accuracy of the biodegradation process of groundwater pollutants, reduces the computational cost, and enhances the physical consistency of the model and the credibility of parameter inversion.
Smart Images

Figure CN119558171B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of groundwater pollution assessment, and in particular to a method and model for inverting key parameters of a groundwater pollutant biodegradation model based on a neural network algorithm. Background Art
[0002] As one of the important freshwater resources on Earth, groundwater is essential for supporting ecosystems, agricultural irrigation, and human daily life. However, with the acceleration of industrialization and the expansion of agricultural activities, groundwater resources are facing serious pollution threats. Therefore, the development of effective numerical simulation optimization methods for groundwater pollutants is of great significance for assessing groundwater pollution risks and formulating subsequent remediation plans. Traditional groundwater solute transport models are usually based on physical and chemical principles, and use numerical simulation methods to predict the migration and transformation behavior of pollutants in groundwater. These models are challenging in parameter determination and model validation, especially when data is scarce or parameters are highly uncertain. The biodegradation process of groundwater pollutants poses significant challenges to the parameter accuracy of numerical models due to its dynamics and complex interactions with groundwater chemical components. In addition, the high-dimensional parameter space and parameter heterogeneity involved in the biodegradation model lead to a significant increase in computational costs. These challenges highlight the urgent need to develop new methods to improve the efficiency and accuracy of biodegradation process simulations.
[0003] In response to the challenges of the dynamics, complexity and high dimensionality of biodegradation model parameters, this paper proposes a method for inverting key parameters in a numerical model of groundwater pollutant biodegradation based on a neural network algorithm. This method uses deep learning technology, especially neural networks, to simulate and predict the biodegradation process of pollutants in groundwater. Through this method, key parameters affecting pollutant biodegradation, such as microbial degradation rate, half-saturation constant, etc., can be inverted and determined more accurately. Summary of the invention
[0004] The invention overcomes the shortcomings of the prior art and provides a method for inverting key parameters in a numerical model of biodegradation of groundwater pollutants based on a neural network algorithm.
[0005] To achieve the above-mentioned purpose, the technical solution adopted by the present invention is: a method for inverting key parameters in a numerical model of biodegradation of groundwater pollutants based on a neural network algorithm, comprising:
[0006] 1) Effectively process high-dimensional and multi-scale data by preprocessing and refining the dimensionality reduction of input data from different sources;
[0007] 2) Build a multi-layer neural network architecture under the framework of physical information neural network, embed relevant reaction equations of the natural attenuation process of pollutants, and couple biochemical reaction mechanisms and data-driven models;
[0008] 3) Train the neural network model used to invert the key parameters of the pollutant biodegradation numerical model, calculate the errors in the training process, and use the Adam optimization algorithm to improve the accuracy of the inversion results;
[0009] 4) Integrate observational data with model prediction results, determine the final parameter settings of the numerical model, and conduct physical quantification of groundwater damage.
[0010] Furthermore, in the present invention, the step 1) pre-processes the input data from different sources to perform refined dimensionality reduction on the high-dimensional data, and effectively processes multi-dimensional and multi-scale data, specifically including:
[0011] Collect and organize input data from multiple sources at multiple scales and from multiple sources;
[0012] Perform data cleaning and standardization, perform feature engineering, and normalize the data;
[0013] Apply kernel principal component analysis to reduce the dimensionality of high-dimensional data;
[0014] Furthermore, in the present invention, the step 2) builds a multi-layer neural network architecture under the framework of physical information neural network, embeds the relevant reaction equations of the natural attenuation process of pollutants, and couples the biochemical reaction mechanism and data-driven model, specifically including:
[0015] In the framework of physical information neural network, based on the convolutional neural network architecture, a neural network parameter inversion model combining the groundwater pollutant degradation mechanism with the data-driven model is constructed;
[0016] Define the input layer and output layer, build a neural network structure with multiple hidden layers, and capture the global dynamic changes of model parameters over a long time scale;
[0017] A comprehensive loss function is designed to quantify the difference between the output results of the neural network model and the groundwater numerical model, and the biodegradation-related equations of pollutants are embedded in the loss function in the form of residuals to achieve the unification of reaction mechanism and data-driven.
[0018] Furthermore, in the present invention, the step 3) trains a neural network model for inverting key parameters of the pollutant biodegradation numerical model, calculates errors during the training process, and uses the Adam optimization algorithm to improve the accuracy of the inversion results, specifically including:
[0019] Calculate the output of the network model from the input layer and generate prediction results;
[0020] Use the loss function to calculate the loss value of the model, and adjust the network output and the weights in the loss function by minimizing the loss function;
[0021] Starting from the output layer, calculate the gradient of the loss function with respect to the output, apply the chain rule to pass it back to each layer, and select the Adam optimization algorithm to update the network parameters;
[0022] Monitor the loss function value and performance indicators during training to avoid overfitting problems.
[0023] Furthermore, in the present invention, the step 4) integrates the observation data and the model prediction results, determines the final parameter setting of the numerical model, and carries out the physical quantification of groundwater damage, which specifically includes:
[0024] Construct an ensemble Kalman filter to integrate observation data with model predictions and iteratively update ensemble members;
[0025] Fine-tune the model on the validation set, using the output of the ensemble Kalman filter to guide the adjustment of model parameters;
[0026] Evaluate the generalization ability of the model on the test set, and evaluate the prediction accuracy and reliability of the model;
[0027] Carry out comparative verification of the neural network parameter inversion model to ensure the feasibility and accuracy of the parameter inversion method.
[0028] The present invention has the following beneficial effects:
[0029] The network input and output data were determined by preprocessing the input data from different sources to ensure the quality and consistency of the data. The kernel principal component analysis method was applied to reduce the dimensionality of high-dimensional parameters, so that the model could capture the key properties of groundwater contaminants more accurately.
[0030] The present invention designs a convolutional long short-term memory network architecture within the framework of physical information neural network to adapt to the spatiotemporal characteristics of groundwater solute migration, enhances the physical consistency of the model, and ensures that the network can effectively process multi-dimensional input data.
[0031] The present invention designs a loss function that introduces a regularization term and uses the Adam algorithm to optimize the loss function, embedding the microbial degradation kinetic equation and the interaction equation into it in the form of residuals, allowing the model to consider the kinetics and interactions of groundwater pollutant biodegradation during the training process, thereby improving the credibility and accuracy of the inversion of groundwater pollutant biodegradation parameters.
[0032] The present invention uses the ensemble Kalman filtering method to integrate observation data and model predictions, divides the data set for cross-validation and tests the generalization ability of the model, and performs the performance of key parameters; uses the trained model for prediction, combines the physical constraints of the physical information neural network, accurately inverts the key parameters, and improves the practical application value of the model.
[0033] By using field observation data, traditional numerical simulation models and neural network models for inverting biodegradation parameters, it was verified that the neural network parameter inversion method and model can improve the accuracy of groundwater pollutant distribution prediction and ensure the accuracy and reliability of the physical quantification of groundwater ecological damage. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, drawings of other embodiments can be obtained based on these drawings without paying creative work.
[0035] Figure 1 Schematic diagram of the overall process;
[0036] Figure 2 Data preprocessing flow chart;
[0037] Figure 3 Neural network structure diagram;
[0038] Figure 4 Schematic diagram of the loss function content;
[0039] Figure 5 Schematic diagram of the effect of parameter inversion method: a is actual observation; b is parameter inversion; c is parameter-free inversion. DETAILED DESCRIPTION
[0040] In order to more clearly understand the above-mentioned objects, features and advantages of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific implementation methods. It should be noted that the embodiments of the present application and the features in the embodiments can be combined with each other without conflict.
[0041] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited to the specific embodiments disclosed below.
[0042] like Figure 1 As shown, the present invention provides a method for inverting key parameters in a numerical model of groundwater pollutant biodegradation based on a neural network algorithm, comprising the following steps:
[0043] S101: Effectively process high-dimensional and multi-scale data by preprocessing and fine-tuning the dimensionality reduction of input data from different sources; S102: Build a multi-layer neural network architecture under the framework of physical information neural network, embed relevant reaction equations of the natural attenuation process of pollutants, and couple biochemical reaction mechanisms and data-driven models;
[0044] S103: Train the neural network model used to invert the key parameters of the pollutant biodegradation numerical model, calculate the errors in the training process, and use the Adam optimization algorithm to improve the accuracy of the inversion results;
[0045] S104: Integrate observational data with model prediction results, determine the final parameter settings of the numerical model, and conduct physical quantification of groundwater damage.
[0046] The present invention solves the problem that it is difficult to accurately invert the key parameters of the numerical simulation model of the biodegradation of pollutants in groundwater under complex geological conditions, and realizes the accurate inversion of the key parameters of the numerical simulation model of the biodegradation process of groundwater pollutants with a long period and a wide range, effectively improving the simulation prediction accuracy of groundwater pollutant migration, and providing a basis for the subsequent groundwater restoration of contaminated sites.
[0047] S101, in one embodiment of the present invention, high-dimensional multi-scale data is effectively processed by preprocessing and refining the dimensionality reduction of input data from different sources, such as Figure 2 As shown, specifically including:
[0048] Collect and organize input data from multiple sources at multiple scales and from multiple sources;
[0049] It should be noted that the multi-source and multi-scale input data include groundwater monitoring related data, including physical parameters such as groundwater permeability, porosity, flow rate, chemical parameters such as dissolved oxygen, pH value, heavy metal ion concentration, biological parameters such as the number and activity of microbial populations, microbial growth rate, and spatiotemporal data such as the change data of groundwater pollutant concentration at different time and space locations. The biodegradation process of groundwater pollutants is affected by many factors, including physical, chemical and biological characteristics. Data sources may include groundwater samples from monitoring stations, laboratory test results, and remote sensing or geographic information system (GIS) data. These data may have different spatial and temporal resolutions and collection methods, and need to be formatted and processed in a unified format;
[0050] Perform data cleaning and filling, carry out standardization processing, perform feature engineering, and perform data normalization;
[0051] It should be noted that data cleaning refers to identifying and processing outliers in the data. Before starting processing, remove obviously erroneous data points and invalid data rows to ensure that the measured values of each observation point are logically consistent. For example, there will be no extreme jumps in temperature, and changes in pollutant concentrations should be in line with common sense. Outlier identification uses the Z-score method to identify outliers in the data. Z-score calculates the degree of abnormality of each data point based on the standard deviation, and its calculation formula is: Where x is the sample value, μ is the feature mean, and σ is the standard deviation. If the Z value of a data point exceeds 3 (that is, the data point is outside three standard deviations of the mean), the data point is considered an outlier. This method can effectively identify potential abnormal data points and ensure the overall quality of the data. For missing values that need to be filled, the K nearest neighbor (K-NN) algorithm is used for filling. This is a distance-based filling method. First, the Euclidean distance between each sample is calculated, the K samples with the closest distance are selected, and their average value is used to fill the missing value. The calculation formula for the Euclidean distance is: Among them, x i and i are the values of the i-th feature of the two samples respectively. For the processing of missing values, select the K value (usually 5) and process all samples with missing values to ensure the integrity of the data.
[0052] It should be noted that data standardization is to eliminate the dimensional differences between different features so that the values of all features are distributed in a similar range, so that the neural network model can converge more easily. The standardization formula is still the Z-score formula. After standardization, the mean of all features will be 0 and the standard deviation will be 1. This processing method helps the optimization algorithm avoid the gradient vanishing problem caused by the excessive value of some features during training.
[0053] It should be noted that the way to perform feature engineering on data is to evaluate the correlation between different features and the target variable through statistical methods (Pearson correlation coefficient). The calculation formula of Pearson correlation coefficient is:
[0054] Among them, X i and Y i are the characteristic value and target value of the sample, respectively. and are their means respectively. By calculating the correlation coefficient of each feature, the feature most relevant to the biodegradation process is selected. In order to improve the expressiveness of the model, nonlinear relationships can also be captured by constructing polynomial features (such as quadratic terms and cubic terms). Assuming that the original features are x1, x2, the constructed polynomial features include: x1, x2, x1x2. This nonlinear combination feature can help the model better fit the complex groundwater pollutant degradation process.
[0055] It should be noted that data normalization helps speed up the training process of the neural network and ensures that different features contribute roughly the same to the model. Data normalization can be performed using the minimum-maximum normalization method to scale the feature values to the range of [0,1], and the formula is: Among them, min(x) and max(x) are the minimum and maximum values of the feature respectively.
[0056] The kernel principal component analysis method is used to reduce the dimension of high-dimensional parameters. The principal components are selected based on the size of the eigenvalues, and the data are projected onto the selected principal components to complete the dimensionality reduction.
[0057] It should be noted that in order to reduce the dimensionality of high-dimensional data and extract the most important features, kernel principal component analysis (KPCA) is applied. Unlike traditional principal component analysis (PCA), KPCA can handle nonlinear data and map the data to a high-dimensional space through a Gaussian kernel function. The formula of the Gaussian kernel function is:
[0058]
[0059] Among them, σ is the bandwidth parameter, which controls the similarity between data points. KPCA reflects the relationship between samples in high-dimensional space by calculating the kernel matrix, and then performs eigenvalue decomposition to select the eigenvector corresponding to the maximum eigenvalue for the data representation after dimensionality reduction. After completing data preprocessing and dimensionality reduction, the data set is divided into training set, validation set and test set in a ratio of 70%, 15%, and 15%, maintaining the independence of data during model training and evaluation.
[0060] S102, in one embodiment of the present invention, a multi-layer neural network architecture is constructed under the framework of a physical information neural network, and the relevant reaction equations of the natural attenuation process of pollutants are embedded. The coupling biochemical reaction mechanism and the data-driven model specifically include:
[0061] In the framework of physical information neural network, based on the convolutional long-term and short-term neural network architecture, a neural network parameter inversion model combining the groundwater pollutant degradation mechanism with the data-driven model is constructed;
[0062] It should be noted that the physical information neural network framework (PINNs) reduces the demand for data volume of neural networks by introducing physical information, and introduces automatic differentiation technology to simplify the processing of nonlinear terms, which reduces the dimensionality difficulty to a certain extent. The principle diagram of PINNs is to embed prior knowledge in physics into neural networks, constrain the output results of neural networks through the physical laws summarized by predecessors, and combine sparse data and physical information (such as differential equations and their physical parameters) to predict physical processes. The physical equations in PINNs can also be biochemical reaction kinetic equations, etc. This patent proposes a new method and model for predicting the distribution of groundwater pollution plumes under the framework of physical information neural networks, which can realize the effective integration of the laws of biochemical degradation reaction mechanisms of pollutants and monitoring data. Select convolutional long short-term memory neural network
[0063] (ConvLSTM) architecture combines the advantages of convolutional neural network (CNN) and long short-term memory network (LSTM). CNN is responsible for extracting the spatial features of data, and LSTM is used to capture dynamic information in time series. This architecture is particularly suitable for dealing with the spatial and temporal characteristics of groundwater pollutants that change over time. During the training process, the biochemical reaction kinetic equation is embedded in the model, the Monod equation in the form of Logistics and other biochemical reaction interaction equations are used as mathematical equations for microbial degradation kinetics, and the loss function is introduced in the form of residuals to achieve the unity of mechanism and data-driven.
[0064] Define the input layer and output layer, build a neural network structure with multiple hidden layers, and capture the global dynamic changes of model parameters over a long time scale;
[0065] It should be noted that the input layer is designed to receive groundwater monitoring data, which includes multi-dimensional, multi-scale spatiotemporal data such as pollutant concentration, dissolved oxygen, pH value, temperature, and other known model parameters to ensure that the model can effectively process data from different sources and different time resolutions. Before this, a data tensor should be constructed to integrate the spatiotemporal data into a tensor format of: Among them, T is the time step, H and W are the spatial dimensions, and C is the feature dimension. The data shape of the input layer corresponds to the number of time steps and features, ensuring that the network can process time series features.
[0066] It should be noted that when constructing the hidden layer, a multi-level neural network structure is set up. Convolutional layers are constructed to extract local spatial features of pollutants, and LSTM layers are constructed to extract the dynamic characteristics of time series data. The conversion layer based on the self-attention mechanism is connected to describe the global dynamic changes of model parameters on a long time scale. The convolutional layer captures the local spatial characteristics of pollutants, while the LSTM layer can better handle time series characteristics, especially the dynamic changes of pollutants over time. The convolutional layer uses local perception fields and shared weights to detect local features in the pollutant concentration distribution, and then convolves the input data by applying filters (convolution kernels) to generate feature maps. The formula for applying a convolution kernel to the input tensor is:
[0067] Among them, X i,j is the input matrix, W is the convolution kernel, b is the bias, Z i,j is the convolution output. The LSTM layer is constructed to capture the dynamic characteristics of time series data and maintain long-term dependencies. Define the LSTM layer, set the number of LSTM units, and use the convolution output as the input of the LSTM layer. The gating mechanism of the LSTM layer controls the flow of information through three gates. The formula is:
[0068] f t =σ(W f ·[h t-1 , x t ]+b f )
[0069] i t =σ(W i ·[h t-1 , x t ]+b i )
[0070] o t =σ(W o ·[h t-1 , x t ]+b o )
[0071] C t =f t ·C t-1 +i t tanh(W C ·[h t-1 , x t ]+b C )
[0072] h t =o t tanh(C t )
[0073] Among them, f t It is the forget gate, t is the input gate, o t is the output gate, C t is the cell state, h t is a hidden state. Set the next layer as a conversion layer to capture long-term dependencies. The conversion layer identifies dependencies between all positions in the modeling sequence under the self-attention mechanism, which is suitable for dealing with the long-term migration process of groundwater pollutants. Although the LSTM layer can handle time dependencies, it may not perform well for longer time series, and the global self-attention of the conversion layer can effectively make up for this. Add a conversion layer based on the local spatiotemporal feature map extracted by ConvLSTM. Through the global self-attention mechanism, the conversion layer can capture the diffusion pattern and global dynamics of pollutants on a long time scale, especially the global characteristics of microbial biodegradation. The feature map of the data processed by ConvLSTM is represented as the input matrix of the conversion layer. The pixel values in each feature map are used as the input of the conversion layer. Design the conversion layer to generate Q, K, and V matrices, and generate query (Q), key (K), and value (V) matrices from the output of LSTM:
[0074] Q=W Q ·h
[0075] K=W K ·h
[0076] V=W V ·h
[0077] Among them, W Q , W K , W V is the weight matrix, and h is the hidden state of the LSTM. The attention weight is calculated using the dot product:
[0078] Among them, dk is the dimension of the key, which is used to scale the dot product. The conversion layer uses the self-attention mechanism to capture the long-term dependencies and global features in the process of pollutant migration. The ConvLSTM layer captures the dynamic changes of pollutants in the short-term time and local space, including the impact of factors such as porosity and flow rate on pollutant diffusion. Based on the local features extracted by ConvLSTM, the conversion layer further captures the pollutant migration pattern and global dynamic characteristics in the long term through the global self-attention mechanism. In this way, the short-term and long-term pollutant migration laws can be comprehensively considered in the model, effectively improving the prediction accuracy of the biodegradation process, thereby achieving accurate inversion of key parameters.
[0079] It should be noted that the design goal of the output layer is to invert the key parameters of biodegradation of pollutants in groundwater. The output data mainly includes two parts: the key parameters of the numerical model of biochemical degradation of pollutants in groundwater and the spatiotemporal distribution of the pollution plume during the natural attenuation of pollutants in groundwater. The number of output nodes is designed according to the number of predicted parameters, and the output layer is used to predict the pollutant concentration at future times, which provides strong support for the prediction of the migration path and concentration distribution of pollutants. Key parameters such as the maximum degradation rate of microorganisms (μ max ) and the half-saturation constant (K s ) can be used for regression prediction through a linear activation function: Among them, W is the weight matrix, h is the feature after convolution and LSTM processing, b is the bias, is the predicted degradation parameter. The difference between the output of the neural network model and the groundwater numerical model is quantified by the loss function. By minimizing the loss function, the key parameters of biochemical degradation are accurately inverted, thereby more accurately predicting the natural attenuation process of pollutants in groundwater.
[0080] Design a comprehensive loss function to quantify the difference between the output of the neural network model and the groundwater numerical model, embed the kinetic equations related to the biodegradation of pollutants into the loss function in the form of residuals, and achieve the unification of reaction mechanism and data-driven;
[0081] It should be noted that if Figure 4 As shown in Figure 1, a comprehensive loss function including data error term, physical error term and regularization term is designed. The main part of the loss function is the data error term, namely the mean square error (MSE), which is used to measure the difference between the model prediction value and the true value. The formula is: in, is the model prediction value, y i is the true value and N is the number of samples. By comparing the actual degradation rate with the rate predicted by the Monod equation, the physical error term can be defined: Among them, r i is the theoretical degradation rate calculated by the Monod equation, The model predicts the rate. To prevent the model from overfitting, add the L2 regularization term: Among them, λ is the regularization coefficient, θ i is the model parameter. The final total loss function combines the above three parts as follows: L = αL data +βL physics +γL reg Among them, α, β, and γ are important hyperparameters for balancing various losses; L data is the data error term, which represents the difference between the model prediction and the observed data; L physicsis the physical error term, indicating whether the model satisfies the physical equation; L reg is a regularization term to prevent overfitting.
[0082] It should be noted that when constructing the physical error term, in order to ensure that the model is consistent with the law of biodegradation of groundwater pollutants, the Monod equation in the form of Logistics and the biochemical reaction interaction equation are added to the physical constraint term. The Monod equation in the form of Logistics is a nonlinear kinetic equation used to describe the degradation rate of pollutants by microorganisms, which combines the growth characteristics of microorganisms in the biodegradation process. First, the Monod equation describes the degradation rate of pollutants by microorganisms: Among them, r is the degradation rate of pollutants by microorganisms, μ max is the maximum specific growth rate (the maximum growth rate of microorganisms under optimal conditions), S is the concentration of the pollutant, and K s is the half-saturation constant, which indicates the pollutant concentration when the degradation rate reaches half of the maximum rate. In order to incorporate the Logistics growth form into the Monod equation, the degradation rate can be made to present an S-shaped curve by introducing factors such as environmental capacity or resource limitations. The common Logistics form is as follows: In this equation, S max is the maximum allowable concentration of pollutants in the system. This form takes the upper limit of pollutant concentration into consideration, so that at high pollutant concentrations, the degradation rate no longer increases indefinitely, but gradually tends to a stable value.
[0083] It should be noted that in the natural attenuation process of pollutants, in addition to the microbial degradation dynamics, there are also a variety of biochemical reactions related to pollutants, microorganisms and environmental factors that degrade pollutants. These reactions can be described by corresponding equations and can also be embedded in the loss function of the neural network to ensure that the model can not only be trained based on data, but also follow the real physical and chemical laws. The following takes the relationship between dissolved oxygen (DO) and nitrate (NO3) and pollutant biodegradation and its application in the loss function as an example. Dissolved oxygen (DO) is one of the important electron acceptors for microbial degradation of pollutants. When the oxygen concentration is insufficient, the degradation rate of microorganisms will decrease. Therefore, the change in dissolved oxygen can be described by the Michaelis-Menten equation: Among them, r DO is the effect of oxygen on the degradation rate, k DO is the effect coefficient of oxygen on microbial degradation rate, DO is the dissolved oxygen concentration, K DO is the half-saturation constant of dissolved oxygen. In the loss function, the physical error term can be defined by the difference between the dissolved oxygen impact rate predicted by the model and the actual measured degradation rate:
[0084] Under anoxic conditions, nitrate may become an electron acceptor, affecting the degradation rate of pollutants. The nitrate reduction reaction can be described by a similar Monod equation: Among them, r NO3 is the effect of nitrate on the degradation rate, k NO3 is the coefficient of nitrate reduction on degradation rate, NO3 is the nitrate concentration, K NO3 is the half-saturation constant of nitrate. The nitrate influence term in the loss function can be written as:
[0085] It should be noted that in the degradation of groundwater pollutants, multiple reactions may occur simultaneously, such as the competitive utilization of oxygen, nitrate and other electron acceptors. At this time, the kinetic equations of different biochemical reactions can be combined to generate a comprehensive degradation rate equation. For the reactions of other electron acceptors, their reduction rate r can be similarly defined. other Taking into account the competitive utilization of multiple reactions, the kinetic equations of these reactions can be combined to generate a comprehensive degradation rate equation. If there are multiple reactions, including microbial degradation and other biochemical reactions (such as oxygen, nitrate, etc.), their rates can be combined in a weighted manner. Let w i is the weight of the ith reaction, and the comprehensive degradation rate r total It can be expressed as: Among them, w i is the weight of the ith reaction, r i is the rate of the ith reaction (e.g., microbial degradation rate, oxygen impact rate, nitrate reduction rate, etc.), and m is the total number of reactions considered. i The competition between reactions can be based on the relative importance of the reaction rates based on the concentration of the reaction substrates. For example, the weights can be calculated by normalizing the substrate concentrations: The rate of each reaction can be calculated by their corresponding equations, and they jointly affect the total degradation rate in a cumulative form. In the loss function of the neural network, the physical error term becomes the sum of the residuals of all reaction rates. Set the predicted and actual values of multiple reaction rates to r pred,i and r true,i , where i represents the i-th reaction. The physical error term L physical It is the sum of the residuals between the predicted values and the actual values of all reaction rates and can be expressed as: Among them, r pred,i is the predicted rate of the ith reaction, r true,iis the actual rate of the ith reaction, and m is the total number of reactions considered (e.g., microbial degradation, oxygen consumption, nitrate reduction, etc.). The final loss function will include multiple error terms, reflecting the combined effects of data error, physical error, and regularization. i is the weight of the ith reaction. i is the rate of the ith reaction: Among them, DL = Data Loss (data error term), PL = Physical Loss (physical error term), PL Monod =Monod equation error, PL DO = Dissolved oxygen (DO) related error, PL NO3 = Nitrate (NO3-) related error, other biochemical reaction errors: Reg = Regularization Loss (regularization term).
[0086] It should be noted that the design of the loss function usually needs to consider all processes related to the transport and degradation of groundwater pollutants to ensure that the model can accurately reflect the real behavior of pollutants. In addition to using biodegradation equations such as the Monod equation, the following physical equations related to pollutant transport can also be embedded in the loss function. Pollutants are diluted in groundwater due to flow. This effect can be simulated by adding a model for the dilution process, such as the convection-diffusion equation:
[0087] Where v is the velocity vector, D is the diffusion coefficient, and C is the contaminant concentration. Contaminants may volatilize from groundwater into the gas phase. Modeling of volatilization processes can be achieved by including relevant mass transfer equations, such as Henry's law: Where q is the adsorption amount, Q max is the maximum adsorption capacity, K a is the adsorption constant and C is the pollutant concentration. The adsorption of pollutants on soil particles affects their migration in groundwater. This process can be described by an adsorption model, such as the Langmuir adsorption equation:
[0088] q=K f C n
[0089] Where q is the adsorption amount, Q max is the maximum adsorption capacity, K a is the adsorption constant and C is the pollutant concentration.
[0090] S103, in one embodiment of the present invention, a neural network model for inverting key parameters of a pollutant biodegradation numerical model is trained, errors in the training process are calculated, and the accuracy of the inversion results is improved using the Adam optimization algorithm, specifically including:
[0091] Calculate the output of the network model from the input layer and generate prediction results;
[0092] It should be noted that the forward propagation process starts from the input layer, passes the input data layer by layer, passes through the convolution layer, LSTM layer and conversion layer, and finally outputs the prediction result. Starting from the input layer, the output of the network is calculated through the weights and activation functions of each layer. For the activation value calculation of the lth layer:
[0093] a (l) =f(W (l) ·a (l-1) +b (l) ) Among them, a (l) is the activation value of layer l, W (l) is the weight matrix of the lth layer, b (l) is the bias of the lth layer, and f is the activation function, such as ReLU or Sigmoid. This step requires processing the input data (such as spatiotemporal characteristics such as pollutant concentration and groundwater flow rate) layer by layer to generate the predicted output of the model.
[0094] It should be noted that the input layer accepts multi-dimensional inputs from the groundwater pollution system, including pollutant concentration, flow rate, temperature, porosity, dissolved oxygen concentration, etc., as well as initial parameters of physical constraints (such as the maximum microbial degradation rate μ max and the half-saturation constant K s ). The input data is first processed by the convolution layer, and the role of the convolution kernel is to extract the local features of the pollutants in space. The convolution operation operates on the input matrix through a sliding window to generate a new feature map. The LSTM layer processes the feature map from the convolution layer and captures the dynamic characteristics of the time series data. The core steps of the convolutional LSTM are the calculation of the forget gate, input gate, hidden gate, cell state, and hidden gate. The LSTM layer models the dynamic changes of pollutants over time. The LSTM layer processes the long-term and short-term dependencies in the time series data through the gating mechanism (forget gate, input gate, output gate), where the forget gate determines which information is discarded, the input gate determines which new information is stored in the cell state, and the output gate determines which information is output.
[0095] It should be noted that to integrate the conversion layer into the forward propagation, the features must first be mapped to the input embedding space of the conversion layer. For time series data, the input is usually the feature map extracted by ConvLSTM or the hidden state H of LSTM. LSTMWhere E is the feature matrix after embedding. Since the transformation layer structure does not have built-in time information, position encoding must be introduced to help the model understand the position relationship in the time series. Position encoding can be expressed as: E = Embedding (H LSTM ) where t is the time step, d is the embedding dimension, and Pt is the position encoding vector. The formula for adding the position encoding to the embedding feature changes to: t =[sin(t / 10000 2i / d ), cos(t / 10000 2i / d )] for i = 0, 1, ..., d / 2 to calculate the self-attention matrix. The self-attention mechanism in the conversion layer is used to capture the dependencies between different time steps: The conversion layer usually uses a multi-head attention mechanism to calculate multiple self-attention heads in parallel and then concatenate the results: MultiHeadAttention (E pos )=Concat(head1, head2,..., head h )·W0, where the calculation of each head is: After the attention mechanism, the data is further processed through the feedforward neural network layer: FeedForward(X) = max(0, X·W1+b1)·W2+b2 where W1 and W2 are the weights of the feedforward network and b1 and b2 are the bias terms. After each attention and feedforward network layer, layer normalization and residual connections are used to optimize the training: LayerNorm(x+Sublayer(x)) where Sublayer(X) represents the layer that has passed through the attention or feedforward network. The output Y of the transformation layer is transformed and combined with the output of the other network layers. In the final stage of the forward propagation, the output of the transformation is used as the input to the fully connected layer for the final prediction: Y final =W fc ·Y Transformer +b fc For specific prediction tasks (such as biodegradation parameter inversion), an appropriate activation function (e.g., linear activation function) is applied at the output layer to generate the final prediction value: In the forward propagation of the convolutional LSTM network with the conversion layer added, the data first passes through the convolutional layer to extract spatial features, and then the ConvLSTM processes the time series dynamic features. After that, the conversion layer captures long-term dependencies and global features through the self-attention mechanism, and finally generates the prediction results. In this way, the model can simultaneously utilize local spatial information and global temporal dependencies to more accurately capture the spatiotemporal changes and biodegradation dynamics of groundwater pollutants. The final output layer is responsible for predicting the target parameters of the model, such as the maximum microbial degradation rate μ max and the half-saturation constant K s, which is the key to parameter inversion. The output is and They are compared with the actual values and thus used to calculate the loss.
[0096] Use the loss function to calculate the loss value of the model. By minimizing the loss function, adjust the weights in the network output and the loss function to balance the importance of the data error term and the physical error term.
[0097] It should be noted that the loss function is used to calculate the difference between the model output and the actual label to obtain the loss value. The choice of loss function depends on the specific task. The mean square error (MSE) is used to measure the difference between the predicted value and the true value, and the cross entropy loss is used to evaluate the accuracy of the predicted category. By minimizing the loss function, the network output is adjusted to conform to the physical laws of biodegradation kinetics and chemical component interactions. The weights in the loss function are adjusted to balance the influence of the data error term and the physical error term, thereby optimizing the prediction performance of the model. The predicted value of the model is obtained from the output layer of the neural network. Assume that the predicted value of the output layer is Among them, is the prediction result of the network for each sample, with a shape of (B, N), where B is the batch size and N is the number of predicted values (usually the number of single predicted values). Get the actual label value from the dataset. The actual label Y is the target value of the network, which is The shape of is the same, that is, (B, N). Calculate the square of the difference between the predicted value and the actual value of each sample: Among them, SquaredError i is the square error of the ith sample. The square errors of all samples are averaged to get the mean square error of the batch: Where B is the batch size and N is the number of predictions. During training, the MSE is calculated once for each batch. For the entire training set, the MSE of all batches can be averaged to obtain the final loss value.
[0098] Starting from the output layer, the gradient of the loss function with respect to the output is calculated, and then passed back to each layer. The chain rule is applied to complete the gradient accumulation, and the Adam optimization algorithm is used to update the network parameters.
[0099] It should be noted that back propagation updates weights layer by layer by calculating the gradient of the loss function to the network parameters. First, the gradient of the loss function to the activation value of the output layer is calculated. For the mean square error loss function:
[0100] Among them, δ (L) is the error of the output layer, is the predicted value, and y is the true value. Then calculate the gradient of the loss function to the output layer weight: Among them, a(L-1) is the activation value of the previous layer of the output layer. For the hidden layer, the error is propagated forward layer by layer. The gradient of each layer is calculated as follows: (l) =(W (l+1) ) T ·δ (l+1) ·f′(z (l) ) where δ (l) is the error W of the lth layer (l+1) The weight matrix from layer l to layer l+1; f′(z (l) ) is the derivative of the activation function of the lth layer, z (l) is the weighted input of layer l. Then calculate the gradient of the weight of layer l: Using the calculated gradient, the network weights are updated using the Adam optimization algorithm. The Adam algorithm uses first-order momentum (mean) and second-order momentum (variance) to adjust the learning rate: and Among them, m t is the first-order momentum of the gradient, v t is the second-order momentum of the gradient, β1 and β2 are the decay rates of the momentum. Correct the calculation deviation of the first-order momentum and the second-order momentum, the formula is as follows: and Update the weights according to the corrected momentum, the formula is as follows: Here, η is the learning rate and ∈ is a constant to prevent division by zero errors.
[0101] Monitor the loss function value and performance indicators during training to avoid overfitting problems.
[0102] It should be noted that the change of loss value is monitored in real time during training to ensure that the loss function gradually decreases. The possible overfitting of the model can be judged by the increase of the loss value of the validation set. In addition to the loss function, it is also necessary to monitor the performance indicators of the model, such as gradient, recall rate, etc. Monitor the gradient value and regularly check and monitor the change of gradient. Focus on the mean and standard deviation of the gradient, draw a histogram or box plot of the gradient, check whether the gradient is evenly distributed, whether there are outliers, and track the change trend of the gradient. Adjust the training strategy according to the gradient monitoring results. If gradient explosion (too large) or gradient disappearance (too small) is observed, these problems can be alleviated by adjusting the learning rate (for example, using a learning rate scheduler or a learning rate decay strategy). In the case of gradient explosion, gradient clipping can be applied to limit the maximum value of the gradient to prevent the parameter update from being too large. If the gradient problem occurs frequently, it may be necessary to check and adjust the weight initialization strategy of the model. Record the gradient monitoring data and analyze it to identify potential problems or improvement points. Generate a visual chart of the training process, such as a gradient change curve, to facilitate in-depth analysis of the model training process. According to the analysis results, necessary adjustments are made to the model training process and hyperparameter settings (such as learning rate, optimizer, etc.) to ensure the stability of the training process and the effectiveness of the model.
[0103] S104, in one embodiment of the present invention, integrating the observed data and the model prediction results, determining the final parameter settings of the numerical model, and conducting the physical quantification of groundwater damage, specifically includes:
[0104] Construct an ensemble Kalman filter, integrate observation data with model predictions, and iteratively update ensemble members.
[0105] It should be noted that the ensemble Kalman filter is an advanced data assimilation method that combines observed data with model prediction results. During the inversion and verification process, the ensemble Kalman filter (EnKF) is used to continuously update the model state, combine the observed data with the numerical simulation results, and optimize the estimation of key parameters. First, a state prediction is performed, and the concentration distribution and change trend of pollutants in the groundwater system are predicted using a numerical simulation model. Then the state is updated, and the prediction results of the model are updated through the ensemble Kalman filter in combination with the field observation data to optimize the estimation of key parameters. To construct an ensemble Kalman filter, the set must be initialized first. Generate a set of initial state samples (set members) by perturbing the initial conditions. Suppose the model state is x, then the initialization set is: {x1 0 ,x2 0 ,…,x N 0}, each member x i 0is the state sampled from the initial distribution. Then each member of the set is predicted through the model to obtain the predicted state: Where f(·) is the state transition function. Calculate the mean of the prediction set: Calculate the covariance matrix P of the forecast set pred : Then calculate the mean and covariance matrix of the observations respectively: and Then calculate the filter gain matrix K: K = P pred H T (HP pred H T +R) -1 Where H is the observation operation matrix and R is the observation noise covariance matrix. Update the set members: Among them, y is the observed data.
[0106] Fine-tune the model on the validation set, using the output of the ensemble Kalman filter to guide the adjustment of model parameters;
[0107] It should be noted that the current model is used to predict the validation set data, and the error between the prediction result and the actual validation set observation data is calculated: Among them, y valid is the actual observed data of the validation set. The model parameters are adjusted based on the statistical characteristics of the error (such as mean square error). Optimization algorithms (such as gradient descent) can be used to minimize the loss function to update the model parameters. The state updated by EnKF is used as the initial state of the model to better adapt to the validation set data.
[0108] Evaluate the generalization ability of the model on the test set, and evaluate the prediction accuracy and reliability of the model;
[0109] It should be noted that the trained model is used to predict the test set data, the model's prediction results are obtained, and the error between the prediction results and the actual test set observation data is calculated. Common evaluation indicators include mean square error (MSE), root mean square error (RMSE) and mean absolute error (MAE). Analyze the values of the evaluation indicators to determine the performance of the model on unseen data.
[0110] Carry out comparative verification of neural network parameter inversion models to ensure the feasibility and accuracy of parameter inversion methods;
[0111] It should be noted that the performance of the model on unknown datasets is ensured by cross-validating different datasets. Cross-validation involves dividing the dataset into a training set and a validation set, training the model on the training set, and evaluating the generalization ability of the model on the validation set. Divide the dataset into multiple folds. For example, 5-fold cross-validation divides the data into 5 subsets. For each fold, use one of the subsets as the validation set and the remaining subsets as the training set. Train the model and evaluate the performance on the validation set. Record the evaluation indicators (such as MSE, RMSE, MAE) for each validation. Calculate the mean and standard deviation of the evaluation indicators for all folds to evaluate the average performance and stability of the model. Based on the results of cross-validation, select the parameter settings that perform best in all folds. These parameter settings should result in good performance of the model on both the training set and the validation set. Perform a final model evaluation and evaluate the final selected model parameters on the full test set to ensure the generalization ability of the model in practical applications.
[0112] It should be noted that the groundwater numerical model after calibration parameters is run to simulate the natural attenuation process of pollutants in groundwater, and the simulated pollutant concentration distribution is compared with the actual observation data to verify the feasibility and accuracy of the parameter inversion method. Combined with the actual site restoration case, the pollutant concentrations of each monitoring well after 150 days from the date of monitoring are sorted and analyzed to draw the concentration distribution map of groundwater pollutants at the site, that is, the pollution plume obtained from the actual observation data (such as Figure 5 a) is used as a reference object for verification. According to the traditional simulation method of solute transport of groundwater pollutants, the biodegradation coefficient obtained by experimental or empirical methods is input into the numerical model, and numerical simulation programs such as MODFLOW and RT3D are run. The distribution of pollutants in groundwater is predicted by solving the mathematical model of the groundwater system, and the known parameters such as model parameters and boundary conditions are input into the model to simulate the spatiotemporal distribution of the pollution plume. The spatiotemporal distribution of the pollution plume obtained by forward simulation (such as Figure 5 b) As the basic data set, it is input into the parameter inversion model based on the neural network. The back propagation algorithm is used to establish a reverse mapping network model from output to input. The values of the biodegradation parameters are continuously adjusted and updated, and the updated values of the variables to be calculated are substituted into the numerical simulation model to obtain the corresponding pollution plume distribution (such as Figure 5c) until the error with the actual pollutant concentration distribution satisfies the minimum convergence rule. By analyzing the results of field observation, simulation and simulation-optimization of actual contaminated sites, compared with the numerical simulation results of the parameter-free inversion model, the accuracy of groundwater contamination plume distribution is improved by about 20% when the biodegradation parameters are calibrated with the neural network parameter inversion model, which is closer to the actual observed field data. Using field observation data, traditional numerical simulation models and neural network models for inverting biodegradation parameters, it is verified that the neural network parameter inversion method and model can improve the accuracy of groundwater pollutant distribution prediction and ensure the accuracy and reliability of the physical quantification link of groundwater ecological damage.
[0113] It should be noted that predicting the distribution of groundwater pollutant concentrations at future sites is a key step in conducting physical quantification of groundwater damage. The key parameters obtained by inversion (such as microbial degradation rate, etc.) are substituted into traditional groundwater numerical simulation models, such as MODFLOW, MT3D or TOUGH2, for model calibration. By adjusting these parameters, the pollutant migration path and degradation rate simulated by the model are consistent with the field observation data. The parameters obtained by inversion are compared with the actual observation data to verify the accuracy of the model prediction. The test data can come from long-term observation records of groundwater monitoring wells or field experimental data. The prediction ability of the model is ensured by comparing the concentration distribution of the simulation results with the actual observation data. The key parameters obtained by neural network inversion (such as degradation rate, reaction kinetic parameters, etc.) are integrated into the traditional groundwater numerical simulation model. For example, the transmission rate and reaction rate parameters are adjusted in MODFLOW, and the pollutant migration and diffusion coefficients are adjusted in RT3D. Through the coupling of the neural network model and the numerical simulation model, a complete groundwater pollutant simulation system is formed. The neural network provides a preliminary estimate of the key parameters, and the numerical simulation is used to predict the spatial and temporal distribution of pollutants. According to the results of the numerical simulation of groundwater pollutants, relevant data can be obtained to quantify the physical damage to the ecological environment. Some formulas are as follows:
[0114] Where H is the amount of damage during the period; t is any year from the occurrence of ecological environmental damage to the restoration to the baseline (t0-t n t0 represents the starting year, which is the year when the ecological environment damage occurs; t n is the termination year, the year when the ecological environmental damage is restored to the baseline; T—base year, which is generally the year in which the ecological environmental damage appraisal and assessment is carried out as the base year; R t—The number of ecological and environmental service functions in the damaged area in year t. For resources, this parameter may be the number of individuals, biomass, life span, resource quantity, energy, productivity or other measures that have important impacts on organisms or ecosystems; for services, this parameter may be the area of affected habitat (hectares), or the length of the river or the area of other habitats, etc. t —The ratio of the ecological environment service function of the damaged area in the tth year to the baseline loss. This ratio changes over time and takes a value of 0-1. r—Discount factor, with a recommended value of 2%-5%. The groundwater pollutant simulation system realizes the effective operation of the physical quantification link of groundwater ecological damage through a perfect and accurate simulation optimization mechanism, and provides a basis and method for restoration feasibility assessment and restoration strategy formulation for groundwater contaminated sites that use monitoring of natural attenuation as a restoration plan.
[0115] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art who is familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A method for inverting key parameters in a numerical model of groundwater pollutant biodegradation based on a neural network algorithm, characterized in that: The following steps are involved: 1) Effectively process high-dimensional and multi-scale data by preprocessing and refining the dimensionality reduction of input data from different sources; 2) Build a multi-layer neural network architecture under the framework of physical information neural network, embed relevant reaction equations of the natural attenuation process of pollutants, and couple biochemical reaction mechanisms and data-driven models; 3) Train the neural network model used to invert the key parameters of the pollutant biodegradation numerical model, calculate the errors in the training process, and use the Adam optimization algorithm to improve the accuracy of the inversion results; 4) Integrate observational data with model prediction results, determine the final parameter settings of the numerical model, and conduct physical quantification of groundwater damage; The step 2) builds a multi-layer neural network architecture under the framework of physical information neural network, embeds the relevant reaction equations of the natural attenuation process of pollutants, and couples the biochemical reaction mechanism and data-driven model, as follows: In the framework of physical information neural network, based on the convolutional neural network architecture, a neural network parameter inversion model combining the groundwater pollutant degradation mechanism with the data-driven model is constructed; Define the input layer and output layer, build a neural network structure with multiple hidden layers, and capture the global dynamic changes of model parameters over a long time scale; Design a comprehensive loss function to quantify the difference between the output of the neural network model and the groundwater numerical model, embed the biodegradation equation of pollutants into the loss function in the form of residuals, and achieve the unification of reaction mechanism and data-driven; The step 3) trains a neural network model for inverting key parameters of the pollutant biodegradation numerical model, calculates the error during the training process, and uses the Adam optimization algorithm to improve the accuracy of the inversion result, as follows: Calculate the output of the network model from the input layer and generate prediction results; Use the loss function to calculate the loss value of the model, and adjust the network output and the weights in the loss function by minimizing the loss function; Starting from the output layer, calculate the gradient of the loss function with respect to the output, apply the chain rule to pass it back to each layer, and select the Adam optimization algorithm to update the network parameters; Monitor the loss function value and performance indicators during training to avoid overfitting problems.
2. The method for inverting key parameters in the numerical model of groundwater pollutant biodegradation based on neural network algorithm according to claim 1 is characterized in that: The step 1) effectively processes high-dimensional multi-scale data by preprocessing and refining the dimensionality reduction of input data from different sources, as follows: Collect and organize input data from multiple sources and scales; Perform data cleaning and standardization, perform feature engineering, and normalize the data; Kernel principal component analysis is used to reduce the dimensionality of high-dimensional data.
3. The method for inverting key parameters in the numerical model of groundwater pollutant biodegradation based on neural network algorithm according to claim 1 is characterized in that: The step 4) integrates the observation data and the model prediction results, determines the final parameter settings of the numerical model, and conducts the physical quantification of groundwater damage, specifically including: Construct an ensemble Kalman filter to integrate observation data with model predictions and iteratively update ensemble members; Fine-tune the model on the validation set, using the output of the ensemble Kalman filter to guide the adjustment of model parameters; Evaluate the generalization ability of the model on the test set, and evaluate the prediction accuracy and reliability of the model; Carry out comparative verification of neural network parameter inversion models to ensure the feasibility and accuracy of parameter inversion methods.
Citation Information
Patent Citations
Soil heavy metal content inversion model generation method and system, storage medium and inversion method
CN110991064A
River sudden water pollution early warning traceability method and system, terminal and medium
CN111898691A