Long short-term memory neural network system identification method based on square root unscented Kalman filtering
By introducing a long short-term memory neural network system identification method based on square root unscented Kalman filtering, the problems of inaccurate models and easy getting trapped in local optima in traditional sewage treatment are solved, and high-precision and robust prediction of sewage treatment process is achieved.
Patent Information
- Application Number
- CN202511002086.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-11-25
AI Technical Summary
In traditional wastewater treatment processes, decision variables for the reaction process are artificially adjusted, making it difficult to establish accurate system models. Furthermore, neural network models are prone to getting trapped in local optima and are difficult to adapt to complex and ever-changing wastewater treatment scenarios.
A long short-term memory neural network system identification method based on square root unscented Kalman filtering is adopted. Combined with the mechanism of nitrogen removal process in wastewater treatment, the adaptive ability and robustness of the model are improved by data preprocessing and iterative updating of model parameters.
It effectively avoids neural networks getting stuck in local optima, improves the prediction accuracy and robustness of key indicators in the wastewater treatment process, and adapts to dynamic changes under complex working conditions.
Smart Images

Figure CN121010033A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of wastewater treatment process system identification, and particularly relates to a long short-term memory neural network system identification method based on square root unscented Kalman filtering. BACKGROUND
[0002] Wastewater treatment is usually through a series of biochemical reactions to remove nitrogen and phosphorus-containing organic matter in wastewater, and the traditional wastewater treatment process often relies on manual adjustment of the decision variables of the reaction process, such as oxygen content, water inflow, etc., and the input of these indicators often depends on the changes of output indicators such as output ammonia nitrogen and total nitrogen. Therefore, it is of great significance to propose an effective and self-adaptive system identification scheme to model the entire wastewater treatment process to prevent in advance and for automatic control system.
[0003] In the field of wastewater treatment, it is often necessary to predict water indicators according to some key variables collected by sensors to ensure that problems can be prevented in advance. In addition, when controlling the denitrification process of wastewater treatment, an accurate system model is also needed. Therefore, it is extremely important to propose an accurate and reliable system identification method. The traditional modeling method based on the reaction mechanism of the wastewater treatment process is difficult to obtain an accurate system model due to the strong coupling of system state variables and too many uncertainties. With the development of machine learning, more and more adaptive methods have been proposed. These schemes rely on data from actual production processes and can obtain an accurate system model after certain data processing and training. They have more obvious advantages in dealing with complex systems.
[0004] Traditional modeling schemes often rely on the mechanism of the wastewater reaction process to model. To some extent, they can clearly reflect the physical causal relationship of some key variables. However, in actual production processes, the process is often too complex and there are many uncertain factors, so it is often difficult to obtain an accurate mechanism model. Moreover, the calculation process of a high-precision mechanism model is often too complex, which greatly increases the calculation cost.
[0005] Neural network models based on machine learning can often obtain better system models due to their strong adaptability when dealing with complex and variable situations such as wastewater treatment. However, these neural network models also have some drawbacks, such as being easily trapped in local optimal solutions, gradient explosion, and the accuracy of parameters cannot be guaranteed. When the model is trapped in a local optimal solution, on the one hand, it will be difficult for the model to jump out of the current optimization path and obtain an ideal model. On the other hand, when the model is trapped in a local optimal solution, the accuracy of the model parameter adjustment may not be guaranteed. SUMMARY
[0006] In order to overcome the defects and deficiencies existing in the prior art, the application provides a long short-term memory neural network system identification method based on square root unscented Kalman filtering, which can fully utilize the excellent modeling capability of the long short-term memory neural network and improve the adaptive capability and robustness of the model in combination with the Kalman filtering algorithm.
[0007] In order to achieve the above purpose, the application adopts the following technical solutions:
[0008] The application provides a long short-term memory neural network system identification method based on square root unscented Kalman filtering, which comprises the following steps:
[0009] Based on the mechanism relationship of the wastewater treatment denitrification process, variables are selected, effluent index data of the wastewater treatment denitrification process is collected, and a to-be-identified model is constructed;
[0010] The effluent index data of the wastewater treatment denitrification process is subjected to data preprocessing;
[0011] A long short-term memory neural network model is constructed as the structure of the to-be-identified model, and square root unscented Kalman filtering algorithm parameters are initialized;
[0012] The long short-term memory neural network model is subjected to parameter updating through cyclic iteration of square root unscented Kalman filtering states;
[0013] Based on the model training result, the structure and parameters of the to-be-identified model are obtained;
[0014] The selected variables are input into the long short-term memory neural network model to output ammonia nitrogen and total nitrogen prediction values.
[0015] As a preferred technical solution, the effluent index data of the wastewater treatment denitrification process is collected, the to-be-identified model is constructed, and specifically comprises:
[0016] The to-be-identified model is expressed as:
[0017] y k =f(U k )
[0018] U k =[y(k-1),y(k-2),…,y(k-l y ),u(k-1),u(k-2),…,u(k-l u )]
[0019] y=[S SNHout ,TN out ] T
[0020] u=[S NHin ,TN in ,COD inQ in ] T
[0021] wherein y k represents the output of the model to be identified, U k represents the input of the model to be identified, k represents the kth acquisition data, l y is the order of the system output, l u is the order of the system input, represents the effluent ammonia nitrogen amount, TN out represents the effluent total nitrogen amount, S NHin represents the influent ammonia nitrogen amount, TN in represents the influent total nitrogen amount, COD in represents the influent biological oxygen demand, Q in represents the influent flow rate.
[0022] As a preferred technical solution, the effluent index data of the wastewater treatment denitrification process are subjected to data preprocessing, specifically including:
[0023] The missing values and abnormal values in the effluent index data are detected, the missing values are supplemented by linear interpolation, and the abnormal values are removed, and the effluent index data are subjected to moving average and standardization processing.
[0024] As a preferred technical solution, a long short-term memory neural network model is constructed, represented as:
[0025] f t = σ (W xf · x t + W hg · h t-1 + b f )
[0026] i t = σ (W xi · x t + W hi · h t-1 + b i )
[0027]
[0028] o t = σ (W xo · x t + W ho · h t-1 + b o )
[0029] h t = o t ⊙ tanh (c t )
[0030] yt = W hy · h t + b y
[0031] wherein f t represents the amount of data information remaining after the previous LSTM neuron passes through the forgetting gate, x t represents the current input vector, W xf , W xi , W xc , W xo represent the input-related weight vector, h t-1 represents the state vector of the previous neural unit hidden layer, W hf , W hi , W nc , W ho , W hy represent the weight of the state vector, b f , b i , b c , b o , b y are the bias vectors of the relevant gates, respectively, and sigma and tanh represent the activation function, i t represents the amount of data information input of the current LSTM neuron, is the candidate state value of the unit, c t is the data information output of the LSTM neuron, c t-1 is the data information output of the previous LSTM neuron, h t is the hidden layer output of the current LSTM neuron, y t is the LSTM neuron output, and is the Hadamard product.
[0032] As a preferred technical solution, the square root unscented Kalman filter algorithm parameters are initialized and represented as:
[0033]
[0034] wherein x0 is the initial value of the state variable, x is the state vector, represents the predicted value of the initial state, E(·) represents the expectation, represents the initial state prediction error covariance matrix, and cholupdate{·} represents the Cholesky update factor.
[0035] As a preferred technical solution, the long short-term memory neural network model is updated by iteratively updating the state of the square root unscented Kalman filter, and specifically includes:
[0036] Start the loop iteration, update the prior state vector when the kth iteration of the square root unscented Kalman filter is performed on the state variable
[0037] Calculate the sigma sample point χ k|k-1 :
[0038]
[0039] λ=α 2 (L+κ)-L
[0040] Wherein, L is the dimension of the state variable, κ, α is an optional parameter, k is used to ensure the semi-definiteness of the covariance matrix, and α is used to control the size of the sigma sample point;
[0041] The sigma sample point and the external input U of the model to be identified are combined k As the input of the long short-term memory neural network model, the measurement value is updated as follows:
[0042]
[0043] Wherein, is the long short-term memory neural network model output corresponding to the current sigma sample point, is the long short-term memory neural network model output corresponding to the i-th sigma sample point, is the prior estimate value of the NARX model output, is the weight when calculating the mean of the sample point;
[0044] Calculate the innovation covariance matrix As follows:
[0045]
[0046] Wherein, qr represents QR decomposition, is the innovation, is the square root form of the measurement noise covariance Q k , is the square root form of the innovation covariance matrix, is the Cholesky factor of , respectively, represent the weights of sample points 0 and 1 when calculating the covariance, respectively represent the prediction vector set based on 2L sigma points and the prediction value based on the initial sigma point;
[0047] Calculate the cross-covariance matrix:
[0048]
[0049] where χ i,k|k-1 , are the i-th sigma sample point and the predicted value corresponding to the i-th sigma sample point, respectively, β represents a non-negative weighting term, denotes the weight of the sample point when calculating the covariance;
[0050] The Kalman gain K k is calculated, and the posteriori estimation value of the state variable is obtained is expressed as:
[0051]
[0052] The state prediction covariance square root is updated, and is expressed as:
[0053]
[0054] where H k is the observation matrix, is the cross-covariance matrix of the state vector and the NARX model output, is the state variable priori estimation covariance matrix, is the Cholesky decomposition factor of the state variable priori covariance matrix, is the updated value of ;
[0055] The adaptive noise estimation is expressed as:
[0056]
[0057] where, is the adaptive measurement noise covariance matrix estimation value corresponding to time step k and time step k-1, is the adaptive process noise covariance matrix estimation value corresponding to time step k and time step k-1, is the measurement noise covariance matrix estimation value, is the i-th diagonal element of the matrix , is the process noise covariance matrix estimation value, and b is the forgetting factor;
[0058] The outlier identification is expressed as:
[0059]
[0060] where α k is the statistic of the residual, R k is the process noise covariance matrix, P(·) represents the probability, χ 2 is a random variable of chi-square distribution, denotes the critical value of chi-square distribution under the significance level α and the degree of freedom s, and αχ represents the maximum tolerance probability of rejecting the null hypothesis, H0 represents the null hypothesis, i.e., for any k, the statistic k does not exceed H1 is an alternative hypothesis, i.e., there exists k such that When the null hypothesis is true, there is no anomaly, and when the alternative hypothesis is true, it means that the residual is too large, and an anomaly occurs;
[0061] According to the parameter adjustment of the outlier, it is represented as:
[0062]
[0063] Where, ρ k is a scaling factor for adjusting the measurement noise covariance matrix and the innovation covariance matrix, when the alternative hypothesis H1 is true, i.e., an outlier is detected, ρ k amplifies the observation error covariance, and reduces the Kalman gain, is the covariance matrix between the predicted observation values based on the information at the previous moment at time step k, N represents the number of samples, is the residual vector at time step k-j.
[0064] The application also provides a long short-term memory neural network system identification system based on a square root unscented Kalman filter, comprising: a to-be-identified model construction module, a data preprocessing module, a long short-term memory neural network model construction module, an initialization module, a parameter updating module, and a prediction module.
[0065] The to-be-identified model construction module is used for selecting variables based on the mechanism relationship of the wastewater treatment denitrification process, collecting effluent index data of the wastewater treatment denitrification process, and constructing a to-be-identified model.
[0066] The data preprocessing module is used for data preprocessing of the effluent index data of the wastewater treatment denitrification process.
[0067] The long short-term memory neural network model construction module is used for constructing a long short-term memory neural network model as the structure of the to-be-identified model.
[0068] The initialization module is used for initializing the parameters of the square root unscented Kalman filter algorithm.
[0069] The parameter updating module is used for updating the parameters of the long short-term memory neural network model through loop iteration of the square root unscented Kalman filter state.
[0070] The prediction module is used for obtaining the structure and parameters of the to-be-identified model based on the model training result, inputting the selected variables into the long short-term memory neural network model, and outputting the predicted values of ammonia nitrogen and total nitrogen.
[0071] As a preferred technical solution, the long short-term memory neural network model construction module is used to construct a long short-term memory neural network model, denoted as:
[0072] f t = σ (W xf · x t + W hf · h t-1 + b f )
[0073] i t = σ (W xi · x t + W hi · h t-1 + b i )
[0074]
[0075] o t = σ (W xo · x t + W ho · h t-1 + b o )
[0076] h t = o t ⊙ tanh (c t )
[0077] y t = W hy · h t + b y
[0078] Wherein, f t represents the amount of data information remaining after the previous LSTM neuron passes through the forgetting gate, x t represents the current input vector, W xf , W xi , W xc , W xo represent the input related weight vector, h t-1 represents the state vector of the previous neural unit hidden layer, W hf , W hi , W hc , W ho , W hy represent the weight of the state vector, b f , b i , b c , b o , b y are the bias vectors of the related gate, respectively, σ and tanh represent the activation function, i tData information input quantity of current LSTM neuron, is a candidate state value of the unit, c t Data information output quantity of LSTM neuron, c t-1 Data information output quantity of previous LSTM neuron, h t Current LSTM neuron hidden layer output, y t LSTM neuron output, is the Hadamard product.
[0079] As a preferred technical solution, the initialization module is used to initialize the square root unscented Kalman filter algorithm parameters, denoted as:
[0080]
[0081]
[0082] Wherein, x0 is the initial value of the state variable, x is the state vector, The predicted value of the initial state, E(·) represents the expectation, The initial state prediction error covariance matrix, cholupdate{·} represents the Cholesky update factor.
[0083] As a preferred technical solution, the parameter updating module is used to update the parameters of the long short-term memory neural network model by iteratively updating the state of the square root unscented Kalman filter, specifically including:
[0084] Start loop iteration, when the square root unscented Kalman filter is iterated for the kth time, update the prior state vector
[0085] Calculate the sigma sample point χ k|k-1 :
[0086]
[0087] λ=α 2 (L+κ)-L
[0088] Wherein, L is the dimension of the state variable, κ and α are optional parameters, κ is used to ensure the positive semi-definiteness of the covariance matrix, and α is used to control the size of the sigma sample point;
[0089] The sigma sample point and the external input U k As the input of the long short-term memory neural network model, the measurement value is updated as follows:
[0090]
[0091] Wherein, is the output of the long short-term memory neural network model corresponding to the current sigma sample point, is the output of the long short-term memory neural network model corresponding to the i-th sigma sample point, is the prior estimate of the NARX model output, is the weight when calculating the mean of the sample points;
[0092] The innovation covariance matrix is calculated as follows:
[0093]
[0094] wherein qr represents QR decomposition, is the innovation, is the square root form of the measurement noise covariance Q k , is the square root form of the innovation covariance matrix, is the Cholesky factor of , respectively represent the weights of the sample points 0 and 1 when calculating the covariance, respectively represent the prediction vector set based on 2L sigma points and the prediction value based on the initial sigma point;
[0095] The cross-covariance matrix is calculated:
[0096]
[0097] wherein χ i,k|k-1 , respectively represent the i-th sigma sample point and the prediction value corresponding to the i-th sigma sample point, and β represents a non-negative weighting term, represent the weight of the sample point when calculating the covariance;
[0098] The Kalman gain K is calculated k and the posterior estimate of the state variable is obtained is represented as:
[0099]
[0100] The state prediction covariance square root is updated, and is represented as:
[0101]
[0102] wherein H k is the observation matrix, is the cross-covariance matrix of the state vector and the NARX model output, is the prior estimate covariance matrix of the state variable, Cholesky factorization of the state variable prior covariance matrix, is updated as
[0103] Adaptive noise estimation is denoted as
[0104]
[0105] where is the adaptive measurement noise covariance matrix estimate corresponding to time step k and time step k-1, is the adaptive process noise covariance matrix estimate corresponding to time step k and time step k-1, is the measurement noise covariance matrix estimate, is the i-th diagonal element of the matrix is the process noise covariance matrix estimate, b is the forgetting factor;
[0106] Outlier detection is denoted as
[0107]
[0108] where α k is the statistic of the residual, R k is the process noise covariance matrix, P(·) denotes probability, χ 2 is a random variable following chi-square distribution, denotes the critical value of chi-square distribution with significance level α and degree of freedom s, α χ denotes the maximum tolerated probability of rejecting the null hypothesis, H0 denotes the null hypothesis, i.e., for any k, the statistic α k does not exceed H1 is the alternative hypothesis, i.e., there exists k such that When the null hypothesis is true, there is no anomaly, and when the alternative hypothesis is true, it means that the residual is too large, and an anomaly has occurred;
[0109] Parameter adjustment according to the outlier is denoted as
[0110]
[0111]
[0112] where ρ k is a scale factor used to adjust the measurement noise covariance matrix and the innovation covariance matrix, when the alternative hypothesis H1 is true, i.e., an abnormal value is detected, ρ k amplifies the observation error covariance, and reduces the Kalman gain, is a covariance matrix between observation values predicted based on previous time information at time step k, N represents a sample number, is a residual vector at time step k-j.
[0113] Compared with the prior art, the present application has the following advantages and beneficial effects:
[0114] (1) The present application introduces an adaptive estimation strategy for process noise and measurement noise in the filtering process, overcoming the problem that the traditional method needs to preset a fixed noise covariance matrix and is difficult to cope with environmental dynamic changes. Existing systems generally rely on experience to set or statically estimate noise parameters, and have poor robustness and generalization ability.
[0115] (2) The present application uses the square root unscented Kalman filter (SR-UKF) algorithm for long short-term memory neural network (LSTM) training, and applies it to dynamic modeling and key indicator prediction of nonlinear and strongly coupled wastewater treatment systems. Unlike existing LSTM training methods using BPTT (backpropagation through time) or momentum descent, this method can better avoid local optimization and improve prediction accuracy. Existing neural network modeling methods lack accuracy and robustness under complex working conditions, and are difficult to adapt to the uncertainty requirements of industrial processes. BRIEF DESCRIPTION OF DRAWINGS
[0116] Figure 1 is an implementation process architecture schematic diagram of the present application based on the square root unscented Kalman filter long short-term memory neural network system identification method.
[0117] Figure 2 is an implementation process architecture schematic diagram of the present application based on the square root unscented Kalman filter long short-term memory neural network system identification method. DETAILED DESCRIPTION
[0118] In order to make the purpose, technical scheme and advantages of the present application clearer and more apparent, the present application will be further described in detail below in combination with the drawings and examples. It should be understood that the specific examples described herein are only used to explain the present application and do not limit the present application.
[0119] Example 1
[0120] As shown in Figure 1 , Figure 2 , the present embodiment provides a long short-term memory neural network system identification method based on square root unscented Kalman filter, comprising the following steps:
[0121] S1: selecting relevant variables based on the mechanism relationship of the wastewater treatment denitrification process, and acquiring a data set by collecting data through a sensor, while determining the structure of the model to be identified, specifically including:
[0122] S11: Simulation data is acquired from the international benchmark simulation model BSM1 and actual factory sensors, and the data is organized into the dataset required for model training.
[0123] S12: Based on the key effluent indicators of wastewater denitrification treatment, namely effluent ammonia nitrogen. Total nitrogen (TN) in effluent out ) and influent ammonia nitrogen (S NHin ), total nitrogen in the influent (TN) in ), biological oxygen demand (COD) in ), inflow rate (Q) in The relationship between the following is established using the model to be identified:
[0124] y k =f(U k )
[0125] The model is a nonlinear autoregressive (NARX) model, where the input to the model to be identified is:
[0126] U k =[y(k-1),y(k-2),…,y(kl)] y ),u(k-1),u(k-2),…,u(kl u )];
[0127] y = [S SNHout ,TN out ] T ;
[0128] u = [S NHin ,TN in COD in Q in ] T ;
[0129] Where k represents the kth data collection, l y For the system output order, l u The NARX model, which provides the system with its input order, contains information about the structure and parameters of the system to be identified.
[0130] S2: Preprocess the collected data, specifically including:
[0131] Missing and outlier values in the collected dataset are detected and filled using linear interpolation. Outliers are removed, and the data is processed by moving average and standardization to ensure the validity and completeness of the dataset.
[0132] S3: Construct a long short-term memory neural network model and initialize the parameters of the square root unscented Kalman filter algorithm, specifically including:
[0133] S31: Establish a long short-term memory neural network model (LSTM), and use the neural network model structure as the structure of the nonlinear autoregressive model in step S12.
[0134] LSTM can effectively capture long-term dependencies in sequences by introducing a set of gate mechanisms to fine control the flow of information. Specifically, the LSTM unit includes three main gate units: input gate, forget gate and output gate, which are used to determine how much new information should be introduced, how much historical state should be forgotten, and how much current state should be output at the current time, and the mathematical expressions are:
[0135] f t =σ(W xf ·x t +W hf ·h t-1 +b f )
[0136] i t =σ(W xi ·x t +W hi ·h t-1 +b i )
[0137]
[0138] o t =σ(W xo ·x t +W ho ·h t-1 +b o )
[0139] h t =o t ⊙tanh(c t )
[0140] y t =W hy ·h t +b y
[0141] Where f t is the amount of data information left by the previous LSTM neuron through the forget gate, i t is the data information input of the current LSTM neuron, is the candidate state value of the unit, and c tThis is the data output of the LSTM neuron. This vector is propagated along the sequence in the time dimension, carrying key information across multiple time steps, thus achieving long-term information retention. t y is the output of the hidden layer of the current LSTM neuron. t This is the output of the LSTM. σ and tanh are the activation functions, ⊙ is the Hadamard product, and x... t W is the current input vector. xf W xi W xc W xo h represents the input-related weight vector. t-1 W is the state vector of the hidden layer of the previous neural unit. hf W hi W hc W ho W hy These are the weights of these state vectors, b f b i b c b o b y These are the bias vectors of the relevant gates;
[0142] S32: The Unscented Kalman Filter (UKF) algorithm is used to train the LSTM neural network model. Specifically, the neural network learning process is regarded as dynamic parameter estimation of a nonlinear system. That is, the weight vector in the LSTM neural network is regarded as the state of the system. The weight parameters are continuously updated iteratively based on the minimum mean square error (MSE) between the actual output and the predicted output. The state space model of the neural network is as follows:
[0143]
[0144] Where, x k The state vector x is composed of a weight matrix and biases. k+1 For the next state vector, U k For the NARX model input in step S12, z k The output of the model is h(·), which is the neural network function, and r is r. k Let the covariance matrix be R k The zero-mean Gaussian process noise vector, q k Let the covariance matrix be Q k The measurement noise vector;
[0145] In addition, considering the sample point generation in the iteration process of the unscented Kalman filter algorithm needs to take the square root of the covariance matrix and the positive definiteness and symmetry of the state error covariance matrix are often lost, leading to the divergence of the Kalman filter algorithm, QR decomposition and Cholesky update algorithm are introduced to effectively avoid the square root operation of the error covariance matrix. The specific expression is as follows:
[0146] P=AA T =R T Q T QR=R T R=SS T
[0147] Wherein, R is the upper triangular matrix obtained by QR decomposition of A T , and S=R T is a lower triangular matrix, which can be updated by propagation to avoid the square root calculation of the error covariance matrix in each iteration; at the same time, in order to update S, Cholesky factor update algorithm is introduced. For the Cholesky factor S of P=AA T , the vector update can be written as:
[0148]
[0149] If u is a matrix instead of a vector, each column of u can be used for continuous update.
[0150] Therefore, the initialization of square root unscented Kalman filter is as follows:
[0151]
[0152] Wherein, x0 is the initial value of state variable, x is the state vector, represents the predicted value of the initial state, E(·) represents expectation, represents the initial state prediction error covariance matrix, cholupdate{·} represents Cholesky update factor.
[0153] S4: Start training with the data set, and update the parameters of the neural network model by iterating the square root unscented Kalman filter state through a loop;
[0154] S41: Start loop iteration, when the square root unscented Kalman filter is iterated for the kth time, update the prior state vector
[0155] S42: Calculate the sigma sample point χ k|k-1 :
[0156]
[0157] λ = a 2 (L + K) - L
[0158] where L is the dimension of the state variable, K, a are optional parameters, K (K≥0) is used to ensure the semi-positive definiteness of the covariance matrix, a (0≤a≤1) controls the size of the sigma sample points.
[0159] S43: The sigma sample points and the NARX model external input U k as the input of the LSTM neural network model, and update the measurements as follows:
[0160]
[0161] where, is the LSTM neural network output corresponding to the current sigma sample point, is the LSTM neural network output corresponding to the i-th sigma sample point, is the prior estimate of the NARX model output, is the weight for calculating the mean of the sample points, and its value is calculated as:
[0162]
[0163] S44: Calculate the innovation covariance matrix as follows:
[0164]
[0165] where qr represents the QR decomposition, is the residual vector (also known as innovation), is the square root form of the measurement noise covariance Q k , is the square root form of the innovation covariance matrix, is the Cholesky factor of , respectively represent the weights of sample points 0, 1 in calculating the covariance, respectively represent the prediction vector set based on 2L sigma points and the prediction value based on the initial sigma point.
[0166] S45: Calculate the cross-covariance matrix:
[0167]
[0168] where x i,k|k-1 , respectively, β (β≥0) represents a non-negative weighting term, represents the weight of the sample point when calculating the covariance;
[0169] S46: Calculate the Kalman gain K k and thus obtain the posterior estimate of the state variable The calculation expression is as follows:
[0170]
[0171] S47: Update the state prediction covariance square root:
[0172]
[0173] wherein H k is the observation matrix, is the cross-covariance matrix of the state vector and the NARX model output, is the state variable prior estimate covariance matrix, is the Cholesky decomposition factor of the state variable prior covariance matrix, is the updated value of
[0174] S48: Adaptive noise estimation, which is prepared for the next round of calculation of the innovation covariance matrix and the like, and the calculation expression is as follows:
[0175]
[0176] wherein is the adaptive measurement noise covariance matrix estimate value corresponding to time step k and time step k-1, is the adaptive process noise covariance matrix estimate value corresponding to time step k and time step k-1, is the measurement noise covariance matrix estimate value, is the i-th diagonal element of the matrix is the process noise covariance matrix estimate value, b is a forgetting factor and 0.95≤b≤0.99;
[0177] S49: Outlier identification, which is represented as:
[0178]
[0179] wherein α k is a statistic of the residual, used to detect abnormal conditions, R k is the process noise covariance matrix, P(·) represents probability, χ 2 is a random variable of chi-square distribution, denotes the critical value of the chi-square distribution with significance level a and degree of freedom s, a χ denotes the maximum tolerated probability of rejecting the null hypothesis. H0 denotes the null hypothesis, i.e., for any k, the statistic a k does not exceed H1 is the alternative hypothesis, i.e., there exists k such that When the null hypothesis is true, there is no anomaly, and when the alternative hypothesis is true, it indicates that the residual is too large, and an anomaly occurs.
[0180] S410: According to the parameter adjustment of the outlier, it is expressed as:
[0181]
[0182]
[0183] where, p k is a scale factor used to adjust the measurement noise covariance matrix and the innovation covariance matrix, when the alternative hypothesis H1 is true, i.e., an abnormal value is detected, p k will amplify the observation error covariance, thereby reducing the Kalman gain and reducing the amplitude of state update, improving the robustness of the filter, is the covariance matrix between the predicted observation values at time step k based on the information at the previous moment, N denotes the number of samples, is the residual vector at time step k-j.
[0184] S5: Obtain the structure and parameters of the to-be-identified model based on the model training result.
[0185] In this embodiment, the state vector posterior estimate value obtained at the end of training is the final weight parameter value of the LSTM neural network, and these weight values constitute the parameters of the NARX model, while the structure of the LSTM neural network serves as the structure of the NARX model. At this point, the NARX model of the effluent ammonia nitrogen and the effluent total nitrogen has been identified, and when the already identified model is used subsequently, only the collected variables need to be input into the LSTM neural network to obtain accurate output ammonia nitrogen and output total nitrogen prediction values.
[0186] In summary, in terms of modeling of nonlinear systems, the SR-UKF algorithm is introduced to replace the traditional Kalman filter algorithm which is only applicable to linear systems, thereby effectively improving the numerical stability and estimation accuracy of the filtering process. Secondly, in view of the unknown and time-varying noise distribution in the sewage treatment process, the adaptive noise estimation mechanism is integrated, so that the system can dynamically adjust the noise covariance parameters, thereby significantly improving the robustness of the model under complex working conditions. Finally, in terms of neural network training, the SR-UKF-based LSTM training method is proposed to replace the traditional backpropagation through time (BPTT) and momentum descent algorithm, thereby solving the problems of easy falling into local optimum and slow convergence, and effectively improving the accuracy and anti-interference ability of the system in the key effluent index prediction task.
[0187] Embodiment 2
[0188] The embodiment provides a long short-term memory neural network system identification system based on square root unscented Kalman filtering, which is used to implement the long short-term memory neural network system identification method based on square root unscented Kalman filtering in the above embodiment 1. The system comprises a to-be-identified model construction module, a data preprocessing module, a long short-term memory neural network model construction module, an initialization module, a parameter updating module and a prediction module.
[0189] In the embodiment, the to-be-identified model construction module is used to select variables based on the mechanism relationship of the sewage treatment denitrification process, collect effluent index data of the sewage treatment denitrification process, and construct a to-be-identified model.
[0190] In the embodiment, the data preprocessing module is used to perform data preprocessing on the effluent index data of the sewage treatment denitrification process.
[0191] In the embodiment, the long short-term memory neural network model construction module is used to construct a long short-term memory neural network model as the structure of the to-be-identified model.
[0192] In the embodiment, the initialization module is used to initialize the parameters of the square root unscented Kalman filter algorithm.
[0193] In the embodiment, the parameter updating module is used to update the parameters of the long short-term memory neural network model by iteratively updating the square root unscented Kalman filter state.
[0194] In the embodiment, the prediction module is used to obtain the structure and parameters of the to-be-identified model based on the model training result, input the selected variables into the long short-term memory neural network model, and output the predicted values of ammonia nitrogen and total nitrogen.
[0195] In the embodiment, the long short-term memory neural network model construction module is used to construct a long short-term memory neural network model, which is represented as:
[0196] f t = σ(W xf · x t + W hf · h t-1 + b f )
[0197] i t = σ(W xi · x t + W hi · h t-1 + b i )
[0198]
[0199] o t = σ(W xo · x t + W ho · h t-1 + b o )
[0200] h t = o t ⊙ tanh(c t )
[0201] y t = W hy · h t + b y
[0202] wherein f t represents the amount of data information remaining after the previous LSTM neuron passes through the forgetting gate, x t represents the current input vector, W xf , W xi , W xc , W xo represent the input-related weight vectors, h t-1 represents the state vector of the previous neural unit hidden layer, W hf , W hi , W nc , W ho , W hy represent the weights of the state vector, b f , b i , b c , b o , b y are the bias vectors of the related gates, respectively, σ and tanh represent the activation functions, i t represents the data information input amount of the current LSTM neuron, is the candidate state value of the unit, c tis the data information output of the LSTM neuron, c t-1 is the data information output of the previous LSTM neuron, h t is the hidden layer output of the current LSTM neuron, y t is the LSTM neuron output, and is the Hadamard product.
[0203] In this embodiment, the initialization module is configured to initialize the parameters of the square root unscented Kalman filter algorithm, denoted as:
[0204]
[0205] where x0 is the initial value of the state variable, x is the state vector, represents the predicted value of the initial state, E(·) represents the expectation, represents the initial state prediction error covariance matrix, and cholupdate{·} represents the Cholesky update factor.
[0206] In this embodiment, the parameter updating module is configured to update the parameters of the long short-term memory neural network model by iteratively performing square root unscented Kalman filtering, and specifically includes:
[0207] Start the loop iteration, and when the k-th iteration of the square root unscented Kalman filter is performed on the state variable, update the prior state vector
[0208] Calculate the sigma sample point χ k|k-1 :
[0209]
[0210] λ=α 2 (L+κ)-L
[0211] where L is the dimension of the state variable, k and a are optional parameters, k is used to ensure the positive semi-definiteness of the covariance matrix, and a is used to control the size of the sigma sample point;
[0212] The sigma sample point and the external input U k of the model to be identified are taken as the input of the long short-term memory neural network model, and the measurement value is updated as follows:
[0213]
[0214] wherein is the output of the long short-term memory neural network model corresponding to the current sigma sample point, is the output of the long short-term memory neural network model corresponding to the i-th sigma sample point, prior estimate of the NARX model output, weight in calculating the mean of the sample points;
[0215] compute the innovation covariance matrix as follows:
[0216]
[0217] where qr denotes the QR decomposition, innovation, measurement noise covariance Q k in square-root form, square-root form of the innovation covariance matrix, Cholesky factor of denote the weights of the sample points 0, 1 in calculating the covariance, respectively, denote the prediction vector set based on 2L sigma points and the prediction value based on the initial sigma point, respectively;
[0218] compute the cross-covariance matrix:
[0219]
[0220] where χ i,k|k-1 , denote the i-th sigma sample point and the prediction value corresponding to the i-th sigma sample point, respectively, and β denotes the non-negative weighting term, denote the weights of the sample points in calculating the covariance;
[0221] compute the Kalman gain K k and obtain the posterior estimate of the state variable denoted as:
[0222]
[0223] update the state prediction covariance square root, denoted as:
[0224]
[0225] where H k is the observation matrix, is the cross-covariance matrix of the state vector and the NARX model output, is the prior estimate covariance matrix of the state variable, is the Cholesky decomposition factor of the prior covariance matrix of the state variable, is the updated value of ; and
[0226] Adaptive noise estimation is expressed as:
[0227]
[0228] in, These are the estimated values of the adaptive measurement noise covariance matrix for time step k and time step k-1. These are the estimated values of the adaptive process noise covariance matrix for time step k and time step k-1. To measure the estimated value of the noise covariance matrix, It is a matrix The i-th diagonal element, is the estimated value of the process noise covariance matrix, and b is the forgetting factor.
[0229] Outlier identification is represented as:
[0230]
[0231]
[0232] Where, α k R is a statistic of the residuals used to detect outliers. k Let χ be the process noise covariance matrix, P(·) denote the probability. 2 Let be a random variable with a chi-square distribution. This represents the significance level α and the critical value of the chi-square distribution with s degrees of freedom. χ H0 represents the maximum tolerance probability for rejecting the null hypothesis. H0 represents the null hypothesis, i.e., for any k, the statistic α... k None more than H1 is the alternative hypothesis, i.e., there exists k such that There is no anomaly when the null hypothesis is true; when the alternative hypothesis is true, it indicates that the residuals are too large and an anomaly has occurred.
[0233] Adjusting the parameters based on outliers is represented as follows:
[0234]
[0235] Where, ρ k It is a scaling factor used to adjust the measurement noise covariance matrix and the innovation covariance matrix. When the alternative hypothesis H1 is true, i.e., outliers are detected, ρ k It amplifies the observation error covariance, thereby reducing the Kalman gain, decreasing the amplitude of state updates, and improving the robustness of the filter. Let N be the covariance matrix between the observations predicted at time step k based on information from the previous time step, and let N represent the sample size. Let be the residual vector at time step kj.
[0236] The above embodiments are the preferred embodiments of the present application, but the embodiments of the present application are not limited to the above embodiments, and any changes, modifications, substitutions, combinations, simplifications, etc. made without departing from the spirit and principles of the present application should be equivalent replacement manners and should be included in the protection scope of the present application.
Claims
1. A system identification method based on square-root unscented Kalman filter long short-term memory neural network, characterized in that, The method comprises the following steps: Based on the mechanism of wastewater treatment denitrification process, the variables are selected, the effluent index data of the wastewater treatment denitrification process are collected, and the to-be-identified model is constructed. The effluent index data of the wastewater treatment denitrification process are preprocessed. A long short-term memory neural network model is constructed as the structure of the to-be-identified model, and the square root unscented Kalman filter algorithm parameters are initialized. The long short-term memory neural network model is updated through cyclic iteration of the square root unscented Kalman filter state. The structure and parameters of the to-be-identified model are obtained based on the model training result. The selected variables are input into the long short-term memory neural network model to output the ammonia nitrogen and total nitrogen prediction values.
2. The square root unscented Kalman filter based long short-term memory neural network system identification method according to claim 1, wherein, The effluent index data of the wastewater treatment denitrification process are collected, and the to-be-identified model is constructed, specifically including: The to-be-identified model is represented as: y k = f(U k ) U k = [y(k - 1), y(k - 2),..., y(k - l y ), u(k - 1), u(k - 2),..., u(k - l u )] y = [S SNHout ,TN out ] T u = [S NHin ,TN in ,COD in ,Q in ] T wherein y k represents the output of the model to be identified, U k represents the input of the model to be identified, k represents the kth acquisition data, l y is the order of the system output, l u is the order of the system input, S NHout represents the effluent ammonia nitrogen amount, TN out represents the effluent total nitrogen amount, S NHin represents the influent ammonia nitrogen amount, TN in represents the influent total nitrogen amount, COD in represents the influent biological oxygen demand, Q in represents the influent flow.
3. The square root unscented Kalman filter based long short-term memory neural network system identification method according to claim 1, wherein, The effluent index data of the wastewater treatment denitrification process are preprocessed, specifically including: The missing values and abnormal values in the effluent index data are detected, the missing values are supplemented through linear interpolation, the abnormal values are removed, the effluent index data are subjected to moving average and standardization processing.
4. The square root unscented Kalman filter based long short-term memory neural network system identification method according to claim 1, wherein, The long short-term memory neural network model is constructed and represented as: f t = σ(W xf · x t + W hf · h t-1 + b f ) i t = σ(W xi · x t + W hi · h t-1 + b i ) o t = σ(W xo · x t + W ho · h t-1 + b o ) h t = o t ⊙ tanh(c t ) y t = W hy · h t + b y wherein f t represents the amount of data information remaining after the previous LSTM neuron passes through the forget gate, x t represents the current input vector, W xf , W xu , W xc , W xo represents the input-related weight vector, h t-1 represents the state vector of the previous neural unit hidden layer, W hf , W hi , W hc , W ho , W hy represents the weight of the state vector, b f , b i , b c , b o , b y are the bias vectors of the relevant gates, respectively, σ and tanh represent the activation functions, i t represents the amount of data information input of the current LSTM neuron, is the candidate state value of the unit, c t is the data information output of the LSTM neuron, c t-1 is the data information output of the previous LSTM neuron, h t is the hidden layer output of the current LSTM neuron, y t is the LSTM neuron output, and is the Hadamard product.
5. The square root unscented Kalman filter based long short-term memory neural network system identification method according to claim 1, wherein, The square root unscented Kalman filter algorithm parameters are initialized and represented as: where x0is the initial value of the state variable, x is the state vector, denotes the prediction of the initial state, E(·) denotes expectation, denotes the initial state prediction error covariance matrix, cholupdate{·} denotes the Cholesky update factor.
6. The square root unscented Kalman filter based long short-term memory neural network system identification method according to claim 1, wherein, The long short-term memory neural network model is updated through cyclic iteration of the square root unscented Kalman filter state, specifically including: beginning a loop iteration, updating the a priori state vector when the kth iteration of the square root unscented Kalman filter is performed on the state variable Compute sigma sample point x k|k-1 : λ = a 2 (L + K) - L Wherein, L is the dimension of the state variable, κ and α are optional parameters, κ is used to ensure the semi-positive definiteness of the covariance matrix, and α is used to control the size of the sigma sampling point; The sigma sample points and the input U outside the model to be recognized k As an input of the long short-term memory neural network model, the update measurement value is as follows: wherein, is the long short-term memory neural network model output corresponding to the current sigma sample point, is the long short-term memory neural network model output corresponding to the i-th sigma sample point, is the prior estimate of the NARX model output, is the weight when calculating the mean of the sample points; Computing the innovation covariance matrix As follows: where qr denotes a QR decomposition, is the innovation, is the measurement noise covariance Q k in square-root form, is the square-root form of the innovation covariance matrix, is the Cholesky factor of denote the weights of the sample points 0, 1, respectively, in the calculation of the covariance, denote the set of prediction vectors based on the 2L sigma points and the prediction value based on the initial sigma point, respectively; The cross-covariance matrix is calculated as: where χ i,k|k-1 , are the i-th sigma sample point and the predicted value corresponding to the i-th sigma sample point, respectively, β represents a non-negative weighting term, denotes the weight of the sample point when calculating the covariance. The Kalman gain K is calculated k and the posterior estimate of the state variable is obtained is expressed as: The state prediction covariance square root is updated and represented as: where H k is the observation matrix, is the state vector and the cross-covariance matrix of the NARX model output, is the state variable prior estimate covariance matrix, is the Cholesky factor of the state variable prior covariance matrix, is the updated value of Adaptive noise estimation is represented as: wherein is an adaptive measurement noise covariance matrix estimate corresponding to time step k and time step k-1, is an adaptive process noise covariance matrix estimate corresponding to time step k and time step k-1, is a measurement noise covariance matrix estimate, is the i-th diagonal element of the matrix is a process noise covariance matrix estimate, and b is a forgetting factor. Outlier identification is represented as: where α k is a statistic of the residuals, R k is the process noise covariance matrix, P(·) denotes probability, χ 2 is a random variable of chi-square distribution, denotes the critical value of chi-square distribution with significance level α and degree of freedom s, α χ denotes the maximum tolerated probability of rejecting the original hypothesis, H0 denotes the original hypothesis, i.e. for any k, the statistic α k does not exceed H1 is the alternative hypothesis, i.e. there exists k such that When the original hypothesis is true, there is no anomaly, and when the alternative hypothesis is true, it indicates that the residuals are too large and an anomaly has occurred; Parameter adjustment according to the outlier is represented as: where ρ k is a scale factor used to adjust the measurement noise covariance matrix and the innovation covariance matrix, when the alternative hypothesis H1 holds, i.e. an outlier is detected, k amplifies the observation error covariance, reducing the Kalman gain, is the covariance matrix between the predicted observations based on the information at the previous time step for time step k, N denotes the number of samples, is the residual vector at time step k-j.
7. A system for system identification based on square root unscented Kalman filter long short-term memory neural network system, characterized in that, It comprises: The to-be-identified model construction module, the data preprocessing module, the long short-term memory neural network model construction module, the initialization module, the parameter updating module and the prediction module. The to-be-identified model construction module is used to select variables based on the mechanism of the wastewater treatment denitrification process, collect the effluent index data of the wastewater treatment denitrification process, and construct the to-be-identified model. The data preprocessing module is used to preprocess the effluent index data of the wastewater treatment denitrification process. The long short-term memory neural network model construction module is used to construct a long short-term memory neural network model as the structure of the to-be-identified model. The initialization module is used to initialize the square root unscented Kalman filter algorithm parameters. The parameter updating module is used to update the long short-term memory neural network model through cyclic iteration of the square root unscented Kalman filter state. The prediction module is used to obtain the structure and parameters of the to-be-identified model based on the model training result, input the selected variables into the long short-term memory neural network model, and output the ammonia nitrogen and total nitrogen prediction values.
8. The square root unscented Kalman filter based long short-term memory neural network system identification system according to claim 7, wherein, The long short-term memory neural network model construction module is used to construct a long short-term memory neural network model, represented as: f t = σ(W xf · x t + W hf · h t-1 + b f ) i t = σ(W xi · x t + W hi · h t-1 + b i ) o t = σ(W xo · x t + W ho · h t-1 + b o ) h t = o t ⊙ tanh(c t ) y t = W hy • h t + b y wherein f t represents the amount of data information remaining after the previous LSTM neuron passes through the forget gate, x t represents the current input vector, W xf , W xi , W xc , W xo represents the input-related weight vector, h t-1 represents the state vector of the previous neural unit hidden layer, W hf , W hi , W hc , W ho , W hy represents the weight of the state vector, b f , b i , b c , b o , b y are the bias vectors of the relevant gates, respectively, σ and tanh represent the activation functions, i t represents the amount of data information input of the current LSTM neuron, is the candidate state value of the unit, c t is the data information output of the LSTM neuron, c t-1 is the data information output of the previous LSTM neuron, h t is the hidden layer output of the current LSTM neuron, y t is the LSTM neuron output, and is the Hadamard product.
9. The square root unscented Kalman filter based long short-term memory neural network system identification system according to claim 7, wherein, The initialization module is used to initialize the square root unscented Kalman filter algorithm parameters, represented as: where x0is the initial value of the state variable, x is the state vector, denotes the prediction of the initial state, E(·) denotes expectation, denotes the initial state prediction error covariance matrix, cholupdate{·} denotes the Cholesky update factor.
10. The square root unscented Kalman filter based long short-term memory neural network system identification system of claim 7, wherein, The parameter updating module is used to update the long short-term memory neural network model through cyclic iteration of the square root unscented Kalman filter state, specifically including: beginning a loop iteration, updating the a priori state vector when the kth iteration of the square root unscented Kalman filter is performed on the state variable Compute sigma sample point x k|k-1 : λ = a 2 (L + K) - L where L is the dimension of the state variable, κ, α are optional parameters, κ is used to ensure the semi-positive definiteness of the covariance matrix, and α is used to control the size of the sigma sampling point; The sigma sample points and the input U outside the model to be recognized k As an input of the long short-term memory neural network model, the update measurement value is as follows: wherein, is the output of the long short-term memory neural network model corresponding to the current sigma sample point, is the output of the long short-term memory neural network model corresponding to the i-th sigma sample point, is the prior estimate of the NARX model output, is the weight when calculating the mean of the sample points; Computing the innovation covariance matrix As follows: where qr denotes a QR decomposition, is the innovation, is the measurement noise covariance Q k in square-root form, is the square-root form of the innovation covariance matrix, is the Cholesky factor of , respectively, denote the weights of the sample points 0, 1, respectively, in the calculation of the covariance, denote the set of prediction vectors based on the 2L sigma points and the prediction value based on the initial sigma point, respectively; Calculate the cross-covariance matrix: where χ i,k|k-1 , are the i-th sigma sample point and the predicted value corresponding to the i-th sigma sample point, respectively, and β represents a non-negative weighting term, denotes the weight of the sample point when calculating the covariance. The Kalman gain K is calculated k and the posterior estimate of the state variable is obtained is expressed as: Update the state prediction covariance square root, denoted as: where H k is the observation matrix, is the state vector and the cross-covariance matrix of the NARX model output, is the state variable prior estimate covariance matrix, is the Cholesky factor of the state variable prior covariance matrix, is the updated value of Adaptive noise estimation, denoted as: wherein is an adaptive measurement noise covariance matrix estimate corresponding to time step k and time step k-1, is an adaptive process noise covariance matrix estimate corresponding to time step k and time step k-1, is a measurement noise covariance matrix estimate, is the i-th diagonal element of the matrix is a process noise covariance matrix estimate, and b is a forgetting factor. Outlier identification, denoted as: where α k is a statistic of the residuals, R k is a process noise covariance matrix, P(·) denotes probability, χ 2 is a random variable of chi-square distribution, denotes the critical value of chi-square distribution with significance level α and degree of freedom s, α χ denotes the maximum tolerated probability of rejecting the original hypothesis, H0 denotes the original hypothesis, i.e. for any k, the statistic α k does not exceed H1 is the alternative hypothesis, i.e. there exists k such that When the original hypothesis is true, there is no anomaly, and when the alternative hypothesis is true, it indicates that the residuals are too large and an anomaly has occurred; Parameter adjustment according to the outlier, denoted as: where ρ k is a scale factor used to adjust the measurement noise covariance matrix and the innovation covariance matrix, when the alternative hypothesis H1 holds, i.e. an outlier is detected, k amplifies the observation error covariance, reducing the Kalman gain, is the covariance matrix between the predicted observations based on the information at the previous time step for time step k, N denotes the number of samples, is the residual vector at time step k-j.