A method for measuring the concentration of ammonia nitrogen in effluent based on neural network and Bayesian optimization
Through the combination of adaptive bidirectional long and short-term memory neural network and improved parallel Bayesian optimization algorithm, the generalization ability and stability of effluent ammonia nitrogen concentration prediction in the existing technology is solved, and high-precision prediction from minute to seasonal level is achieved, which improves the intelligent regulation capabilities of sewage treatment plants.
Patent Information
- Application Number
- CN202510614762.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-05-14
AI Technical Summary
The existing deep learning methods have poor generalization ability, poor stability and low accuracy in the prediction of effluent ammonia nitrogen concentration in the wastewater treatment field. It is difficult to simultaneously capture the fluctuations of ammonia nitrogen concentrations from minute to seasonal level. The optimization algorithm converges slowly when the seasonal data distribution is offset, which cannot meet the real-time control needs.
The effluent ammonia nitrogen concentration measurement method based on adaptive bidirectional long and short-term memory neural network and improved parallel Bayesian optimization algorithm is adopted. Multi-scale features are extracted through the timing convolution layer and the SE attention mechanism, and the input step size is dynamically adjusted. The parameter search efficiency is improved in combination with the parallel Bayesian optimization algorithm to achieve high-precision prediction of different working conditions.
It improves the accuracy of the prediction of ammonia nitrogen concentration in effluent and the generalization ability of the model, and can maintain high prediction accuracy under different operating conditions, provide reliable support for intelligent regulation of sewage treatment plants and reduce operation and maintenance costs.
Smart Images

Figure CN120144969B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of sewage treatment and artificial intelligence, and particularly relates to a method for measuring the effluent ammonia nitrogen concentration based on neural network and Bayesian optimization. Background Art
[0002] With the increasingly strict environmental protection standards and the complexity of sewage treatment processes, the accurate prediction of the effluent ammonia nitrogen concentration has become a key task to ensure the stable operation of sewage treatment plants and the compliance of water quality. Ammonia nitrogen is one of the core indicators of sewage discharge. Exceeding the concentration standard of ammonia nitrogen will not only cause water eutrophication, but also may pose hazards to the ecological system and human health. The traditional method relying on manual sampling and laboratory detection has lag, and it is difficult to provide real-time guidance for process control. Therefore, developing an efficient and accurate ammonia nitrogen concentration prediction model has important engineering application value.
[0003] In recent years, deep learning technology has been introduced into the field of sewage treatment due to its powerful feature extraction ability for the prediction of the effluent ammonia nitrogen concentration. However, existing deep learning methods still have problems such as poor generalization ability, weak stability, and low accuracy when predicting the effluent ammonia nitrogen concentration. Summary of the Invention
[0004] Aiming at the problems existing in the prior art, the purpose of the present invention is to propose a method for predicting the effluent ammonia nitrogen concentration based on an adaptive long short-term memory neural network and an improved parallel Bayesian optimization algorithm to improve the prediction accuracy and generalization ability of the model under different working conditions.
[0005] To achieve the above purpose, the technical solution adopted by the present invention is: a method for measuring the effluent ammonia nitrogen concentration based on neural network and Bayesian optimization, including the following steps:
[0006] Step 1: Construct a basic model with a temporal convolutional layer, an attention mechanism, and a bidirectional long short-term memory neural network, collect feature variables x related to the prediction of the effluent ammonia nitrogen concentration, calculate the Pearson correlation coefficient, retain the feature variables x according to the Pearson correlation coefficient, preprocess the retained feature variables x, and normalize each feature variable x to the interval [0, 1] by using the Min-Max method;
[0007] Step 2: Input the feature variables x into the basic model for feature extraction and enhancement, extract the local dependence relationship in the time series data through the temporal convolutional layer TCL, recalibrate the features through the SE attention mechanism, enhance the response of key features and suppress noise;
[0008] Step 3: Extract the long-term temporal features of the feature sequence from the feature variables x after feature extraction and enhancement through a bidirectional long short-term memory neural network with an adaptive mechanism, dynamically adjust the input step size to adapt to the changes in different time periods in the data, capture the bidirectional time dependencies in the sequence, and output the predicted value;
[0009] Step 4: The parallel Bayesian optimization algorithm dynamically adjusts the acquisition function according to different stages of the optimization process, evaluates multiple candidate points simultaneously, and adaptively acquires the function in parallel to output the measured value.
[0010] For the above method for measuring the ammonia nitrogen concentration in the effluent based on neural network and Bayesian optimization, in Step 1, the calculation formula for the Pearson correlation coefficient is: , where, r is the Pearson correlation coefficient, x i and y i are the i th observed values of the two variables, x m and y m are the means of the variables x and y respectively, n is the number of samples, and the feature variables include the finally selected dissolved oxygen at the end of aerobic stage, total suspended solids at the end of aerobic stage, effluent pH value, effluent oxidation-reduction potential, and effluent nitrate nitrogen.
[0011] For the above method for measuring the ammonia nitrogen concentration in the effluent based on neural network and Bayesian optimization, the preprocessing includes first performing data cleaning on the collected original data to handle missing values and outliers. For the missing values, different filling strategies are adopted according to the data characteristics. For continuous variables, the time series interpolation method is used to reasonably fill the values using the numerical values of adjacent time points before and after; for the outliers, combining statistical methods and business experience, the outlier points are identified by calculating the distribution range of each feature, and the sliding window median replacement method is used for correction.
[0012] For the above method for measuring the ammonia nitrogen concentration in the effluent based on neural network and Bayesian optimization, Step 2 includes:
[0013] Step 2-1: Extract the local dependencies of the time series data through the temporal convolutional layer TCL. The output x c of TCL is calculated by the formula:
[0014] , where, δ (⋅) is the ReLU activation function used to introduce non-linearity, Wk,c,c′ is the weight matrix of the convolutional kernel, x t+k,c is the value of the input data at time step t + k and input channel c ; b c′ is the bias term, b ∈R c′ ;
[0015] Step 2-2: Re-calibrate the feature variable x through the SE attention mechanism to enhance the response of key features and suppress noise, perform channel re-calibration, and apply the obtained weight s after excitation to the original input features.
[0016] The above method for measuring the ammonia nitrogen concentration in the effluent based on neural network and Bayesian optimization, the Step 2-2 includes:
[0017] Step a: Perform compression, use global average pooling to generate a channel descriptor z c′′ , which is used to emphasize the spatial information of the global distribution. The formula is: , where x ijc ′′ represents the feature value of the i -th channel at height j and width c ; H and W represent the height and width of the input channel respectively;
[0018] Step b: The SE attention mechanism reduces the parameters and complexity by using the compression ratio r, and increases the non-linearity by using the ReLU activation function;
[0019] Step c: Calculate the weight of each channel through the sigmoid function to achieve effective re-calibration of the features. The output M 1 and M 2 of the excitation are calculated by the weights of two fully connected layers s as: , where M 1∈R c′′×c′′ / r , M 2∈R c ′′ / r×c′′ , z is composed of c ′′ channel descriptors, z ∈R c′′×1 is the output of the compression step, s ∈R c′′×1 , σ is the sigmoid activation function;
[0020] Step d: Perform recalibration of the channels, and apply the weights obtained after excitation s to the original input features. For each channel, the recalibrated output is: , where s c′′ is the weight of channel c '', x c′′ is the original input feature of channel c '', x s is the output after SE recalibration.
[0021] The above-mentioned method for measuring the ammonia nitrogen concentration in the effluent based on neural network and Bayesian optimization, step 3 includes:
[0022] Step 3-1: The bidirectional long short-term memory neural network ABiLSTM processes the input sequence step by step through LSTM units. The LSTM units include the update of the input gate, forget gate, output gate, and cell state. The algorithm formulas are as follows:
[0023] ,
[0024] ,
[0025] ,
[0026] ,
[0027] ,
[0028] ,
[0029] where x t ( x t ∈R m×1 ) is the received input vector, h t ( h t ∈R m×1 ) is the hidden state at time step t, h t−1 is the hidden state of the previous time step, g t is the new candidate value, f t , i t , o t are the forget gate, input gate, and output gate at tOutput at a moment c t is t the cell state at a moment W f , W O , W i , W C ( W ∈R h×m ) respectively represent t the weight matrices connecting the input x t to different gates b f , b i , b C , b O ( b ∈R h×1 ) is the bias term;
[0030] Step 3-2: Define a window with a minimum length of Δ t min . Dynamically adjust the input step size based on the error, and only process the data within the time steps contained in the window each time. The dynamically adjusted window size for the prediction error is: , where α is the adjustment coefficient, and error t is the mean square error of the current time step;
[0031] Step 3-3: Capture the bidirectional temporal features of the data through forward and backward propagation. For the input x t ∈R m×η , the specific operation process of ABiLSTM can be expressed as:
[0032] ,
[0033] ,
[0034] ,
[0035]
[0036] where H t1 , H t2 , H t respectively represent the forward, backward, and hidden states of ABiLSTM Ot Is the output of ABiLSTM, W f xh , W b xh ∈R h×m , W f hh , W b hh ∈R h×h , W hq ∈R q×2h Represents the weight matrix of the model, b f h , b b h ∈R h×η , b q ∈R q×2η Represents the bias term of the model, q Is the final output dimension.
[0037] The above method for measuring the concentration of ammonia nitrogen in effluent based on neural network and Bayesian optimization, the mean square error error of the current time step t Is: , where: y t Is t The true value at time y p Is t The predicted value at time
[0038] The above method for measuring the concentration of ammonia nitrogen in effluent based on neural network and Bayesian optimization, in the step 4, the parameter combination to be optimized by the model Θ Is: the number of layers and hidden units of the ABiLSTM layer, the compression ratio of SE r And the L2 regularization parameter, by optimizing the L2 regularization parameter to optimize the weight matrix of the ABiLSTM layer and the TCL layer W And the bias matrix b .
[0039] The above method for measuring the concentration of ammonia nitrogen in effluent based on neural network and Bayesian optimization, the step 4 includes:
[0040] Step 4-1: Initialize the prior distribution of the Gaussian process and select the initial candidate point set: , where, m ( Θ p ) is the mean function,Θ p denotes the p th new input matrix, Θ d denotes the set of known input points, k ( Θ p , Θ d ) is the Gaussian kernel function, representing Θ p and Θ d the covariance matrix between each group of data in ;
[0041] Step 4-2: Obtain the observed data by inputting Θ P , and combine with the prior distribution to calculate the formulas for the mean and variance of the posterior distribution of Θ P as follows:
[0042] ,
[0043] ,
[0044] where μ is the mean of the posterior distribution, V is the variance of the posterior distribution, k ( Θ P ) and ( k ( Θ P ) are the observed value vectors calculated by the kernel function, K is the covariance matrix defined by the kernel function K ( Θ P , Θ d ), B is the noise variance, I is the identity matrix, Q is the set of known objective function values;
[0045] Step 4-3: Select multiple parallel evaluation candidate points by maximizing the acquisition function. The parallel adaptive acquisition function is:
[0046] ,
[0047] where D n represents the known observed data set, D n ∈ x , a irepresents the i-th acquisition function, i = 1, 2, 3, k = 1, 2, 3, … n ;
[0048] Step 4 - 4: According to the formulas for the mean and variance of the posterior distribution calculated Θ P update the posterior distribution until the accuracy requirement is met.
[0049] The above - mentioned method for measuring the effluent ammonia - nitrogen concentration based on neural network and Bayesian optimization, the acquisition function includes:
[0050] In the initial stage of optimization, due to limited information of the objective function, the upper confidence bound UCB function is adopted to enhance the exploration ability, quickly cover the parameter space and identify potential advantageous regions. The UCB function is: , where μ ( Θ P ) is the predicted mean of the objective function at the point Θ P , v ( Θ P ) is the predicted standard deviation at the point Θ P , f is a constant, which is a parameter to control the balance between exploration and exploitation;
[0051] In the middle stage of optimization, with the accumulation of information, it turns to the expected improvement EI function to balance exploration and exploitation. The EI function is: , where Θ P + is the input vector corresponding to the currently known optimal parameter;
[0052] In the late stage of optimization, the probability of improvement PI function is adopted, focusing on fine - searching in the identified potential regions to achieve precise adjustment of the optimal solution. The PI function is: , where P is a probability measure, γ is a non - negative parameter to control the exploration intensity.
[0053] The beneficial effects of a method for measuring the concentration of ammonia nitrogen in effluent based on neural network and Bayesian optimization are as follows: By synergistically extracting multi-scale features and strengthening key features through the temporal convolutional layer and SE attention mechanism, the accuracy of predicting the concentration of ammonia nitrogen in effluent is significantly improved; The adaptive adjustment of the long short-term memory neural network is used to dynamically adjust the input step size to adapt to the temporal patterns of different working conditions; The improved parallel adaptive Bayesian optimization algorithm is combined to improve the parameter search efficiency. At the same time, through ABiLSTM, the complete capture of the time scale from minutes to seasons is realized, the generalization ability of the model is improved, and high prediction accuracy can be maintained under different working conditions, providing reliable support for the intelligent regulation of sewage treatment plants. Through normalization processing, it is ensured that features with different dimensions are comparable, and at the same time, the over-influence of some large-value features on model training is avoided. Through the parallel Bayesian optimization algorithm, the acquisition function can be dynamically adjusted according to different stages of the optimization process, improving the coverage rate of the parameter space, enhancing the optimization efficiency, and this method speeds up the exploration speed of the parameter space by evaluating multiple candidate points simultaneously, enhancing the generalization ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 It is a schematic diagram of the TCL-SE-ABiLSTM model structure in an embodiment of the present invention;
[0055] Figure 2 It is a schematic diagram of the ABiLSTM structure in an embodiment of the present invention;
[0056] Figure 3 It is a schematic diagram of the predicted result of ammonia nitrogen in effluent in summer in an embodiment of the present invention;
[0057] Figure 4 It is a schematic diagram of the predicted result of ammonia nitrogen in effluent in winter in an embodiment of the present invention; DETAILED DESCRIPTION OF THE INVENTION
[0058] To enable those skilled in the art to better understand the technical solutions of the present invention, the technical solutions of the present invention will be described below in conjunction with the specific embodiments and the accompanying drawings.
[0059] Embodiment 1
[0060] In the field of predicting the concentration of ammonia nitrogen in the effluent of urban sewage treatment plants, accurately capturing the law of water quality change is a key technical problem, which is specifically manifested as follows.
[0061] (1) Existing models are difficult to synchronously capture the ammonia nitrogen concentration fluctuation characteristics from the minute level to the season level, resulting in a decrease in the prediction accuracy of the long-term trend.
[0062] (2) Traditional methods cannot adaptively adjust the weights of input features, affecting the robustness of the model under extreme weather conditions.
[0063] (3) Existing optimization algorithms have a slow convergence speed when dealing with seasonal data distribution shifts and are difficult to meet the requirements of real-time control.
[0064] (4) The model needs to be retrained under different operating conditions due to differences in data distribution, increasing the operation and maintenance costs.
[0065] To address the above problems, this embodiment proposes an ammonia nitrogen concentration prediction method based on an adaptive bidirectional long short-term memory neural network and an improved parallel Bayesian optimization algorithm. This method collaboratively extracts multi-scale features and strengthens key features through a temporal convolutional layer and an SE attention mechanism; uses an adaptive adjustment long short-term memory neural network to dynamically adjust the input step size to adapt to the temporal patterns of different operating conditions; and combines an improved parallel adaptive Bayesian optimization algorithm to improve the parameter search efficiency. This embodiment can maintain high prediction accuracy under different operating conditions and provide reliable support for the intelligent regulation of wastewater treatment plants.
[0066] As Figure 1 - Figure 2 shown, a method for measuring the effluent ammonia nitrogen concentration based on a neural network and Bayesian optimization includes the following steps.
[0067] Step 1: Construct a basic model with a temporal convolutional layer, an attention mechanism, and a bidirectional long short-term memory neural network, collect feature variables x related to the prediction of the effluent ammonia nitrogen concentration, calculate the Pearson correlation coefficient, retain the feature variables x according to the Pearson correlation coefficient, preprocess the retained feature variables x, and normalize each feature variable x to the interval [0, 1] using the Min-Max method.
[0068] Step 2: Input the feature variables x into the basic model for feature extraction and enhancement. Extract the local dependencies in the time series data through the temporal convolutional layer TCL, and recalibrate the features through the SE attention mechanism to enhance the response of key features and suppress noise.
[0069] Step 3: Pass the feature variables x after feature extraction and enhancement through a bidirectional long short-term memory neural network with an adaptive mechanism to extract the long-term temporal features of the feature sequence, dynamically adjust the input step size to adapt to changes in different time periods in the data, capture the bidirectional time dependencies in the sequence, and output the predicted value.
[0070] Step 4: A parallel Bayesian optimization algorithm dynamically adjusts the acquisition function according to different stages of the optimization process, evaluates multiple candidate points simultaneously, and adaptively acquires the function in parallel to output the measured value.
[0071] Through the feature extraction and enhancement of TCL-SE, the capture of ABiLSTM temporal features, and the optimization of parameters by the Bayesian optimization algorithm, the effluent ammonia nitrogen concentration can be accurately predicted.
[0072] Embodiment 2
[0073] This embodiment is a specific illustration of Embodiment 1.
[0074] As Figure 1 The overall flowchart realizes the real-time prediction of the effluent ammonia nitrogen concentration through TCL-SE combined with ABiLSTM and the adaptive Bayesian optimization algorithm and Figure 2 The detailed structure of ABiLSTM shows the working process of ABiLSTM. A method for measuring the effluent ammonia nitrogen concentration based on neural network and Bayesian optimization includes the following steps.
[0075] Step 1: Collect the characteristic variables related to the prediction of the effluent ammonia nitrogen concentration, and calculate the Pearson correlation coefficient through formula (1), and retain the characteristic variables with |r|>0.6.
[0076] (1),
[0077] where r is the Pearson correlation coefficient, x i and y i are the i th observations of two variables, x m and y m are the means of variables x and y respectively, n is the number of samples. Finally, the dissolved oxygen at the end of aerobic stage, the total suspended solids at the end of aerobic stage, the effluent pH value, the effluent redox potential, and the effluent nitrate nitrogen are selected as the input characteristic variables. The sample data set is x = x 1, x 2,..., x n , and the collected effluent ammonia nitrogen concentration data is expressed as y = y 1, y 2,..., y n , y i is the i-th effluent ammonia nitrogen concentration point, n represents the number of samples of easily measurable variables.
[0078] Step 2: First, perform data cleaning on the collected raw data to handle missing values and outliers. First, comprehensively clean the raw data, focusing on solving the problems of missing values and outliers. For missing values, different filling strategies are adopted according to the data characteristics: for continuous variables such as dissolved oxygen and pH value, the time series interpolation method is used, and the values of adjacent time points before and after are used for reasonable filling; for discrete data, the common values of this feature are used for completion. The detection of outliers combines statistical methods and business experience. By calculating the distribution range (mean ± 3 times the standard deviation) of each feature, the outlier points are identified, and the sliding window median replacement method is used for correction, which not only eliminates the abnormal interference but also maintains the time continuity of the data. After completing the data cleaning, normalize all feature variables, and use the Min - Max method to linearly transform each feature to the interval [0, 1], ensuring the comparability of features with different dimensions and avoiding the excessive influence of some large - value features on model training.
[0079] Step 3: Feature extraction and enhancement. When the feature variables x are input into the model, the model first extracts the local dependencies of the time - series data through the Temporal Convolutional Layer (TCL). The output of the TCL x c can be calculated by the following formula:
[0080] (2),
[0081] where, δ (⋅) is the ReLU activation function, used to introduce non - linear characteristics, W k,c,c′ is the weight matrix of the convolutional kernel, x t+k,c is the value of the input data at time step t + k and input channel c at, b c′ is the bias term, b ∈R c′ .
[0082] Then, recalibrate the features through the SE attention mechanism to enhance the response of key features and suppress noise. Specifically, first perform compression, and use global average pooling to generate a channel descriptor z c′′ , used to emphasize the spatial information of the global distribution. The formula is as follows:
[0083] (3),
[0084] where, xijc '' represents the eigenvalue of the i channel with height j and width c at the H -th W channel. r and M 1 M and M 2 c′′×c′′ / r are the weights of two fully connected layers ([[]] M 1 ∈ R c′′ / r×c′′ ), respectively. Then, the output M 1 M of the excitation calculated by the weights s 1 s and c′′×1 2
[0085] (4) is
[0086] where z is composed of c '' channel descriptors, z ∈ R c′′×1 is the output of the compression step, s ∈ R c′′×1 ), σ and s is the sigmoid activation function. Finally, the recalibration of the channel is performed, and the weight
[0087] obtained after excitation is applied to the original input feature. For each channel, its recalibrated output is
[0088] where s c′′ is the weight of channel c '', x c′′ is the original input feature of channel c '', x s is the output after SE recalibration.
[0089] Step 4: After the data is subjected to feature extraction and enhancement through TCL-SE, the long-term temporal features of the feature sequence are extracted through an adaptive bidirectional long short-term memory neural network (ABiLSTM), and finally the predicted value is output. Specifically, ABiLSTM first processes the input sequence step by step through LSTM units. The core calculations of the LSTM unit include the input gate, forget gate, output gate, and update of the cell state, and its algorithm formulas are as follows:
[0090] (6),
[0091] (7),
[0092] (8),
[0093] (9),
[0094] (10),
[0095] (11),
[0096] Among them, x t ( x t ∈R m×1 ) is the received input vector, h t ( h t ∈R m×1 ) is the hidden state at time t, h t−1 is the hidden state of the previous time step, g t is the new candidate value, f t , i t , o t are the outputs of the forget gate, input gate, and output gate at t time, c t is t the cell state at time, W f , W O , W i , W C ( W ∈Rh×m ) represent respectively t the connection input at a moment x t to the weight matrices of different gates, b f , b i , b C , b O ( b ∈R h×1 ) is the bias term, σ and is the sigmoid activation function.
[0097] Then, by defining a window with a shortest length of Δ t min , the input step size is dynamically adjusted based on the error, and only the data of the time steps contained within the window is processed each time. The window size dynamically adjusted according to the prediction error is:
[0098] (12),
[0099] where, α is the adjustment coefficient, and error t is the mean square error of the current time step. The calculation formula of error t is:
[0100] (13),
[0101] where: y t is t the true value at the moment, y p is t the predicted value at the moment.
[0102] Finally, the bidirectional temporal features of the data are captured through forward and backward propagation. For the input x t ∈R m×η , the specific operation process of ABiLSTM can be expressed as:
[0103] (14),
[0104] (15),
[0105]
[0106] (16),
[0107] where, Ht1 , H t2 , H t represent the forward, backward, and hidden states of ABiLSTM respectively. O t is the output of ABiLSTM. W f xh , W b xh ∈R h×m , W f hh , W b hh ∈R h×h , W hq ∈R q×2h represents the weight matrix of the model. b f h , b b h ∈R h×η , b q ∈R q×2η represents the bias term of the model. q is the final output dimension, and T is the transpose.
[0108] Step 5: Use the improved parallel adaptive Bayesian optimization algorithm for parameter optimization. In this embodiment, an improved parallel Bayesian optimization algorithm with an adaptive acquisition function is proposed. This algorithm can dynamically adjust the acquisition function according to different stages of the optimization process, improve the optimization efficiency, and by simultaneously evaluating multiple candidate points, it speeds up the exploration speed of the parameter space and enhances the generalization ability of the model. The parameter combination to be optimized by the model Θ is: the number of layers and hidden units of the ABiLSTM layer, the compression ratio of SE r and the L2 regularization parameter. By optimizing the L2 regularization parameter, the weight matrices W and bias matrices b of the ABiLSTM layer and the TCL layer are optimized.
[0109] Step 6: Evaluate the model through two indicators: the normalized root mean square error (NRMSE) and the coefficient of determination (R 2 ):
[0110] (24),
[0111] (25),
[0112] Among them, y max is the maximum value in the observed data; y min is the minimum value in the observed data; K is the total number of observed values; y j is the average value of the actual observed values.
[0113] Through the feature extraction and enhancement of TCL-SE, the capture of ABiLSTM time series features, and the optimization of parameters by the Bayesian optimization algorithm, the ammonia nitrogen concentration in the effluent can be accurately predicted.
[0114] Example 3
[0115] As Figure 1 - Figure 4 shown, a method for measuring the ammonia nitrogen concentration in the effluent based on a neural network and Bayesian optimization includes the following steps.
[0116] Step 1: Data collection. The data related to the ammonia nitrogen concentration in the effluent used in this example are all from a sewage treatment plant in Beijing, with time spans from January 16, 2016 to January 22, 2016 and from September 16, 2016 to September 22, 2016 respectively. The data sampling interval is 1h, and there are a total of 1064 samples. Calculate the Pearson correlation coefficient through formula (1), and retain the feature variables with |r|>0.6.
[0117] (1),
[0118] Among them, r is the Pearson correlation coefficient, x i and y i are the i th observed values of two variables, x m and y m are the means of variables x and y respectively, n is the number of samples. Finally, the dissolved oxygen at the end of aerobic treatment, the total suspended solids at the end of aerobic treatment, the effluent pH value, the effluent redox potential, and the effluent nitrate nitrogen are selected as input feature variables. The sample data set is x = x 1, x 2,..., x n , and the collected ammonia nitrogen concentration data in the effluent is expressed as y = y 1, y 2,...,y n , y i is the i-th effluent ammonia nitrogen concentration point, n represents the number of samples of easily measurable variables.
[0119] Step 2: Data processing and feature selection. First, clean the collected raw data to handle missing values and outliers. Secondly, use the Min - Max method to normalize each feature to the interval [0, 1] finally.
[0120] Step 3: Feature extraction and enhancement. When the feature variable x is input into the model, the model first extracts the local dependencies in the time - series data through the Temporal Convolutional Layer (TCL). The Temporal Convolutional Layer can efficiently capture the mutual influence between each time step in the time series through one - dimensional convolution operations, especially able to identify multi - scale features in the sequence. This local feature extraction mechanism enables the model to automatically learn the relationships between each time point when processing complex data with time - correlation, thus better grasping the dynamic changes in the data.
[0121] The output of TCL x c can be calculated by the following formula: (2),
[0122] where, δ (⋅) is the ReLU activation function, used to introduce non - linear characteristics, W k,c,c′ is the weight matrix of the convolutional kernel, x t+k, c is the value of the input data at time step t + k and input channel c at, b c′ is the bias term, b ∈R C′ .
[0123] Then, the model recalibrates the features through the SE attention mechanism. The SE attention mechanism can effectively enhance the response of key features while suppressing those unimportant or noisy features by learning the non - linear relationships between channels. This process, through the compression of global information and the recalibration of channels, enables the model to focus on the features with high influence on the prediction task while ignoring the redundant information with less impact on the result. Specifically, first, perform compression, use global average pooling to generate a channel descriptor zc′′ , which is used to emphasize the spatial information of the global distribution.
[0124] The formula is as follows: (3),
[0125] where, x ijc ′′ represents the eigenvalue of the i width of j and the c ′′th channel at the height of H and W represent the height and width of the input channel respectively. Secondly, SE reduces the parameters and complexity by using the compression ratio r and increases the non-linearity by using the ReLU activation function. Finally, the weight of each channel is calculated by the sigmoid function to achieve effective recalibration of the features. M 1 and M 2 are the weights of the two fully connected layers ([[]] M 1 ∈ R c′′×c′′ / r , M 2 ∈ R c′′ / r×c′′ ), then the output M 1 and M 2 of the two fully connected layers is calculated as the excitation output s ( s ∈ R c′′×1 ) as:
[0126] (4),
[0127] where, z is composed of c ′′ channel descriptors, z ∈ R c′′×1 is the output of the compression step, s ∈ R c′′×1 , σ is the sigmoid activation function. Finally, the recalibration of the channel is performed, and the weight s obtained after excitation is applied to the original input features. For each channel, its recalibrated output is:
[0128] (5),
[0129] where, s c′′ is the weight of channel c ′′, x c′′ is the original input feature of channel c ′′, x s is the output after SE recalibration.
[0130] Step 4: Prediction of ammonia nitrogen concentration in the effluent in different seasons. After the data is subjected to feature extraction and enhancement through TCL-SE, it then enters the adaptive bidirectional long short-term memory neural network (ABiLSTM) for further processing. When ABiLSTM extracts the long-term temporal features of the feature sequence, it first captures the bidirectional temporal dependencies in the sequence through the BiLSTM structure. Different from the traditional unidirectional LSTM, BiLSTM can take into account the context information of the current moment and its previous and subsequent moments simultaneously, thus effectively capturing the more complex and long-term temporal dependencies in the sequence. This bidirectional feature enables ABiLSTM to more accurately understand the long-term trends and changing patterns of time series data.
[0131] In addition, ABiLSTM also introduces an adaptive mechanism to adapt to the changes in different time periods in the data by dynamically adjusting the input step size. This feedback mechanism based on the prediction error enables the model to capture features more flexibly by continuously adjusting the input step size. Specifically, ABiLSTM first processes the input sequence time step by time step through the LSTM unit. The core calculations of the LSTM unit include the update of the input gate, forget gate, output gate, and cell state, and its algorithm formulas are as follows:
[0132] (6),
[0133] (7),
[0134] (8),
[0135] (9),
[0136] (10),
[0137] (11),
[0138] Among them, x t ( x t ∈R m×1 ) is the received input vector, h t ( h t ∈R m×1 ) is the hidden state at time t, h t−1 is the hidden state of the previous time step, g tis a new candidate value, f t , i t , o t are the outputs of the forget gate, input gate, and output gate at t time, c t is t the cell state at time, W f , W O , W i , W C ( W ∈R h×m ) respectively represent t the weight matrices connecting the input at time x t to different gates, b f , b i , b C , b O ( b ∈R h×1 ) are the bias terms, σ is the sigmoid activation function.
[0139] Then, by defining a window with a minimum length of Δ t min , the input step size is dynamically adjusted based on the error, and only the data within the time steps contained in the window is processed each time. The window size dynamically adjusted according to the prediction error is:
[0140] (12),
[0141] where, α is the adjustment coefficient, and error t is the mean squared error at the current time step. The calculation formula for error t is:
[0142] (13),
[0143] where: y t is t the true value at time, y p is t the predicted value at time.
[0144] Finally, capture the bidirectional temporal features of the data through forward and backward propagation. For the input x t ∈R m×η , the specific operation process of ABiLSTM can be expressed as:
[0145] (14),
[0146] (15),
[0147] , (16),
[0148] Among them, H t1 , H t2 , H t respectively represent the hidden states of forward, backward, and ABiLSTM. O t is the output of ABiLSTM. W f xh , W b xh ∈R h×m , W f hh , W b hh ∈R h×h , W hq ∈R q×2h represents the weight matrix of the model. b f h , b b h ∈R h×η , b q ∈R q×2η represents the bias term of the model. q is the final output dimension. Ø is the tanh activation function.
[0149] Step 5: Optimization algorithm. This embodiment proposes an improved parallel Bayesian optimization algorithm for the adaptive acquisition function. This algorithm can dynamically adjust the acquisition function according to different stages of the optimization process, improve the optimization efficiency, and this method speeds up the exploration speed of the parameter space by simultaneously evaluating multiple candidate points and enhances the generalization ability of the model. The parameter combination to be optimized by the model ΘIt is: the number of layers and the number of hidden units of the ABiLSTM layer, the compression ratio of SE r and the L2 regularization parameter. By optimizing the L2 regularization parameter, the weight matrices of the ABiLSTM layer and the TCL layer are optimized W and the bias matrix b .
[0150] Specifically, first initialize the prior distribution of the Gaussian process and select the initial candidate point set, that is:
[0151] (17),
[0152] wherein, m ( Θ p ) is the mean function, k ( Θ p , Θ d ) is the Gaussian kernel function, and is the p th new input matrix Θ p ( p <= n , n is Θ the total number of groups of) and the known input point set Θ d ( Θ d ∈R m×d ) is the covariance matrix between each group of data in the set.
[0153] By inputting Θ P to obtain the observed data, combined with the prior distribution, calculate the mean and variance of the posterior distribution of Θ P :
[0154] (18),
[0155] (19),
[0156] wherein, μ is the mean of the posterior distribution, V is the variance of the posterior distribution, k ( Θ P )( k ( Θ P )∈R d×1 ) is the observed value vector calculated by the kernel function, K is the kernel function K ( ΘP , Θ d ) The defined covariance matrix ( K ∈R d×d ), B is the noise variance, which is used to simulate the random noise in the observed values of the objective function. I is the identity matrix ( I ∈R d×d ), Q ( Q ∈R d×1 ) is the known set of objective function values. By inputting Θ P to obtain the observed data, combined with the prior distribution, calculate Θ P the mean and variance of the posterior distribution.
[0157] Then, multiple parallel evaluation candidate points are selected by maximizing the acquisition function. The parallel adaptive acquisition function is:
[0158] (20),
[0159] where, D n ∈ x , i = 1, 2, 3, k = 1, 2, 3, … n . By using the acquisition function, the points for which the mean and variance are to be calculated next are selected, and the points that meet the accuracy are finally determined by continuously iterating the acquisition function. Equation (20) means that the left side of the equation represents the entire acquisition function, which is divided into three types, and the right side is divided into three acquisition functions a1, a2, and a3. Specifically, the acquisition functions at different stages are as follows.
[0160] (1) In the initial stage of optimization, due to limited information about the objective function, the upper confidence bound (UCB) function is used to enhance the exploration ability, quickly cover the parameter space, and identify potential advantageous regions. The UCB function is:
[0161] (21),
[0162] where, μ ( Θ P ) is the predicted mean of the objective function at the point Θ P , v ( Θ P ) is the predicted standard deviation at the point Θ P , f is a constant, which is a parameter that controls the balance between exploration and exploitation.
[0163] (2) As information accumulates, shift to the expected improvement (EI) function to balance exploration and exploitation. The EI function is:
[0164] (22),
[0165] where, Θ P + is the input vector corresponding to the currently known optimal parameters.
[0166] (3) In the later stage of optimization, adopt the probability improvement (PI) function, focus on the identified potential areas for fine search to achieve precise adjustment of the optimal solution. The PI function is:
[0167] (23),
[0168] where, P is the probability measure, γ is a non - negative parameter controlling the exploration intensity. Finally, repeat the above steps to obtain the new ([[]] Θ P , f ( Θ P ))). According to equations (18) and (19), update the posterior distribution until the accuracy requirement is met.
[0169] Step 6: Model evaluation. Evaluate the model through two indicators: the normalized root - mean - square error (NRMSE) and the coefficient of determination (R 2 ):
[0170] (24),
[0171] (25),
[0172] where, y max is the maximum value in the observed data; y min is the minimum value in the observed data; K is the total number of observed values; y j is the average value of the actual observed values.
[0173] Table 1 and Table 2 are respectively the comparison of the prediction results of different methods for the effluent ammonia - nitrogen concentration in summer and winter. From Table 1 and Table 2, it can be concluded that the R2 and NRMSE in this embodiment have better prediction effects than other methods both in summer and winter.
[0174] Table 1 Prediction results of different models for the effluent ammonia - nitrogen concentration in summer:
[0175] 。
[0176] Table 2 Prediction results of different models for ammonia nitrogen concentration in winter effluent:
[0177] 。
[0178] Through Figure 3 The online prediction result graph of ammonia nitrogen concentration in summer effluent and Figure 4 The online prediction result graph of ammonia nitrogen concentration in winter effluent demonstrate the high accuracy and generalization ability of this embodiment in predicting ammonia nitrogen concentration in effluent.
[0179] The above embodiments are only for explaining the structural concept and characteristics of the present invention, aiming to enable ordinary technicians in the field to understand the content of the present invention and implement it accordingly, and shall not be used to limit the protection scope of the present invention. Any equivalent changes or modifications made according to the essence of the content of the present invention should be covered within the protection scope of the present invention.
Claims
1. A method for measuring the concentration of ammonia nitrogen in effluent based on neural network and Bayesian optimization, characterized in that, It includes the following steps: Step 1: Construct a basic model with a temporal convolutional layer, an attention mechanism, and a bidirectional long short-term memory neural network. Collect feature variables x related to the prediction of the effluent ammonia nitrogen concentration, calculate the Pearson correlation coefficient, retain the feature variables x according to the Pearson correlation coefficient, preprocess the retained feature variables x, and normalize each feature variable x to the interval [0, 1] using the Min-Max method; Step 2: Input the feature variables x into the basic model for feature extraction and enhancement. Extract local dependencies in the time series data through the temporal convolutional layer TCL, and recalibrate the features through the SE attention mechanism to enhance the response of key features and suppress noise, including: Step 2-1: Extract the local dependencies of the time series data through the Temporal Convolutional Layer (TCL). The output of the TCL x c The calculation formula is: , where δ (⋅) is the ReLU activation function, which is used to introduce non-linearity W k,c,c′ is the weight matrix of the convolutional kernel x t+k,c is the value of the input data at time step t + k and input channel c ; b c′ is the bias term b ∈R c′ ; Step 2-2: Recalibrate the feature variables x through the SE attention mechanism to enhance the response of key features and suppress noise. Perform channel recalibration, and apply the obtained weight s after excitation to the original input features, including: Step a: Perform compression and use global average pooling to generate a channel descriptor z c′′ , which is used to emphasize the spatial information of the global distribution. The formula is: , where x ijc ′′ represents the eigenvalue of the i -th channel at the position with height j and width c ; H and W represent the height and width of the input channel respectively; Step b: The SE attention mechanism reduces the parameters and complexity by using a compression ratio r, and increases the non-linearity by using the ReLU activation function; Step c: Calculate the weights of each channel through the sigmoid function to achieve effective recalibration of features, and calculate the output of the excitation through the weights of two fully connected layers M 1 and M 2 calculate the output of the excitation s as: , where M 1 ∈ R c′′×c′′ / r , M 2 ∈ R c ′′ / r×c′′ , z consists of c '' channel descriptors, z ∈ R c′′×1 is the output of the compression step, s ∈ R c′′×1 , σ is the sigmoid activation function; Step d: Perform channel recalibration and apply the weights obtained after excitation s to the original input features. For each channel, the recalibrated output is: , where s c′′ is the weight of channel c '', x c′′ is the original input feature of channel c '', x s is the output after SE recalibration; Step 3: Extract the long-term temporal features of the feature sequence from the feature variables x after feature extraction and enhancement through a bidirectional long short-term memory neural network with an adaptive mechanism, dynamically adjust the input step size to adapt to changes in different time periods in the data, capture the bidirectional time dependencies in the sequence, and output the predicted value; Step 4: Parallel Bayesian optimization algorithm, dynamically adjust the acquisition function according to different stages of the optimization process, evaluate multiple candidate points simultaneously, and parallel adaptive acquisition function to output the measured value.
2. The method for measuring the effluent ammonia nitrogen concentration based on neural network and Bayesian optimization according to claim 1, wherein In step 1, the calculation formula for the Pearson correlation coefficient is as follows: , where r is the Pearson correlation coefficient, x i and y i are the i th observed values of two variables, x m and y m are the means of variables x and y respectively, n is the number of samples, and the characteristic variables include the finally selected dissolved oxygen at the aerobic end, total suspended solids at the aerobic end, effluent pH value, effluent oxidation-reduction potential, and effluent nitrate nitrogen.
3. The method for measuring the ammonia nitrogen concentration in the effluent based on neural network and Bayesian optimization according to claim 1, wherein, The preprocessing includes first performing data cleaning on the collected raw data to handle missing values and outliers. For the missing values, different filling strategies are adopted according to the data characteristics. For continuous variables, the time series interpolation method is used to reasonably fill with the values of adjacent time points before and after; for the outliers, combined with statistical methods and business experience, the outliers are identified by calculating the distribution range of each feature, and corrected by replacing with the median of the sliding window.
4. The method for measuring the ammonia nitrogen concentration in the effluent based on neural network and Bayesian optimization according to claim 1, characterized in that, The said Step 3 includes: Step 3-1: The bidirectional long short-term memory neural network ABiLSTM processes the input sequence step by step through the LSTM unit. The LSTM unit includes the update of the input gate, forget gate, output gate, and cell state. The algorithm formula is as follows: , , , , , , Among them, x t ( x t ∈R m×1 ) is the received input vector, h t ( h t ∈R m×1 ) is the hidden state at time t, h t−1 is the hidden state of the previous time step, g t is the new candidate value, f t , i t , o t are the outputs of the forget gate, input gate, and output gate at t time step respectively, c t is t the cell state at time step, W f , W O , W i , W C ( W ∈R h×m ) respectively represent t the weight matrices connecting the input x t to different gates at time step, b f , b i , b C , b O ( b ∈R h×1 ) are the bias terms; Step 3-2: Define a window with a minimum length of Δ t min , dynamically adjust the input step size based on the error, and only process the data within the time steps contained in the window each time. The dynamically adjusted window size for the prediction error is: , where α is the adjustment coefficient, and error t is the mean square error of the current time step; Step 3-3: Capture the bidirectional temporal features of the data through forward and backward propagation. For the input x t ∈R m×η , the specific operation process of ABiLSTM can be expressed as: , , , , where H t1 , H t2 , H t represent the forward, backward, and hidden states of the ABiLSTM respectively, O t is the output of the ABiLSTM, W f xh , W b xh ∈R h×m , W f hh , W b hh ∈R h×h , W hq ∈R q×2h denotes the weight matrix of the model, b f h , b b h ∈R h×η , b q ∈R q×2η denotes the bias term of the model, q is the final output dimension.
5. The method for measuring the ammonia nitrogen concentration in the effluent based on neural network and Bayesian optimization according to claim 4, characterized in that, The mean square error error at the current time step t is as follows: , where: y t is t the true value at time y p is t the predicted value at time 6. The method for measuring the ammonia nitrogen concentration in the effluent based on neural network and Bayesian optimization according to claim 1, characterized in that, In the step 4, the parameter combinations to be optimized by the model Θ are: the number of layers and hidden units of the ABiLSTM layer, the compression ratio of the SE r and the L2 regularization parameter. The weight matrices of the ABiLSTM layer and the TCL layer are optimized by optimizing the L2 regularization parameter W and the bias matrices b .
7. The method for measuring the ammonia nitrogen concentration in the effluent based on neural network and Bayesian optimization according to claim 6, characterized in that, The said Step 4 includes: Step 4-1: Initialize the prior distribution of the Gaussian process and select the initial candidate point set: , where m ( Θ p ) is the mean function, Θ p represents the p -th new input matrix, Θ d represents the set of known input points, k ( Θ p , Θ d ) is the Gaussian kernel function, representing Θ p and Θ d the covariance matrix between each pair of data in; Step 4-2: By inputting Θ P obtain the observed data, and combine with the prior distribution to calculate Θ P The formulas for the mean and variance of the posterior distribution of are respectively: , , Among them, μ is the mean of the posterior distribution, V is the variance of the posterior distribution, k ( Θ P ) and ( k ( Θ P ) are the vectors of observed values calculated through the kernel function, K is the covariance matrix defined by the kernel function K ( Θ P , Θ d ), B is the noise variance, I is the identity matrix, Q is the set of known objective function values; Step 4-3: Select multiple parallel evaluation candidate points by maximizing the acquisition function. The parallel adaptive acquisition function is: , wherein, D n represents a known observed data set, D n ∈ x , a i represents the i-th acquisition function, i i = 1, 2, 3, k i = 1, 2, 3, … n ; Step 4-4: Update the posterior distribution according to the formulas for the mean and variance of the posterior distribution obtained by calculation Θ P until the accuracy requirement is met.
8. The method for measuring the ammonia nitrogen concentration in the effluent based on neural network and Bayesian optimization according to claim 7, characterized in that, The said acquisition function includes: In the initial stage of optimization, due to limited information about the objective function, the Upper Confidence Bound (UCB) function is adopted to enhance the exploration ability, quickly cover the parameter space, and identify potential advantageous regions. The UCB function is as follows: , where μ ( Θ P ) is the predicted mean of the objective function at the point Θ P , v ( Θ P ) is the predicted standard deviation at the point Θ P , f is a constant parameter that controls the balance between exploration and exploitation; In the mid-term of optimization, as information accumulates, the EI function is expected to be improved to balance exploration and exploitation. The EI function is: , where Θ P + is the input vector corresponding to the currently known optimal parameters; In the later stage of optimization, the Probability Improvement (PI) function is adopted to focus on the identified potential areas for fine search to achieve precise adjustment of the optimal solution. The PI function is as follows: , where P is the probability measure, γ is a non - negative parameter that controls the exploration intensity.
Citation Information
Patent Citations
BiLSTM voltage deviation prediction method based on Bayesian optimization
CN113554148A
Convolutional layer and self-attention mechanism-based effluent ammonia nitrogen concentration measurement method
CN118228766A